A high-throughput machine learning pipeline designed to predict molecular properties, including drug toxicity, with high accuracy and granular SHAP interpretability.
Interpret model predictions with TreeSHAP visualisations for direct insights into molecular substructure contributions.
Handle large chemical compound datasets efficiently with memory-optimised Parquet / PyArrow partitioning.
Achieved RMSE 0.1577 and R² 0.99 on validation benchmarks.
Feature attribution analysis isolates top molecular descriptors responsible for toxic chemical pathways.