Artificial Intelligence in Material Science
artificial-intelligence-in-material-science
Syllabus
Faculty: Dr. Satyanarayan Dhal
Code(Credit) : CUTM 3164(0-2-2)
Course Objectives:
1. To impart hands-on proficiency in Python-based workflows for acquiring, cleaning, exploring, and featurizing materials data from open scientific databases (Materials Project, AFLOW, NIST-JANAF).
2. To develop the ability to build, optimize, evaluate, and interpret classical machine learning and deep learning models (Ridge, RF, XGB, MLP, CNN, GNN, GPR, Bayesian Optimization) for materials property prediction and design.
3. To cultivate research-grade professional practices — reproducibility, version control, explainability, and scientific communication — through a structured capstone project delivered as a documented GitHub portfolio.
Upon successful completion of this course, students will be able to:
Course Syllabus
Environment Setup: Anaconda & Virtual Environments; GitHub Initialization & Python Warm-up; First Visualization & Git Workflow; Advanced pandas for Materials Data; Data Types, Encoding & Git Branching; Repository Documentation & Peer Review
Materials Project API: Setup & Authentication; Querying & Exploring Binary Oxides; Structured Queries, Merging & Export; Data Profiling & Missing-Value Strategies; Duplicates & Outlier Detection; Domain-Driven Cleaning & Correlation Analysis;
NIST-JANAF Cp Dataset: Loading & Cleaning;
Temperature Dependence & Group Analysis;
Publication-Quality Visualization
Module 3:
Feature Engineering & Descriptors
Why Features? Elemental Property Statistics; CBFV Generation & Inspection
Feature Correlations & Chemical Similarity; Matminer Featurization (Magpie & Deml); Feature Selection I: Variance & Correlation Filters; Feature Selection II: RFE & Importance; Structural Descriptors; Sine Coulomb Matrix & Custom Features; Combined Feature Matrix & Quick Benchmark
Baseline Linear Regression & Ridge; Lasso, ElasticNet & Sparsity; Diagnostics: Parity & Residual Plots; Support Vector Regression Basics; SVR Hyperparameter Tuning; Feature Scaling & Model Trade-offs; Decision Trees: Build & Interpret
Overfitting & Tree Feature Importance; Random Forests & OOB Validation; Gradient Boosting; XGBoost & Optuna Optimization; Grand Model Comparison & PBL Sprint 1
Cross-Validation Strategies; Data Leakage & sklearn Pipelines; Learning Curves & GroupKFold; SHAP: Global Explanations; SHAP: Local Explanations; LIME & Physical Interpretation
Neural Network Fundamentals & Data Pipeline; Building & Training an MLP
Regularization, Ablation & Benchmark; Microstructure Images & Augmentation
CNN from Scratch; Transfer Learning & Grad-CAM; Graphs & CGCNN Concepts
Building Graphs & a Toy GCN; GCN on Materials Project Data
Gaussian Process Regression
Gaussian Process Regression; Uncertainty Bands & Kernel Choice; Calibration & the Active-Learning Link; Bayesian Optimization: Theory & Setup; Running BO with the Ax Platform; BO vs. Random Search & Realism Check; Reproducibility Standards & Repo Audit
Portfolio Reports & Master README
Peer Portfolio Review & Reflection
Description:
The project component now comprises 30 supervised 1-hour sessions (15 classes of 2 hours), organized into 10 phases of 3 sessions each. All hours are faculty-supervised execution time; independent reading, data sourcing, and report drafting continue between sessions and are evidenced in the logbook.
Code(Credit) : CUTM 3164(0-2-2)
Course Objectives:
1. To impart hands-on proficiency in Python-based workflows for acquiring, cleaning, exploring, and featurizing materials data from open scientific databases (Materials Project, AFLOW, NIST-JANAF).
2. To develop the ability to build, optimize, evaluate, and interpret classical machine learning and deep learning models (Ridge, RF, XGB, MLP, CNN, GNN, GPR, Bayesian Optimization) for materials property prediction and design.
3. To cultivate research-grade professional practices — reproducibility, version control, explainability, and scientific communication — through a structured capstone project delivered as a documented GitHub portfolio.
Course Outcomes (COs)
Upon successful completion of this course, students will be able to:
CO1:
Explain the theoretical foundations of supervised, unsupervised, and deep learning as applied to materials property prediction.
CO2:
Acquire, clean, and explore materials datasets from open databases (Materials Project, AFLOW, NIST-JANAF) and generate, validate, and select compositional and structural feature vectors using CBFV and Matminer.
CO3:
Build, train, and optimize classical ML (Ridge, RF, XGB) and deep learning (MLP, CNN, GNN) models for materials property regression and classification.
CO4:
Evaluate models using rigorous cross-validation, interpret predictions with SHAP/LIME, and apply Gaussian Process Regression and Bayesian Optimization to materials design problems.
CO5:
Document, version-control, and present AI-driven materials research results in reproducible GitHub repositories, reports, and presentations.
Course Syllabus
Module 1:
Foundations & Environment
:
Environment Setup: Anaconda & Virtual Environments; GitHub Initialization & Python Warm-up; First Visualization & Git Workflow; Advanced pandas for Materials Data; Data Types, Encoding & Git Branching; Repository Documentation & Peer Review
Module 2:
Materials Data Acquisition & EDA
Materials Project API: Setup & Authentication; Querying & Exploring Binary Oxides; Structured Queries, Merging & Export; Data Profiling & Missing-Value Strategies; Duplicates & Outlier Detection; Domain-Driven Cleaning & Correlation Analysis;
NIST-JANAF Cp Dataset: Loading & Cleaning;
Temperature Dependence & Group Analysis;
Publication-Quality Visualization
Module 3:
Feature Engineering & Descriptors
Why Features? Elemental Property Statistics; CBFV Generation & Inspection
Feature Correlations & Chemical Similarity; Matminer Featurization (Magpie & Deml); Feature Selection I: Variance & Correlation Filters; Feature Selection II: RFE & Importance; Structural Descriptors; Sine Coulomb Matrix & Custom Features; Combined Feature Matrix & Quick Benchmark
Module 4:
Classical & Ensemble ML Models
Baseline Linear Regression & Ridge; Lasso, ElasticNet & Sparsity; Diagnostics: Parity & Residual Plots; Support Vector Regression Basics; SVR Hyperparameter Tuning; Feature Scaling & Model Trade-offs; Decision Trees: Build & Interpret
Overfitting & Tree Feature Importance; Random Forests & OOB Validation; Gradient Boosting; XGBoost & Optuna Optimization; Grand Model Comparison & PBL Sprint 1
Module 5:
Model Evaluation & Explainability
Cross-Validation Strategies; Data Leakage & sklearn Pipelines; Learning Curves & GroupKFold; SHAP: Global Explanations; SHAP: Local Explanations; LIME & Physical Interpretation
Module 6:
Deep Learning for Materials
Neural Network Fundamentals & Data Pipeline; Building & Training an MLP
Regularization, Ablation & Benchmark; Microstructure Images & Augmentation
CNN from Scratch; Transfer Learning & Grad-CAM; Graphs & CGCNN Concepts
Building Graphs & a Toy GCN; GCN on Materials Project Data
Module 7:
Advanced Methods: UQ, Bayesian Optimization
Gaussian Process Regression
Gaussian Process Regression; Uncertainty Bands & Kernel Choice; Calibration & the Active-Learning Link; Bayesian Optimization: Theory & Setup; Running BO with the Ax Platform; BO vs. Random Search & Realism Check; Reproducibility Standards & Repo Audit
Portfolio Reports & Master README
Peer Portfolio Review & Reflection
Project Work (20 Hours)
Description:
The project component now comprises 30 supervised 1-hour sessions (15 classes of 2 hours), organized into 10 phases of 3 sessions each. All hours are faculty-supervised execution time; independent reading, data sourcing, and report drafting continue between sessions and are evidenced in the logbook.
5.1 Project Tracks (Student Choice)
Track | Description | Deliverable |
Track A: Property Prediction | Build and compare ML models to predict a materials property (bandgap, formation energy, bulk modulus, hardness) using a real database. | Full pipeline: data → features → 3+ models → evaluation → SHAP analysis → GitHub repo. |
Track B: Classification & Phase Prediction | Classify material phases, stability, crystal system, or microstructure category from compositional/structural features or images. | Classification pipeline with ROC curves, confusion matrix, feature importance, SHAP, GitHub repo. |
Track C: Advanced / Research-Grade | Implement an advanced method (GNN, Bayesian Optimization, GPR, active learning) on a materials problem of choice. For high-performing students. | Advanced implementation, ablation study, classical baseline comparison, full documentation, GitHub repo. |
Resources & References
8.1 Textbooks
ID | Title | Author(s) | Publisher / Year |
T1 | Machine Learning: A Probabilistic Perspective | Kevin Murphy | MIT Press, 2012 |
T2 | Deep Learning | Goodfellow, Bengio, Courville | MIT Press, 2016 |
T3 | Hands-On Machine Learning with Scikit-Learn, Keras & TensorFlow | Aurélien Géron | O'Reilly, 3rd Ed. |
T4 | Materials Modelling using Density Functional Theory | Feliciano Giustino | Oxford UP, 2014 |
T5 | Interpretable Machine Learning | Christoph Molnar | Online (free), 2023 |
Session plan & materials
No materials published yet.
