Presentations

2027

TMS 2027
AlloyBot self-driving laboratory to develop high temperature alloys 100x faster
2027 TMS Annual Meeting & Exhibition, Symposium: AI-Enabled Materials Processing, 2027
Abstract
We present our AlloyBot self-driving laboratory currently being built at UW-Madison to accelerate the development of high temperature alloys by 100x. AlloyBot uses robotics, automation, and AI to experimentally test and optimize new alloy materials with minimal human assistance. As first building block, we previously presented the automatic arc-melting synthesis module, which can produce over 150 bulk alloys per week. Here, we present subsequent modules which enable automatic sample preparation, processing, microstructure analysis, and performance testing. Artificial Intelligence is integrated to enable real-time data processing and closed-loop alloy design.
TMS 2027
AlloyBot automated SEM-XRD pipeline for high-throughput microstructure characterization
2027 TMS Annual Meeting & Exhibition, Symposium: AI-Enabled Materials Processing, 2027
Abstract
Accelerated alloy development depends not only on rapid synthesis but also on characterization throughput that can keep pace. Conventional microstructure analysis remains largely manual, creating a bottleneck for data-driven design in self-driving laboratories. Following our AlloyBot automatic arc-melting synthesis, as next building block we here present a pipeline automating acquisition and analysis through X-ray diffraction (XRD) integrated with scanning electron microscopy (SEM) and energy-dispersive X-ray spectroscopy (EDS). An efficient human-in-the-loop learning framework is used, in which expert oversight initially guides microstructure interpretation and is reduced as confidence in the system is established. Further, the workflow couples diffraction-based phase information with spatially resolved compositional and microstructural data, such that each modality refines the interpretation of the other. This yields consistent, machine-readable outputs suited for model training and rapid processing-structure-property assessment. We discuss the architecture of the pipeline and its role as a characterization backbone for increasingly autonomous experimental workflows.

2026

AIChE 2026
Adaptive semantic segmentation of alloy microstructures in scanning electron microscopy via foundation model embeddings and uncertainty-guided annotation
2026 AIChE Annual Meeting, 10E Interactive Session: Data and Information Systems, 2026 Poster, presenting author
Abstract
Scanning electron microscopy (SEM) is central to alloy microstructure characterization and is routinely used to study phase structure, morphology, and processing-induced microstructural variation in metallic systems [1,2]. However, automating segmentation of SEM images remains difficult in exploratory materials workflows. Conventional supervised segmentation models such as U-Net and related task-specific deep networks can perform well when large labeled datasets are available, but they typically require substantial annotation effort, assume a fixed set of predefined classes, and are not well suited to settings in which newly acquired images contain previously unseen microstructural appearances [3]. More recent foundation-model-based segmentation tools such as Segment Anything [4] and SAM 2 [5] provide powerful “promptable” and “zero-shot” segmentation capabilities, while interactive microscopy tools such as Cellpose 2.0 [6] and μSAM [7] have helped reduce manual effort in related imaging settings. Yet an important challenge remains unresolved for alloy characterization, i.e., how can we efficiently learn and refine scientifically meaningful semantic labels over time as new SEM images arrive, without repeatedly retraining a large segmentation model whenever new visual variants appear? To address this challenge, we propose a human-in-the-loop segmentation framework for SEM microstructure analysis that combines (large-scale) pretrained vision foundation model embeddings with uncertainty-guided expert annotation. Our approach uses dense patch-level embeddings extracted from DINOv3 [8] to provide transferable visual representations for SEM images. A multi-class classifier is then trained on embeddings extracted from a small set of expert-annotated regions corresponding to the semantic microstructural concepts of interest. Predictive entropy from the classifier is used to identify uncertain regions and focus additional expert labeling on the most informative image locations. This creates an iterative workflow in which expert effort is targeted rather than exhaustive, and previously labeled examples are retained and reused over time through a growing feature bank. A key aspect of the framework is that adaptation occurs in the pretrained embedding space rather than through repeated fine-tuning of the underlying vision model. As a result, the workflow remains computationally lightweight while still allowing the segmentation behavior to evolve as new examples are observed. This is especially important for exploratory alloy characterization, where objects belonging to the same scientifically meaningful class may vary substantially in contrast, brightness, texture, or geometry across images. In our framework, segmentation is defined primarily by the semantic concepts conveyed through expert-provided examples rather than by fixed assumptions tied to low-level image appearance alone. We demonstrate the approach on sequential alloy SEM images, where new images are processed over time and the labeled feature bank grows as expert corrections are incorporated. In this setting, the framework becomes progressively more effective as more images are analyzed: annotation effort decreases when previously seen microstructure classes reappear, while the method remains flexible enough to incorporate newly observed classes when needed. More broadly, these results suggest that lightweight supervised adaptation of dense foundation-model embeddings can progressively learn to reproduce expert microstructure labeling behavior and provide a practical path toward efficient, robust, and adaptive SEM segmentation under limited annotation budgets. The overall framework is general and may be useful well beyond alloy microscopy, including in other scientific imaging workflows where semantic classes evolve over time and expert-guided refinement is essential. References [1] Dong, Zhichao, et al. "Microstructural evolution and characterization of AlSi10Mg alloy manufactured by selective laser melting." Journal of Materials Research and Technology 17 (2022): 2343-2354. [2] Yuan, Xu, et al. "Microstructures and mechanical properties of cast Al-Mg-Si alloy with combined addition of Sc and Zr." Journal of Materials Research and Technology 35 (2025): 5097-5106. [3] Ronneberger, Olaf, Philipp Fischer, and Thomas Brox. "U-net: Convolutional networks for biomedical image segmentation." International Conference on Medical image computing and computer-assisted intervention. Cham: Springer international publishing, 2015. [4] Kirillov, Alexander, et al. "Segment anything." Proceedings of the IEEE/CVF international conference on computer vision. 2023. [5] Ravi, Nikhila, et al. "SAM 2: Segment anything in images and videos." arXiv preprint arXiv:2408.00714 (2024). [6] Pachitariu, Marius, and Carsen Stringer. "Cellpose 2.0: how to train your own model." Nature methods 19.12 (2022): 1634-1641. [7] Archit, Anwai, et al. "Segment anything for microscopy." Nature methods 22.3 (2025): 579-591. [8] Siméoni, Oriane, et al. "Dinov3." arXiv preprint arXiv:2508.10104 (2025).
AIChE 2026
Beyond optimization in molecular dynamics: goal-directed Bayesian design with trajectory-based noise modeling
2026 AIChE Annual Meeting, 10D: Advances in Computational Methods and Numerical Analysis, 2026 Presenting author
Abstract
Molecular dynamics (MD) simulations are a powerful tool for probing molecular-scale structure and dynamics, but systematically exploring simulation inputs such as composition, temperature, or solvent conditions quickly becomes computationally prohibitive. This challenge is amplified when the target observables are estimated from finite, autocorrelated trajectories, so the magnitude of the statistical noise varies across the input space. In addition, in many MD settings, the scientific objective is not simply to optimize a single scalar quantity, but rather to identify broader system-level features such as threshold regions, level sets, and transition boundaries. To address these challenges, we present MD-BAX [1], a goal-directed Bayesian design framework for molecular simulation campaigns that builds on Bayesian Algorithm Execution (BAX) [2] and its posterior-sampling variant [3], while explicitly accounting for input-dependent noise through trajectory statistics. Rather than assuming constant observation noise, MD-BAX estimates the uncertainty at each simulated input directly from the underlying MD trajectory using autocorrelation-based analysis and incorporates these estimates into a fixed-noise Gaussian process surrogate (see, e.g., [4] for more details on input-dependent Gaussian process models). The resulting surrogate provides more reliable uncertainty quantification for guiding adaptive sampling toward the scientific objective of interest. We demonstrate the framework on coarse-grained simulations of amphiphilic block copolymers, using the radius of gyration as the primary observable. In particular, we consider tasks involving identification of threshold-exceeding regions and mapping coil-to-globule transition behavior as a function of polymer composition [5] and solvent quality [6]. In both cases, MD-BAX concentrates simulations in the most informative regions of parameter space and recovers physically meaningful structures such as level-set boundaries and transition manifolds with substantially fewer simulations than non-adaptive sampling. These gains are closely tied to improved uncertainty calibration, consistent with recent work emphasizing the importance of evaluating predictive uncertainty rather than mean accuracy alone [7]. Overall, MD-BAX provides a practical route to closed-loop, goal-directed design in molecular simulation. More broadly, this work shows that observation-noise modeling is not merely a technical detail in adaptive MD workflows; when the objective is to learn system behavior from noisy trajectories, propagating trajectory-derived uncertainty through the design loop can materially improve where computational effort is spent. Because the framework is modular, it can be extended readily to a broad range of molecular simulation settings in which observables are estimated from noisy trajectories and the goal is to infer “scientifically meaningful” system behavior rather than identify a single optimal condition. References [1] Tan, Tianhong, et al. "MD-BAX: A general-purpose Bayesian design framework for molecular dynamics simulations with input-dependent noise." The Journal of Chemical Physics 164.8 (2026). [2] Neiswanger, Willie, Ke Alexander Wang, and Stefano Ermon. "Bayesian algorithm execution: Estimating computable properties of black-box functions using mutual information." International Conference on Machine Learning. PMLR, 2021. [3] Cheng, Chu X., et al. "Practical bayesian algorithm execution via posterior sampling." Advances in Neural Information Processing Systems 37 (2024): 135186-135207. [4] Goldberg, Paul, Christopher Williams, and Christopher Bishop. "Regression with input-dependent noise: A Gaussian process treatment." Advances in neural information processing systems 10 (1997). [5] Van den Oever, J. M. P., et al. "Coil-globule transition for regular, random, and specially designed copolymers: Monte Carlo simulation and self-consistent field theory." Physical Review E 65.4 (2002): 041708. [6] Spaeth, Justin R., Ioannis G. Kevrekidis, and Athanassios Z. Panagiotopoulos. "A comparison of implicit-and explicit-solvent simulations of self-assembly in block copolymer and solute systems." The Journal of chemical physics 134.16 (2011). [7] Rasmussen, Maria H., et al. "Uncertain of uncertainties? A comparison of uncertainty quantification metrics for chemical data sets." Journal of Cheminformatics 15.1 (2023): 121.
AIChE 2026
A modular generate-then-optimize framework for sample-efficient multi-objective molecular design
2026 AIChE Annual Meeting, 10E Advances in Machine Learning and Intelligent Systems I, 2026
Abstract
Designing molecules that satisfy multiple competing objectives remains a central challenge in molecular discovery, particularly when each evaluation requires expensive simulation or experiment testing [1]. This difficulty is amplified by the enormous size of chemical space, which makes exhaustive search impossible even for relatively restricted molecular families. Bayesian optimization (BO) provides a principled framework for sample-efficient search, while generative models offer a route to move beyond fixed molecular libraries. However, many existing methods tightly couple generation and optimization through continuous latent spaces [2], which can create architectural entanglement, invalid or redundant decoded structures, and practical difficulties in performing scalable batch acquisition. In addition, popular parallel multi-objective acquisition strategies such as batch expected hypervolume improvement (qEHVI) can become increasingly expensive to optimize as batch size grows [3]. In this work, we present a modular “generate-then-optimize” framework for de novo multi-objective molecular design that decouples candidate generation from uncertainty-aware acquisition [4]. In the first stage, any generative engine (such as a genetic algorithm, variational autoencoder, diffusion model, or reinforcement learning policy) can be used to construct a large and diverse pool of candidate molecules. In the second stage, we introduce a new batch acquisition function, qPMHI (batch probability of maximum hypervolume improvement), that selects the batch of candidates most likely to produce the maximum Pareto hypervolume expansion. The key mathematical result is that qPMHI decomposes additively across candidates, so the optimal batch can be found exactly by ranking candidate-wise probabilities estimated via posterior sampling from the surrogate model. This avoids the combinatorial search typically associated with large-batch selection and makes it practical to optimize directly over large discrete candidate pools. A major advantage of this framework is modularity. Because generation and selection are separated, the method can flexibly combine different molecular generators, probabilistic surrogate models, and molecular representations without redesigning the overall workflow. This makes it possible to pair recent generative AI tools with rigorous uncertainty-aware optimization methods in a simple and scalable manner. More broadly, the framework provides a practical strategy for combining expressive candidate generation with principled decision-making in large combinatorial spaces. We benchmark the approach on two case studies. First, on a standard drug-like molecule design benchmark involving two common molecular properties, the proposed method achieves higher final hypervolume than the strongest baseline and identifies a substantially broader Pareto frontier. Second, we apply the framework to de novo design of quinone-based organic electrode materials for aqueous redox flow batteries. In this setting, the method consistently uncovers novel, diverse, and high-performing candidates and achieves larger Pareto front improvements than strong optimization and fragment-recombination baselines. These results demonstrate that the proposed framework can effectively combine generative modeling and BO to accelerate multi-objective discovery in chemically realistic spaces. Although our focus here is de novo molecular design, the underlying idea is considerably more general. Any problem involving a large discrete and/or combinatorial design space, an inexpensive proposal mechanism, and an expensive evaluation pipeline could potentially benefit from the same generate-then-optimize strategy. Thus, beyond the specific molecular applications considered here, this work highlights a broader route for combining generative AI with optimization in data-efficient scientific discovery. References [1] Janet, Jon Paul, et al. "Accurate multiobjective design in a space of millions of transition metal complexes with neural-network-driven efficient global optimization." ACS central science 6.4 (2020): 513-524. [2] Gómez-Bombarelli, Rafael, et al. "Automatic chemical design using a data-driven continuous representation of molecules." ACS central science 4.2 (2018): 268-276. [3] Daulton, Samuel, Maximilian Balandat, and Eytan Bakshy. "Differentiable expected hypervolume improvement for parallel multi-objective Bayesian optimization." Advances in neural information processing systems 33 (2020): 9851-9864. [4] Muthyala, Madhav R., et al. "Generative Multiobjective Bayesian Optimization with Scalable Batch Evaluations for Sample-Efficient De Novo Molecular Design." Industrial & Engineering Chemistry Research 65.1 (2025): 628-642.

2025

AIChE 2025
Deep learning founded alternative to high throughput screening for noncanonical amino acid incorporation
2025 AIChE Annual Meeting, Advances in Experimental Methods for Protein Engineering, 2025
Abstract
Proteins serve as an attractive solution platform to challenge issues in many fields ranging from energy to healthcare. This capacity is expanded through the introduction of non-canonical amino acids (NCAAs). NCAAs add upon the preexisting library of 20 canonically encoded amino acids, carrying along with newfound, reliable chemistries. Combining this with the relationship between structure and function found in biology results in a system that can perform novel chemistry at intentional locations. This manifests in many ways including better drug delivery through conjugation (i.e., pharmaceuticals) to safer sequestering of rare earth metals (i.e., sustainability). However, improper incorporation of these non-canonicals may perturb the structure to where wild-type functionality is reduced or even destroyed. Due to the lack of guidelines, many incorporation attempts result in trivial results, often requiring high throughput screening methods, costing countless hours and experimental expenses. By providing intuition of incorporation sites, NCAAs become more accessible as a technology platform, resulting in the overall expansion of synthetic biology. We seek to algorithmically discover these guidelines, creating a workflow in silico in which conjugation sites can be rationally selected based off a variety of parameters. By this introduction of computational methods, we quickly screen and select promising conjugation site candidates for more informed mutagenesis rounds. The project explores (1) the selection of the decision space from which the model will train off and (2) the model architecture itself. Due to the size of proteins and the data-dependent nature of deep learning models, we increased the density of our decision space through rationally selecting which residues to train off and the number of NCAAs explored (200+ unique NCAAs). Amino acid surface exposure and distance from the binding complementary determining region were screened and mutated independently to one another, with combination of screens tested afterwards. In the end, 60 total sites (20 from each screening process) were selected and mutated using the NCAA database. After designing our decision space, we moved into training our model. The model views proteins through the lens of a language, where each amino acid functions as a word to build out the sequence. Through understanding this protein language, the model can identify sensical locations to incorporate non-canonicals based on its characteristics. We then characterize the function of resultant nanobodies through examination of cell viability and target receptor binding affinity in vitro using relevant disease model cells. Upon collection of data, we plan to revisit the original computational model, leveraging statistical analysis to identify key features that have the highest correlation to both the binding interaction and intercellular activity, with the end goal of creating a predictive model accounting for necessary parameters. Through this, we can not only improve the accessibility of engineering NCAA containing proteins but also distill a deeper understanding of a protein's resiliency to structural perturbations.
AIChE 2025
Generative multi-objective Bayesian optimization with scalable batch evaluations for sample-efficient de novo molecular design
2025 AIChE Annual Meeting, Machine Learning for Soft and Hard Materials I: Soft Matter, 2025
Abstract
Designing molecules that must satisfy multiple, often conflicting objectives is a central challenge in molecular discovery. The enormous size of chemical space and the cost of high-fidelity simulations have driven the development of machine-learning-guided strategies for accelerating design with limited data. Among these, Bayesian optimization (BO) offers a principled framework for sample-efficient search, while generative models provide a mechanism to propose novel, diverse candidates beyond fixed libraries. However, existing methods that couple the two often rely on continuous latent spaces, which introduces both architectural entanglement and scalability challenges. This work introduces an alternative, modular “generate-then-optimize” framework for de novo multi-objective molecular design. At each iteration, a generative model constructs a large, diverse pool of candidate molecules, after which a novel acquisition function, qPMHI (multi-point Probability of Maximum Hypervolume Improvement) , selects a batch of candidates most likely to expand the Pareto front. The key insight is that qPMHI decomposes additively, enabling exact, scalable batch selection through simple probability ranking estimated via Monte Carlo sampling. We benchmark the framework against state-of-the-art latent-space and discrete molecular-optimization methods, demonstrating significant improvements across synthetic benchmarks and a case study in sustainable energy storage, where our approach rapidly identifies novel, diverse, and high-performing quinone-based cathode materials for aqueous redox-flow batteries.

2024

AIChE 2024
Computational discovery of plastic-binding peptides for remediating microplastic pollution
2024 AIChE Annual Meeting, Microplastics Contamination: Research, Treatment and Mitigation, 2024
Abstract
Methods are needed to remediate microplastic (MP) pollution so as to limit potential harm to the environment and human health. One promising approach is to use polypeptides to detect and/or capture MP pollution, as they can adsorb strongly to micro- and nanomaterial sized objects. Short peptides that bind to plastics are particularly interesting because their small size means that they are more economical to manufacture than proteins. However, peptides cannot currently be applied to MP remediation because there are no known peptides that bind to the many common plastics. In this work, we aim to fill this gap by using biophysical modeling and machine learning to discover 12-residue plastic-binding peptides (PBPs) that bind strongly to the common plastics polyethylene, polystyrene, polypropylene, and PET. We have been designed PBPs using the Pep tide B inding D esign (PepBD) algorithm. PepBD searches for PBPs by using Metropolis Monte Carlo sampling to randomly change the amino acid sequence or structure of a starting peptide adsorbed to a plastic surface. The changes are accepted or rejected based on the change in the peptide score, a measure of the peptide-plastic interaction energy and the stability (i.e., internal energy) of the peptide in the adsorbed conformation. Over the course of sampling millions of sequences and structures, PepBD identifies peptides with high affinity to the plastic. The best PBPs designs were found in molecular dynamics (MD) simulations to bind much more strongly to the target plastic than random amino acid sequences. Experimental testing supports our simulation results for the polyethylene design; designs for the other plastics are currently being evaluated. We next designed PBPs that have affinity for plastic as well as two other important features: high solubility in water, which facilitates MP remediation in aqueous environments; and preferential binding to a target plastic over others, which can help characterize or separate components of MP waste. The predicted solubilities were significantly improved by including the CamSol solubility score in the peptide score. The binding preference of a peptide for one plastic over others was moderately increased by training a long short-term memory (LSTM) network on PepBD data to predict the peptide’s score given the amino acid sequence, and then searching for sequences with large score differences between a target plastic and an off-target plastic. By identifying PBPs for many common plastics, we hope that peptide-based methods for remediating MP pollution can begin to be developed.