Chemistry, Chemical Engineering, and Materials Science
Track Chairs
Sanat K. Kumar, Bykhovsky Professor of Chemical Engineering, Columbia Engineering
Simon Billinge, Professor, Materials Science and Applied Physics and Applied Mathematics, Columbia Engineering
Yevgeny Rakita, Postdoctoral Research Scientist, Data Science Institute, Columbia University
Track Session Program - All times EST
DAY 1
Monday, December 14
12:30 PM - 2:00 PM: Track Session 2
Martin L. Green, Leader, Materials for Energy and Sustainable Development Group, National Institute of Standards and Technology (NIST)
Talk Title: Autonomous (AI-Driven) Experimental Materials Science
Abstract: Several critical technologies are currently materials-limited, awaiting novel materials solutions for advancement, e.g., transportation (light-weight, high-strength alloys, and corrosion resistant coatings), sustainable energy (earth-abundant catalysts for solar fuel production), and nanoelectronics (low-energy/operation integrated circuit devices). A fundamental obstacle to efficient discovery of novel materials and processing schemes is the large number of candidate materials, and their associated processing parameters. Modern materials scientists train for decades in the fundamentals of the field to develop the knowledge and intuition required to discover new materials. They must then plan and conduct a series of experiments to test their hypotheses, analyze the resultant data, generate new hypotheses based on the knowledge gained, and plan the next series of experiments. They are constrained to study relatively simple materials (e.g., those that may contain ≤ 4 elements and involve ≤ 2 processing parameters). This is because while the number of variable parameters (n) increases linearly, the number of candidate materials/processing schemes increases exponentially (~102n). I will discuss the development of a novel NIST platform which will place artificial intelligence in charge of materials optimization and discovery, thereby transforming “dumb” instrumentation into “intelligent” systems capable of planning, executing, analyzing, and extracting knowledge from experimental and modeling data. The marriage of artificial intelligence and materials science will result in a sea-change, enabling a 102 – 104 acceleration of materials discovery and process optimization.
Alex Hexemer, Senior Scientist and Computing Program Lead Computing, Advanced Light Source, Lawrence Berkeley National Laboratory
Talk Title: Machine Learning for Characterization Techniques
Abstract: The materials discovery cycle contains many different components, including synthesis, characterization and data analysis and interpretation. In the past few decades, automatic synthesis pipelines have been established for many chemistry and materials systems. For characterization, many advanced techniques, such as X-ray scattering and NMR crystallography, have enabled the structure identification of various chemical, biological and materials systems, including polymers, inorganic materials, and proteins. These techniques have been developed and improved substantially over the past few decades, which brings high-throughput experimental discovery into reach. Meanwhile, these breakthroughs produce enormous data amounts. However, the process of understanding the structural features from data is still very labor-intensive. It requires many man-hours of work by highly specialized and trained scientific staff to interpret the data and identify the structure correctly. Recently, machine learning, a branch of artificial intelligence, has demonstrated the potential to tackle and accelerate the analysis of common techniques such as tomography, scattering, and spectroscopy.
Poster Lightning Talks
Himaghna Bhattacharjee, University of Delaware: Thermochemical Data Fusion Using Graph Representation Learning
Gilchan Park, Brookhaven National Laboratory: Predicting Precursors of Inorganic Synthesis Reactions using Literature Mining
Jordan Winetrout, University of Colorado Boulder: Machine Learning Framework for Predicting Elastic Performance of Carbon Nanotubes and Elucidating the Influence of Defects
Emil Thyge Skaaning Kjær, University of Copenhagen: Characterising the Atomic Structure of Mono-Metallic Nanoparticles From X-Ray Scattering Data using Conditional Generative Models
Hung Vuong, Columbia University: Using a Machine Learning Approach to Determine the Space Group of a Structure From the Atomic Pair Distribution Function
Kyle Sherman, Binghamton University: Machine Learning the Mysterious Long Time Dynamics of Spin Ice
4:00 PM - 5:30 PM: Track Session 4
Keith A. Brown, Assistant Professor (ME, MSE, Physics), Department of Mechanical Engineering, Boston University
Talk Title: Learning to Design Hierarchical Materials with Autonomous Researchers
Abstract: Nature has taught us that intricate structure from the molecular scale to the macroscale can lead to exquisite material properties. While inspiring, the extraordinarily vast number of permutations of composition, processing conditions, and structures spanning these scales makes brute force exploration of this parameter space intractable. Further, simulation cannot accurately and rapidly predict many important material properties such as non-linear mechanical properties such as toughness, indicating that experiments are necessary to explore this space. Autonomous researchers present a unique opportunity to address the complexity presented by hierarchical materials through their combination of automation to perform experiments rapidly and their use of machine learning to select experiments to achieve specific goals. In this presentation, we describe the emergence of autonomous research systems and efforts in the materials community to use them to overcome critical bottlenecks in experimental materials development. We explore the acceleration in research possible with autonomous systems using the Bayesian experimental autonomous researcher (BEAR), a system that combines additive manufacturing, robotics, and mechanical testing to design and evaluate mechanical components without human intervention. By guiding this system using Bayesian optimization-based active learning, we report a ~60 fold reduction in the number of experiments needed to achieve a target performance. Finally, we explore the further acceleration possible when simulation is combined with the BEAR to realize a physics-informed system that provides a further ten-fold acceleration in terms of the number of experiments. In addition to providing a path for the design of high performance hierarchical materials, autonomous research systems represent a unique convergence of the learning community and the materials research community that hold the promise to accelerate experimental materials development.
Nongnuch Artrith, Research Scientist, Department of Chemical Engineering, Columbia University
Talk Title: (Machine) Learning What Makes Catalysts Good
Abstract: Machine learning (ML) has proven a powerful tool for accelerating the computational characterization of energy materials [1-3]. There is a growing number of case studies identifying descriptors of catalytic performance using ML instead of physical intuition. ML is ideally suited for the pattern detection in large uniform data sets, but consistent experimental data sets on catalyst studies are often small. Here we demonstrate how a combination of machine learning and first-principles calculations can be used to extract knowledge from a relatively small set of experimental data [4]. The approach is based on combining a complex machine-learning model trained on a computational library of transition-state energies with simple linear regression models of experimental catalytic activities and selectivities from the literature. Using the combined model, we identify the key C-C bond-scission reactions involved in ethanol reforming and perform a computational screening for ethanol reforming on monolayer bimetallic catalysts with architectures TM-Pt-Pt(111) and Pt-TM-Pt(111) (TM = 3d transition metals). The model also predicts four promising catalyst compositions for future experimental studies. The approach is not limited to ethanol reforming but is of general use for the interpretation of experimental observations as well as for the computational discovery of catalytic materials.
Andrew Ferguson, Associate Professor of Molecular Engineering; and Deputy Dean of Equity, Diversity, and Inclusion, Pritzker School of Molecular Engineering, University of Chicago
Talk Title: Reconstructing Protein Folding Trajectories from Experimentally Measurable Observables
Abstract: Proteins are molecular machines. Understanding their dynamics and mechanisms of operation is a grand challenge in molecular biophysics and a pre-requisite to the rational design and engineering of proteins with desired structure and function. Sophisticated single molecule techniques have enabled the resolution of protein structure to within a few Angstroms, but no techniques are currently available to follow their real-time dynamical evolution with similarly atomistic resolution. By integrating concepts and tools from statistical mechanics, dynamical systems theory, manifold learning, and deep learning, we have developed an approach to recover atomistic protein folding trajectories from one-dimensional time series in experimentally measurable coarse-grained observables such as a radius of gyration or distance between two fluorescent probes. In a computational application to the artificial mini-protein chignolin we demonstrate recovery accuracies better than 2 Å and lay the theoretical and algorithmic foundations to apply this technique to real experimental data.
DAY 2
Tuesday, December 15
11:00 AM - 12:30 PM: Track Session 5
Eun-Ah Kim, Professor, Department of Physics, Cornell University
Talk Title: Interpretable Machine Learning of Quantum Matter Data
Abstract: Decades of efforts in improving computing power and experimental instrumentation were driven by our desire to better understand the complex problem of quantum emergence. However, the increasing volume and variety of data made available to us today present new challenges. Employment of machine learning could embrace these challenges and turn them into opportunities. However, the rigorous framework for scientific understanding requires the interpretability of any machine learning essential. I will discuss our recent results using machine learning approaches designed to be interpretable from the outset. Specifically, I will present discovering order parameters and its fluctuations in voluminous X-ray diffraction data and discovering signature correlations in quantum gas microscopy data.
John Wright, Associate Professor, Department of Electrical Engineering, Columbia Engineering
Talk Title: Nonconvex Learning with Physical Data
Abstract: Many problems in sensing and modeling the physical world can be naturally cast as numerical optimization problems, in which we seek models that are consistent with observed data and satisfy physical hypotheses such as sparsity. In domains such as chemistry, physics, and neuroscience, this kind of intuitive data modeling typically leads to nonconvex optimization problems. While worst-case nonconvex optimization is impossible in general, physical problems are typically structured. In this talk, we show through examples how this structure can be leveraged to formulate problems that can be solved globally using efficient methods. We discuss applications of these ideas to fast chemical imaging, where they suggest new, more efficient sensors, and to the analysis of electron microscopy data, where they suggest new data analysis strategies.
Includes joint work with Yuqian Zhang (Rutgers), Han-Wen Kuo, Dan Esposito (Columbia) and Abhay Pasupathy (Columbia)
Gerbrand Ceder, Professor, Department of Materials Science and Engineering, University of California, Berkeley
Talk Title: Information Extraction and Learning by Large-scale Text-Mining of the Scientific Literature
Abstract: The overwhelming majority of scientific knowledge is published as text, which is difficult to analyze by either traditional statistical analysis or modern machine learning methods. In contrast, the main source of machine-interpretable data for the materials research community has come from structured property databaseswhich encompass only a small fraction of the knowledge present in the research literature. Beyond property values, publications contain valuable knowledge regarding the connections and relationships between the data items as interpreted by the authors. I will demonstrate the extraction of codified synthesis recipes from text using a combination of neural networks, transformer models, and grammar tree parsing. Extraction the details of synthesis, including precursor compounds, synthesis operations and their numerical details, requires a very high precision of information extraction, and a tolerance to deal with imprecise and non-standard language. This effort has led to a new large data set of codified solid-state synthesis reactions which can be queried to obtain interesting information on choice of synthesis operations and precursors. In more recent work we have integrated this approach with the extraction of data from figures so that synthesis and property outcomes can ultimately be related.
Poster Lightning Talks
Junhui Huang, Columbia University: Synthesis of Novel Antibacterials
Jacob Monroe, National Institute of Standards and Technology (NIST): Variational Autoencoders As A Unifying Framework for Molecular Coarse Graining, Back-mapping, and On-The-Fly Generation of Efficient Monte-Carlo Moves
Samichhya Paudel, Howard University: Pose Filter-Based Machine Learning Tools to Enhance Structure-Based Drug Discovery for G Protein-Coupled Receptors
David Sheen, National Institute of Standards and Technology (NIST): Automated Consistency Analysis of 2D Nuclear Magnetic Resonance Spectra for Machine Learning Applications in Quality Control
Valentin Stanev, University of Maryland, College Park: Predicting The Absorption Spectra Of Azobenzene Dyes
Zhiping (Peter) Zhang, Cornell University: Metabolic Pathway Design Using Deep Learning
Yevgeny Rakita, Columbia University: Towards Predictive Synthesis Pathways of Materials with Desired Properties
2:30 PM - 4:00 PM: Track Session 7
Krishna Rajan, Erich Bloch Chair, Empire Innovation Professor, Department of Materials Design and Innovation, University at Buffalo
Talk Title: Machine Learning for the Study of Structural Motifs in Complex Crystal Chemistries
Abstract: In this presentation, we describe the application of machine learning strategies that are sensitive to the details of the local neighborhood and coordination geometry of atoms that govern global properties of materials. We use the well-established computational chemistry formalism of Hirshfeld Surfaces and show how it serves as a powerful machine readable fingerprint for the classification and prediction of materials properties.
Sergei V. Kalinin, Corporate Fellow, The Center for Nanophase Materials Sciences, Oak Ridge National Laboratory
Talk Title: Can (Almost) Unsupervised Machine Learning Learn (and Control) Chemistry and Physics from Atomically-Resolved Imaging Data?
Abstract: Rich functionalities of quantum and strongly correlated materials emerge from the interplay between the electronic, orbital, lattice, and spin degrees of freedom that often lead to complex structural and electronic phenomena spanning atomic to mesoscopic scales. In many cases, these phenomena are associated with translational symmetry breaking, local frozen disorder, or strongly correlated disorder. However, the relevant mechanisms and roles of individual subsystems often remain unknown. Over the last decade, Scanning Transmission Electron Microscopy has emerged as a powerful quantitative probe of materials structure and functionality on the atomic level, providing high veracity information on local chemical bonding, composition, and symmetry breaking distortions. We aim to harness the power of machine learning methods to build a comprehensive picture of the chemistry and physics of quantum materials from these observations. In this presentation, I will illustrate the application of rotationally-invariant variational autoencoders (rVAE) towards the effective exploration of the chemical evolution of the system based on local structural changes, effectively discovering molecular building blocks and chemical reactions pathways in unsupervised manner. I will further illustrate the extension of this approach in encoder-decoder architectures to establish the parsimonious structure-property relationships in complex materials on an example of plasmonic nanostrucutres. These allow to question such as (a) what responses are possible in a given materials systems, and (b) what local geometries are required to enable them. Finally, on an example of the direct electron beam manipulation of plasmonic structures I illustrate the approach for practical implementation of these concepts.
This research is supported by the by the U.S. Department of Energy, Basic Energy Sciences, Materials Sciences and Engineering Division and the Center for Nanophase Materials Sciences, which is sponsored at Oak Ridge National Laboratory by the Scientific User Facilities Division, BES DOE.
Andrew Gordon Wilson, Assistant Professor, Courant Institute of Mathematical Sciences and Center for Data Science, New York University
Talk Title: Bayesian deep learning and probabilistic model construction
Abstract: To answer scientific questions, and reason about data, we must build models and perform inference within those models. But how should we approach model construction and inference to make the most successful predictions? How do we represent uncertainty and prior knowledge? How flexible should our models be? Should we use a single model, or multiple different models? Should we follow a different procedure depending on how much data are available?
In this talk I will present a philosophy for model construction, grounded in probability theory. I will exemplify this approach with methods that exploit loss surface geometry for scalable and practical Bayesian deep learning. The talk will primarily be based on https://arxiv.org/abs/2002.08791 (NeurIPS 2020).
The MLSE 2020 Chemistry, Chemical Engineering, and Materials Science Track is sponsored by the Department of Chemical Engineering, Northeastern University
Participating Speakers & Track Chairs
Additional Opportunity: Machine Learning in Materials Science Tutorial Workshop
3 day workshop in parallel to the MLSE 2020 Chemistry & Materials Science Track: December 13 - 15, 2020 from 9:00 AM - 11:00 AM EST
Tools from data science and machine learning are increasingly adopted in materials science, but most resources for beginners were originally developed with computer-science applications in mind. This workshop addresses the need to train materials scientists in this rapidly developing area by introducing state-of-the-art machine-learning methods for concrete materials science applications.
After a general introduction lecture, different machine-learning approaches will be introduced in tutorials using actual materials data either from experiments or from simulations. Workshop participants will be able to work interactively on their own laptops.
Research Submissions
Submission Deadline: October 15, 2020
Acceptance Decision: November 15, 2020
Accepted Poster Upload Deadline: December 1, 2020
Poster Presentations & Event Date: December 14-15, 2020 (see the full program here) - posters will be online, with Zoom and text-based chatting.
The Chemistry, Chemical Engineering, and Materials Science track will first collect abstracts, which should be received by the submission deadline. Please see this guide here for optimal formatting of your abstract.
Following review, posters from accepted researchers must be uploaded by December 1, 2020. Posters will be linked, with a 30-second video explanation by the lead author(s), to the conference website. There will be a Zoom URL and text-based chatting capabilities for discussion. Posters must be a single page with text and graphics fully visible on a small screen. Any single-page format is acceptable. In case it is helpful, we provide a template here.
