Earth and Environmental Sciences

 Track Chairs

  • Pierre Gentine, Associate Professor, Earth and Environmental Engineering, Columbia Engineering

  • Laure Zanna, Professor, Center for Atmosphere Ocean Science, Department of Mathematics, Courant Institute, New York University

Track Session Program - All times EST

DAY 1
Monday, December 14

11:00 AM - 12:30 PM: Track Session 1

Elizabeth A. Barnes, Associate Professor, Department of Atmospheric Science, Colorado State University

Talk Title: Leveraging Interpretable Neural Networks for Earth Science Discovery

Abstract: Recent advances in machine learning have yielded many breakthroughs in commercial applications, and these techniques hold enormous promise for scientific discovery. While exciting advances with these tools have already been seen in other scientific disciplines, e.g. life sciences, they have been more slowly embraced by the geoscience community. One possible explanation for this is the perceived “black box” that outputs an answer without any explanation as to “why?” or “how?”. Here, I will discuss how the field can make the most of machine learning interpretation techniques (i.e. “explainable AI”) to open the black box and push the bounds of scientific discovery. This has profound implications for machine learning use in science, as it not only increases trust in the output, but also allows us to learn new science from the decision making process of the algorithm itself. I will discuss applications in climate science, including seasonal prediction and the atmospheric response to climate change. While these examples are focused on climate problems, the tools and the approach are widely applicable and offer an exciting path for the future of geoscientific research.

Galen McKinley, Professor of Earth and Environmental Sciences, Columbia University and the Lamont-Doherty Earth Observatory

Talk Title: Quantifying the Ccean Carbon Sink with Sparse Data, Physical Models and Machine Learning

Abstract: The ocean plays an important role in modulating climate change by sequestering human-emitted carbon from the atmosphere; 40% of fossil fuel emissions since the industrial revolution are now stored in the ocean. While overall air-sea CO2 flux is well estimated, regional patterns are more challenging to capture but just as important to understand. Quantifying the CO2 flux across the air-sea interface requires monthly maps of surface ocean partial pressure of CO2 (pCO2), which is typically estimated with physical models of the ocean circulation and biogeochemistry, or by extrapolating very sparse (2% coverage) direct observations to full coverage using auxiliary driver data. In this work, we combine the strengths of physical models and machine learning extrapolations to assess uncertainty and develop an improved extrapolation algorithm. First, we use a suite of physical models as a testbed with which to examine the skill of machine learning extrapolations from sparse data. Neural networks and Extreme Gradient Boosting (XGB) algorithms reconstruct pCO2 seasonality with high fidelity, but fail to accurately reconstruct lower frequency variability, particularly in the Southern Ocean. Data sparsity is the key limiting factor. Second, we use physical models as prior estimates of the pCO2 distribution, and then implement an XGB algorithm to correct these priors toward observations. With this approach, physical models provide spatial relationships not fully captured in the data. In comparison to other extrapolation products, the resulting pCO2 estimates have the best fit to independent data.

Authors: Galen A. McKinley, Lucas Gloege, Jake Stamell, Rea Rustagi, and Tian Zheng

George Karniadakis, Charles Pitts Robinson and John Palmer Barstow Professor of Applied Mathematics, Brown University; and Research Scientist, MIT

Talk Title: Approximating Functions, Functionals, and Operators using Deep Neural Networks for Diverse Applications

Abstract: We will present a new approach to develop a data-driven, learning-based framework for predicting outcomes of physical and geophysical systems and for discovering hidden physics from noisy data. We will introduce a deep learning approach based on neural networks (NNs) and generative adversarial networks (GANs). Unlike other approaches that rely on big data, here we “learn” from small data by exploiting the information provided by the physical conservation laws, which are used to obtain informative priors or regularize the neural networks. We will also make connections between Gauss Process Regression and NNs, and discuss the new powerful concept of meta-learning. We will demonstrate the power of PINNs for several inverse problems, and we will demonstrate how we can use multi-fidelity modeling in monitoring ocean acidification levels in the Massachusetts Bay. We will also introduce new NNs that learn functionals and nonlinear operators from functions and corresponding responses for system identification. The universal approximation theorem of operators is suggestive of the potential of NNs in learning from scattered data any continuous operator or complex system. We first generalize the theorem to deep neural networks, and subsequently we apply it to design a new composite NN with small generalization error, the deep operator network (DeepONet), consisting of a NN for encoding the discrete input function space (branch net) and another NN for encoding the domain of the output functions (trunk net). We demonstrate that DeepONet can learn various explicit operators, e.g., integrals, Laplace transforms and fractional Laplacians, as well as implicit operators that represent deterministic and stochastic differential equations. More generally, it can learn multiscale operators spanning across many scales and trained by diverse sources of data simultaneously. There are many versions of PINNs, e.g., variational (VPINNs), stochastic (sPINNs), conservative (cPINNs), nonlocal (nPINNs), generalized (xPINNs), etc, and we will provide some highlights. In addition, we will present our recent theoretical results on the convergence and generalization of PINNs.

2:30 PM - 4:00 PM: Track Session 3

Gustau Camps-Valls, Full Professor, Image Processing Laboratory, Universitat de València

Talk Title: How to Surf the Physics and Machine Learning Interplay

Abstract: Do you still think that data, deep learning (DL) and computers are enough for actual learning? Maybe for cats and dogs discrimination in photographic images they suffice. However, most problems in Earth sciences aim to do inferences about the system, where accurate predictions are just one tiny bit of the whole problem. Inferences mean understanding variables relations, deriving models that are physically interpretable, causally explanatory, parsimonious and tractable. DL models alone are excellent approximators, but very often do not respect the most elementary laws of physics, like mass or energy conservation, so consistency and confidence is compromised. In this talk I will introduce five ways to live in the Physics and machine learning interplay. Modern machine learning methods that can learn differential equations from data, encode physics priors and constraints from algorithmic fairness principles, improve parameterizations by variational forward-inverse modeling, emulate physical models, and blend data-driven and process-based models. This is a collective long-term AI agenda towards developing and applying algorithms capable of discovering knowledge from Earth data.

Pedram Hassanzadeh, Assistant Professor of Mechanical Engineering, Rice University

Talk Title: Data-driven Modeling of Subgrid-Scale Processes using Deep Neural Nets and Transfer Learning

Abstract: Resolving all the relevant length and time scales in simulations of the turbulent flows, e.g., in the atmospheric and oceanic circulations, remains out of reach due to computational constraints. In practice, low-resolution models resolve the large-scale processes while the effects of the small-scale processes are often parameterized in terms of the large-scale variables. More recently, super-parameterization (SP), which involves solving for small-scale processes on a high-resolution grid embedded within the low-resolution grid, has attracted attention and shown advantages over parameterization, but SP’s applicability remains limited to its high computational cost. In the past few years, data-driven parameterization (DD-P) using deep learning has shown promising results, but numerical stability (in online or a posteriori simulations) and generalization (i.e., extrapolation) have remained as challenging and important issues to address.  In this talk, using a multi-scale Lorenz 96 system and 2D turbulent flow as testbeds, we 1) Introduce a data-driven SP (DD-SP) framework in which the equations of small-scales are integrated data-drivenly using deep learning to reduce the cost, 2) Show the promises of DD-P in representing subgrid-scale effects, in particular by capturing energy backscattering, while remaining numerically stable, and 3) Demonstrate how transfer learning enables DD-P and DD-SP to generalize to more chaotic systems or flows with higher Reynolds numbers. 

Rose Yu, Assistant Professor, Department of Computer Science and Engineering, UC San Diego

Talk Title: Physics Guided Deep Learning for Spatiotemporal Dynamics

Abstract: While deep learning has shown tremendous success in many domains, it remains a grand challenge to incorporate physical principles to such models for applications in physical sciences. In this talk, I will discuss (1) Turbulent-Flow Net: a hybrid approach for predicting turbulent flow by marrying well-established computational fluid dynamics techniques with deep learning (2) Equivariant Net: a systematic approach to improve generalization of spatiotemporal models by incorporating symmetries into deep neural networks. I will demonstrate the advantage of our approaches to a variety of physical systems including fluid and traffic dynamics.

Yoo-Geun Ham, Associate Professor, Chonnam National University (South Korea)

Talk Title: Deep Learning for multi-year ENSO forecasts

Abstract: Variations in the El Niño-Southern Oscillation (ENSO) are associated with a wide array of regional climate extremes and ecosystem impacts. Robust, long-lead forecasts would therefore be valuable for managing policy responses. But despite decades of effort, forecasting ENSO events at lead times more than one year remains problematic. Here we show that a statistical forecast model employing a deep learning approach produces skillful ENSO forecast for lead times of up to one and a half years. To circumvent the limited amount of observation data, we use transfer learning to train a Convolutional Neural Network (CNN) first on the Coupled Model Intercomparison Project phase 5 (CMIP5) historical simulations and subsequently on reanalysis from 1871 to 1973. During the validation period from 1984 to 2017, the all-season correlation skill of the Nino3.4 index of the CNN model is significantly higher than those of current state-of-the-art dynamical forecast systems. Also, the CNN model is better at predicting the detailed zonal distribution of sea surface temperatures, overcoming a weakness of dynamical forecast models. A heatmap analysis indicates that the CNN model predicts ENSO events based on physically reasonable precursors. In addition, we developed the All-season Convolutional Neural Network (A_CNN) model to account for ENSO seasonality. The correlation skill of the ENSO was particularly improved for forecasts of the boreal spring, which is the most challenging season to predict. More importantly, heat map values indicated a clear time evolution with an increasing forecast lead time. This enhances the role of the A_CNN model as a diagnostic tool by revealing the comprehensive influence of various climate precursors on ENSO behavior.


DAY 2
Tuesday, December 15

12:30 PM - 2:00 PM: Track Session 6

Aleyda Trevino, Graduate Student, Department of Earth and Planetary Sciences, Harvard University

Talk Title: Nonlinear Growth Relationship Indicated by a Neural Network Reconstruction of Drought using Tree-Ring Widths

Abstract: Climate reconstructions using proxies and machine learning methods are at times faced with issues of historical climatic data availability and proxy resolution. We navigate the issue of the limited temporal overlap between tree ring data and the instrumental record in order to understand the underlying relationship between climate and tree growth. A two-layer neural network is explored for purposes of reconstructing summertime self-calibrated Palmer Drought Severity Index (scPDSI) across the contiguous United States. Reconstructions using neural networks are more skillful than a linear approach at 75% of the sites if evaluated by the coefficient of efficiency and at 54% when using the Pearson correlation coefficient. The increased reconstruction skill is associated with the ability of the network to capture nonlinear climate-growth relationships. In the Southwest, in particular, a nonlinear response function captures a diminishing sensitivity of growth to moisture under wetter conditions, consistent with alleviation of moisture stress. This work allows us to understand climate-growth relationships more fully, and potentially open up more horizons for the usage of machine learning methods in questions that might have limited data and resolution

Craig Connolly, Postdoctoral Researcher, Environmental Health Sciences, Lamont-Doherty Earth Observatory, Columbia University

Talk Title: Predicting Arsenic Contamination and Heterogeneity Across Scales Using Environmental Geospatial Analysis and Machine Learning

Extended Abstract: Chronic exposure to arsenic (As) in groundwater and rice is a staggering public health crisis and threat to food security worldwide. Levels of As in groundwater and rice often range from safe to very dangerous over small spatial scales, making it challenging to understand the environmental conditions and their interactions that govern As mobility and toxicity. Despite decades of research on the origin of As contamination, we are still unable to sufficiently predict As concentrations in groundwater and rice in most environments and with enough confidence to predict exposure levels or to make effective management decisions that reduce the risk of chronic As exposure. This deficiency is in part because prior predictive models lack clear mechanistic linkages and/or high-resolution geospatial information at scales consistent with those at which extreme spatial heterogeneity in As are observed. We recently used an extensive dataset (>400,000 sites) of remotely-sensed estimates of flooding as well as other relevant geomorphic variables in South and Southeast Asia with a machine learning (Random Forest) technique to demonstrate that flooding is a master variable that regulates the key environmental conditions responsible for As levels in heterogeneous aquifers. Our modeling and mechanistic framework also lends itself to evaluating the relationship between flooding and rice-As levels, as well as how climate and anthropogenic impacts on flooding will change levels of As in groundwater and rice in the future. As larger and more complex datasets are added to these efforts, new and innovative types of data analysis will be essential.

Names and Affiliations of Co-Authors

  • Melinda L Erikson (Upper Midwest Water Science Center, USGS)

  • Ana Navas-Acien (Environmental Health Sciences, Columbia University)

  • Steven N Chillrud (Lamont-Doherty Earth Observatory, Columbia University)

  • Beck DeYoung (Union College)

  • Athena Anh-Thu Nghiem (Lamont-Doherty Earth Observatory, Columbia University)

  • Mason O Stahl (Union College)

  • Benjamin C Bostick (Lamont-Doherty Earth Observatory, Columbia University)

Jing Gao, Assistant Professor of Geospatial Data Science, Department of Geography, University of Delaware

Talk Title: Mapping Global Urban Land for the 21st Century with Data-Driven Simulations and Shared Socioeconomic Pathways

Abstract: Urban land expansion is one of the most visible, irreversible, and rapid types of land cover/land use change in contemporary human history, and is a key driver for many environmental and societal changes across scales. Yet spatial projections of how much and where it may occur are often limited to short-term futures and small geographic areas. Here we produce a first empirically-grounded set of global, spatial urban land projections over the 21st century. We use a data-science approach exploiting 15 diverse datasets, including a newly available 40-year global time series of fine-spatial-resolution remote sensing observations. We find the global total amount of urban land could increase by a factor of 1.8–5.9, and the per capita amount by a factor of 1.1–4.9, across different socioeconomic scenarios over the century. Though the fastest urban land expansion occurs in Africa and Asia, the developed world experiences a similarly large amount of new development. [Published on 08 May 2020, in Nature Communications, http://doi.org/10.1038/s41467-020-15788-7]

Names and Affiliations of Co-Authors

  • Brian O'Neill, Joint Global Change Research Institute, College Park, MD

Yi-Ming Quin, PhD Student, Environmental Science and Engineering, Harvard University

Talk Title: Assessing the Non-Linear Effect of Atmospheric Variables on Organic Aerosol Particle Concentration Using Machine Learning Methods

Abstract: Atmospheric particulate matter (PM) consists of a large organic fraction. The concentration of organic aerosol particles (OA) is highly variable in the atmosphere, depending on a wide variety of factors, such as emission flux, atmospheric oxidants, and temperature. Due to the complex interaction of the numerous parameters, accurate estimation of the effect of target parameters on the OA concentration is often challenging. The study herein used a random forest machine-learning algorithm to predict the particle concentrations of primary organic aerosol (POA) and oxygenated organic aerosol (OOA) and access their influence by different atmospheric conditions at an urban site and a rural site in Hong Kong. The model was trained with archived POA and OOA concentrations, gaseous pollutant concentrations, and the associated meteorological parameters. The random forest model was observed to be able to explain more than 80% of the observed traffic-POA, cooking-POA, and OOA. Whereas, a multiple linear regression model can only explain 30–50% of these OA concentrations. The dependencies of the OA concentration on different atmospheric conditions (e.g., NO2, O3, and other conditions) were calculated via the accumulated local effect (ALE) algorithm within the random forest. The ALE calculates the predicted response on the target parameters while isolating the changes in the cofounding factors. The efficient handling of non-linear and interaction effects makes the technique flexible and suitable for assessing the non-linear effect of atmospheric conditions on the organic aerosol particle concentration.

4:00 PM - 5:30 PM: Track Session 8

J. Emmanuel Johnson, Researcher & PhD Candidate, University of Valencia

Talk Title: Gaussianization for Earth - An Information Theoretic Perspective

Abstract: Remote sensing data often exhibit characteristics that are difficult to tackle with machine learning including heterogeneity, multivariate and multi-source. One of the biggest challenges of all is dealing with the curse of dimensionality; a common characteristic of Earth Observation (EO) data. Gaussianization is a class of machine learning approaches that is effective in computing density estimates of your data. This framework uses a sequence of composite invertible transformations which transform data from its original domain to a base Gaussian domain. In addition to this transformation via Gaussianization, we can also compute information theory measures (ITMs), which are particularly relevant for the analysis of Earth system data. The mean, variance and correlation provide first and second order measures and are typically used in analysis but ITMs can give higher order measures capturing more complexity and hence providing more insight on the problem at hand. We showcase how Gaussianization is useful in a selection of Earth observation data analysis problems including: synthesizing new data from Earth observation data, quantifying the information content across various Earth observation data, and computing similarity metrics on key land surface variables relevant for drought detection.

Rea Rustagi, Undergraduate Student (3rd year), Major in Applied Math, Columbia Engineering

Talk Title: Strengths and Weaknesses of Three Machine Learning Methods for pCO2 Interpolation

Talk Abstract: The ocean plays an important role in sequestering carbon dioxide (CO2) from the atmosphere, reducing the effects of climate change now and in the future. Quantifying CO2 flux across the air-sea interface requires time-dependent maps of surface ocean partial pressure of CO2 (pCO2). However, direct measurements of ocean pCO2 are sparse in space and time, making direct quantification of CO2 flux challenging. Various machine learning approaches have been used to upscale these sparse measurements and create gap-free maps of pCO2, including feed forward neural network (NN), extreme gradient boosting (XGB), and random forest (RF). Here we evaluate the ability of these three Machine Learning (ML) approaches to statistically reconstruct full-coverage surface ocean pCO2 from sparse in situ data. We use 100 members across four Earth system large ensemble models, where each member represents a different reality. 20% of the data is withheld as a test set with the remaining data being used to train and evaluate each ML approach. First, we sample each member’s full-field model pCO2 as real-world pCO2. We then use each ML approach to reconstruct the pCO2 field from each member. Finally, we evaluate the reconstruction performance by comparing back to the original un-sampled member. This is done for all 100 members. We find XGB and RF perform better than NN based on a suite of regression metrics. However, NN generalized well to regions with no observations. Overall, XGB outperforms NN and RF, with lowest mean bias and consistent performance across the 100 members.

Names and Affiliations of Co-Authors

  • Jake Stamell (1st author) - Masters Student, Data Science Institute, Columbia University

  • Luke Gloege - Post-Doctoral Student in the Department of Earth and Environmental Engineering

  • Galen A. McKinley - Professor of Earth and Environmental Sciences

Yudi Wu, PhD Candidate, College of Engineering, Florida Agricultural and Mechanical University

Talk Title: Water Quality Deterioration near Culverts in Appalachia National Forest

Abstract: Near the culverts in Appalachia National Forest, iron concentration not attaining the standard has been commonly observed. Three iron release mechanisms, i.e., iron pyrite decomposition (IRM I), organic decomposition coupled with Fe(III) reduction (IRM II) and element iron corrosion (IRM III) were identified and were found to be responsible for ferrous iron release. The soil and water samples were collected from the eleven culvert sites in Appalachia National Forest and were analyzed. Various statistic methods were used to identify the correlation of iron release mechanisms with measured parameters. Using the principle component analysis, five principle components (PCs) were found to capture the variances that significantly contributed to the elevated iron concentration, among which PC1 and PC2 were the two dominating contributors and were associated with IRM I and IRM II. PC 3 accounted for 9.18% of the variance and was attributed to IRM III. Based on IRM I, ferrous iron was released from pyrite decomposition, which was correlated with sulfate elevated concentration in the water. The soil samples analyzed by X-ray powder diffraction (XRD) further evidenced that sulfate-related mineral contributed to this process.


Participating Speakers & Track Chairs


Research Submissions

Submission Deadline: October 15, 2020
Acceptance Decision: November 15, 2020

The Earth and Environmental Sciences track will accept 4 types of submissions:

Research Highlights abstract (250 words, maximum) Highlight abstracts highlight recently published work  and contextualize it for the broader MLSE audience.  

Extended abstract (2 pages, maximum) Extended abstracts describe unpublished original research. These abstracts are non archival and any posting of them by MLSE is only optional. 

Poster (1 page, maximum) Posters describe unpublished original research. We welcome submissions covering late-breaking results as well as works in progress.

Panel (2 pages, maximum) Panels should aim to assemble a knowledgeable group to lead an active discussion with the audience around a topic of broad interest to the MLSE audience.