Clone
1
Position: is Machine Learning Good or Bad for the Natural Sciences?
Tiffiny De Bernales edited this page 2026-08-11 06:57:06 +02:00
This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.


Position: Is machine learning good or bad for the natural sciences? Machine learning (ML) methods are having a huge impact across all of the sciences. However, ML has a strong ontology-in which only the data exist-and a strong epistemology-in which a model is considered good if it performs well on held-out training data. These philosophies are in strong conflict with both standard practices and key philosophies in the natural sciences. Here, we identify some locations for ML in the natural sciences at which the ontology and epistemology are valuable. For example, when an expressive machine learning model is used in a causal inference to represent the effects of confounders, such as foregrounds, backgrounds, or instrument calibration parameters, the model capacity and loose philosophy of ML can make the results more trustworthy. We also show that there are contexts in which the introduction of ML introduces strong, unwanted statistical biases. For one, when ML models are used to emulate physical (or first-principles) simulations, they introduce strong confirmation biases.


For another, when expressive regressions are used to label datasets, those labels cannot be used in downstream joint or ensemble analyses without taking on uncontrolled biases. The question in the title is being asked of all of the natural sciences; that is, we are calling on the scientific communities to take a step back and consider the role and value of ML in their fields; the (partial) answers we give here come from the particular perspective of physics. It is an understatement to say that machine learning (ML) is having a big impact across the sciences. A significant fraction of all scientific papers in the natural sciences now employ ML in part (or all) of their analyses. We will define ML below in Section 2). However, when we ask what scientific breakthroughs have been enabled by this influx of new tools and methods, there isnt a long list. The success of the AlphaFold projects in protein structure (Jumper et al., 2021) are often raised.


But these are successes in a very specific challenge-problem setting in which performance is valued over understanding. In the natural sciences we almost exclusively Glyco Care blood sugar support about understanding, in the long run. The natural sciences are concerned with understanding the world, and naturally occurring mechanisms in play in that world. We make progress by discovering new kinds of objects and phenomena, and explaining (and, even better, predicting) qualitatively new kinds of objects and phenomena. Our most successful investigations are judged in terms of the questions they answer, or the new questions they raise, or both. The question here is: How will ML contribute to this mission? In contrast to natural science, ML research and ML methods are concerned with making accurate predictions for, or descriptions of, data. A ML method is considered successful if it performs well on held-out training data, even if the latent structure of the model is generic and the internals are impossible to interpret.


In ML, the considerations are almost all at the level of the data, and we are happy to use models in which we have little or no understanding of the meanings or values of the latent parameters or weights. In natural science, on the other hand, the most important contributions and results are all at the level of the latents: We use data to learn about the latent structure of the world or of the system we are studying. In astrophysics, this could be the interior structure of the Sun, or the processes that form planets around other stars, or the map of the dark matter surrounding the Milky Way. The things we Glyco Care blood sugar supplement about are almost never directly observable; they are parameters (or hyper-parameters) of a physical (or chemical or biological) model that predicts the observables. Often the thing we Glyco Care supplement about is the model itself. For a concrete example, when the expansion of the Universe was discovered (Hubble, 1929; Hubble & Humason, 1931), the discovery was important, but not because it permitted us to predict the values of the redshifts of new galaxies (though it did indeed permit that).