Causal Machine Learning: A Deductive-Inductive Framework for Sociological Research
Kölner Zeitschrift für Soziologie & Sozialpsychologie (Cologne Journal of Sociology & Social Psychology), 78(3): 1089–1123. (Special Issue on Explanation and Causality in Sociology)
Causal explanation is central to sociological research, shaping both theoretical development and empirical inquiry. This paper argues that causal machine learning, a technique that integrates deductive identification strategies with inductive estimation techniques, offers an analytical approach for modeling complex, nonlinear social processes within the potential outcomes framework. We argue that causal machine learning operates through an iterative feedback loop: Theoretical assumptions guide flexible estimation, which inductively uncovers complex heterogeneities and nonlinearities, and these discoveries subsequently refine and expand sociological knowledge. Drawing on a systematic review of recent sociological research (2014–2024), we highlight how causal machine learning is advancing work in three key areas: causal effect heterogeneity, causal mediation analysis, and time-varying causal inference. These developments expand the methodological tool kit available to sociologists and strengthen the discipline's ability to test, refine, and extend theories of social explanation. We conclude by outlining emerging directions, including high-dimensional causal inference and generative artificial intelligence, that are opening new methodological frontiers in causal machine learning for sociology.
Causal Machine Learning: From Identification to Estimation
Causal inference in the social sciences centers on counterfactual reasoning, with a focus on estimating the "effects of causes." For example, we may ask about the effects of college education: What do life outcomes look like if an individual chooses to attend college versus not attending? Even with detailed data on college completion and later-life outcomes, we cannot answer this question directly because we do not observe the counterfactual: what would have happened if the same individual had made a different choice of attending versus not attending. This challenge motivates the potential outcomes framework, which formally distinguishes between observed and unobserved outcomes.
The standard identification assumptions—consistency, conditional ignorability, and positivity—allow the ATE to be identified from the observed data distribution. This means the ATE can be expressed as a statistical quantity that depends only on observed variables.

Causal machine learning can yield more robust estimates in settings with high-dimensional covariates and complex or non-linear relationships within the covariate space—features often arising from social processes. Yet causal machine learning is not merely an advanced estimation strategy. It is best understood as a mode of reasoning that integrates deductive design with inductive modeling. Once this causal framework is in place, the task becomes inductive: estimating the identified statistical quantity using flexible, data-driven algorithms. As shown in Fig. 1, the causal machine learning workflow integrates deductive reasoning (used to define and identify causal estimands) with inductive modeling (used to estimate them with flexible, data-driven modeling).
Sociologists routinely construct theoretical models that imply causal relationships and work with observational data to study complex, nonlinear social processes. The causal machine learning workflow mirrors the structure of much sociological research. Rather than replacing existing methods, causal machine learning offers an alternative approach, providing new tools to engage enduring sociological questions with greater flexibility.
A Review of Causal Machine Learning in Sociology
Through our literature review, we identified three major themes that reflect how causal machine learning has been incorporated into sociological research: (1) causal effect heterogeneity, (2) causal mediation analysis, and (3) time-varying causal inference. All three domains involve complex and dynamic social processes—challenges that machine learning is particularly well suited to model.
While our review focuses on three core domains, we also note the growing interest in emerging areas such as causal inference with high-dimensional data and the use of generative artificial intelligence for causal reasoning. Although still relatively rare in sociological publications, these developments reflect broader shifts in the computational social sciences and are likely to shape future methodological innovation in the field.
In each area, machine learning expands the capacity of existing causal inference strategies by modeling nonlinear, dynamic, and socially structured processes. For studying effect heterogeneity, methods such as causal forests and metalearners uncover variation in treatment effects that would be difficult to prespecify in conventional regression models.
These developments demonstrate how causal machine learning can expand the scope of causal inference in sociology. Importantly, causal machine learning complements rather than replaces the deductive logic of causal inference. At its core, causal inference depends on clearly defined estimands and identification assumptions that allow researchers to interpret estimates as causal effects. Once these assumptions are articulated, machine learning provides flexible tools for estimation.
