← Back to publications

Recent Developments in Causal Inference and Machine Learning

By Jennie E. Brand, Xiang Zhou, Yu XieApril 26, 2023

Annual Review of Sociology, 49, 81–110.

This article reviews recent advances in causal inference relevant to sociology. We focus on a selective subset of contributions aligning with four broad topics: causal effect identification and estimation in general, causal effect heterogeneity, causal effect mediation, and temporal and spatial interference. We describe how machine learning, as an estimation strategy, can be effectively combined with causal inference, which has been traditionally concerned with identification. The incorporation of machine learning in causal inference enables researchers to better address potential biases in estimating causal effects and uncover heterogeneous causal effects. Uncovering sources of effect heterogeneity is key for generalizing to populations beyond those under study. While sociology has long emphasized the importance of causal mechanisms, historical and life-cycle variation, and social contexts involving network interactions, recent conceptual and computational advances facilitate more principled estimation of causal effects under these settings. We encourage sociologists to incorporate these insights into their empirical research.

Many important questions in the social sciences, and everyday life, are causal questions. For example, we want to know how parental divorce affects children, how attending college affects job prospects, or how moving to a new neighborhood affects children’s academic performance. We ask what would happen if individuals did or did not experience an event, like divorcing or attending college. Since reviews in sociology by Winship & Morgan (1999) and Gangl (2010), the literature on causal inference has developed several new promising directions. Some of the most exciting areas of development lie at the intersection of causal inference with machine learning. This review describes several key identification strategies for causal inference and how machine learning methods can enhance our estimation of causal effects. Throughout our review, we describe some empirical applications of these methods in sociology.

We emphasize four main principles in our review. First, the plausibility of the assumptions underlying different research designs and identification strategies varies by application. Second, causal effect heterogeneity is the norm, and it complicates extrapolation. Third, when evaluating social mechanisms in sociological research, we need to attend to confounding along the causal pathway, i.e., for the treatment–outcome relationship and the treatment–mediator and mediator–outcome relationships. Fourth, temporal and spatial interference, typical in social settings, complicate the definition, identification, and estimation of causal effects.

Causal Effect Identification and Estimation

Empirical work can be descriptive, such that we establish facts through associations between observables. For example, we might observe that college graduates earn higher wages than non– college graduates. But to evaluate causal effects, we draw on counterfactuals, i.e., we ask how much college-educated individuals would have earned without a college degree. The potential-outcomes framework offers a conceptual apparatus for defining causal effects.

Recent developments in randomized experiments include adaptive designs for evaluating optimal treatment assignment. For example, multi-armed bandits tailor treatments to individuals when they need treatment. The design aims to balance the goals of exploration (i.e., evaluating the effects of different treatment conditions) and exploitation (i.e., assigning units to treatment conditions with higher payoffs). For example, consider an online setting where treatment is assigned sequentially to different units, and the outcome for each unit is measured quickly after treatment assignment. A multi-armed bandit assigns treatment conditions based on information learned up to the point of the assignment, thus allowing researchers or policy makers to assign more units to conditions with higher payoffs. Sociological applications of multi-armed bandits remain scarce, but it is a promising approach for future studies. Researchers have adapted machine learning methods to estimate causal parameters to mitigate these and other concerns central to causal inference. First, to adapt machine learning to the

Causal Effect Heterogeneity

Individuals differ not only in pretreatment characteristics (i.e., pretreatment heterogeneity) but also in how they respond to a common treatment (i.e., treatment effect heterogeneity). Analyses that estimate heterogeneous treatment effects can yield insights into how scarce social resources are distributed in an unequal society and how events differentially impact populations with different expectations of their occurrence.

In some cases, we may hypothesize that an event has significant consequences for some subgroups but less or no effect among others. Scholars may aim to identify the most responsive subgroups to determine which individuals benefit most from treatment so that policy makers can better assign different treatments to balance competing objectives, such as reducing costs and maximizing outcomes for targeted groups. An important feature of the potential-outcomes framework is that it allows for general heterogeneity in treatment effects from the outset. Attending to treatment effect heterogeneity can also help extrapolate findings to diverse populations and contexts.

Causal Effect Mediation

While traditional sociological approaches to mediation analysis relied on parametric structural equation models to define and estimate direct and indirect effects, a large body of research has emerged within the causal inference literature that disentangles the tasks of causal definition, identification, and estimation. Causal mediation analysis seeks to uncover whether and how a treatment affects an outcome by quantifying the pathways through which a causal effect operates. Building upon the potential-outcomes framework and graphical causal models, a new body of research has provided model-free definitions of direct and indirect effects, established the assumptions needed for identifying these effects, and developed an array of estimation strategies. These tools can help researchers discover mechanistic explanations, build theories, and design policy interventions. Sociologists would do well to consider these conceptual and computational tools in studies involving mechanisms. This section briefly reviews the causal approach to mediation analysis and its recent developments.

Temporal and Spatial Interference

Many sociological questions involve the study of effects over time or interactions within networks. Indeed, historical or life-cycle variation and network interactions lie at the center of sociological inquiry. But these settings complicate the definition and identification of causal effects. Just as sociologists studying temporal variation or network settings should consider causal processes, causal inference scholars should consider the complications involved in allowing treatments and effects to vary over time and interference between units under study. SUTVA posits that one unit’s outcome is not affected by the treatment status of other units in the population. However, we often face temporal or spatial interference that renders SUTVA untenable. This section briefly reviews causal inference methods developed to study temporal and spatial interference.

Spatial interference may arise in settings where units under consideration are not isolated but are connected by a common physical or social space, such as schools, neighborhoods, and friendship networks, leading to spillover effects. In such settings, one unit’s potential outcome is a function of not only its treatment status but also the treatment status of other related units.

In many social settings, people interact with each other through multiple channels and networks, such as friends, family, neighbors, and others. It is important to estimate the spillover effects that arise through each network; however, those network interactions are often unobserved, rendering unbiased estimation of spillover effects difficult. Egami (2021) develops sensitivity analysis methods for assessing the potential influence of unobserved networks on causal findings. Relatedly, An (2018) emphasizes the importance of collecting data on treatment diffusion to measure treatment interference properly and then to estimate the direct treatment effect, treatment interference effect, and treatment effect on interference.

Read the full article →