Skip to Content
BayesiaLabVideos, Tutorials, Examples, & Case StudiesCase Study: Vehicle Size, Weight, and Injury Risk

Case Study: Vehicle Size, Weight, and Injury Risk

Introduction

Objective

This paper aims to illustrate how Bayesian networks and BayesiaLab can help overcome certain limitations of traditional statistical methods in high-dimensional problem domains. We consider the vehicle safety discussion in the recent Final Rule on future CAFE standards issued by the Environmental Protection Agency (EPA) and the National Highway Traffic Safety Administration (NHTSA) to be an ideal topic for our demonstration purposes.

When we reference the EPA/NHTSA Final Rule, we are specifically referring to the version of the document that was signed on August 28, 2012, and subsequently submitted to the Federal Register. However, when discussing the overall rationale presented in the Final Rule, we also implicitly include all the supporting studies that informed it. The Corporate Average Fuel Economy (CAFE) regulation was enacted by the U.S. Congress in 1975 with the goal of improving the average fuel economy of passenger cars and light trucks.

This paper focuses on technique rather than the subject matter itself, but our findings will undoubtedly yield new insights. We do not intend to challenge the conclusions of the EPA/NHTSA Final Rule. Instead, we aim to examine the overall problem domain independently while considering the rationale outlined in the Final Rule. Rather than simply replicating existing analyses with different tools, we will incorporate a broader set of variables and employ alternative methods to provide a complementary perspective on certain aspects of this issue. By moving beyond the traditional parametric methods used in EPA/NHTSA studies, we intend to demonstrate how Bayesian networks can serve as a robust framework for forecasting the impacts of regulatory interventions. Ultimately, our goal is to utilize Bayesian networks to evaluate the consequences of actions that have yet to be taken.

We will restate several original research questions to better align with our explanatory goals. While the EPA/NHTSA required a macro view of this domain, focusing on societal costs and benefits, we believe that Bayesian networks are particularly effective for understanding high-dimensional dynamics at a micro level. Therefore, we will examine this area with greater detail by incorporating additional accident attributes and using finer measurement scales.

For the sake of clarity, we will also limit our study to a more narrowly defined context, specifically vehicle-to-vehicle collisions rather than all types of motor vehicle accidents. It is important to emphasize that our analysis will focus solely on vehicle safety. We will not address any of the environmental justifications presented in the EPA/NHTSA Final Rule. Thus, our focus will be on a small segment of the overall problem domain.

This paper demonstrates a typical research workflow by presenting a sequence of alternating questions and answers. Throughout this discussion, we will gradually introduce various concepts specific to Bayesian networks, addressing each topic as it arises. In the initial chapters, we focus on providing extensive detail, including step-by-step instructions and numerous screenshots for using BayesiaLab. As we progress to later chapters, we will begin to simplify some of the technical aspects to emphasize the broader perspective of Bayesian networks as a powerful framework for reasoning.

Background

In October 2012, the Environmental Protection Agency (EPA) and the National Highway Traffic Safety Administration (NHTSA) issued the Final Rule, “2017 and Later Model Year Light-Duty Vehicle Greenhouse Gas Emissions and Corporate Average Fuel Economy Standards.”

One of the most important concerns in the Final Rule was its potential impact on vehicle safety. This should not be surprising as it is a commonly held notion that larger and heavier vehicles, which are less fuel-efficient, are generally safer in accidents. This belief is supported by the principle of conservation of linear momentum and Newton’s well-known laws of motion. In collisions of two objects of different mass, the deceleration force acting on the heavier object is smaller. Secondly, larger vehicles typically have longer crumple zones that extend the time over which the velocity change occurs, thus reducing the deceleration. Vehicle manufacturers and independent organizations have observed this many times in crash tests under controlled laboratory conditions.

It is also known that vehicle size and weight are key factors for fuel economy (please note that “Weight” and “mass” are used interchangeably throughout this paper). More specifically, the energy required to propel a vehicle over any given distance is a linear function of the vehicle’s frontal area and mass. Thus, a reduction in mass directly translates into a reduced energy requirement, i.e. lower fuel consumption.

Therefore, at least in theory, a conflict of objectives arises between vehicle safety and fuel economy. The question is, what does the real world look like? Are smaller, lighter cars really putting passengers at substantially greater risk of injury or death? One could hypothesize that so many other factors influence the probability and severity of injuries, including highly advanced restraint systems, that vehicle size may ultimately not determine life or death.

Given that the government, both at the state and the federal level, has collected records regarding hundreds of thousands of accidents over decades, one would imagine that modern data analysis can produce an in-depth understanding of injury risk in real-world vehicle crashes.

This is precisely what EPA and NHTSA did in order to estimate the societal costs and benefits of the proposed new CAFE rule. In fact, a large portion of the 1994-page Final Rule is devoted to discussing vehicle safety. Based on their technical and statistical analyses, they conclude that there is a safety-neutral compliance path with the new CAFE standards that includes mass reduction (we refer to the version of the document signed on August 28, 2012, which was submitted to the Federal Register. Page numbers refer to this version only).

General Considerations

To provide motivation and context for our proposed workflow, we will briefly discuss a number of initial thoughts regarding the EPA/NHTSA Final Rule. As an introduction to the technical discussion, we will first bring up a number of general considerations about the problem domain that will influence our approach.

Active Versus Passive Safety

The EPA/NHTSA studies have used “fatalities by estimated vehicle miles traveled (VMT)” as the principal dependent variable. This measure thus reflects all contributing as well as mitigating factors with regard to fatality risk. This includes human characteristics and behavior, environmental conditions, and vehicle characteristics, and behavior (e.g. small passenger car with ABS and ESP). In fact, the fatality risk is a function of one’s own attributes as well as the attributes of any other participant in the accident. In order to model the impact of vehicle weight reduction at the society-level, one would naturally have to take all of the above into account.

As opposed to a society-level analysis, we are approaching this domain more narrowly by looking at the risk of injury only as a function of vehicle characteristics and accident attributes. We believe that this approach helps to isolate vehicle crashworthiness, i.e. a vehicle’s passive safety performance, as opposed to performing a joint analysis of crash propensity and crashworthiness. This implies that we omit the potential relevance of vehicle attributes and occupant characteristics with regard to preventing an accident, i.e. active safety performance. It would be quite reasonable to include the role of vehicle weight in the context of active safety. For instance, the braking distance of a vehicle is, among other things, a function of vehicle mass. Similarly, occupant characteristics most certainly affect the probability of accidents, with younger drivers being a well-known high-risk group.

As a result of drivers’ characteristics and vehicles’ behavior, at least a portion of victims (and their vehicles) “self-select” themselves through their actions to “participate” in an accident. Speaking in epidemiological terms, our study may thus be subject to a self-selection bias. This would indeed be an issue that would have to be addressed for society-level inference. However, this potential self-selection bias should not interfere with our demonstration of the workflow while exclusively focusing on passive safety performance.

Dependent Variable

The EPA/NHTSA studies use a binary response variable, i.e. fatal vs. non-fatal, in order to measure accident outcomes. In the narrower context of our study, we believe that a binary response variable may not be comprehensive enough to characterize the passive safety performance of a vehicle.

Also, survival is not only a function of the passive safety performance of a vehicle during an accident, but it is also influenced by the quality of the medical care provided to the accident victim after the accident.

While it is a widely held belief among experts that vehicle safety has much improved over the last decade, the recent study by Glance et al. (2012) reports that, given the same injury level, there has also been a significant reduction in mortality of trauma patients since 2002.

“In-hospital mortality and major complications for adult trauma patients admitted to level I or level II trauma centers declined by 30% between 2000 and 2009. After stratifying patients by injury severity, the mortality rate for patients presenting with moderate or severe injuries declined by 40% to 50%, whereas mortality rates remained unchanged in patients with the least severe or the most severe injuries.”

Given that the fatality data that was used to inform the EPA/NHTSA Final Rule was collected between 2002 and 2008, we speculate that identical injuries could have had different outcomes, i.e. fatal versus non-fatal, as a function of the year when the injury occurred. Thus, we find it important to use an outcome variable that characterizes the severity of injuries sustained during the accident, as opposed to only counting fatalities.

Covariates

Similar to the binary fatal/non-fatal classification, other key variables in the EPA/NHTSA studies are also binned into two states, e.g., two weight classes Cars<2,950lbs.,Cars>2,950lbs.{Cars<2,950 lbs., Cars>2,950 lbs.} While the discretization of variables will also become necessary in our approach with Bayesian networks, we hypothesize that using two bins may be too “coarse” as a starting point. By using two intervals only, we would implicitly make the assumption of linearity in estimating the effect of vehicle weight on the dependent variable.

Furthermore, we speculate that a number of potentially relevant covariates can be added to provide a richer description of the accident dynamics. For instance, in a collision between two vehicles, we presume the angle of impact to be relevant, e.g. whether an accident is a frontal collision or a side impact. Also, specifically for two-vehicle collisions, we consider that the mass of both vehicles is important, as opposed to measuring this variable for one vehicle only. We will attempt to address these points with our selection of data sources and variables.

Consumer Response

The “law of unintended consequences” has become an idiomatic warning that an intervention in a complex system often creates unanticipated and undesirable outcomes. One such unintended consequence might be the consumers’ response to the new CAFE rule.

The EPA/NHTSA Final Rule notes that all statistical models suggest a mass reduction in small cars would be harmful or, at best, close to neutral and that the consumer choice behavior given price increases is unknown. Also, the EPA/NHTSA Final Rule has put great emphasis on preventing vehicle manufacturers from “downsizing” vehicles as a result of the CAFE rule: “in the agencies’ judgment, footprint-based standards [for manufacturers] discourage vehicle downsizing that might compromise occupant protection.” (EPA Final Rule, p. 214).

However, the EPA/NHTSA Final Rule does not provide an impact assessment with regard to future consumer choices in response to the new standards. Given that the Final Rule states that vehicle prices for consumers will rise significantly, “between $1,461 and $1,616 per vehicle in MY 2025” (EPA Final Rule, p. 123) as a direct consequence of the CAFE rule, one can reasonably speculate that consumers might downsize their vehicles.

Rather, the Final Rule states: “Because the agencies have not yet developed sufficient confidence in their vehicle choice modeling efforts, we believe it is premature to use them in this rulemaking.” (EPA Final Rule, p. 310). We speculate that this may limit one’s ability to draw conclusions with regard to the overall societal cost.

Unfortunately, we currently lack the appropriate data to build a consumer response model that would address this question within our framework. However, in terms of the methodology, we have presented a vehicle choice modeling approach in our white paper, Modeling Vehicle Choice and Simulating Market Share.

Technical Considerations

Assumption of Functional Forms

Given the familiar laws of physics that are applicable to collisions, one could hypothesize about certain functional forms for modeling the mechanisms that cause injuries of vehicle passengers. However, a priori, we cannot know whether any such assumptions are justified. Because this is a common challenge in many parametric statistical analyses, one would typically require a discussion regarding the choice of functional form, e.g. justifying the assumption of linearity.

We are not in a position to reexamine the choice of functional forms in the EPA/NHTSA studies. However, our proposed approach, learning Bayesian networks with BayesiaLab, has the advantage that no specification of any functional forms is required at all. Rather, BayesiaLab’s knowledge discovery algorithms use information-theoretic measures to search for any kind of probabilistic relationships between variables. As we will demonstrate later, we can capture the relationship between injury severity and angle of impact, which is clearly nonlinear.

Interactions and Collinearity

All of the studies supporting the EPA/NHTSA Final Rule use a broad set of control variables in their regression models. However, none of the studies use any interaction effects between these covariates. As such, an assumption is implicitly made that the covariates are all independent. However, examining the relationships between the covariates reveals that strong correlations do indeed exist, which violates the assumption of independence. In fact, collinearity is highlighted numerous times, e.g. “NHTSA considered the near multicollinearity of mass and footprint to be a major issue in the 2010 report and voiced concern about inaccurately estimated regression coefficients.” (Kahane, p. xi).

The nature of learning a Bayesian network does automatically take into account a multitude of potential relationships between all variables and can even include collinear relationships without a problem. We will see that countless relevant interactions between covariates exist, which are essential to capture the dynamics of the domain.

Causality

This last point is perhaps the most challenging one among the technical issues. The EPA/NHTSA studies use statistical models for purposes of causal inference. Statistical (or observational) inference, as in “given that we observe,” is not the same as causal inference, as in “given that we do.” Only under strict conditions, and with many additional assumptions, can we move from the former to the latter.9 Admittedly, causal inference from observational data is challenging and can be controversial. All the more it is important to clearly state the assumptions and why they might be justified.

With Bayesian networks we want to present a framework that allows researchers to explore this domain in a “causally correct” way, i.e., allowing, with the help of human domain knowledge, to disentangle “statistical correlation” and “causal effects.”

Sections

This case study is presented in five parts:

Summary

  1. This paper was developed as a case study for exhibiting the analytics and reasoning capabilities of the Bayesian network framework and the BayesiaLab software platform.
  2. Our study examined a subset of accidents with the objective of understanding injury drivers at a detail level, which could not be fully explored with the data and techniques used in the context of the EPA/NHTSA Final Rule.
  3. For an in-depth understanding of the dynamics of this problem domain, it was of great importance to capture the multitude of high-dimensional interactions between variables. We achieved this by learning Bayesian networks with BayesiaLab.
  4. On the basis of the Bayesian networks learned from data, BayesiaLab’s Likelihood Matching was used to estimate the exclusive Direct Effects of individual variables on the outcome variable. The estimated effects were generally consistent with prior domain knowledge and the laws of physics.
  5. Simulating domain interventions, e.g. the impact of regulatory action required carrying out causal inference, using Likelihood Matching. We emphasized the distinction between observational and causal inference in this context.
  6. With regard to this particular collision type, we conclude that injury risk remains a function of both mass and size. More specifically, as a result of decomposing the individual effects of vehicle size and weight, the notion of “mass reduction being safety-neutral given a fixed footprint” could not be supported in this specific context at this time.

Conclusion

Learning Bayesian networks from historical accident data using the BayesiaLab software platform allows us to comprehensively and compactly capture the complex dynamics of real-world vehicle crashes. With the domain encoded as a Bayesian network, we can “embrace” the high-dimensional interactions and leverage them for performing observational and causal inference. By employing Bayesian networks, we provide an improved framework for reasoning about vehicle size, weight, injury risk, and ultimately about the consequences of regulatory intervention.

References

2017 and Later Model Year Light-Duty Vehicle Greenhouse Gas Emissions and Corporate Average Fuel Economy Standards. Final Rule. Washington, D.C.: Department of Transportation, Environmental Protection Agency, National Highway Traffic Safety Administration, October 15, 2012. https://federalregister.gov/a/2012-21972 .

2017 and Later Model Year Light-Duty Vehicle Greenhouse Gas Emissions and Corporate Average Fuel Economy Standards. Department of Transportation, Environmental Protection Agency, National Highway Traffic Safety Administration, August 28, 2012.

A Dismissal of Safety, Choice, and Cost: The Obama Administration’s New Auto Regulations. Staff Report. Washington, D.C.: U.S. House of Representatives Committee on Oversight and Government Reform, August 10, 2012.

Bastani, Parisa, John B. Heywood, and Chris Hope. U.S. CAFE Standards - Potential for Meeting Light-duty Vehicle Fuel Economy Targets, 2016-2025. MIT Energy Initiative Report. Massachusetts Institute of Technology, January 2012.

http://web.mit.edu/sloan-auto-lab/research/beforeh2/files/CAFE_2012.pdf
PDF

Chen, T. Donna, and Kara M. Kockelman. “THE ROLES OF VEHICLE FOOTPRINT, HEIGHT, AND WEIGHT IN CRASH OUTCOMES: APPLICATION OF A HETEROSCEDASTIC ORDERED PRO-BIT MODEL.” In Transportation Research Board 91st Annual Meeting, 2012. http://www.ce.utexas.edu/prof/kockelman/public_html/TRB12CrashFootprint.pdf .

“Compliance Question - Will Automakers Build Bigger Trucks to Get Around New CAFE Regulations?” Autoweek. Accessed September 9, 2012. http://www.autoweek.com/article/20060407/free/60403023. 

Conrady, Stefan, and Lionel Jouffe. “Causal Inference and Direct Effects - Pearl’s Graph Surgery and Jouffe’s Likelihood Matching Illustrated with Simpson’s Paradox and a Marketing Mix Model,” September 15, 2011.

Conrady, Stefan, and Lionel Jouffe. Case Study: Modeling Vehicle Choice and Simulating Market Share

“Crashworthiness Data System - 2009 Coding and Editing Manual.” U.S. Department of Transportation National Highway Traffic Safety Administration, January 2009.

“Crashworthiness Data System - 2010 Coding and Editing Manual,” n.d.