Friday, January 10, 2020

Predicting the Infant Mortality Rate: Panel Data Regression of U.S. States from 2011-2014

In the spring semester of my junior year at Long Beach State, I created a linear regression based final project for a human development class. Despite being assigned as a qualitative research project, I explained to the professor that I had just renewed my Stata license and wanted an excuse to get value out of it. She gave me permission to implement a linear model on the condition that the qualitative aspect would not be compromised, and that it would not contain unaccessible econometrics jargon. This would be the basis of the paper that would go on to be named "Explaining Interstate Differences in the Infant Mortality Rate." In this article, I am going to use my insights and additional experience gained from almost two years of additional econometrics practice to critique and iterate upon the paper.

This update has been quite a long time in the making. The regression never felt like it modeled reality that well, and over the summer break following that semester I read some more literature and thought of potential avenues to explore. I was very fortunate to have an environmental economics professor in the fall semester who was fresh off her thesis on the effects of urban pollution on infant mortality rates. She provided me with great direction, and many of the adjustments I make in this paper are based on the notes I took during that meeting.

There are two reasons why it has taken me so long to update this model, after all, that meeting was in mid September 2018. First, my computer skills at the time were not at the level to create replicable longitudinal research. The panel data creation process was done in Excel, and was messy and overwhelming for me at the time. Second, some variables that I used for the original paper, in particular, physicians per capita, were only published publicly for every other year, and the alternative, health care expenditure per capita, hasn't been updated since 2014. This was a massive problem, since with data expansion, the new angle I wanted to take was focusing on the effects that the Affordable Care Act's Medicaid expansion had on states that passed the legislature. But since 2014 was the first time period in which states began to pass the expansion, the effects are surely to be lagged. But it's finally here, fully replicable and more mature.

The Good

  • Well Researched and Cited
    • I demonstrated good attention to detail for citations and providing support for claims made in the paper.
    • With 12 references, this is one of my papers with the most citations. I was very careful to cross all of my t's and dot the i's, especially since my professor criticized my project proposal for making claims without providing sufficient evidence.
  • Qualitative Research and Exploration of the Literature
    • I think that the introduction and literature review are the high water marks of the paper. Both sections clearly explain the problems I am attempting to explain, and provide interesting and thought provoking support for the thesis.
      • I really liked contextualizing individual state's mortality rates by comparing them with other nations. The United States is an extremely wealthy country, but there is quite a lot of heteroskedasticity beneath the surface.
  • Transforming Regression Analysis into Accessible Results
    • This paper meets the criteria laid out by my human development professor by avoiding insular statistics and econometrics jargon and creating a paper that is easy to understand the results of. However, do keep an eye out for a very similar bullet point in the "Ugly" section.

The Bad

  • A Snapshot of a Single Point in Time
    • This is a point that I am not particularly embarrassed by; if I had an "The Ehh" category, this would be a better fit. Ultimately, with one year of data, it is impossible to draw conclusions on a subject such as this due to missing out on time trends and other time series factors.
    • While this model would have benefited from more time periods, any panel data attempt would have been a black box. With a cross sectional model, I was aware of the potential pitfalls in creating a non biased estimation, but the same could not be said for panel data.
  • To Log, or not to Log
    • In the original paper, I expressed uncertainty as to proper specification of household income in the regression. Ultimately, the stance i ended up with was not taking the natural log of household income, but acknowledging the merits behind using a log. My environmental economics professor reinforced my position and said that in this case the data require a log. The chart below is a histogram from 2011 through 2014, which highlights the quite normal distribution of household income across states.
  • Omitted Variables
    • While there are certainly variables to be included that have a purpose in the model, there is not a glaring omission. 
    • Referring back to my notes from the advice I received, racial demographics were recommended by my professor to further provide support for interstate differences. Racial demographics are not even distributed across states, most notably, in the case of the Great Migration following the American Civil War and Reconstruction era.
    • Smoking rate and insurance coverage are two more health indicators that have relevant correlation with the health of a mother. Nicotine is a brutal habit to kick, and I would expect that states that have higher rates of tobacco dependence would have higher rates of pregnant smokers. Insurance coverage could be a good indicator of the propensity that a pregnant women has to avoid medical appointments due to cost.
Kilpinen, Jon T., 2014. "Percent Black or African American, 2010." United States Map Gallery. Map 01.04. http://scholar.valpo.edu/usmaps/9
Data source: United States Census Bureau

The Ugly

  • Gini Coefficient is a Linear Combination of Household Income
    • This is easily the biggest mistake I made in the modeling process. The Gini Coefficient is based on the Lorenz Curve, which models the distribution of cumulative income across the population. Since the sum of all household income is the cumulative income of a country, including both the Gini Coefficient and household income per capita biases the results of the model.
  • General Misuse of Statistics and Interpretation of OLS Models
    • Sound familiar? Overall, this paper compromises the theory at the benefit of being accessible. The interpretation of the regression results are wrong, in equal parts due to a lack of experience and a desire to avoid jargon like "statistical significance."

Atonement

Data

With nearly two years of additional statistical programming and data cleaning experience, I can comfortably create the replicable research required for this project. When organizing the individual csv files, I even implemented some automation to speed up the process and reduce the lines of code. By using the "lapply" or listapply function, I put all 20 of the individual data frames into a list and wrote code that then applied across each data frame in the list. The code, datasets, and graphics created for use in this paper are published on my GitHub.

As mentioned earlier in the paper, I elected to replace physicians per capita with healthcare expenditures per capita (HCE) due to the ease of access to health expenditures data. The tradeoff with this decision is having the data pressed between 2011 and 2014. Going further back in time with this dataset would result in having a lot more exogenous factors relating to the financial crisis such as unemployment and stress level. If the Centers for Medicare & Medicare Services would ever update their interstate data, the panel data could be extended out towards 2018. Another shortcoming I would like to acknowledge is the nominal dollars used in HCE and personal income per capita. A major advantage of creating this project within R is that I can add in data to control for inflation upstream, but I did not think of that during the research process. While inflation in the time periods following the financial crisis were historically low, this would still be a pareto improvement to the model.

Other variables added to the original estimation are: racial demographics for the four most populous races in the U.S., proportion of Medicaid users, proportion of the uninsured, emissions per capita, and the smoking rate in each state. The expected coefficients of Medicaid recipients and the uninsured are

Table 1 highlights the summary statistics of the dataset, the states with the lowest and highest infant mortality rates in the dataset are Alaska in 2011 and Mississippi in 2013, respectively. Alaska's result is quite interesting, while their IMR was lowest in 2011, it has steadily increased to 5.10 deaths per 1000 in 2012, then 5.77, to 6.67 in 2014, a 74% increase just over this four year period. While this may seem like a very convincing trend, Alaska's IMR seems to have a lot of variance. The years preceding 2011 bounce around from mid 3 deaths to high 6 deaths, and the years following 2014 average out to 6.1 deaths. The answer could just be the variance of a low, rural population like Alaska, or it could be the result of policy that manages to extend healthcare data tracking to the disconnected, native populations.

Theoretical Model

I will be using a fixed effects model, this choice is influenced by two factors: first, in the discussion I had with my environmental economics professor, she recommended a fixed effects model. Second, when exploring the data, I used the Hausman Test to test for consistency between the fixed and random effects models. With a p-value of .009, the null hypothesis that both models are consistent is rejected, which suggests that the fixed effects model is a better fit. While I do not want to rely on a one-size-fits-all test like the Hausman Test as justification, these two points in favor of the fixed effects model lead me to choose it as the main model.

A downside of using fixed effects is the lack of measurement of time-invariant factors; this means that the original intent of measuring differences between groups of states will not be possible. Instead, this model will focus on the changes within states that affect the IMR.

In addition to a fixed effects model, I also recreated the original OLS model used in my final paper, with a few differences. Instead of being based in 2016, the model is based in 2014 due to the data availability issues referenced earlier. Physicians per capita is replaced with HCE, and the South dummy variable is changed to Southeast, based along the Bureau of Economic Analysis' regions. This change is quite minor, but since I am doing economic research, I figured the delineations that the BEA makes is better suited for this project.

The expectations for the magnitude of the coefficients remains the same from the original paper. I expect personal income and health care expenditures to have inverse relationships with IMR, since one would expect that as personal income and HCE rises, the quality of healthcare obtained rises. Obesity rate is expected to be positively correlated with IMR, as states with higher concentrations of obesity would be expected to have higher rates of obesity during pregnancy, which is linked to a wide variety of health issues, as stated by the Mayo Clinic.

With the updated model, Medicaid and uninsured rates would be expected to be positively correlated to IMR. According to this report published by the Kaiser Family Foundation, while Medicaid is not perfect, adults on Medicaid increase their access to preventative and regular medical care, especially when compared to uninsured adults.

Based solely along household income demographics, we would expect states with higher proportions of whites and asians to have lower infant mortality rates. Conversely, states with higher proportions of hispanics and african americans would be expected to have higher rates of infant mortality. However, only accounting for household income would be a very reductionist view of the relationship between race and health indicators like IMR. In fact, the argument of the effects of race on health outcomes versus socioeconomic status on health outcomes has continued for hundreds of years. As noted by Kawachi, Daniels, and Robinson, this debate in the United States predates the Civil War, with 





Summary Statistics
StatisticNMeanSt. Dev.MinPctl(25)Pctl(75)Max
IMR2046.1231.1943.8405.1076.8939.600
Medicaid.Only2040.1270.0310.0500.1100.1500.220
Uninsured2040.1280.0420.0300.1000.1600.230
HCE2047,905.8531,228.9785,3417,019.58,759.511,944
pi_pc20444,422.5807,758.27232,16338,541.848,335.871,334
ObesityRate20428.3523.3572026.130.736
Black2040.1080.1080.0040.0300.1500.500
SmokingRate20419.7163.6179.70017.20022.00029.000
emi_pc20421.92618.5084.20011.85024.625117.200

Linear Regressions of Infant Mortality Rate
Cross Sectional IMR Fixed Effects IMR
(Intercept) 0.398 (1.805)      
HCE1000 0.082 (0.130) -0.355 (0.287)
pi_pc1000 -0.006 (0.024) -0.022 (0.048)
ObesityRate 0.171 *** (0.045) -0.049 (0.047)
southeast 0.896 ** (0.335)      
Medicaid.Only       0.202 (5.085)
Uninsured       -2.765 (4.577)
Black       19.762 (13.697)
emi_pc       -0.174 *** (0.046)
SmokingRate       -0.017 (0.046)
N 51      204     
R2 0.514  0.128 
logLik -60.943       
AIC 133.885       
*** p < 0.01; ** p < 0.05; * p < 0.1.