Friday, May 31, 2019

How Does a Coach's Tenure Affect Team Performance?

Introduction:

Over the past ten seasons, the Phoenix Suns of the NBA have had 8 head coaches. It's no secret that coaches of professional sports teams have some of the most tenuous jobs in all of America. One bad season, or even a few unfortunate injuries, and coaches are usually the first ones to pack their bags. Each front office has a different approach to deciding when to move on from a hire: some coaches are fired without a chance to implement their systems, some are fired simply due to plain incompetence, and some are given leeway to organize their systems, only to have no success. What measures do front offices use to determine when a coach is ineffective?

 
All 6 Phoenix Suns non-interim head coaches from 2008-2018
 From left to right (Mike D'antoni 2003-2008, Terry Porter 2008-2009, Alvin Gentry 2009-2013, 
Jeff Hornacek 2013-2016, Earl Watson 2016-2017, Igor Kokoškov 2018-2019)

The easiest example of the "coaching carousel" would be the Cleveland Browns of the NFL. Similar to the Suns, the Browns have also had 8 coaches in the past decade. After 4 years of consecutive head coach firings, the new General Manager of the Browns, Sashi Brown, entered with a different plan. Brown hired Hue Jackson, a longstanding assistant coach in the NFL. Jackson was given much more leeway to implement his system than his predecessors, yet still failed spectacularly. Over the course of  two-and-a-half seasons, Jackson totaled a record of 3-36-1, the worst stretch by a single head coach in the NFL history. This dismal streak included a completely win-less season, the 2nd time that had ever been accomplished in the NFL and a staggering 0-20 record on the road. Jackson's tenure raises many questions, for starters, How important is coach tenure? If it is important, then was Jackson simply the wrong guy at the wrong time? Furthermore, at what point is a front office able to differentiate between growing pains and incompetence? Through regression analysis, I hope to provide analytic insight that can augment front office decision making to more efficiently move from hire to hire.

Despite an NFL coaching resume stretching over 15 years, Hue Jackson was not the right fit for the Browns.
(Photo: David Dermer, USA TODAY Sports)

Theoretical Framework:

The choice between panel data models was a point of weakness in my knowledge and an area I put a lot of thought and research into. My belief is that the unobserved effect on winning varies between team to team, this would seem to make the Random Effects model a better fit for the data. However, Jeffrey M. Wooldridge, author of the econometrics textbook I used as a guideline, emphasizes that correlation (or lack thereof) between the unobserved effect and the dependent variables is the most important deciding factor on choosing between Random and Fixed Effects. Since I don't believe that I can justify the covariance of a_i and x_it being zero, I will instead use the Fixed Effects model, however I will include the Random Effects as a robustness measure to check if coefficients are drastically different between the two.

OFFRTG and DEFRTG are advanced analytics ratings created by basketball analyst Dean Oliver, they combine many statistics to create a value that rates how well a team performs on offense or how effective a team is at preventing points on defense. Since these formulas include a wide variety of variables, I will omit the basic counting statistics such as points or rebounds to avoid any potential multicollinearity problems. The expected result over additional years of coaching experience would be that offensive rating would increase and defensive rating would decrease as players have more familiarity with the implemented systems.

While offensive and defensive rating measure on the court performance: coach tenure, turnover percent, total cap spending, and average age represent managerial and front office decision making. The effect of coach tenure on wins is presumed to be positive; as a coach is able to train players and players have better understandings of their role within the system, teams should exhibit better results. Furthermore, additional years allow front offices to tailor drafted players or free agents to a coaches' strengths. On the subject of front offices, total cap spending and average age are variables I chose to represent the current stage of competition that the team is in. The reasoning is that teams with championship ambitions will certainly have one or more players being paid the maximum contract. In addition, teams competing for a title tend to devalue their draft picks and young players as they tend to not align with the their time frame for winning, as such, younger teams should be expected to win less games.

Ownership spending is an area that is certainly relevant to the success of a team, but does not have publicly available data. One would expect a team to perform better when given access to state-of-the-art training facilities, or a robust analytics staff. But, for the purposes of this paper, we will have to rely on publicly available front office data to test our assumptions.

According to correlation matrix and variance inflation factor below, none of the dependent variables exhibit a problematic amount of multicollinearity. The correlation matrix highlights the previously mentioned relationship between cap spending and success as 61% of total cap can be explained through offensive rating. In addition, the matrix highlights the detrimental bivariate relationship that turnovers have on every metric in the model. R does not have a command that runs a VIF test on Random or Fixed effects models, the recommended solution for this online was to use a pooling model for the test. With these concerns of the Classical Assumptions assuaged, the next section will focus on exploratory data analysis.


Data:

Salary cap data was pulled from Spotrac, coaching tenure data was pulled from Basketball-reference, and advanced analytics data was pulled from the NBA's stats page.

Starting off, it's important to get a sense of what the current coaching situation looks like. The histogram below represents the number of years each team has had their current coach for the 2018-2019 season. There were four mid-season coaching changes, With a mean tenure of just over 3.9 years, one data point on the far reaches of the graph is hard to ignore.

As of 2019, Gregg Popovich of the San Antonio Spurs has notched his 23rd season as their head coach. By many, Popovich is considered the greatest head coach of all time, and one glance at his accomplishments adds legitimacy to that claim. During his tenure, Popovich has: won 5 NBA championships, set the record for most NBA wins, and has not missed the playoffs since the 1996-1997 season. By sporting standards, the Spurs' have sustained success for a lifetime. To put this in perspective: the New Orleans Pelicans did not exist as a franchise the last time the Spurs missed the playoffs, the Memphis Grizzlies were the Vancouver Grizzlies and had only existed for one year the last time the Spurs missed the playoffs.

Gregg Popovich and the success of the Spurs are monolithic, and serve as the poster child of how great coaching and management can affect a franchise. It should be noted that Popovich's coaching stardom did not follow the traditional trajectory. Popovich was originally the General Manager of the Spurs but fired his head coach, Bob Hill, and moved himself to head coach. Popovich's first season was a disaster, with a final record of 20-62, due to a large variety of injuries and an aging roster. Fortunately for Popovich, his team won the first pick in the NBA draft and selected Tim Duncan, who would go on to win 5 NBA championships with the Spurs and be considered one of the all time greats. Looking back in hindsight, the whole situation turned out great for Popovich, but had he not have been the GM, and instead just an interim head coach, he would not have had the leeway to execute such a plan.


Below is a graph demonstrating the relationship between offensive rating and coach tenure from 2015 to 2019. I originally used a static graph, but when pouring through the data I noticed some quite drastic swings in Offensive Rating. In just a few years, the results of teams designing their offenses around efficiency is apparent. This basketball strategy is sometimes called "Moreyball" named after Houston Rockets General Manager, Daryl Morey. Moreyball relies on two tenants: efficiency and pace. In the modern NBA, efficiency means emphasizing shot selection, favoring heavy three point shooting and avoiding mid-range shots. Moreyball takes this philosophy to the extreme, with a seemingly midrange-phobic strategy. According to FiveThirtyEight, 80% of the Houston Rockets shots occur either beyond the three point line or at the basket.


Empirical Results: 

As stated in the theoretical framework section, the model of focus is fixed effects, with the random effects model used as a robustness measure. The dependent variable in each specification is games won over the course of the 82 game regular season. The first and third specifications are the model as represented in the original framework. The magnitude of each independent variable is what was predicted in the theoretical framework section. Log of Total Cap, Average Age, Coach Tenure, Offensive Rating are all have a positive effect on Wins, while higher Defensive Rating and Turnovers have a negative.

Coach Tenure, the independent variable that comprises most of the focus of this paper, is very weakly positive, for each additional year a coach has with a given team, this model estimates that the team would win one tenth more games. I do not believe that the model in its current form models a coach's agency very well. Sure, turnovers are generally an indication of the discipline of a team, and a coach is responsible for instilling disciplined habits into their players, but there needs to be metrics that account for a coach's individual strategy or unique decision making. Of course, so much of a coach's decision making is private, either behind closed doors, or in a team huddle, that trying to model decision making is working with very incomplete information.

Both offensive rating and defensive rating are extremely correlated with wins, I believe these variables could be dominant variables and causing severe imperfect multicollinearity problems. Wooldridge's Introductory Econometrics makes no note of dominant variables in its text, but Studenmund's Using Econometrics defines a dominant variable as, "...so highly correlated with the dependent variable that it completely masks the effects of all other independent variables in the equation." Does Offensive Rating and Defensive Rating mask the effects of all other independent variables? The second and fourth specifications are the fixed effects and random effects models without Offensive and Defensive Rating, with those independent variables omitted, the Adjusted R-squared drops from 86% to just 19%. In addition, the other independent variables in specifications two and four change quite drastically in response to dropping the ratings. Average Age and Turnovers demonstrate a significantly discernible effect on the Wins, and the coefficient of Log of Total Cap changes magnitude from positive to negative. Although these are all quite drastic changes, this does not indicate that the rating variables are dominant. The reasoning for this is due to omitted variable bias, without Offensive Rating or Defensive Rating in the model, Turnovers per 100 is the only variable that denotes on-the-court performance. This gap in the model violates Classical Assumption three, which assumes that independent variables are uncorrelated with the error term. I would describe Offensive and Defensive Rating as highly important independent variables, rather than explicitly dominant.




Regression Results
Wins
FixedRTGFixedRandomRTGRandom
(1)(2)(3)(4)
Log of Total Cap0.207-3.0971.528-1.143
(1.855)(3.223)(1.580)(3.156)
Average Age0.3854.582***0.393*4.278***
(0.326)(0.640)(0.207)(0.596)
Coach Tenure0.0100.728-0.0940.377
(0.191)(0.454)(0.071)(0.299)
Offensive Rating2.336***2.459***
(0.134)(0.094)
Defensive Rating-2.478***-2.579***
(0.144)(0.102)
Turnovers per 100-0.272-1.962**0.111-1.644**
(0.384)(0.883)(0.257)(0.821)
Constant14.101-27.268
(23.742)(67.308)
Observations150150150150
R20.8930.3700.9410.335
Adjusted R20.8600.1900.9390.317
F Statistic158.553*** (df = 6; 114)17.000*** (df = 4; 116)2,299.846***73.034***

Notes:***Significant at the 1 percent level.
**Significant at the 5 percent level.
*Significant at the 10 percent level.


Conclusions and Future Research:

Although the testing failed to reject the null hypothesis that coaching tenure has no effect on winning percentage, this experiment yielded some interesting results and served as good practice in R. What this model does do well is highlight the importance of on-the-court success. There is an age old sports adage that goes, "Winning cures everything" and this research serves as another piece of proof to add to the pile.

I believe that this framework has applications for future research in fields beyond sports. The question of how upper management should evaluate lower management is the generalized problem this paper is trying to solve; this framework could be easily transformed into a model that tests how effective management experience is on improving quality control in a factory or any broader application of oversight on management. In the model's current state, however, there is not any measure of the quality of the decisions that a coach makes, only the quality of the on-the-court product. With Offensive and Defensive Ratings included, the model runs into multicollinearity concerns, but without them there is omitted variable bias representing on-the-court performance.

Future research into this topic would need to find a way to represent coaching decisions in a way that isn't reliant on counting statistics. Shot distribution percentages, how teams handle transition opportunities, and pick-and-roll frequency would be good starting points for modeling team strategy. One tradeoff of these metrics is they are relatively new concepts and only have a few years of data. By using these, the size of the dataset would decrease by 20%, going from 5 years to 4 years of data. For a model that I feel could use a larger sample, these new metrics may not result in a pareto improvement of the model. These are decisions I hope to explore further in a future papers.