Wednesday, May 15, 2013

The Winning Formula: Balance

It may be the trendiest topic in the NBA: balance. Do you need it? Do you want it? What really is it? Every NBA GM, analyst or fan may have a different opinion. Should a team be built around one (or more) superstar(s) or should it be evenly assembled with complimentary parts? Not only has it become popular for superstar players to pair up (Heat) or teams go all-out to obtain a superstar (Rockets), it has become commonly accepted that this is the best way to build a successful team. However, when well-balanced teams such as this year’s Nuggets showed us inequality might be overrated, I began wondering which is the better way to build a team. The truth is, there really is no right answer; there have been successful teams built with both symmetric and skewed distributions. However, by analyzing certain performance statistics, Win Shares and salary data from the past 11 NBA seasons, I tried to determine which approach was more reliable.

Read more after the jump



Basic Statistics

I wanted to find a way to measure this balance, or imbalance, and analyze whether it was to a team's benefit, or detriment. A perfect way to do this was using the statistic of standard deviation. Standard deviation measures spread from the mean, so if a team has a higher standard deviation of points, they had more spread out scoring. If they have a lower standard deviation of points, they had more balanced scoring. However, this measure is greatly influenced by the actual amount of points scored. Teams with who score more points are naturally going to have a higher standard deviation of points. To control for this, I divided each players scoring total by the total number of team points, resulting in a percentage that estimates each player's share of the team’s points scored in a given season. I then took the standard deviation of those percentages, resulting in the Adjusted standard deviation of points (Adjusted SD). A team with a higher Adjusted SD of points had a majority of its points come from a few players, while a team with a lower Adjusted SD of points had a balanced scoring attack. The same obviously applies for assists, rebounds, 3-pointers, steals, blocks, and turnovers.
           
It turns out that the way these stats are distributed amongst the team has a significant impact on the season total of the stat, the team’s offensive rating (Points/100 possessions), and their overall winning percentage. Using single regressions, I found that having more spread out scoring contributions (a higher Adjusted SD of points) leads to more points overall, a higher offensive rating and more wins. The same applies to assists, rebounds, steals and blocks. In other words, uneven assists (or rebounds, etc) leads to more overall assists, a higher offensive rating, and more wins. The same logic even applies to turnovers, as having a wide spread of turnovers leads to more on court success, probably because you want your turnovers limited to your primary ball handler. Interestingly enough, unlike all the previous stats mentioned, having a wider spread of turnovers doesn’t predict having more overall turnovers.

This is the same thing Nima Shaahinfar found in his analysis, summarized here. His results differed from mine in that he found that rebounds should be evenly dispersed amongst a team, as it creates a more efficient offensive and defensive unit. Shaahinfar reasoned this means better offense and defense arecreated if everybody crashes the boards. He used lineup statistics, whereas I used season aggregate data. He accounted for the stats of players in relation to the lineup they played with, while I used season totals, combining all of a team’s lineups into one data set. My method may be less precise, but I still feel that looking at how a team’s points, rebounds, etc, were allocated across an entire season is a valuable exercise and the results still have significance.

Using my data in multiple regression models yielded some note-worthy results. The most interesting one is displayed below in regression model 1, predicting offensive rating. The model predicts that, while controlling for how well a team shoots overall (FG%) and how well it shoots the three (3P%), the spread of those 3-point shot attempts and overall shot attempts is significant. With a positive coefficient on Adjusted SD of Field Goal Attempts and a negative coefficient on Adjusted SD of 3-Point Attempts, the model tell us that teams benefit from an uneven distribution of overall shots, but an even distribution of 3-point attempts. This shows us two things. First, successful offensive teams have many guys taking threes. Second, taking the Adjusted SD of 3PA into account, the fact FGA should be unbalanced means that our 2-point shots should have an uneven distribution as well. These results coincide with Shaahinfar’s results displayed in his blog.


Click to Enlarge

  Note: the analysis here is kind of tricky. If you have a lower Adjusted SD of 3PA, or a more even distribution of people shooting threes, you probably have just more guys who can shoot threes. Every team has at least three players who can hit a three, but when your guys 5, 6, and 7 are shooting threes, they are shooting them for a reason; they’re probably good at them. So it makes sense that a team with a low Adjusted SD of 3PA has a higher Offensive Rating, they have a lot of guys who can hit threes. The practical advice: load up on effective players who can shoot threes.

All of these results must be taken with a grain of salt. For example, the Warriors this past season had the most spread out 3PA (highest Adjusted SD of 3PA) mostly due to the Splash Brothers. The numbers show that offensive efficiency increases when those measures are lower, but this wasn’t the case for the Warriors and no team is going to pass up on Reggie Miller 2.0 and Reggie Miller 2.5 to keep their spreads as even as possible. The point is, over the past eleven seasons, the trend is that better offensive teams have had a balanced 3-point attack, but clearly every team has their own formula.

An interesting way to look at this concept of balance and imbalance is shot attempts and points. I ran single regressions on both the Adjusted SD of Field Goals Attempted and points. The models predicted an increase in the adjusted spread of both shots taken and points scored is better for your team. Did this hold true this past season? Well, listed below are the top and bottom 10 teams in spreading out both shot attempts and points. As you can see, this is an indicator of success.

 
Click to Enlarge


Win Shares

Another interesting stat to look at is Win Shares, basketball’s version of Wins Above Replacement(WAR). Win Shares relies heavily on Offensive and Defensive ratings. Developed by Dean Oliver, Offensive and Defensive ratings mostly use basic box score statisticssuch as PTS, FTA, REB, 3PM, AST, STL, BLK, and TOV as the inputs. By encompassing how effective a player was on the offensive and defensive end, a Win Share essentially represents the number of wins contributed by a given player. The sum of all the Win Shares on a team results in a number very close to their actual win total. So, by dividing each Win Share by the total number of Win Shares, the resulting percentages represent the proportion of a single win that can be credited to a given player. Looking at the standard deviation of these percentages (Adjusted SD of Win Shares) shows the distribution of a team’s overall contributions. How spread out or balanced does a team want its players’ individual impacts to be? Interestingly enough, it is better to be as balanced as possible. When put into a regression model while controlling for previous winning percentage, the Adjusted SD of Win Shares had a negative coefficient, meaning more spread out Win Shares is damaging for a team. In regression 2 below, I added in the Adjusted SD of specific stats that are part of the Win Share formula and modeled the non-independence amongst teams by including a random intercept. The results are similar.

Click to Enlarge
-->I also performed a similar regression, but added in the totals of all the individual stats as well (total points, assists, etc) and the adjusted SD of these stats all remained significant. This tell us that both the quantity and distribution of these stats are important.

Wait how does this make sense? You want more spread-out scoring, passing, rebounding and defending statistics, yet a balanced distribution of Win Shares, a statistic that is essentially a direct measure of a players scoring, passing, rebounding and defending? Yes, this is true and it stresses an important point. Players must have roles. Successful teams have players playing to their strengths in certain areas. They have players who score, other players who rebound, others who assist; all ideally contributing to a balanced distribution of Win Shares.

The takeaway here is that the public perception of players is skewed towards individual stats over team impact. I claimed earlier that an uneven dispersion of points, or having a player dominate the scoring, predicted more overall scoring. However, the correct way the phrase it is that having a player whose role is to score leads to a more effective offense. The same goes for having a player whose role it is to pass and get more assists, etc.

Other Measures

Want one more way to look at balance on a team? How about the allocation of minutes and funds? The results from a multiple regression predicting winning percentage from the standard deviation of minutes and Adjusted SD of Salary (because some teams differ a lot in their overall salary) are below.
 
Click to Enlarge

When regressing against winning percentage while controlling for previous year’s success, the spread of a team’s payroll has a positive relationship. Essentially, more successful teams have an unbalanced distribution of their salaries. The same results were found for minutes played; a more uneven allocation of minutes predicts more success. The reasoning behind regression 3’s results are much more obvious. If a team is evenly distributing its minutes and not having just a few players play the majority of the time, it probably doesn’t have any great players. And if a team’s salary is very balanced, it probably doesn’t have a great player on the team who is worth a big contract.

Applying this same logic to the previously mentioned measures of spread like points and rebounds, it makes sense why bad teams tend to have more balanced scoring. They have nobody who can score in bunches and differentiate himself from the rest of the pack. That’s why teams pay a premium for production, why one-dimensional scorers like Carmelo Anthony get max-deals, and even why irrational scorers like Michael Beasley, JR Smith, and Jamal Crawford have a place in this league (Thank God). But what we learned earlier with the application of Win Shares into the equation is that if you have a scorer, leave him to scoring. Surround him with other guys who can defend, rebound, pass and hit threes. Everybody else should ideally contribute an even share of these other statistics.

Conclusion

I started this post talking about balance. Do NBA teams really want balance? Well the answer is yes, in some respects, and no, in others. It depends on your team. An optimal roster should have unbalanced salaries, but this doesn’t mean all players shouldn’t contribute. Successful teams have had a more even distribution of Win Shares, meaning they receive significant contributions from everyone. You want players with specific roles, and players who know those roles. Does that mean you don’t want a player like LeBron James, who can lead his team in rebounds, points and assists and any given night? Of course not, he’s the best player in the league. No team isn’t going to sign LeBron because he will skew their allocation of Win Shares. Certain players tend to ruin all the analytics done in the NBA. They usually have LeBron or Durant in their name. Thanks a lot guys.


Contact Joey Shampain at joseph.shampain@gmail.com with any questions

Also check out part one!

Labels: , , ,

Wednesday, May 8, 2013

The Winning Formula: Age and Experience


Though the analytics first broke into the sports world in baseball thanks to Bill James, sabremetrics, and Brad Pitt (Billy Beane), basketball is now riding the wave as well, and arguably even higher. Like the MLB (and most other major sports), NBA front offices now have a statistical and analytical focus. Teams and coaches are viewing players through different, more diagnostic, lenses. Writers such as Zach Loweand Kevin Peltonare doing what Jonah Hill (Peter Brand) did for sabremetrics, translating the complexities front office and coaches are analyzing on a daily basis into layman’s terms. Being an avid NBA fan and typical nerd, I have grown to admire the analytic work done by the basketball community, so much so I decided to give it a try myself…

As a Statistical Science major here at Cornell, I am conducting an independent study, with ILR Organizational Behavior Professor Emily Zitek, looking at the effects of NBA team composition and performance on its overall success. Essentially, what is “The Winning Formula”? In a series of blog posts within the next few weeks, I will discuss many factors that contribute to a team’s success, such as experience, age, Pythagorean Wins, Win Shares, standard deviation of performance statistics and Dean Oliver’s Four Factors. Prerequisite knowledge is not needed, nor is a degree in statistics, only a curiosity as to why your favorite NBA teams perform the way they do. Though I will not make any groundbreaking conclusions, I hope to paint clearer picture as to why certain teams are successful and others are not.

Two of the most frequently cited factors that determine success are age and experience. That team is too old (Lakers). That team is too young (Bobcats). That team is too inexperienced (Rockets). That team is so veteran-savvy (Spurs). Does all of this theoretical hypothesizing done by the media have factual merit? I tried to take a statistical approach to answer that question by analyzing team data from the past 11 seasons, from the 2002-2003 season until the just-finished 2012-2013 season.

First, it is important to understand that experience and age, though correlated, are two very different things. Chris Copeland was a 29-year-old rookie this year, with 0 years of NBA experience. John Wall and Rookie Damian Lillard are the same age, though John Wall was drafted 3 years ago.

Most citations of “youngest” and “oldest” team have to do with team averages. However, when looking at the age and experience of a team it doesn’t make sense to calculate the average age or experience amongst its players. Why should Grant Hill (Age 40, 17 years of NBA experience) and his 437 total minutes played this season have the same impact on Clippers’ age and experience as Eric Bledsoe (Age 23, 2 years of NBA experience), who played more than three times as many minutes (1553) as Hill? Therefore, by weighting a player’s age and experience by the percentage of his team’s minutes he played during the season, I created two new variables, Weighted Age and Weighted Experience. In this example, Eric Bledsoe’s spry 23 years get three times as much weight as Grant Hill’s ancient 40 years. Essentially, Weighted Age represents the average age of all 5 players on the floor at any given time throughout the season. The results of the Weighted Age and Weighted Experience calculations for Playoff and Non-Playoff teams in the last 11 seasons can be seen below.

As you can see, playoff teams, on average, give minutes to players 1.42 years older and with 1.47 more years of experience
Click to Enlarge

Weighted Age also allows us to look at another interesting team characteristic, the Age Premium. This is the difference the between the Weighted Age and Average Age of a team, or essentially the emphasis they put on age. If a team has a positive Age Premium, they are playing their older players more minutes. The same reasoning obviously applies for the Experience Premium. Playoff teams have higher Premiums that Non-Playoff Teams, illustrating that not only do Playoff Teams play older and more experienced players, they give their elderly players a greater share of the minutes than Non-Playoff Teams.

For the entire NBA, the negative Age Premium(-0.05) and positive Experience Premium(0.378) illustrates an interesting trend. Over the past 11 seasons, teams, on average, are giving playing time to players with more experience, yet slightly less age. This tells us front offices and coaches value greater NBA experience, but not greater player age. Why? As I will demonstrate, experience has been proven to be better indicator or success.

I ran simple and multiple regressions on the above variables to try to predict team winning percentages. I used a lagged dependent variable of the previous season’s winning percentage to take into consideration how good the team was last year and included a random intercept to model the non-independence amongst teams. Average Age and Average Experience were not used because, as previously stated, how is a 40-year old Grant Hill sitting on the bench for most of the season going to help you win? The below table gives the regression coefficients (standard errors in parentheses) with regular season winning percentage as the dependent variable.

Click to Enlarge

It can be seen that both Weighted Age and Weighted Experience are significant indicators of success considering last year’s performance. Due to the standard deviation of Weighted Experience being 1.6 years, regression 1 predicts that a team one standard deviation above the mean of Weighted Experience (6.698) can expect to win about 10.17% more of its games (8.34 games in an 82-game season) than a team one standard deviation below the mean (3.498). Age has a similar, but not as significant impact. Given the 1.68 year standard deviation of Weighted Age, regression 2 suggests that a team one standard deviation above the mean of Weighted Age (28.41) can expect to win about 9% more of its games (7.38 games in an 82-game regular season) than a team one standard deviation below the mean (25.05).

The strength of experience over age is stressed again when both are included in multiple regression model 3. Due to the correlated nature of Weighted Age and Weighted Experience, I can’t make a claim about an increase in wins predicted by a given increase in Weighted Experience, like in the previous paragraph. The coefficients can no longer be interpreted literally, yet they still show important concepts. After taking into account Weighted Age, Weighted Experience is still significant, however, after taking into account Weighted Experience, Weighted Age is insiginificant. Essentially, once you account for a team’s Weighted Experience, Weighted Age doesn’t explain any more of the variability remaining in team winning percentages. However, even if given Weighted Age and the previous year’s winning percentage, Weighted Experience can still add a lot to the picture.

Regression 4 mathematically proves an observable concept. Given how good a team was last year, their Experience Premium (the emphasis they put on age) is a significant predictor of success. Essentially, if a team is playing its players with more experience, it is more likely to be a better team. If a team is playing its younger players the majority of its minutes, it is more likely to be a bad team. This makes intuitive sense; teams with negative experience premiums are probably in their development stage, hence giving playing time to younger players, and teams with positive experience premiums are most likely built to win now. However, this is not true for all teams (this year’s Thunder have a negative Experience Premium for example), and as with most regressions, this is meant to show a trend, not the rule.

We’ve figured out that both having older, more experienced players on the floor tends to improve your chances of winning in the regular season, but for most NBA franchises (spare the Bobcats), the real goal isn’t to succeed in the regular season and make the playoffs, it's to win a championship. I ran the same four regressions as earlier, but used Playoff Wins, a perfect barometer of playoff success, as the dependent variable. This time, none of the independent variables besides previous winning percentage were found to be statistically significant in predicting playoff success. What does this tell us? Having an older, more experienced team is beneficial in the regular season, but doesn’t seem to have much of an effect in the playoffs. In fact, I even ran the regressions trying to predict regular season wins while restricting my sample to only teams who made the playoffs, and neither Weighted Age, Weighted Experience, nor the Experience Premium has a significant impact. Age and Experience are what separates the good teams from the bad teams, but doesn’t pull a leader out of the front of the pack.

These findings coincide with the conclusions made by James Tarlow in a paper titled "Experience and Winning in the National Basketball Association" presented at the 2012 MIT Sloan Sports Analytics Conference . However, Tarlow went even further, looking at team chemistry and coaching regular season and postseason experience. He found, unlike player experience, those variables are significant predictors of playoff success. I suggest giving the article a read.

It shouldn’t be shocking that experience matters more than age in the NBA, at least with regards to regular season performance. In fact, I would hypothesize that it matters in every sport and professional endeavor. NBA GM’s know this and build their rosters accordingly. A player’s old age is a detriment when entering the NBA Draft. Age for a player like 23 year-old Gorgui Dieng entering the draft this year is a clear disadvantage. A player like Dieng needs to develop, so by the time he has adequate experience he may only have a few years left in his physical prime.  On the other hand, Nerlens Noel is 19 years-old. NBA teams would love to get their hands on him and have his peak physical performance coincide with a sufficient level of NBA experience in his mid-20’s.

The takeaway here though, is that this is a trend, not the rule. You can be a very young and inexperienced team and make the playoffs, look at the Rockets this year. However, as James Tarlow argued, lack of coaching playoff experience and team chemistry will be a detriment. But, just for fun, here are the least experienced and youngest teams to get in the playoffs in the past 11 seasons

 

As you can see, the Thunder tend buck the trend. In fact, the 2011 Thunder won the most (9) postseason games out of all the above teams.  So what’s the real solution? Get Kevin Durant on your team. In all seriousness, this points to a significant point. Though you may have just wasted ten minutes of your life reading this post, there are many other factors besides age and experience influence winning and losing in the NBA. In fact, the R2 for best model (Regression 3) is .393, meaning only 39.3% of the variability in regular season winning percentages can be explained by a team’s Weighted Age, Weighted Experience and their previous winning percentage. So even though a team with “savvy veterans” may do better over the course of the regular season due to their experience, please realize this is not the only reason why. They may have Lebron James.

Contact Joey Shampain at joseph.shampain@gmail.com with any questions



Labels: , , ,