Showing posts with label competition. Show all posts
Showing posts with label competition. Show all posts

Tuesday, February 15, 2011

NBA League Size and Competitiveness, Part the Last: Calculating Competition

Today we embark on a journey through the perilous land of inventing your own statistics.  As a wrap-up for this series on NBA competitiveness and league size, I wanted to create a kind of competitiveness index based upon the regular season results in any given season.  That has proved to be a more difficult task then I originally anticipated, for reasons which will become clear.

First off, though, why would I want to do something like this?  Well, as I discussed in Part One, I'm reading The Book of Basketball, by Bill Simmons, and was struck by how definitive his assessment of which NBA seasons were competitive and which weren't is.  In particular, some of the early seasons in the NBA sparked comments like, "Everyone had a good team back then."  I wanted to try to figure out if he was right because, as sabermetrics has taught us, often people who are passionate and well-informed fans of a sport still don't really understand what's going on.

For example, in baseball it was long believed that carrying a .300 batting average alone was sufficient to make you a good hitter.  "A .300 hitter" was - still is - an honorable appellate, as well as a since qua non of baseball success.  What about a player like Juan Pierre, though, whose career .298 average puts him close enough to be called a .300 hitter? Is he really any good?  Old-time baseball wisdom would say yes.  He's fast, he hits for a high average, and he's the kind of guy that people assume is a good fielder, whether he is or not.  But even offensively, you can dive deeper into his batting lines and see that he's a deeply, deeply flawed player.

You see, Juan Pierre does not really draw walks.  Nor does he hit for power.  So despite a career .300 average, he sports a .347 OBP - not bad, but not good enough for someone who aspires to be an integral part of a team's success.  Moreover, his .366 career slugging percentage means that he's a singles hitter.  Those many hits he does generate aren't in the gaps or over the fence (as evidenced by his 14 career homers in almost 1600 games).  Now, the traditional baseball viewpoint would be that all of Pierre's singles are made up for by his stolen bases...  Which is fair, except he has lead the league in getting caught stealing 6 times, and stolen bases only 3 times.

 Which is all to say that being a .300 hitter alone used to look great, and still looks great.  But looks can be deceiving.  No one should confuse Juan Pierre with a great hitter.  Similarly, sometimes a league might look competitive without actually being competitive.  And so I embarked on this little blog-project to prove Simmons right and/or wrong.

In Part Two I found and discussed that, while defining competition - let alone assessing it - might be very difficult, we can at least see that, as the league gets larger, so too does the standard deviation of winning percentage.  From one perspective that means that the league is getting less competitive - in the sense that teams are less jumbled together - but from another it means the league is getting more competitive - in the sense that there are more elite teams in any given season.  And that's exactly what my work for today's post shows.

What I did was develop a formula for "competitiveness," using the number of above-.500 teams, the mean of their winning percentages, and the standard deviation of their winning percentages.  My reasoning was this: if a higher percentage of teams are above .500 in a given season, the league is more competitive.  Similarly, the higher the average winning percentage of those teams, the more competitive the league is.  Lastly, the more condensed those winning percentages are (the lower the standard deviation), the more competitive the league is.  The advantage of this approach, of course, is that we can completely ignore any team that finished .500 or worse.  Those teams, I reasoned, don't really count (even if many of them do make the playoffs, thanks to the NBA's "everybody makes it" attitude towards the postseason).  The disadvantage, as the statistically acute among you will see, is that the components I have selected here are all closely related to standard deviation of winning percentage league wide.

What does that mean?  Well, let me show you.

X - Number of teams, Y - "Competitiveness"
My formula for competitiveness is messy, but worth sharing.  Brackets indicate the separate components, which I tried to normalize so that 1 was more or less "average":

[(0.5 + Percentage of teams above .500)] x [10 x (mean above .500 - .5)] x [(stdev above .500 - mean above .500) / (stdev above .500 + mean above .500)] x 50

I multiplied the whole thing by 50 just to pull it up into a more readable and intuitive range.  Basically, 50 is normal (as you can see, the trendline above is close to, though not quite at, 50), while anything above 50 is a particularly competitive season, and anything below 40 is uncompetitive.

Now this graph alone doesn't show you anything problematic.  Like our graph from Part Two, it has a weak, but present trend going upwards, and...  Wait.  It looks very similar to that graph.

So I graphed "competitiveness" by year, and made the following line graph:

"Competitiveness" (Y) by season (X)

 I then did the same with Standard Deviation of winning percentage:

STDEV of Wpct (Y) by season (X)
Now you may notice that these two graphs look almost exactly the same.  With a sinking feeling - starting to realize the folly of my ways - I graphed the two against each other:


Competitiveness (Y) against STDEV of Wpct (X)
 The result is unambiguous.  My "Competitiveness" ranking basically tells me that when the standard deviation of the winning percentages league wide are high, the competitiveness is also high.  Which, of course, is the opposite of what I was suggesting in Part Three.  Yeah.

The result is hardly surprising, as I said, because of the components in my formula.  While the percentage of better-than-.500 teams may not have much bearing on standard deviation of winning percentage, obviously when the mean of the winning percentage of teams above .500 is higher, so too will be the standard deviation of winning percentage of all teams.  Meanwhile, the final component of my formula - accounting for standard deviation of above-.500 winning percentages - will be inversely related to standard deviation of winning percentages league wide, but not enough, obviously, to disrupt the high correlation between "competitiveness" and SD of winning percentage league wide.

But really, this only goes to show that "competitiveness" is a highly ambiguous term.  Where one fan might think the most competitive season is the one where all of the teams are bunched together, another might prefer the one with five great teams and five terrible ones.  It's really a matter of perspective.

Where Simmons makes his determination, then, is probably the best place: skill of players in the league.  While you do have to be careful here - because all evaluation of player skill is heavily influenced by the relative skills of his contemporaries, and things like changing league sizes mess with our understanding of what is good and what is great - probably the best way to assess the competitiveness of the league at any point is to assess the overall skill of the players in the league at a given time.  That's a much more challenging project, but I can imagine going through players and seeing where great careers overlap, and figuring out when talent has been at its apex and nadir.  Of course, Simmons does that kind of thing for a living - though without relying too much on numerical analysis and going more with his perception, a more-than-fair, if perilous, approach.  John Hollinger also does that for his living, relying absolutely on numbers.  So between the two of them, you can probably get a good sense of what's going on.

Finally, if you want to see the nuts and bolts of my work - messy as it is - I've posted my workbook to GoogleDocs.  Do with it what you will.

Friday, February 11, 2011

NBA League Size and Competitiveness, Part Three: Outliers

Before I dive into more statistical analysis in an effort to answer the question as to whether a larger NBA leads to stiffer competition or not, I want to take a brief (ha!) interlude to consider a few outlier seasons.  In Part Two we saw this graph:

Again, X is number of teams, Y is standard deviation of winning percentage
 There are, as you can see, a handful of data points here that are particularly far from the trendline, in both directions.  From the ultra-competitive mid-1950s to the mess that was the last two seasons of the ABA, I've picked the eight most notable points on either side of the trendline to discuss in this post.  I'll be grouping these by era.

The Early Days - 1950s

The 50s were an interesting time for the NBA.  The league was small, there was no three point line, and the shot clock didn't come to the league until the 1954-55 season.  Moreover, the league - and the country - had not quite worked out a number of racial issues, and so the league was dominated by white players.  What's more, because basketball was still new, there were not throngs of kids who grew up playing the game (football and especially baseball were the sports of the time in America), meaning it was harder to find talented athletes.

The 1952-53 season was one of the "least competitive" - at least by standard deviation of winning percentage - in the history of the NBA.  With a SD of .198, the league was both top and bottom heavy.  Interestingly, no team won more than 70% of their games, but the SD is so high because only one team won between 40 and 60% (the Fort Wayne Pistons).  At the high end, the New York Knicks went 47-23, the Syracuse Nationals went 47-24, the Boston Celtics went 46-25, the Minneapolis Lakers (which makes much more sense than the Los Angeles Lakers) went 48-22, and the Rochester Royals went 44-26.  On the other hand, the Baltimore Bullets and Philadelphia Warriors went 16-54 and 12-57 respectively.

Despite the disparity in winning percentages in 1952, point differentials were much smaller.  The league's high-scorers from Rochester averaged 86.3, while the Indianapolis Olympians averaged 74.6 points per game.  Both are astoundingly low by modern standards, but more remarkable is the gap - or lack thereof - between the two.  Consider the 2009-2010 NBA, in which the Phoenix Suns averaged 110.2 ppg, while the New Jersey Nets put up only 92.4.  As for point differential, the Milwaukee Hawks went 27-44, but were outscored, on average, only 77.4 to 75.9.  One gets the sense that they fell behind, and then had the clock milked against them.

The shot clock changed everything in the NBA, and it's no accident that three of the most competitive season in NBA history were 1954-55, 1955-56, and 1956-57.  Those were the first three seasons of the shot clock, and the clumping of W-L records alone shows that teams were really struggling to understand how to play in a transforming league.  SDs for those years were .087, .061, and .054(!).

That the league was changing was obvious.  In 1954-55 the Boston Celtics averaged over 100 points per game (on both defense and offense), both NBA firsts.  No team stood out in those three seasons, however, despite the complete transformation of the game.  In part this was due to a small league, in part a lack of standout talent, in part a shorter season, and in part, of course, a drastic rule change.  It took until the 1957-58 season for some team to start to pull away from the pack, some team to start to "get it" in this new era of no-running-out-the-clock basketball.  That team?  The Boston Celtics.  Not surprisingly, they had been the best offensive team in the league before the shot clock, and they continued to be for years afterwards.  What catapulted them to dominance, however, was defense.  They kept scoring, and added Bill Russell as a rookie in 1956-57.  Even that year they won their first NBA championship and finished 44-28, but the league as a whole was still bunched together (the Western Division featured no less than three 34-38 teams).

As Russell's career took off, the competitiveness of the league disappeared.  The Celtics became such a dominant force that by the 1959-60 season, the pendulum had swung so far in the other direction that a 59-16 Celtics team, combined with a 19-56 Cincinnati squad, led to a SD of .188, the second highest in league history.  Of course, 1959 is one of those seasons where we have to wonder about what "competitiveness" really means.  The uneven talent distribution did mean that Cincinnati, Minneapolis, New York, and Detroit got hammered far more often than not, but the other four teams were all about .600.  Boston's 59-16 was countered by Philadelphia's 49-26, and Syracuse's 45-30.  The St. Louis Hawks finished 46-29 in the West, as well.

By now scoring was way, way up as well.  The Celtics, in 1959, averaged 124.5 on offense, and 116.2 on defense.  Average!  No wonder Wilt Chamberlain was able to put up 37.6 a game, or Bob Cousy averaged almost 10 assists.  Anyway, this is one of those seasons, ironically, where "everyone had a good team" in the opinion of Bill Simmons, and he has a point.  Everyone may not have had a good team, but four teams definitely did.  Consider the following key players (win shares in parentheses):

Boston -  Bill Russell (13.8), Bob Cousy (7.9), Bill Sharman (7.8), Tom Heinsohn (7.7)
Philadelphia - Wilt Chamberlain (17.0), Tom Gola (9.9), Paul Arizin (9.2)
Syracuse - Dolph Shayes (9.5), George Yardley (9.0), Larry Costello (8.0)
St. Louis - Cliff Hagan (11.8), Bob Pettit (11.5), Clyde Lovellette (9.0)

Who cares if Minneapolis's Elgin Baylor (11.5, with no support from anyone else on the team) was the only other really good player in the league, the best teams were all loaded with talent.  Which, of course, only begs the question: which is better, a league with 4 great and 4 terrible teams, or a league with 8 teams who all have a change to beat each other?

The End of the ABA, and the Merger - 1970s

The early and mid 1970s were an interesting time for the NBA.  The ABA - which included three-pointers and lots more slam dunks - emerged as a competitor, but also hemorrhaged money and was responsible for seasons that were, by any measure, extremely uncompetitive.  It was, in short, more of a show league, but it put pressure on the NBA just the same.  1972 was a pivotal year, in particular, because the NBA switched TV partners, moving to CBS from ABC after the season.  It was also, by standard deviation, the single least competitive season in NBA history.

On the plus side, a 68-14 Boston Celtics team romped through its division, while 60-22 Milwaukee and 60-22 Los Angeles led the way in the Western Conference.  The New York Knicks ended up upsetting Boston in the Conference Finals and went on to crush Los Angeles in five games in the Finals.  But, like 1959, 1972 was marked by a handful of great teams and a handful of truly awful ones.

On the minus side was Philadelphia, finishing an unheard of 9-73, a full 59 games behind Boston.  Buffalo - from Philadelphia and Boston's division, went 21-61, meaning the Atlantic had two of the best and two of the worst teams in the NBA.  Portland, cellar-dwellers in the West, was also an uninspiring 21-61, and their division mates, the Sonics, went 26-56.  In all, it was a year of extremes, in a league waiting for a merger (which would bring, among others, Julius Erving to the NBA), with its TV deal caught in limbo, and yet with all of the attention that a third New York - Los Angeles finals brought to the league.  Throw in Wilt Chamberlain and Kareem Abdul-Jabaar, and the NBA was making inroads in mainstream America.

Meanwhile, the ABA was falling apart.  The 1974-75 and 1975-76 seasons were, well, horrible.  With the merger looming, and teams facing bankruptcy, competitive balance suffered.  Posting SDs of .191 and .190, the last two seasons of the ABA were more or less a joke.  Already down to 10 teams in 1974, after a 27-57 season Memphis closed up shop, and San Diego and Utah both played fewer than 20 games in 1975, finishing 3-8 and 4-12 respectively before calling it quits.  The final season of the ABA, then, featured only 7 teams including a Virginia squad that finished 15-69 for the second year in a row.  Ultimately, in a top heavy league, it made sense for the New York Nets, the Denver Nuggets, the San Antonio Spurs, and the Indiana Pacers to make the jump to the NBA as the ABA finally closed its doors.

The 1975-76 NBA season had been, in stark contrast to the ABA, highly competitive (SD of .105).  At 54-28, Boston led the Eastern Conference, while a 59-23 Golden State team led the West.  Other than those two teams, however, no one finished higher than .600, while only one team - 24-58 Chicago - finished below .300.  Parity was the word, and so the addition of four good teams from the ABA, along with a redistribution of talent from the folding ABA teams, meant that 1976-77 would be one of the most competitive ever in the NBA.

With a SD of .098, 1976-77 is the most competitive season since the merger.  It's no accident that it happened the first year after the merger, for reasons discussed above.  For the second straight season, only one NBA team finished below .300, the New York Nets.  Meanwhile, the other ABA transfers did better, with San Antonio and Denver posting solid above .500 seasons (Denver, in fact, won their division), and Indiana finishing 36-46.  No one really stood out in 1976, however, with the 53-29 Lakers the class of the league.

This was no diluted league, however, and while the lack of great teams might frustrate some, there was no shortage of great players.  Take a look at some of the leader boards to see what I mean:

Points:
1) Pete Maravich - New Orleans
2) Kareem Abdul-Jabaar - Los Angeles
3) David Thompson - Denver
4) Billy Knight - Indiana
5) Elvin Hayes - Washington

Rebounds:
1) Kareem - Los Angeles
2) Moses Malone - Houston
3) Artis Gillmore - Chicago
4) Elvin Hayes - Washington
5) Bill Walton - Portland

Win Shares:
1) Kareem - Los Angeles
2) Gilmore - Chicago
3) Hayes - Washington
4) Dr. J - Philadelphia
5) Bobby Jones - Denver

Not even mentioned in those statistical categories are first team All-NBA-er Paul Westphal and 2nd teamers George Gervin, Geroge McGinnis, and Jo Jo White.  Rookie of the Year Adrian Dantley, All-Stars Dan Issel, Bob Lanier, Rick Barry, Dave Cowens, John Havlicek, Bob McAdoo, Rudy Tomjanovich, and Earl Monroe are also worth mentioning.  In short, it was a banner year, talent-wise, for the NBA.  It just so happened that few of those players were teammates (Issel, Bobby Jones, and David Thompson with Denver were possibly the best trio in the league, but not good enough to get past Bill Walton's eventual champion Portland).

The NBA continued to see parity in 1977-78 (SD of .111) and 1978-79 (SD of .103), but eventually we settled into a happy medium as Larry Bird and Moses Malone came into their own in the early 80s, followed, of course, by MJ.

Modern Era - 1990s and Beyond 

There's not as much to say about the NBA since the merger.  As the three-point line became and accepted part of the game, and as free agency settled in and the draft became what it is today, changes to the league structure have become much smaller.  As you can see in the graph, there's a much narrower range of variability from one season to another in the modern NBA, and that's probably a better measure of competitiveness than anything else.  The modern NBA has struck a balance, for the most part, between too few and too many teams being in contention each season, with room for the occasional extreme.

A couple noteworthy seasons include 1983-84 and 2006-07 (with SDs of .115 and .132, respectively), which were both on the "competitive" end of our SD spectrum.  1983 was Bird's first MVP season, featuring a - guess who? - Los Angeles vs. Boston finals.  While no team really pulled away, both Boston and Los Angeles were excellent, and the SD is so low mainly because no team was truly awful.  The 27-55 Chicago Bulls, of course, were one of the league's worst, and bad enough to land none other than Michael Jordan in the draft the next season (at #3 overall).*

* This was a crazy, crazy draft.  Check out some of the picks, here:
1) Hakeem Olajuwon - Houston
2) Greg Oden - Portland.  Oops, I meant Sam Bowie.  It's just, they're exactly the same player.  And, just like with Kevin Durant, the next guy was maybe a little better.
3) Michael Jordan - Chicago
4) Sam Perkins - Dallas 
5) Charles Barkley - Philadelphia 
7) Alvin Robertson - San Antonio
9) Otis Thorpe - Kansas City 
11) Kevin Willis - Atlanta
16) John Stockton - Utah

2006-07, meanwhile, was a parity hodge-podge, both for teams and for players.  Dirk Nowitzki won the MVP because, hey, why not?  And then his Dallas team proceeded to get dismantled by the eight-seed Golden State Warriors in the first round.  So that was maybe a bad choice.  Meanwhile the Spurs and boring Tim Duncan coasted through the regular season (finishing 58-24, which is pretty good for coasting) only to absolutely dominate the playoffs, beating Denver 4-1, Phoenix 4-2, Utah 4-1, and sweeping LeBron's Cavaliers in the finals.  Much as I hate the Spurs, Tim Duncan is, you know, really really good, and was at his best (or close) in 2006-07.

Our last two notable seasons are 1996-97 and 1997-98, responsible for two of the higher SDs in league history at .191 and .189.  These were the last two seasons of the Jordan Bulls, whose dominance in the East was matched by - in the regular season anyway - the Stockton-Malone-Ostertag (joke) Jazz in the West.  The NBA was kind of on cruise control in the late 90s.  Parity was at an all-time low - at least since the merger - but no one seemed to mind that Utah, Miami, Chicago, Seattle, Los Angeles, and Houston were winning at or above 60 games a season while the rest of the league was mediocre or terrible.  The NBA brass had to be happy, at least, because New York was at least decent, for most of the decade, meaning the huge media markets of Chicago, LA, and New York were drawing viewership, while Utah was a nice wrinkle and a good foil to the Bulls.  That is, they were good enough to win a game or two, but not good enough to really challenge for the title as long as Jordan had at least one leg.

Conclusion

Diving into individual seasons and eras tells more about the methodology I've been using than the results.  The fact is, competitiveness is subjective, and while some people will prefer parity, others prefer leagues like those of the late 90s, when parity is non-existent because a small handful of teams dominate every year.  The fact is, either type of league can be successful.  In Europe, the Premiership and other soccer leagues are routinely extremely top-heavy.  When was the last time someone not named Chelsea, Manchester United, or Arsenal won the EPL?  Answer: Blackburn Rovers in 1994-95 (during Alan Shearer's prime)*.  Yeah.  And yet, people keep watching even though the same three teams are at the top every year.

* And no, I won't apologize to other Americans for knowing who Alan Shearer is.

I would argue, though, that the EPL is supremely competitive for exactly that reason: there are a small set of teams that must get a result basically every match.  Then there are a lot of other teams that are also well-matched, and while they're not fighting for the league title, they are trying to avoid relegation, or climb high enough to qualify for the Euro Cup, if not the European Champions League.  The same is more or less true in the NBA.  While, realistically, it's hard to win a seven game series against superior opposition, there's still plenty of incentive - in revenue, principally - to make the playoffs, and for better teams there's the incentive of home-court advantage that drives competition throughout the regular season.

If anything, the biggest flaw in the NBA's competitiveness in the modern era - the real cause of the larger SDs we see now - is not dilution, league size, free agency, or the salary cap.  I'm sure those things contribute, but I think the biggest culprit is incentive: there's undeniably incentive for good teams to win games, but there's also incentive for bad teams to lose games.  Because the NBA draft lottery is designed to give inferior teams a better chance at higher picks, tanking is all-too common, and tanking has as much effect on standard deviations of winning percentage as title chasing does.  If only the NBA had a relegation system!  But that's a post for another time.

Tuesday, February 8, 2011

NBA League Size and Competitiveness, Part Two: The Data

"Errors using inadequate data are much less than those using no data at all." - Charles Babbage

That is the spirit with which this post will proceed.  The question we're trying to get at is whether or not a larger league in the NBA (or, really, in any sport) leads to a more "diluted" product.  I've reinterpreted this question to be the following: is the league more or less competitive when there are more teams?  As discussed in Part One, it's not easy to really tell where the overall talent level of a league is because all the statistics players compile - all the games they play, the championships team win, and so on - are contextual.  The best player from the 1950s was still the best player from the 1950s, and looks great in retrospect, even if he wouldn't even make an NBA roster today.

Therefore, I transformed "diluted" into "competitive," because it seems to me that the one stands for the other.  That is, when we think the league is diluted, what we're really saying is that the league is not competitive, that there are too many players who are not good enough to hang with the few good ones, and that the few good ones are causing a small set of teams to dominate.  In a non-diluted league - in a competitive league - "everyone has a good team," to use Bill Simmons's language from my last post.  The result, no one - or few people - have a team that just trounces everyone else.

So today we're going to make a first pass at the data.  That first pass?  Looking, simply, at the standard deviation of winning percentage for each year in NBA (and ABA) history.  When was the league the most competitive (smallest standard deviation), when was it the least competitive (largest standard deviation), and is there any discernible trend as the league expands (does a larger league tend to be more or less competitive)?  Without further ado, here's a graph of our results.  X-axis is number of teams, Y-axis is standard deviation of winning percentage.

X is number of teams, Y is standard deviation of winning percentage

As you can see, there is a slight upwards trend here, but it's pretty small.  The R value (R-squard is on the graph) is about 0.17, which is not really significant unless you're doing social sciences research.  Nevertheless, a big reason why our correlation is so small is how spread out the standard deviations of winning percentage were when there were only eight to ten teams in the league.  As you can see, the left-most data points are much, much more spread out than the rightmost, with values ranging from barely over 0.05 all the way to almost 0.20.  What does this mean?  It means that, in the league's most competitive season a mere 5% separated average teams from good teams, and almost everyone was within 10%.  Put in sports-fan friendly terms, the best winning percentage in the league was about .600, while the worst was about .400.  That's baseball territory.  On the other hand, in the least competitive season, the separation led to a best team with an .800 winning percentage and a worst team with a .200 winning percentage.  Now, that's not precise (we'll get into the exact numbers shortly), but that's roughly what standard deviation tells you.

So, back when there were only eight professional teams, there was a huge variety from year to year.  Of course, with only eight teams we expect more variety in standard deviation, because, hey, fewer data points means each data point has more influence.  Thus, if one team wins 90% of their games one season in an eight team league, that value is going to skew the overall standard deviation much more than if one team wins 90% of their games in a 30 team league.  And, indeed, the 95-96 Chicago Bulls (who went 72-10) did not make 95-96 anywhere close to one of the least competitive seasons in NBA history.  Had that happened in 1960, the story would have been different.  For example, one of the "least competitive" - by standard deviation of winning percentage - seasons in NBA history was 1952, when a 12-57 Philadelphia team joined a 16-54 Baltimore team to drag the whole league down.

If we take only the seasons since the merger - that is, only seasons in which the league has more than 20 teams, we get the following instead:


Now we have an R value of .46, which is getting much closer to significant.  Indeed, while random variation obviously plays a huge roll - as do hard-to-quantify things like player skill and pre-NBA training, as well as injury management and so on - there seems to be little doubt that larger leagues are at least somewhat less competitive, according to standard deviation of winning percentage.  Consider that there have been only five seasons in which the SD of winning percentage was under 0.15 since the league went to 25 teams in 1988, whereas there were over 20 such seasons in the 40 years before then.

What this really shows, though, is not competitiveness or dilution, but talent distribution.  That is, regardless of the level of talent in the league at any given time, the more spread out that talent is, the lower the standard deviation of winning percentage will be.  The more concentrated, conversely, the higher the standard deviation of winning percentage will be.  Whether this is a measure of dilution is up for debate.  Also, while in a smaller league a larger standard deviation here might mean less competitiveness (one or maybe two dominant teams), in a 30 team league it might be exactly what we want (5 or 6 really good teams).

For example, this season the Miami Heat have Dwayne Wade, LeBron James, and Chris Bosh.  The Los Angeles Lakers have Lamar Odom, Kobe Bryant, and Pau Gasol.  The Boston Celtics have Kevin Garnett, Paul Pierce, Ray Allen, and Rajon Rondo.  Now, any of those ten players would be the best or second best player on most other teams.  Whether because of finances, smarts, collusion, or some combination of factors, we're currently watching a league where talent has conglomerated onto a small set of teams that routinely beat up on inferior opposition.  While I didn't run the (still changing) numbers from this season's NBA, so far there's a team with an .840 winning percentage (San Antonio), three teams above .700 (Dallas, Miami, Boston), and six more teams above .600 (Chicago, Atlanta, Orlando, Oklahoma City, Los Angeles, New Orleans).  On the other side of the coin, there's a .154 Cleveland team, plus five other teams below .300 (Sacramento, Minnesota, Washington, Toronto, and New Jersey).

Is this season's NBA competitive or not?  There are, at this point, ten legitimately good teams, any of whom - given the right breaks - could win the NBA Finals.  That sounds extremely competitive to me.  On the other hand, there are also at least six teams that are flat out awful, meaning that a large portion of each day's games are over before they start.  Is Cleveland really going to beat Miami?  Does Minnesota stand a chance against Oklahoma City?  Even though upsets happen, I doubt any circumstance would arise where a fan would feel like one of those inferior teams really deserved to beat one of the top ones.  They would need lots of lucky breaks.  And that, I think, is a mark of an uncompetitive league, when the bottom third of the league stands little to no chance against the top third.

But wait.  Does that mean the league is uncompetitive, or does it just mean that talent is distributed unevenly?  The latter is certainly true.  The former is more a question of taste and perspective.  We could say the same about the league being diluted.  When it comes down to it, if you make the league smaller, players who seemed great in a 30 team league will look good, and players who seemed good will look average.  Does that mean the league is better or worse?  Or does it mean that our perspective changes?

Consider a historical example.  When the ABA and the NBA merged for the 1976-1977 season, the NBA had one of its least competitive seasons ever, with a standard deviation of winning percentage under .100.  Bill Simmons says, in The Book of Basketball, that this one time the league actually under-expanded, as the 18 teams from the NBA and the 8 remaining viable teams from the ABA became 22 instead of 26.  That meant that a lot of ABA talent got redistributed to (mostly) bad NBA teams, meaning that talent was distributed about as evenly as ever in the history of the league.  The 53-29 Lakers lead the NBA that season, with a 50-32 Denver and 50-32 Philadelphia on their heels.  Those were the only three teams that won 60% of their games or more.

So where does 1976-77 sit in terms of dilution, competitiveness, and distribution of talent?  Really, we can only answer the ladder.  Talent was widely distributed.  Was the league diluted?  Was it competitive?  That's a matter of opinion.  Because talent was widely distributed, it was certainly competitive in a broad sense, but many fans would rather see great teams (and, by extension, terrible teams) than good ones (and merely bad ones).  As far as dilution, that's a more complicated question still.

See, when the league goes from 8 teams to, say, 12 teams, we'll tend to think of it as diluted because players who weren't previously good enough now are.  Similarly, contraction seems to eliminate dilution, because suddenly all of those marginal players are gone.  But give it ten years after expansion and contraction, and we no longer feel that way, because it's a matter of perspective and perception.  If the NBA cut 10 teams this offseason, the result would definitely be a short-term feeling of "raising the level" of the league, and probably increased competitiveness in the sense of a smaller standard deviation of winning percentages.  However, after 10 seasons of the new, 20 team NBA, we'd get used to seeing guys who had previously been their team's #1 as role players on the "deeper" teams in the smaller league.  New draft picks who once would have been franchise guys for bad teams would suddenly never be at the top of the league.  These new would-be stars, however, would never be thought of as franchise players who turned into role players.  We'd just consider them role players.  Suddenly, over time, the league would start to look a lot like it does now, only with fewer teams.

The same goes in the other direction.  A more diluted league is all well and good to talk about, but no one talks about how diluted NCAA Division I college basketball is, despite the fact that it adds new teams almost every year.  Sure, talent distribution is pretty extreme in the NCAA, but even so there are usually a good 20 or so teams that have a legitimate chance to win the Tourney every season (if you think this is an exaggeration, consider Butler), given the right breaks, and a favorable series of match-ups in March Madness.  Is the NCAA diluted?  Maybe, in some sense, but in another sense it's almost a crazy question to ask.

I would argue the same is true in the NBA.  Is the NBA diluted or not?  That's not really a good question, because being diluted is relative, and the stats players compile are relative, and even wins and losses are relative.  Bill Simmons believes the NBA is diluted because he grew up watching a league with a dozen teams in it.  I don't, because I grew up watching a league with 27-30 teams in it.  NCAA fans are used to the 400-something Division I teams, so there's never really a discussion.

What we do have, however, are some interesting measures of talent distribution and, in some sense, competitiveness.  We've already seen the trend - as the league expands, there is both a tightening up of standard deviations of winning percentage (less variety from season to season), and a slight (very slight) upwards trend.  Next time, we'll dive a little deeper into that data and look at some of the outlier seasons.  Moreover, we haven't given up on the competitiveness question - when has the league been most competitive? With more teams or with fewer? - so we're going to tease out only the good teams and run the same analysis here on them (that is, how many good teams are there in a season, and how good are they).  Stay tuned.

Friday, February 4, 2011

NBA League Size and Competitiveness, Part One: Introduction

Thanks to a generous friend, I've recently begun reading Bill Simmons's colossal The Book of Basketball.  I say colossal because, as you may not be aware, the book is about as long as Anna Karenina.  It's long, it's big, and so far, anyway, it's extremely entertaining.  A significant portion of the book seems to be Simmons - perhaps better known simply as "The Sports Guy" - taking digs at Vince Carter, Kareem Abdul-Jabar, and Wilt Chamberlain, whilst trumpeting (who else, as a kid who grew up in Boston?) anyone who played for the Celtics, and especially Bill Russel.  Which is all very fun.

Anyway, in the first few chapters, Simmons has already made a point - well, he's made many points - with which I disagree.  He is a firm believer, it seems, that expansion has diluted the NBA, and that the league's competitiveness was much higher when he was a kid.  "Back in my day," he never says, but might as well, "Basketball players had to try harder, because every night they played against teams filled with All-Stars."  Now, "my day," in this case, refers to the Russel era of the late 50s and early 60s, when the Boston Celtics - despite the huge competitive balance of the NBA at the time - won eight championships in a row (and nine out of ten, and ten out of thirteen).  If only we could have that again!

In all seriousness, though, Simmons does an excellent job describing what makes for success in the NBA, and, frankly, he knows way more about it than I do.  He points out - and rightly so - that basketball statistics are deeply flawed, because they don't capture the magical things that allow teams to win games.  I would say that Simmons is right: points scored, assists, rebounds, blocks, and steals do not accurately measure a player's contribution to his team.  Not even close.  That doesn't mean statistics have no place in basketball, it just means that basketball statistics have to get better and, what's more, that might be impossible to do because unlike in baseball, a team's success in basketball has more to do with how teammates work together than with how individuals perform.  (Inhales).  The linchpin of Simmons argument, here, is that Wilt Chamberlain - for all his statistical dominance - was a terrible teammate who's teams rarely won championships, while Bill Russel was actually a better and more valuable player, as evidenced by his bevy of MVP awards and Championship rings.  And you know what, I buy it.

I still don't buy, however, that the modern NBA is somehow watered down compared to the NBA of the 60s, and I do think that statistics can demonstrate why.  A while back I explored how NBA rosters are constructed, using Win Shares, and discovered that teams, as a whole, follow a highly predictable model.  That model, to rehash, is that the average "best player" on a team accumulates 9.3 Win Shares in a season, and each subsequent player accumulates less in a logarithmic way.  The stunning result was, at least for the season I looked at, a correlation coefficient of exactly one.  League wide, there's a very strong trend towards a regular distribution of success on the court.

What does this have to do with competitive balance?  Not much, but I want to point out that it jives well with the qualitative description of successful NBA teams that Simmons gives in his book.  He argues that teams need a great player, followed by a couple all-stars, followed by some key role players.  If you look at my old post and the graph with the Lakers and Celtics, it's easy to see that their model fits well with that description.

Now, I bring this up because Simmons points out how many All-Stars were on the Celtics and their rival Lakers and Warriors back in the 60s.  The teams were stacked, he tells you, replete with great talent.  Not like todays teams, where many teams are lucky to have even one All-Star.

Of course, the easiest hole to poke in this argument - that teams had more All-Stars back in the 60s - is a direct result of a smaller league.  Not because the talent level was necessarily higher, but because there were fewer players from which to draw an All-Star team.  Of course the Celtics had a bunch of All-Stars in the 60s, because the league only had eight teams.  That means that, even if every team was equal, an all-star roster of 12 would mean taking three players from each team in both of the four-team divisions!  Since the Celtics were also the best team in the league, it's only reasonable that they would have four or five All-Stars in any given season.

Compare that to today's league.  With 30 teams in the league, it's hard to have even two All-Stars from the same team, because an individual player has to out-shine so many others.  What's more, a second (or third) best player on a given team is going to have an even harder time, because he has to look better than the best player on many other teams, not easy to do given the limited and flawed statistics available in the modern NBA.  I realize that may be a bit opaque, so let me clarify using the simplest example: points.

Consider two teams that score 100 points per game.  On Team A, King Star scores 25 a game, whilst his brother Duke Star scores 20 a game.  Thing is, Duke takes way fewer shots, because he's more accurate, and is generally just a more efficient player than King Star, despite King's gaudy numbers.  Now, King Star is a perennial All-Star and fan favorite, and he's still plenty good, so he's going to the All-Star Game no matter what.  Duke is on the cusp, especially because Team B - which also scores 100 a game - features Selfish McGee (also known as Allen Iverson), a player who plays the same position as Duke, but scores 30 points a game in twice as many shots, thanks to a higher-paced offense and a team that has no other reliable scorers.  So Duke Star, in order to make it to the All-Star game, has to outplay either King or Selfish in the eyes of the people who make these decisions.

Of course, that's no different now than it was 40 years ago.  What's different now, instead, is that Duke is up against his equivalent on 14 other teams (whether they be like Duke, like Selfish, or like the heretofore unmentioned Crappy Sullivan), instead of 3.  Suddenly Duke, who's just as good - maybe even better - than the number two guy the Celtics had back in 1962, doesn't even make the All-Star game, while he would have been a shoe-in, at least as a backup, back in the 60s.

Phew.  OK, all that out of the way, let's actually get to the point of the post, which is how to actually assess whether the NBA is more or less competitive now.  How do we do this?  Is it best to look at players or teams?  What statistics should we use, in order to compare against eras?  In fact, there are many ways we could study the question, but the easiest and most intuitive, to me anyway, is simply to look at wins and losses.  I'm struck by a sentence in The Book of Basketball, which goes something like this: "I'm telling you, everyone had a good team back then."  Now I know what Simmons means is "All the good teams had good teams back then," because he knows that, even then, there were cellar-dwellers.  The reality is, every game played has a winner and a loser, and one of the constants in all sports is that, league wide, the average winning percentage is always exactly .500.  It goes without saying that, in order to get to .500, there will always be some teams that are much better and some teams that are much worse, and some teams that are right about in the middle.

What do we make of the claim, then, that everyone had a good team?  Well, what I think Simmons means is, there were more - a larger group of teams - well above average, and fewer really bad ones (and, likely, fewer really great ones).  He might phrase that as "more great teams, fewer average ones," but that's just a perceptual thing.  We live in a time where "average," in sports, has come to mean "bad," and "mediocre" has come to mean "absolutely terrible."  Ironically, "terrible" is something we don't actually dislike: the Timberwolves are terrible, but in a lovable kind of way.  It's mediocre teams we can't stand.

Anyway, how do we test whether or not everyone had a good team, given that we think it means that there was better competitive balance, that fewer teams were terrible, and fewer were so good that the games weren't even worth playing?  Well, there's a pretty easy - if tedious - way, that does not require digging into the deeply flawed player statistics of the 60s.  We can, in fact, compare across eras and leagues easily - as Simmons does when he says that the modern NBA is watered down compared to the old NBA - using wins and losses.  It's simple, really.  We just need to look at standard deviations of Win-Loss records throughout NBA history, and we'll see when the NBA has been at its most competitive.  In short, smaller standard deviations means the league is more competitive, while larger ones mean that the league is less competitive (more top and/or bottom heavy).

Now, there are some concerns here.  First off, those old leagues were so small that our sample size is going to be tiny.  Standard Deviations don't mean a lot when you're talking about 8 data points.  That is, they don't mean a lot if you're trying to be predictive based on only 8 data points.  But, in this case, I think we'll be fine, because we're just trying to deduce how "spread out" the quality of teams has been throughout NBA history.  Standard Deviation is exactly the statistic we want to use.  Since we'll be able to get a broad view of competitiveness, we'll be able to take the first steps towards assessing the competitiveness or watered-down-ness of the NBA across eras, irregardless of silly things like small league sizes making it easier to win championships (because, hey, fewer opponents) or make it to the All-Star game.

I honestly don't know what I'll find in doing this, even though my hypothesis is that the modern NBA is, if anything, more competitive than the NBA of the 60s.  I might be wrong.

As an extra outlet (for both me and Simmons), I'll also calculate the mean and standard deviation of the smaller set of "good" teams in the league.  I haven't yet decided how to draw this line, but I'm initially thinking that anyone above .500 makes the cut.  Basically, if we find a relatively constant standard deviation across time, we'll still want to test is maybe, in certain eras, the "good teams" are more evenly balanced with each other.  Now, this will be built into our bigger SD calculation, but we'll be cutting out noise like a team or two that finishes with a winning percentage of .130, and thereby makes the whole league's SD look way bigger than it is.  Indeed, I think Simmons would agree that the occasional really really bad team shouldn't count against any assessment of the competitiveness of the league as a whole, and so we'll do a parallel calculation that cuts out those really bad teams.

So to recap, here's the method: I'll be going through every season of professional basketball on basketball-reference.com (oh the wonders of being unemployed), and putting every team's W-L record into a spreadsheet.  From there, it's easy to calculate mean (which will always be half the games in the season) and standard deviation of wins per league per year.  The lower that SD, the more competitive the league.  I'll also pull out just the above .500 teams, and run the same calculations, to see if maybe there was more competitiveness amongst the good teams than in the league as a whole.  Finally, I'll do a smaller cut of outliers, removing just the really really bad teams (teams more than 2 SDs from the mean), and recalculate the league without their nefarious influence.

What will I find?  You'll have to come back to my next epically long blog post to find out, because I don't know yet.

Saturday, December 25, 2010

LeBron James Misunderstands Contraction

In honor of the NBA's insane practice of playing fifty games on Christmas Day, here's a basketball post for all of you yule-tide guards, forwards, and centers who are skipping out on dinner to watch hoops and surf the web.

I read an interesting piece last night on ESPN.com in which LeBron James says the following: "Hopefully the league can figure out one way where it can go back to the '80s where you had three or four All-Stars, three or four superstars, three or four Hall of Famers on the same team," James said. "The league was great. It wasn't as watered down as it is [now]."  LeBron goes on to propose that taking good players off of bad teams and then putting them onto better teams would improve the NBA, principally because - we can assume - such a move would increase the overall talent pool of the league, and thereby improve competitiveness.

Leaving aside the many financial reasons why the NBA would never even begin to consider contracting franchises, let's take a quick look at the classic argument that fewer teams equals more competition.  The claim that any league will become "watered down" with the addition of new teams is a surprising one to me, given the large number of contrary examples available across sports.  Without even beginning to do quantitative or statistical analysis, you can point to college football, college basketball, or European soccer as prime examples of sports where more teams hardly thwarts competitive balance.  While certainly the overall talent level of the NFL is higher than the NCAA, no one (that I've heard, anyway) really argues that the NCAA should trim the FCS down to its 30 best programs, in large part because you'd have a lot of argument about which programs belonged and which didn't.

Similarly, while European soccer leagues are notoriously top-heavy (Rangers and Celtic win the Scottish league every year; Roma, Inter, and AC Milan duke it out in Italy; Barcelona and Real Madrid are the only two Spanish contenders; Manchester, Arsenal, and Chelsea usually top the EPL; and so on), there's not nearly so much to separate those league champions and runners-up from each other, as the Champions League routinely demonstrates.  And, what's more, the there are enough fans to support not only Premiership (or equivalent) teams around Europe, there are enough to support second, third, fourth, and sometimes even fifth tier professional clubs.  Far from decreasing the competitiveness or intrigue of the leagues, that there are well over 100 progessional soccer clubs in England alone seems to improve the overall quality of the game, because more people (hence, more talent) actually have an opportunity to make a living playing the sport.

This last point is especially important when you consider player development.  European soccer stars often hone their skills at a young age by playing for lower-division teams on loan.  Instead of riding the bench and just participating in practice, a future phenom has a chance to participate in real competitive matches, real promotion races, and real cup games with other players who - while not of Premiership quality - are also professionals, are experienced and knowledgeable, and are trying really hard because, hey, it's their job and their passion.  Compare that to American sports, where development almost always takes place in leagues dominated by - indeed, almost exclusively composed of - younger players.  Perhaps it's a minor difference, but the point is that more professional teams seems not only to help soccer in Europe financially, but competitively as well.

Returning to the NBA, LeBron is right in assuming that the overall level of talent in the NBA would increase with contraction.  It stands to reason that, if you take the current 150 starters (from 30 teams) and trim that down to, say, 100 starters (for 20 teams), you're generally going to improve the average overall quality of starters league-wide.  Likewise down the roster, where the roughly 360 NBA players (assuming a roster of 12) would be cut substantially to 240.  Since the NBA is the premier basketball league in the world, it's safe to assume that you'd be going from very close to the best 360 basketball players in the world to the best 240.  Of course there's some fudge-room at the edges, where evaluating talent and meeting team needs might mean that the 250th best player makes it onto a roster before the 230th best does, but roughly you're going to be in that range.

But does decreasing the number of players, and therefore increasing the overall average quality of players league-wide, really improve the league?  That really depends upon what you want to see, as a fan, and what the league is trying to accomplish from a competition standpoint.  Certainly you're likely to see a higher quality of basketball in a smaller league, but not by very much.  Given that basketball talent, like talent in most areas, is almost certianly normally distributed, some rough math (assuming about 10,000 professional / aspiring professional basketball players in the world; a very rough guess) tells us this: in real terms, the difference between the 240th best player and the 360th best player in the world is about the same as the difference between the best player in basketball and the second best.  Those 120 players in-between, in other words, are pretty close to each other in skill.

What LeBron James misunderstands, then, is two-fold.  1) Success in the NBA - like in any sport at the highest level - is only partially the result of talent. 2) Improving overall average talent does not necessarily improve competitiveness.

The second of these misunderstandings first.  When you cut a league's size, even by a substantial number like 10 NBA teams, you don't change the talent pool all that much, as we see above.  But even more importantly, in changing the average level of talent, you do little to change the distribution of that talent.  Sure, the overall quality of play league-wide might improve by some small measure, but you're still liable to have a small set of dominant teams, a bigger set of middling teams, and another small set of poor teams.  While sometimes the league will skew one direction or another, it would be foolish to forget that team quality, just like player talent, tends to be normally distributed.  Lowering the league size will make that distribution less obviously normal, but it doesn't change that some teams will still be stacked while others are terrible.

Now, it is still right to say that the worst team in the league will be better in a 20 team league than in a 30 team leauge.  However, the point is that the best team will also be better, and probably by roughly the same amount.  Which means that LeBron's Heat would not be playing a star-studded opponent every night.  Far from it; they'd be heavy favorites just as often, if not more often, in a smaller league.

Consider, as an example to bring the point home, your fantasy league.  Instead of distributing the NBA's 360 players amongst 30 teams, you've split them up between 10 or 12 (or something).  Now, in pure talent terms, the team you put together featuring the best players from 5 different NBA teams is pretty awesome, but the other guys in your league have done the same thing.  Suddenly, in that setup, everyone has great players on their team, and winning becomes a matter of getting the best of the best.  Odds are someone (or sometwo) dominates your fantasy league, a bunch of other teams are middle-of-the-pack, and a couple teams - probably abandoned - really suck.  Similarly, in a smaller NBA some of those stars on bad teams might be united with stars from good teams to form all-star rosters, but stardom is relative, and that guy who looks like a stud in a league of 30 might suddenly be average in a smaller league.  Just like in fantasy sports, smaller leagues tend to make what used to be great players look good, and good players look average.  That's what happens when you shift the average talent level upwards.

If anything, LeBron should advocate for expansion if he wants to see teams with more great players.  It's a lot easier to be three standard deviations (or more) above average when the average is lower, after all.

As for the other misunderstanding I mentioned above, even in a league with a smaller distribution of talent, there's not likely to be a substantive change in competitive balance because talent is only a part of what makes a basketball team successful.  Is it a big part?  Of course.  But - and especially as you decrease talent disparity - things like how well players do their jobs, how well the coaches game-plan, and, of course, luck play huge roles in determining outcomes.

Consider college basketball.  In the NCAA, it is enough for Duke to simply show up and beat most teams in the country because they have more talent.  Not so in the NBA.  Sure, the Lakers will probably beat the Timberwolves 9 times out of 10 - or even 49 times out of 50 - on talent alone.  But Duke will beat Bethune-Cookman or Denver University 999 times out of 1000 because they are that much more talented.  Even the worst NBA team is still composed of 12 of the top 400ish basketball players in the world (which is far from true in college basketball).  That alone is enough to allow them to compete against anyone they might play, even if they have a disadvantage.  A gameplanned and prepped worst team (by talent) in the NBA would probably beat a completely unprepared best team (by talent) fairly often.

What separates great teams in the NBA, then, is not merely talent, but how that talent is used, and how the coaching staff decides to employ that talent.  Indeed, it is easy for us to confuse talent for preparation at the elite echelons of sport, especially because the ability to fit into a system often wins out over talent in determining who should get a roster spot.  Regardless, it's silly to think that a smaller league would change the impact of game-planning and roster-construction on NBA success.

I want to close with a final observation that seems almost too obvious.  I wonder whether LeBron realizes that, if you cut the number of teams in the NBA, you also cut the total number of points scored, the total number of assists, the total number of rebounds, and so on.  Even if you keep the schedule at 82 games a team, those two or four or ten missing teams don't have players amassing points and minutes, and as the Heat are discovering this year, it's a lot harder to be the guy who scores 30 points per game when you've got 3 guys capable of scoring 30 per game.

A team, in a smaller league, has a couple options: it can try to distribute the ball evenly to its many star players, in which case none of them look quite as starry as before, or it can anoint a leader, and given him the ball most, and make everyone else a role player (or somewhere in between those two options).  In other words, teams still have to make the exact same decisions in a smaller league that they do now.  And, in the end, most of them would choose to take current "stars" and turn them into role-players to even better stars, just like the successful USA Olympics basketball team in 2008 did with Kobe Bryant (defensive specialist) and Carlos Boozer (rebounder), just like the Lakers do, and just like the current Heat are starting to do.

So to any of you who believes that contraction is the way to a better league in any sport, I challenge you to think again.  After all, if smaller leagues were really more compelling, we'd all watch the Harlem Globetrotters instead of March Madness.