Showing posts with label big ideas. Show all posts
Showing posts with label big ideas. Show all posts

Tuesday, June 24, 2008

What's an Assist Really Worth?

Part 1: Getting to the Root of Assists

A couple weeks ago I stumbled onto the Numbers Guy blog at the Wall Street Journal, and added it to my Google Reader. One of the posts in the archive reminded me of something I've been meaning to write about for quite some time: how I deal with assists in the calculation of PAPER.

Since PAPER isn't a linear weights system in which each event has a positive or negative value, I've never been too much concerned with the point-value of an assist. Knowing the difference between the likelihood of making an assisted shot versus a non-assisted one, however, is important. To that end, I compared Field Goal Percentage and Assists:Attempts for all regular-season conference games in the four-season period between 2003 and 2006 and came up with the following (click on the graph to see a larger version):

In the average ACC game in that period, there would be about 6 additional field goals scored for every 10 assists a team racked up. Note that I didn't say the assists led to the baskets being scored; since on a team level, assists are just a subset of field goals made, the assists don't cause the extra baskets to be made. Something else does--something that isn't listed in the box score--and it's something we've got to guess at.

We know the field goal percentages, and we know how many assists were recorded. The fourth critical piece of information if we're going to have any hope of valuing assists properly is the number of times shots were taken on which an assist would have been recorded had the shot been made, regardless of whether or not the shot went into the basket. I call this fourth number the Setup. If we can determine how many Setups occurred, we can find two different field goal percentages for each team: Setup FG% (Assists divided by Setups), and Solo FG% (Field Goals Made minus Assists, divided by Field Goal Attempts minus Setups).

Using the information above, we can see that a team with an Assist-to-Attempt ratio of 0.10 would have an expected field goal percentage of .357; improving that ratio to 0.50 would give an expected field goal percentage of .600. Assuming that both Setup and Solo FG% remain constant, their respective values must be .755 and .297.

Obviously, though, that's an average. Shooting prowess varies from team to team, so those are hardly one-size-fits-all numbers. What's needed is a framework by which we can infer the Setup rates of each team. While it would be tempting to use the ratio of Setup to Solo FG%, which is about 2.5-to-1, that would give an impossible Setup FG% greater than 1 for any Solo FG% greater than .400. Likewise, simply applying the difference of about .46 would break down at Solo FG% greater than .540.

The solution I came up with was to estimate Setup FG% to be the cube-root of Solo FG%, and the red line on the above chart represents the cube-root best-fit line for league-average in the study. While not quite perfect, using the cube-root method has the advantage of never breaking down for any FG% between 0 and 1. Using the cubic formula, we can estimate the number of Setups (S) for any team as long as we know how many field goals they made (M) and attempted (T) and how many assists (A) were recorded as follows:

That looks like an awful lot to keep track of, but it's a fairly simple equation to build into a database or spreadsheet function.

Coming up in Part 2: What does all this allow us to do?

Thursday, February 28, 2008

TAPE: Raising the Bar

Up to this point, a team's TAPE has been a representation of their expected winning percentage against an average team. The problem is, with 341 teams in Division I, there are between 160 and 170 teams that are better than average at any given moment. Average teams, unless they're lucky enough to play in the Ivy League or the SWAC or the Patriot League or the MEAC, aren't even among the best teams in their own conference. Average teams are nowhere near the postseason radar. Using an average team as a reference point in a Division I ratings system just doesn't make any sense at all.

So, if average is out, what is in? Every year around this time, bubble talk begins in earnest. While an increasing number of mid-majors are getting bubble attention, most of the focus rightfully is on the teams from the six BCS conferences. (Yes, I do realize that it's silly to use a football term in a basketball context, but, really, what other nomenclature would work? "Power conference" sounds like a meeting of energy executives, and a "high major" is a marching band leader with an illegal smile, so, as much as I hate the sport with the funny-shaped ball, I'll stick with the BCS.) When it comes to assessing those teams, the most important bit of information seems to be whether or not they have a winning conference record.

With that in mind, TAPE has now been changed to reflect how each team would fare not against the average Division I team, but what their record might look like if they played their entire schedule against BCS-level competition. The number of teams with ratings better than .500 has dropped from 165 to 48. It just so happens that if the NCAA field was built using TAPE, the cut line for at-large bids would be .500. Neat, huh?

Saturday, July 14, 2007

Hurricane Forecast

Miami's projections are ready.

I'm going to take a few days off from generating the projections. Half of them are up, and with the season still almost 4 months away I don't feel any real rush to get the rest out there. Instead I'm going to spend some time on a new project.

In developing the similarity scores used to generate the player projections, it's really jumped out at me just how close a correlation there is between a player's size and the statistics that he generates. Amazingly enough, height and weight are not used at all in generating the similarity scores, yet the top comps are almost always dead ringers, from a body-type perspective, for the base player.

I've been using player size as a proxy for position in PAPER, but that leads to a few different kinds of problems. First, sometimes there are guys--Ishmael Smith and Will Bowers, to cite one from either extreme off the top of my head--who are so far outside the normal range of players that the regressions that tell me what the "league average" player of their size should be doing just don't have enough good data to work with. They wind up with funky results (a player T.J. Bannister's size in 2005, for example, should have blocked -0.006 shots per defensive possession according to the model) that I've either got to just roll with or jury rig out somehow. In the grand scheme of things, it's not a huge deal, or one that affects the bottom line number much, if at all, but I've never cared for it.

The second problem is that it treats players who are big or small for their position differently than they probably should be. There's an adjustment using BMI that refines the raw height number into something a little bit closer to the truth, essentially adding an inch or so to the height of the stouter players while shaving one off of the scrawnier ones, but PAPER still treats Mamadi Diane and Greivis Vasquez exactly the same, even though they have vastly different responsibilities on the court.

Getting to the point, I think there may be a way to make PAPER better by scrapping the size adjustments and using positional adjustments instead. I'm already making what I think is a pretty safe assumption: players of similar size generally tend to play the same position. But if I can break that down, to find some markers in the numbers that say, "This guy is a shooting guard," or, "That guy is a defensive specialist," it will only make the system better.

I don't know if this will go anywhere or not, but I'm going to think on it for a few days and see what happens. In the meantime, if there's a player on a team that hasn't been posted yet, and you're just dying to see what the Magic Spreadsheet says he's going to do next year, drop me an email. Running the individual projections takes like 5 seconds, and I do them out of curiosity all the time; it's just the formatting everything into a nice little package that takes a couple hours for each team.