Is there a mathematician in the house?
Unlike most of my arguments with K, this one may have a definitive answer. A little help, please.
E: "...took two freakin' hours to get home on the bus tonight! And another thing, what is up with these commentators who cite polls on hypothetical McCain-Obama and McCain-Clinton match-ups, concede that all the results are within the margin of error, and then go on to "analyze" the results, as if they'd just conceded that Obama had a cat named Tinkerbell?"
[Bit of transcript redacted in which K and E actually agree that any such pre-general election poll is ultimately meaningless, given that Clinton's numbers are, as previously mentioned, "pre-shrunk," while Obama's are innocent of any meaningful encounters with organized Republican reptilianism.]
K: "But if a series of such polls are conducted, each saying essentially the same thing, can't we be more confident in saying that his electability [qualified by the previous redaction; K and I are here just arguing for the fun of it] is higher than hers?"
E: "No, you're just replicating a meaningless result. Nothing times n is still nothing."
K: "But doesn't probability play into it? The confidence of getting result X in a single poll is 95%. But the confidence of getting result X and result Y...?"
E: "But these aren't linked events. The margin of error is just the sampling error between that particular sample and the population; it resets as soon as you draw a new sample."
K: "Ok. Say you have ten polls that each suggest Obama would garner 46 percent of the vote, with a three-point margin of error, so you know the actual percentage would range from 43 to 49. If you have multiple such polls, aren't you establishing that the true value is likely to be in the center, and, in fact, 46?"
E : "No. The margin of error tells you that you don't know the exact shape -- the tails, mainly -- of your distribution. For each sample, the chance that you have that outer limit falling in the right slot, 43 through 49, is equal for every value within that range, each time. Getting ten polls that give you a 46 is no weirder than getting heads on each of ten coin tosses -- and no more of a basis for predicting what the 11th would be."
K: "But, assuming the questions are the same each time, aren't you, in a way, aggregating your sample? If the population of interest is the American voter, isn't each draw just a subset of the same sample? Say that each of these unanimous samples had an n of 100. Your aggregate then has an n of 1000 and thus does have a lower margin of error."
E: "Shhh. I'm blogging."
-----------
Honestly, I think he's just a shade shy of correct. Essentially what we're talking about is meta-analysis, right? I can't remember how to actually do those, but I'm guessing the computation takes into account exactly these issues. Can anyone give a quick-and-dirty explanation for an innumerate such as me (and sometimes K)?
And yes, I suppose I ought to ease up my criticism if I can't even get it right myself. But my original irritation remains: in the majority of cases, the analyst is talking about one poll in which all candidates are within the margin. And each discrete instance of this occurs many, many times. I still maintain: that's neither news reporting nor analysis; it's news manufacture, on a gross and disturbing scale.
E: "...took two freakin' hours to get home on the bus tonight! And another thing, what is up with these commentators who cite polls on hypothetical McCain-Obama and McCain-Clinton match-ups, concede that all the results are within the margin of error, and then go on to "analyze" the results, as if they'd just conceded that Obama had a cat named Tinkerbell?"
[Bit of transcript redacted in which K and E actually agree that any such pre-general election poll is ultimately meaningless, given that Clinton's numbers are, as previously mentioned, "pre-shrunk," while Obama's are innocent of any meaningful encounters with organized Republican reptilianism.]
K: "But if a series of such polls are conducted, each saying essentially the same thing, can't we be more confident in saying that his electability [qualified by the previous redaction; K and I are here just arguing for the fun of it] is higher than hers?"
E: "No, you're just replicating a meaningless result. Nothing times n is still nothing."
K: "But doesn't probability play into it? The confidence of getting result X in a single poll is 95%. But the confidence of getting result X and result Y...?"
E: "But these aren't linked events. The margin of error is just the sampling error between that particular sample and the population; it resets as soon as you draw a new sample."
K: "Ok. Say you have ten polls that each suggest Obama would garner 46 percent of the vote, with a three-point margin of error, so you know the actual percentage would range from 43 to 49. If you have multiple such polls, aren't you establishing that the true value is likely to be in the center, and, in fact, 46?"
E : "No. The margin of error tells you that you don't know the exact shape -- the tails, mainly -- of your distribution. For each sample, the chance that you have that outer limit falling in the right slot, 43 through 49, is equal for every value within that range, each time. Getting ten polls that give you a 46 is no weirder than getting heads on each of ten coin tosses -- and no more of a basis for predicting what the 11th would be."
K: "But, assuming the questions are the same each time, aren't you, in a way, aggregating your sample? If the population of interest is the American voter, isn't each draw just a subset of the same sample? Say that each of these unanimous samples had an n of 100. Your aggregate then has an n of 1000 and thus does have a lower margin of error."
E: "Shhh. I'm blogging."
-----------
Honestly, I think he's just a shade shy of correct. Essentially what we're talking about is meta-analysis, right? I can't remember how to actually do those, but I'm guessing the computation takes into account exactly these issues. Can anyone give a quick-and-dirty explanation for an innumerate such as me (and sometimes K)?
And yes, I suppose I ought to ease up my criticism if I can't even get it right myself. But my original irritation remains: in the majority of cases, the analyst is talking about one poll in which all candidates are within the margin. And each discrete instance of this occurs many, many times. I still maintain: that's neither news reporting nor analysis; it's news manufacture, on a gross and disturbing scale.
2 Comments:
-
Ted Frank said...
-
- 11:21 AM
-
Ted Frank said...
-
- 11:22 AM
Post a CommentKevin's right in the abstract that, ceteris paribus, multiple samples should correspond to a larger sample with a smaller margin of error, but I've seen pollster-bloggers claim that that such aggregations are not appropriate when one is talking different sampling methods, but I haven't seen the mathematical argument for their claim.
You're correct that the question is meaningless because the Hillary apple is a different animal than the Obama orange, and Obama's poll numbers are artificially high given that voters haven't seriously considered the question. (Huckabee's numbers are artificially low for the same reason, but that's a different story and, one hopes, a moot one.)
P.S. How's that three-week deadline going?
<< Home