Showing posts with label statistical malpractice. Show all posts
Showing posts with label statistical malpractice. Show all posts

Tuesday, August 5, 2008

More statistical malpractice from Tweed: Joel Klein and his claims do not measure up

Check out Elizabeth Green’s article today in the NY Sun, in which she asked three academics to scrutinize the validity of the administration’s claim of narrowing the achievement gap between ethnic and racial groups.

See, for example, Bloomberg’s recent testimony before Congress, in which he said that “over the past six years, we’ve done everything possible to narrow the achievement gap – and we have. In some cases, we’ve reduced it by half.”

Yet the evidence for this is weak to non-existent. On the national tests called the NAEPs, there has been no narrowing of the achievement gap in any area since the Bloomberg/Klein reforms were instituted:

An analysis by the National Center for Education Statistics, the research arm of the federal Education Department, concludes that no achievement gaps have narrowed at all in New York City between 2003 and 2007. The only gap that moved in any significant direction is the one between poor students and the rest of the population, which widened slightly, that analysis said. The National Center for Education Statistics also concludes that upward trends in the reading scores of black and Hispanic fourth-graders lauded by Mr. Klein are not statistically significant.

In the article, Joel Klein reveals his statistical illiteracy:

“Those are just confidence levels. Nobody is saying this is a science," Mr. Klein said. He added: "If three points is flat, and four points is statistically significant, then what you're doing is, you're playing something of a game."

Chief press officer David Cantor called the memo from NCES "a politicized gloss.”
Instead, it is the DOE who insists on playing games – and politicizing the issue, by continuing to slander experts as somehow biased when they provide objective evidence that the non-stop PR spin issuing from Tweed has no basis in reality.
This is hardly the first time the DOE has revealed such statistical malpractice. Jim Liebman, law professor and head of the DOE accountability office, is a repeat offender. I recall one episode in particular when Liebman, testifying before the City Council, insisted that he wasn't basing school grades primarily on the results of a few tests, since each test was really "multiple assessments" given out over "multiple days," resulting in "multiple measures" of proficiency.
Another instance of this was the presentation of Jennifer Bell-Elwanger, head of testing for DOE, who was giving a power point presentation to the Panel for Educational Policy in late November, following the release of the NAEP results. She continually pointed out gains that, according to the NCES, were not statistically significant. Patrick Sullivan, Manhattan rep to the PEP and fellow blogger here, who seems to understand data better than anyone currently employed by the DOE, questioned her closely, saying, "But by definition these are insignificant gains, no?" Which she, of course, fervently denied.
Indeed, the memo from the NCES was prepared in response to a highly misleading email that Klein sent to nearly every NYC resident after NAEP results were first reported. In his email, he falsely claimed that the results showed “good progress that is consistent with the overall picture” and said that the NAEP showed a narrowing of the gap in nearly all areas, when the data itself revealed quite the opposite.

According to NYC’s results on the state exams, the situation is more complicated. The achievement gap is narrowing in some areas when one looks at “proficiency” levels, that is whether a student is at a level 1, 2, 3, etc., but not in terms of the actual scale scores.

Some testing experts consider proficiency levels less meaningful than scale scores, as they can be arbitrary, subjective and easy to manipulate. Daniel Koretz, a professor at Harvard and a national expert on testing, has just published the must-read book of the summer, Measuring Up: What Educational Testing Really Tells Us. Here is what Koretz has to say about proficiency levels:

….the percents deemed proficient [on state tests] are largely unrelated to states’ actual levels of student achievement…are also often inconsistent across grades or among subjects in a grade....[This system] obscures a great deal of information [because] of the coarseness of the resulting scale.
Even more importantly, why should we trust the state scores at all, when we know that the gains that they purport to show are not reflected in more trustworthy measures like the NAEPs?

The best thing about the Koretz’ book is his lucid explanation of why “test score inflation” inevitably occurs when you attach high-stakes to exams, and how this undermines the integrity and validity of the results; this has increasingly been the case throughout the nation as a result of NCLB, but even more here in NYC, as a result of the increasingly high-stakes policies of the Bloomberg/Klein administration.

Steve Koss has written about this eloquently on our blog, in relation to Campbell’s Law: “The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor.”

I have tried to explain this phenomenon to many elected officials, staff, and reporters over the years, apparently with little success. I certainly don’t know a single NYC media outlet that has ever mentioned it, though Campbell’s Law was cited in some recent letters to the NY Times in response to the administration’s experiment to pay students for high scores.

I recall a lengthy discussion of this issue several years back with NY Times reporter David Herszenhorn while he was still on the education beat, to explain my opposition to the Mayor’s newly-proposed 3rd grade retention policy. One of the reasons I so vociferously opposed this policy, and still do, was not just that it was unfair to the student to base such a life-altering decision on the basis of one single, fallible test score, with such large margins of error; and not just that retention has been shown to have a racially-disparate impact and hurt rather than help most low-performing students.

My opposition was also due to the fact that the more significant consequences are attached to any test, the less its results can be trusted as a reliable gauge of real learning.

Since then, of course, the administration has piled on more and more high-stakes consequences -- for students, teachers, and schools – by adding fifth and seventh grade retention, awarding principals, teachers and students monetary rewards for high scores, and threatening to close down schools if scores don’t improve fast enough. The scores themselves have been rendered entirely meaningless as a result, as excessive test prep, teaching to the test, cheating, and other strategies to “game” the system has totally overtaken our schools.
Here is what Koretz has to say about this:

"One might expect that with the huge increase in the amount of testing in recent years, we would know more…Ironically, the reverse is true. While we have far more data now than we did twenty of thirty years ago, we have fewer sources of data that we can trust. The reason is simple: the increasing in testing has been accompanied by a dramatic upsurge in the consequences attached to scores. This is turn has created incentives to take shortcuts --- various forms of inappropriate test preparation, including outright cheating – that can substantially inflate test scores, rending trends seriously misleading or even meaningless.”
To the administration, all this appears acceptable, because as long as test scores go up, this justifies their policies; in truth, they don’t seem to care if the increases are meaningful or not.

Their laissez-faire attitude is revealed by the total lack of interest evinced in following-up on even well-documented cases of cheating. (See for example this story in the NY Sun, which though it says the DOE is “investigating” this will likely lead nowhere, as such stories have in the past.)

This “anything goes” attitude is also reflected in Klein’s remarks in a recent interview in the NY Post:

Q: What about complaints about the report-card grades for schools?

A: The report cards were probably one of the noisy periods. But . . . I can't tell you how many principals said to me, 'You know, chancellor, I didn't get the right grade but I promise you I won't get the same one next year,' so I think that had a big impact.

Here, Klein implicitly acknowledged that even while principals do not accept the fairness of these grades, based primarily on one-year gains in scores, he is content as long as they guarantee that they will get these scores to rise in the future, by any means possible.

Wednesday, April 9, 2008

Use test scores for tenure? Not a good idea, with these bumblers

So the NY State budget finally was decided, with all the proposed cuts to NYC schools restored, and the promise of CFE maintained, yet all the newspaper editorial boards and bloggers can do is to blather about a provision in the budget that prohibits the use of student test score data in making teacher tenure decisions.

Eduwonkette provides some of the links to the bloggers who are so outraged as to contend that this is the end of the civilized world.

Actually, the final language in the budget bill was a reasonable compromise, in which it was agreed that there will be a two year moratorium while a commission considers how best this information can be utilized to inform tenure decisions.

Evaluating a teacher’s competence on standardized test scores alone is not sufficient, since the gains or losses that any class achieves in scores is often highly erratic from year to year, is partly based on factors such as class size which is quite variable across NYC schools, and the background of students in each class.

Actually, research shows that its not just the current class size that helps determine the rate of learning, but a student's past class sizes, which can change the entire trajectory of his or her academic career.

And what are they going to do about the fact that many of the tests are given in the middle of the year? The DOE's proposed solution is to give last year's teacher half the credit, but that assumes equal effectiveness of all teachers -- which is contrary to the whole point of this exercise - that some teachers are more effective than others.

Moreover, test scores do not tell the whole story. Other evidence of a teacher's skills and value are equally if not more important, including her ability to motivate students, keep them engaged, and guide them in their writing, their projects and all other types of creative learning that cannot be assessed by test scores alone.

Most importantly, it is by now abundantly clear that this statistically illiterate administration cannot be trusted to use this data carefully and intelligently, with a grain of salt and in relation to other critical factors, given their record on merit pay and school grades.

In both cases, they chose to base the results primarily (85%) on test scores, with more than half based upon the essentially unreliable gains or losses in scores over one year alone.

Tying tenure to test scores could have very destructive effects, discouraging teachers from taking on classes of struggling or special ed students, and lead to a further loss of morale, with even more test prep replacing real learning.

A hiatus of two years is a terrific idea since whatever is decided will be implemented by a new administration that will hopefully be more trustworthy with the use of such data. We know that the bunch of bumbling amateurs in charge of our schools now would never be able to figure out how to balance all these factors in an intelligent, humane and constructive fashion.

For more on this issue, including comments from Chancellor Klein, Randi Weingarten and me, see the Channel 2 report here.

Sunday, December 2, 2007

"Negative learning" and statistical malpractice at the Panel on Educational Policy

At last week’s meeting of the Panel on Education Policy at Tweed, Jim Liebman’s performance in attempting to defend the indefensible – the school grading system that he designed -- was breathtaking in its ignorance.

Liebman, the current DOE accountability “czar,” is a former criminal attorney, currently on leave from the Columbia law school, with no training or experience in education policy, statistics or testing, and yet the entire educational focus of the DOE is now based upon his faulty theories and expensive initiatives, including the $80 million supercomputer called ARIS, assigning letter grades to all schools primarily on the basis of one year’s worth of test scores, devoting millions of more dollars and hours of precious classroom time to interim standardized assessments, and the creation of “data inquiry teams” in all schools – all in the effort to “differentiate instruction” which in the end will be impossible without smaller classes.

At the PEP meeting, in order to justify the school grading system, he fastened on the “F” that PS 35 in Staten Island received, a school in which 98% of its students are on grade level in math, and 86% in ELA. Why did this exemplary school receive an “F”? Because last year, only 35% of its students improved their scores over the year before in reading, and only 23% in math – though research shows that a large part of annual variations in test scores are based on chance alone and are statistically unreliable. (For more on this, see my Daily News oped and a previous posting, Ten reasons to distrust the new accountability system.)

During the discussion, Liebman compared PS 35 to one of its “peer” schools – the Anderson school, a citywide Gifted and Talented school that accepts students on the basis of their high IQ and high test scores. When Patrick Sullivan pointed out the unfairness of comparing PS 35 to a selective school like Anderson, Liebman said it didn’t matter how the kids got there, they should all make the same annual gains. He failed to mention, however, that elementary schools are grouped with other schools according to only the roughest measures of demography –and that no statistician would compare the performance of a school that selects its students on the basis of test scores with a neighborhood school, like PS 35, that has to admit every child in its zone.

There was an abundance of statistical malpractice on display that night -- between Liebman’s presentation and the talk given by the DOE testing “expert”, Jennifer Bell-Elwanger, who tried to convince the panel that the city’s lack of significant progress on the NAEPs since 2003 was indeed real progress. Both of these individuals would have flunked an elementary course in statistics if they had tried to make these arguments in a college exam.

When asked wouldn’t it better to have separate grades for achievement and progress, rather than collapse all these categories into one grade, even if he were convinced that the lack of one year’s progress in test scores was significant (which it isn’t) Liebman replied that the good thing about giving a single grade is that it gets people’s attention (or something like that.) One could say the same about threatening to cut off the hands of someone accused of theft, or even capital punishment, which doesn’t mean it’s a remotely fair practice or even useful.

More recently, in response to questions about class size from parents in Manhattan and Queens, Liebman has insisted that the reason the DOE refuses to reduce class size is that classes would have to shrink to below 15 students to improve instruction and/or student achievement. In other words, lowering class size from 30 to 20 would make absolutely no difference.

Not only is such a statement absurd to anyone who has actually spent any time teaching in the public schools or observing classrooms, it is completely unsupported by research. Instead, it is simply another lame excuse that opponents of reducing class size like to throw up as a smokescreen in order to discourage such efforts.

Here is a comment sent to me from Chuck Achilles, a principal investigator of the famed STAR experiment in Tennessee and a professor of at Eastern Michigan University and at Seton Hall University. Chuck is also one of the premier class size researchers in the world:

“Hi Leonie:

I thought that the “below 15” idea (archaic) had faded. Anyone who says that is uninformed and ought to be asked (challenged) publicly to defend the assertion. It came once from one meta-analysis (Glass & Smith, 1988) that was very limited in its n of observations (77, of which some were for physical skills like hitting a tennis ball against a wall.) Just in STAR, we had more than 1300 observations in the range of 12-28 students. We typically analyzed reading outcomes, but sometimes we did math (giving us 2600 comparisons) and could have used other academic (test) outcomes… I’ve faxed some pages to show the linear effect: About a correlation of -.35 for each student added to a class. Because STAR used the class average as the unit of analysis, this means (approximately) the addition of each student to a class in the n=12-28 range reduces the class average score (about .1 of a month per year.) Later analyses show that it is cumulative.

Chuck A.”

Here is a fact sheet with numerous citations, showing there is no threshold in terms of reducing class size; and that the increase in achievement in relation to the decrease in class size is roughly linear.

Liebman reminds me of a phenomenon called “negative learning” ---in layman’s terms, a little learning is a dangerous thing. One would think that someone who got his reputation by writing about the high error rate in capital punishment would have a little humility and understand the possibility of human fallibility in making absolute judgments, but no such luck.

Saturday, November 10, 2007

Eduwonkette on the "statistical malpractice" of this administration

Check out the telling critique of the new school grading system from Eduwonkette, an astute new blogger:

"Earlier this week, Mayor Michael Bloomberg flexed his muscles by threatening to close F schools as early as June. He quipped, "Is this a wake-up call for the people who work there? You betcha."

Through analyzing these data, I've concluded that the people in need of a wake-up call work not at F schools, but at the NYC Department of Education....There are five reasons the report cards might kindly be called statistical malpractice".

One statistical anomaly she points to: the grades received by the schools that run from 6-12th grades. Of the 33 schools, 22 have different grades at the middle and high school level -- many of them sharply different. Examples?

"Consider the Academy of Environmental Science - its high school got a C, but its middle school got an F. At Hostos Lincoln Academy of Science, the middle school got a D, but the high school got a B. At the Bronx School for Law, Government, and Justice, the middle school got an F, but the high school got a C."

Clearly the grades these schools received reveal nothing useful about the leadership or the overall quality of the school, as the administration would maintain.

Is this a wake-up call for the people who work there? You betcha.