Showing posts with label growth scores. Show all posts
Showing posts with label growth scores. Show all posts

Friday, December 13, 2019

Why the DOE's analysis of the second year of the literacy coach program provokes more questions than answers

Yesterday, the DOE released an evaluation of the second year of its literacy coach program, created under Chancellor Farina in 2016-2017.  The program is budgeted at $85.7 million this year, for a total of approximately $235 million since its inception. Each literacy coach has a salary of $90,000 to $150,000, and there are approximately 515 of them working with K-2 teachers in selected elementary schools across the  city.  The number of schools and coaches have expanded each year.
In August 2018, the DOE released a brief power point which purportedly contained the sole written evaluation of the first year of the program.  By analyzing the growth scores from October to May of second graders who were administered the Gates-MacGinitie Reading Tests (GMRT) at schools that received literacy coaches, compared to students at similar schools without coaches, they found no positive results in either word decoding, word knowledge, or comprehension.

The results were expressed in grade equivalents, which allowed one to see that rather than progressing, both groups had fallen further behind in terms of grade level.
I submitted a Freedom of Information request for the second year results shortly afterward, since these students' test scores should have been available and could have easily been analyzed by then.  Instead, the DOE waited until December 2018 to tell me that NO such analysis had yet been done. Yet they subsequently proposed an expansion of the program in any case, with little evidence to back it up.

I submitted another FOIL request in August 2019, and on Monday, Dec. 9, 2019 they told me they would delay any substantive response till January 16, 2020, with this excuse: "[4] the need to review records to determine the extent to which they must be disclosed, [and 5] the number of requests received by the agency."
Three days later, on Dec. 12, they released the second year evaluation, at which point there were 236 literacy coaches, serving 298 elementary schools, spending an average of 20 periods with each teacher.  168 reading coaches were assigned to one elementary school only and 68 served two schools.
The study found a slight but significantly larger average growth score in reading comprehension among second graders who were tested in a sample of about one third of the first and second cohort of schools that had access to literacy coaches, as compared to second graders at similar schools that did not.  No significant gains were found overall, or in the other two specific areas of word decoding or word knowledge.  

In several ways, the evaluation  is quite thin.  Here are some of the most glaring omissions:
1- One would expect that the DOE would not merely test 2nd graders in the program, but also analyze the scores of these same students as 3rd graders, to see if the slight gains that they experienced earlier had persisted or disappeared.  They could do this without administering more tests by comparing their 3rd grade state test scores to students at similar schools without coaching. No such analysis is mentioned in the paper.
2- Curiously, the DOE also seemed to switch its methodology in reporting test score gains compared to the first year study.  See this:
 “...in this report we use extended scale scores. In previous reporting, we used grade equivalent scores. Scale scores refer to the continuous scale on which GMRT results are measured, from Pre-Reading to Adult Reading. While grade equivalents are more easily understandable, scale scores are more precise and are used for analyses. Scale scores on the GMRT capture students’ reading ability on a linear scale that is useful for both comparison across grades and for analysis.”
I'm not sure why they made this change, unless the one area of comprehension where there appeared to be a significant difference, the gain was too small to show up in terms of grade equivalents.  Or perhaps they didn't want to reveal that even those students at schools that had achieved significant gains still fell behind in terms of grade level?
3.  The report makes other claims without any data to support them, and overall it is surprisingly sparse in actual statistics.  See this claim for example: "Students of teachers who received ULit coaching grew more than students of teachers who did not."
The study does not include any  data to back this up, and does not actually say that the difference in growth was significant.
What this may suggest is that while the growth scores of students at schools that had access to coaches were significantly greater in the area of comprehension, the same significance did not occur when the scores of the students of teachers that actually received coaching were compared to students of teachers who did not receive coaching.

This passage is similarly hard to understand:
Moreover, the more coaching a teacher received, the more growth the students had, on average. The difference between students whose teachers had more coaching and those who did not was statistically significant.

Again, no data is provided to back up either claim, and no definition of  “more coaching” is offered in which would allow one to evaluate the claim made by the second statement.  How much more coaching was needed for a teacher' students to exhibit statistically significant test score gains? 
4.  The report also doesn't provide the full questions or results of the teacher/coach/principal surveys.  Though it says that almost half of teacher respondents responded that if they had a reading coach the following year, they would like the coach to work with them "one-on-one", it doesn't report if teachers were asked if they wanted coaching at all, or if they believed the funding might be better spent on more classroom teachers to lower class size or to hire intervention teachers who would work directly with struggling readers.

No information is provided either as to whether the coaches believed their training was useful, even though a third of school leaders reported that  "the coach was out of the building too often for professional learning."
5.  We are now in the midst of the fourth year of the program. Why is the DOE just now  releasing second year results, instead of the third year results that should be available - especially given that we are spending nearly $100 million a year on the program? Or do the results have to be carefully sifted and parsed before released, as these appear to be?
All in all, there needs to be a more comprehensive study with much more data provided before one can be assured that the program is providing benefits to kids worth the cost.  And such an analysis would be far more credible coming from an experienced and independent research outfit like RAND.

Tuesday, May 10, 2016

Breaking: the Lederman decision and Gallup poll: the beginning of the end of high-stakes testing?

Today, the court decision in the Sheri Lederman case was issued.  Judge Roger McDonough of the NY State Supreme Court concluded that rating teachers via their students' growth scores on the state exams is "arbitrary and capricious."  He cited a wealth of evidence from affidavits of academic experts such as Linda Darling-Hammond, Sean Corcoran, Aaron Pallas, Carol Burris, Audrey Amrein-Beardsley, Jesse Rothstein  and others, showing that the system of evaluating teachers by means of test scores is unreliable, invalid, unfair and makes no sense.  The full court decision is below.

Here is the message from her attorney (and husband) Bruce Lederman:

"I am very pleased to attach a 13 page decision by J
udge Roger McDonough which concludes that Sheri has “met her high burden and established that Petitioner’s growth score and rating for the school year 2013-2014 are arbitrary and capricious.” The Court declined to make an overall ruling on the rating system in general because of new regulations in effect. However, decision makes (at page 11) important observations that VAM is biased against teachers at both ends of the spectrum, disproportionate effects of small class size, wholly unexplained swings in growths scores, strict use of curve.

The decision should qualify as persuasive authority for other teachers challenging growth scores throughout the County. Court carefully recites all our expert affidavits, and discusses at some length affidavits from Professors Darling-Hammond, Pallas, Amrein-Beardsley, Sean Corcoran and Jesse Rothstein as well as Drs. Burris and Lindell . It is clear that the evidence all of these amazing experts presented was a key factor in winning this case since the Judge repeatedly said both in Court and in the decision that we have a “high burden” to meet in this case. The Court wrote that the court “does not lightly enter into a critical analysis of this matter … [and] is constrained on this record, to conclude that petitioner has met her high burden” ...To my knowledge, this is the first time a judge has set aside an individual teacher’s VAM rating based upon a presentation like we made.

THANKS to all who helped in this endeavor."

At the same time, a national poll was released by the Gallup organization showing how most parents, teachers, students and administrators do not believe state exams are useful:

 Most teachers find their quality of the state exams are only "fair" or "poor":
And most families, whatever their income level, do not believe that these exams improve learning:

Let's hope that together these poll results, along with the Lederman decision, sound the death knell for the obsession with high-stakes testing that has overtaken our schools.

Tuesday, October 2, 2012

Why no one in his right mind should believe the school grades OR the teacher growth scores

UPDATE: Here is Gary Rubinstein's new graph of school ranks. this year compared to last, according to the Progress reports.  The "x" axis is 2011; the "y" axis is 2012. 
The new DOE school grades for elementary and middle schools, euphemistically called “Progress reports,” came out with much fanfare, with 217 schools potentially put on the closing list because they either received failing grades or three “Cs” in a row, more than ever before.  As parent leader Shino Tanikawa pointed out in the New York Post, DOE is using unreliable grades to implement terribly misguided policies. (If you'd like to see your school's grade anyway, you can find it here.)

Last year’s headlines ran like this: School report cards stabilize after years of unpredictability. This year again, reporters cited claims of DOE officials, who “highlighted the stability of this year’s reports.”
It wasn’t true last year and it isn’t this year either.  Here is a figure from Gary Rubinstein’s blog from last year, showing no correlation between the rank order of schools in 2010 and 2011. I suspect the Gary’s illustration will be similar for this year, once he gets around to making it.
As InsideSchools reported, 24 out of the 102 schools that received “D”s or “F”s this year had received top grades of “A” or “B” the year before. Other high-performing schools, such as PS 234 in Tribeca that received an “A” last year, fell precipitously to a “C” with the same principal, same staff and most of the same students.  The school plunged from the 81st to the 4th percentile, which would have meant a “D,” if not for the DOE rule that no school that performs in the top third citywide can receive a grade lower than C, as Michael Markowitz pointed out in a comment on GothamSchools.
According to the DOE formula, 80-85% of school’s grade depends on last year’s test scores on the state exams.  Most of that figure is based on supposed “progress,” i.e. the change in test scores from the year before.  As I have pointed out many times, including in this 2007 Daily News oped, Why parents and teachers should reject the new grades”, experts in statistics have found that the change in test scores at the school level from one year to the next is highly erratic and up to 32-80% random. 
In recognition of this fact, Jim Liebman, who developed the school grading system, originally told skeptical parents that the system would eventually incorporate three years of test scores, which would lessen the huge amount of volatility, a promise that DOE has failed to abide by.  For proof of this promise, see Beth Fertig’s book, Why Can’t U teach me 2 read?:
“…[Liebman] then proceeded to explain how the system would eventually include three years’ worth of data on every school, so the risk of big fluctuations from one year to the next wouldn’t be such a problem. (p.121)”
To make things worse, the state exams last year and the scoring guides were riddled with errors, as many parents, teachers and students noted . Finally, test scores are not a good way to assess school quality, for myriad reasons, even if the tests were perfect and the formula based on multiple years worth of data, as I pointed out in this NYT Room for Debate column
LESSON: Anyone who believes in the accuracy of these school grades is sadly misinformed, and DOE’s attempts to dissuade parents from sending their kids to schools with low grades or to close schools based upon such unreliable system is intellectually and morally bankrupt.
In another untenable move, the teacher “growth scores,” also based on the one year’s change in test scores, but this time at the classroom level, have been released to principals outside NYC.  These growth scores will be incorporated in the new teacher evaluation system to be imposed statewide.  See this excellent column by Carol Burris about how the heedless use of growth scores is likely to damage the education of our kids. 
Here are the detailed results of the statewide survey by principals, including their comments, showing that the vast majority believe that these scores are NOT an accurate reflection of the effectiveness of individual teachers.  More than 70% of principals also said they were “doubtful” or strongly opposed to any use of growth scores in teacher evaluations.  Aside from the annual volatility in these scores, there are many other problems with relying upon such flawed measures of teacher quality:
The growth scores, developed for the state by the consulting company AIR, attempted to adjust only for the following demographic factors: 

  • Economic disadvantage (ED), but without differentiating free lunch or reduced lunch students – very different, with very different expected outcomes;
  • Students with disabilities (SWDs), but not types or severity of disability;
  • English language learners (ELLs).

And even though AIR did attempt to control for the above factors, they still admitted that teachers who work at schools with large numbers of students who were poor or had disabilities tended to have lower growth scores. 
AIR made NO attempt to control for classroom characteristics such as class size, or the racial/ethnic background of students, or any other variable that is known to affect achievement. This is different from the NYC value-added teacher data reports, released last year by DOE, widely derided as unfair and unreliable, which at least attempted to control for many of these factors.
Even so, Bruce Baker, professor at Rutgers, found that while the teacher data reports claimed to control for class size, teachers at schools with larger class sizes were significantly more likely to be found ineffective than those who taught at schools with small classes, and as “class size increases by one student, the likelihood that a teacher in that school gets back to back bad ratings goes up by nearly 8%.”  The situation with these growth scores is yet worse, with no attempt made to control for class size at all.  Is it fair to deny NYC teachers tenure and/or risk losing their jobs because they are saddled with larger classes than teachers in the rest of the state?
AIR also found that teacher growth scores tended to be higher in “schools with students with higher mean levels of ability in both ELA and in Mathematics.” This shows that teachers who work in schools with large numbers of low-achieving students are more likely to be found ineffective. 
Apparently NYC principals have not yet been offered the opportunity to examine the growth scores of their teachers, unlike principals in the rest of the state. As far as I know, DOE has not explained why.  With far larger numbers of struggling students who are economically disadvantaged and crammed into larger classes, it is quite likely that a higher proportion of NYC teachers will be found ineffective than elsewhere in the state– and thus unfairly penalized for teaching disadvantaged students in worse conditions.
Here is the link where you can find the AIR technical manuals on growth scores.