Showing posts with label value added. Show all posts
Showing posts with label value added. Show all posts

Friday, August 05, 2011

letter + response from Joanne Barkan re teacher value-added at Dissent


From Leonie Haimson:

See the critical comment on "Firing Line" from Claire Robertson-Kraft who is the Associate Director of Operation Public Education, a University of Pennsylvania project which has been instrumental in promoting value-added modeling.

And Joanne Barkan's response at this link:.

http://dissentmagazine.org/atw.php?id=521

Barkan has a  great rebuttal you should read; here are some other points of my own:

"Most of the reformers’ systems are not simply focused on measuring performance and dismissing teachers, but also on providing them with more robust mentorship and meaningful ongoing professional development."

Umm, not really! They focus on simplistic and reductionist methods of evaluation to push as many teachers out schools asap.  Hanushek has said that 5-10% of teachers should be fired each year.  There is nothing if the corp reform agenda that really deals with giving teachers more support or a better chance of success.

As I see it, the issue isn’t with the elements of the reforms themselves. Rather, it’s a matter of the process used to design and implement the new systems. Largely driven from the top-down, these initiatives are being rolled out much too quickly without garnering the input of teachers and other stakeholders that would be necessary to sustain the efforts over time.

No, the elements of reforms themselves are essentially destructive and anti-thetical to good teaching.  This is like Joel Klein saying his biggest error is not communicating his policies better to parents.

Rather than engaging in polarizing discussion, we should be focusing our efforts on supporting promising legislation like a recent Illinois education law—which brought multiple stakeholders to the table to work together as partners in reform.

Yes, Jonah Edelman of Stand for Children told the truth about that; how he outfoxed and outgamed the teachers unions – in part by offering millions of dollars to state legislators to push his essentially anti- teacher legislation.

Monday, December 27, 2010

NY Times on Value-Added: Dec. 27, 2010

http://www.nytimes.com/2010/12/27/nyregion/27teachers.html?_r=1&hp

By SHARON OTTERMAN
Published: December 26, 2010

For the past three years, Katie Ward and Melanie McIver have worked as a team at Public School 321 in Park Slope, Brooklyn, teaching a fourth-grade class. But on the reports that rank the city's teachers based on their students' standardized test scores, Ms. Ward's name is nowhere to be found.

Melanie McIver, a teacher at Public School 321 in Park Slope, Brooklyn, with Elizabeth Phillips, background, the school principal. Both women have seen issues related to the city's system of ranking teachers, which is at the heart of a lawsuit in State Supreme Court in Manhattan.
"I feel as though I don't exist," she said last Monday, looking up from playing a vocabulary game with her students.

Down the hall, Deirdre Corcoran, a fifth-grade teacher, received a ranking for a year when she was out on child-care leave. In three other classrooms at this highly ranked school, fourth-grade teachers were ranked among the worst in the city at teaching math, even though their students' average score on the state math exam was close to four, the highest score.

"If I thought they gave accurate information, I would take them more seriously," the principal of P.S. 321, Elizabeth Phillips, said about the rankings. "But some of my best teachers have the absolute worst scores," she said, adding that she had based her assessment of those teachers on "classroom observations, talking to the children and the number of parents begging me to put their kids in their classes."

It is becoming common practice nationally to rank teachers for their effectiveness, or value added, a measure that is defined as how much a teacher contributes to student progress on standardized tests. The practice was strongly supported by President Obama's education grant competition, Race to the Top, and large school districts, including those in Houston, Dallas, Denver, Minneapolis and Washington, have begun to use a form of it.

But the experience in New York City shows just how difficult it can be to come up with a system that gains acceptance as being fair and accurate. The rankings are based on an algorithm that few other than statisticians can understand, and on tests that the state has said were too narrow and predictable. Most teachers' scores fall somewhere in a wide range, with perfection statistically impossible. And the system has also suffered from the everyday problems inherent in managing busy urban schools, like the challenge of using old files and computer databases to ensure that the right teachers are matched to the right students.

All of this was not as important when the teacher rankings were an internal matter that principals could choose to heed or ignore. City officials had pledged to the teachers' union that the rankings would not be used in the evaluation of teachers and that they would resist releasing them to the public.

But over the past several months, the system of teacher rankings has been catapulted to one of the most contentious issues facing the city's 80,000-member teaching force. A new state law, passed this year to help New York win Race to the Top money, pledges that by 2013, 25 percent of a teacher's evaluation be based on a value-added system. The city has begun urging principals to consider rankings when deciding whether to grant tenure. And the city now supports the release of the data to the 12 media organizations, including The New York Times, that have requested it.

The departing schools chancellor, Joel I. Klein, defended the release of the rankings in an e-mail to school staff members, acknowledging that they had limitations but calling them "the fairest systemwide way we have to assess the real impact of teachers on student learning."

"For too long," Mr. Klein wrote, "parents have been left out of the equation, left to pray each year that the teacher greeting their children on the first day of school is truly great, but with no real knowledge of whether that is the case, and with no recourse if it's not."

But the United Federation of Teachers, the city's teachers' union, has sued to keep names in the rankings private, arguing that the data is flawed and would result in unnecessary harm to the reputation of teachers. The matter is now before Justice Cynthia Kern of State Supreme Court in Manhattan.

New York City began ranking teachers in the 2007-8 school year as part of a pilot project intended to improve classroom instruction. The project, which cost $1.3 million, with an additional $2.3 million budgeted over the next 18 months, was expanded in the 2008-9 school year to give rankings to more than 12,000 fourth- through eighth-grade teachers.

In New York City, a curve dictates that each year 50 percent of teachers will receive "average" rankings, 20 percent each "above average" and "below average," and 5 percent each "high" and "low." Teachers get separate rankings for math and English.

In support of the model, Douglas Staiger, an economics professor at Dartmouth College, cites research showing that if a teacher receives a high-performing score one year, there is a modest likelihood that he or she will receive a high-performing score the following year. The correlation is about 0.3, he said, with 1 being perfect, and 0 being no correlation. This means that about one-third of teachers ranked in the top 25 percent would appear among the top quarter of teachers the next year.

While that year-to-year link may seem low, in the budding and messy exercise of trying to quantify what makes students learn, it is one of the strongest predictors of future student performance, along with the reduction of class size. That means that, on average, students placed for a year with a high-value-added teacher will do better than those placed with a low-value-added teacher. Dr. Staiger placed the improvement at about three percentile points on a typical standardized test.

"This information is useful but has to be used with caution," he said. "It's that middle ground. It's not useless, but it's not perfect."

Yet a promising correlation for groups of teachers on the average may be of little help to the individual teacher, who faces, at least for the near future, a notable chance of being misjudged by the ranking system, particularly when it is based on only a few years of scores. One national study published in July by Mathematica Policy Research, conducted for the Department of Education, found that with one year of data, a teacher was likely to be misclassified 35 percent of the time. With three years of data, the error rate was 25 percent. With 10 years of data, the error rate dropped to 12 percent. The city has four years of data.

The most extensive independent study of New York's teacher rankings found similar variability. In math, about a quarter of the lowest-ranking teachers in 2007 ended up among the highest-ranking teachers in 2008. In English, most low performers in 2007 did not remain low performers the next year, said Sean P. Corcoran, the author of the study for the Annenberg Institute for School Reform, who is an assistant professor of educational economics at New York University.

The high margin of error for most scores, something the city refers to as the confidence interval, is another source of uncertainty, Dr. Corcoran said. In math, judging a teacher over three years, the average confidence interval was 34 points, meaning a city teacher who was ranked in the 63rd percentile actually had a score anywhere between the 46th and 80th percentiles, with the 63rd percentile as the most likely score. Even then, the ranking is only 95 percent certain. The result is that half of the city's ranked teachers were statistically indistinguishable.

"The issue is when you try to take this down to the level of the individual teacher, you get very little information," Dr. Corcoran said. The only rankings that people can put any stock in, he said, are those that are "consistently high or low," but even those are imperfect.

"So if you have a teacher consistently in the top 10 percent," he said, "the chances are she is doing something right, and a teacher in the bottom 10 percent needs some attention. Everything in between, you really know nothing."

In New York, the rankings face an additional set of issues. The state tests on which they were based became, over time, too predictable and easy to pass, and this summer the state began to toughen standards. Daniel Koretz, a Harvard professor whose research helped persuade the state to toughen standards, said that as a result it was impossible to know whether rising scores in a classroom were due to inappropriate test preparation or gains in real learning. Rankings that include the tougher standards will not be available until the next academic year.

"It would make sense to wait until the problems with the state test are sorted out, because we are going to get it wrong a lot of the time," Dr. Koretz said.

City officials defended using the state tests as a basis for the rankings, saying that they remained predictive of other outcomes, like graduation rates. Echoing Dr. Corcoran, the officials said they were most interested in identifying teachers at the extremes. "We have read the studies on it, and it is the best quantitative method that we have," said John White, a deputy chancellor. "When used in concert with other pieces of information, it can help us judge teacher effectiveness."

Beyond the formulas and tests, individual errors — like the one that led Ms. Ward to be left out altogether — have generated controversy. The teachers' union claims that it has found at least 200 such errors, including teachers' getting rankings for subjects they did not teach (sometimes they did well, sometimes poorly). Mr. White would not provide an estimate for the error rate, but noted that principals had 18 months to correct mistakes in class lists, starting from when the scores were first distributed.

Mr. White said on Tuesday that before the next round of rankings was released, teachers would be able to review class lists to verify which students they taught, a practice that generally did not happen in the past. Douglas N. Harris, an economist affiliated with the center at the University of Wisconsin that produces the city's rankings, called the science behind them promising, and said that they had jump-started a wider effort to come up with better measures of teacher performance, which was long overdue.

But Dr. Harris urged caution in reading too much into the early crop of rankings, and added, "As a general rule, you should be worried when the people who are producing something are the ones who are most worried about using it."

Sunday, September 05, 2010

Vaule Added: Grading the graders

Sep 3rd 2010, 13:46 by R.M. | WASHINGTON, DC

Description:
http://www.economist.com/sites/default/files/20100904_USP504_290.jpgA FEW
weeks ago I wrote a post about the use of value-added statistical analysis
to evaluate teacher effectiveness
<http://www.economist.com/blogs/democracyinamerica/2010/08/teachers_unions>
. Briefly, I think it's a step in the right direction, but teachers deserve
a more comprehensive evaluation system. Since then the Los Angeles Times,
whose reporting on the subject prompted my initial comment, has published a
database <http://projects.latimes.com/value-added/> with the ratings of
about 6,000 elementary-school teachers based on the paper's value-added
analysis. So, if you are a parent in the Los Angeles area, you can now find
out if your child's teacher is "least effective", "less effective",
"average", "more effective", or "most effective", based on the standardised
test scores of their students. Predictably, the teacher's unions threw a fit
<http://www.latimes.com/news/local/la-me-teacher-react-20100830,0,1507297.st
ory> , with the local union planning a protest in front of the Times
building (having already called for a boycott of the paper).

The unions have resisted using student test scores for evaluation purposes
for some time, though they are finally coming round to idea. It has never
been quite clear why teachers or their unions consider their occupation
uniquely difficult to evaluate. There are many jobs that lack simple metrics
with which to gauge effectiveness, and most make do with some combination of
evaluation procedures. The results are not perfect, but Conor Friedersdorf
captures
<http://andrewsullivan.theatlantic.com/the_daily_dish/2010/09/brief-thoughts
-on-teacher-pay.html> why it is so odd for teachers, especially, to resist
using unavoidably inexact methods, like value-added testing.

[E]very week they read student assignments and use their fallible judgment
to assign a letter grade, often based on opaque, somewhat arbitrary
standards. This process culminates in a report card sent home at the end of
every semester. It typically assesses achievement on an A to F scale that
presumably doesn't capture every nuance of student mastery over a subject.
High school teachers who give out these grades do so knowing that for many
students they'll one day be scrutinized by college admissions officers,
who'll admit or deny applicants largely based on the average of these
somewhat arbitrary grades that don't capture every nuance of a student's
academic abilities.

Despite its imperfections, I haven't many teachers eager to do away with
grades, and while I've seen a lot of teachers complain about being evaluated
based on test scores-a complaint with which I sympathize-I've never seen a
persuasive defense of "masters degrees earned" or "years worked" as a better
metric of quality. Yet teachers unions champion a status quo that relies on
these very measures.

For all their harrumphing, the Los Angeles teachers union and the local
school district have agreed to negotiate a new evaluation system. "Top
district officials have said they want at least 30% of a teacher's review to
be based on value-added," reports the Times. "But they have said the
majority of the evaluations should depend on observations." That's a start.
The Obama administration is pushing for greater transparency
<http://articles.latimes.com/2010/aug/25/local/la-me-ed-grants-20100825>
and more use of value-added analysis across the country. The next step is
agreeing on what to do with teachers who are deemed "least effective",
because heaven forbid we fire any of them.

http://www.economist.com/blogs/democracyinamerica/2010/09/evaluating_teacher


When Does Holding Teachers Accountable Go Too Far?
By DAVID LEONHARDT
The start of the school year brings another one of those nagging, often unquenchable worries of parenthood: How good will my child’s teachers be? Teachers tend to have word-of-mouth reputations, of course. But it is hard to know how well those reputations match up with a teacher’s actual abilities. Schools generally do not allow parents to see any part of a teacher’s past evaluations, for instance. And there is nothing resembling a rigorous, Consumer Reports-like analysis of schools, let alone of individual teachers. For the most part, parents just have to hope for the best.
That, however, may be starting to change. A few months ago, a team of reporters at The Los Angeles Times and an education economist set out to create precisely such a consumer guide to education in Los Angeles. The reporters requested and received seven years of students’ English and math elementary-school test scores from the school district. The economist then used a statistical technique called value-added analysis to see how much progress students had made, from one year to the next, under different third- through fifth-grade teachers. The variation was striking. Under some of the roughly 6,000 teachers, students made great strides year after year. Under others, often at the same school, students did not. The newspaper named a few teachers — both stars and laggards — and announced that it would release the approximate rankings for all teachers, along with their names.
The articles have caused an electric reaction. The president of the Los Angeles teachers union called for a boycott of the newspaper. But the union has also suggested it is willing to discuss whether such scores can become part of teachers’ official evaluations. Meanwhile, more than 1,700 teachers have privately reviewed their scores online, and hundreds have left comments that will accompany them.
It is not difficult to see how such attempts at measurement and accountability may be a part of the future of education. Presumably, other groups will try to repeat the exercise elsewhere. And several states, in their efforts to secure financing from the Obama administration’s Race to the Top program, have committed to using value-added analysis in teacher evaluation. The Washington, D.C., schools chancellor, Michelle Rhee, fired more than 100 teachers this summer based on evaluations from principals and other educators and, when available, value-added scores.
In many respects, this movement is overdue. Given the stakes, why should districts be allowed to pretend that nearly all their teachers are similarly successful? (The same question, by the way, applies to hospitals and doctors.) The argument for measurement is not just about firing the least effective sliver of teachers. It is also about helping decent and good teachers to become better. As Arne Duncan, the secretary of education, has pointed out, the Los Angeles school district has had the test-score data for years but didn’t use it to help teachers improve. When the Times reporters asked one teacher about his weak scores, he replied, “Obviously what I need to do is to look at what I’m doing and take some steps to make sure something changes.”
Yet for the all of the potential benefits of this new accountability, the full story is still not a simple one. You could tell as much by the ambivalent reaction to the Los Angeles imbroglio from education researchers and reform advocates. These are the people who have spent years urging schools to do better. Even so, many reformers were torn about the release of the data. Above all, they worried that although the data didn’t paint a complete picture, it would offer the promise of clear and open accountability — because teachers could be sorted and ranked — and would nonetheless become gospel.
Value-added data is not gospel. Among the limitations, scores can bounce around from year to year for any one teacher, notes Ross Wiener of the Aspen Institute, who is generally a fan of the value-added approach. So a single year of scores — which some states may use for evaluation — can be misleading. In addition, students are not randomly assigned to teachers; indeed, principals may deliberately assign slow learners to certain teachers, unfairly lowering their scores. As for the tests themselves, most do not even try to measure the social skills that are crucial to early learning.
The value-added data probably can identify the best and worst teachers, researchers say, but it may not be very reliable at distinguishing among teachers in the middle of the pack. Joel Klein, New York’s reformist superintendent, told me that he considered the Los Angeles data powerful stuff. He also said, “I wouldn’t try to make big distinctions between the 47th and 55th percentiles.” Yet what parent would not be tempted to?
One way to think about the Los Angeles case is as an understandable overreaction to an unacceptable status quo. For years, school administrators and union leaders have defeated almost any attempt at teacher measurement, partly by pointing to the limitations. Lately, though, the politics of education have changed. Parents know how much teachers matter and know that, just as with musicians or athletes or carpenters or money managers, some teachers are a lot better than others.
Test scores — that is, measuring students’ knowledge and skills — are surely part of the solution, even if the public ranking of teachers is not. Rob Manwaring of the research group Education Sector has suggested that districts release a breakdown of teachers’ value-added scores at every school, without tying the individual scores to teachers’ names. This would avoid humiliating teachers while still giving a principal an incentive to employ good ones. Improving standardized tests and making peer reports part of teacher evaluation, as many states are planning, would help, too.
But there is also another, less technocratic step that is part of building better schools: we will have to acknowledge that no system is perfect. If principals and teachers are allowed to grade themselves, as they long have been, our schools are guaranteed to betray many students. If schools instead try to measure the work of teachers, some will inevitably be misjudged. “On whose behalf do you want to make the mistake — the kids or the teachers?” asks Kati Haycock, president of the Education Trust. “We’ve always erred on behalf of the adults before.”
You may want to keep that in mind if you ever get a chance to look at a list of teachers and their value-added scores. Some teachers, no doubt, are being done a disservice. Then again, so were a whole lot of students.
David Leonhardt is an economics columnist for The Times and a staff writer for the magazine.

Monday, July 26, 2010

important new study about huge error rates in value-added teacher evaluation


New mathematica study for the US DOE’ Institute of Education Sciences, showing that if using value –added teacher test scores to evaluate teachers, the error rate is 35% (by using one year of test score data) and still as high as 25% for three years of test score data.:

“ This paper addresses likely error rates for measuring teacher and school performance in the upper elementary grades using value-added models applied to student test score gain data. Using realistic performance measurement system schemes based on hypothesis testing, we develop error rate formulas based on OLS and Empirical Bayes estimators. Simulation results suggest that value-added estimates are likely to be noisy using the amount of data that are typically used in practice.

Type I and II error rates for comparing a teacher’s performance to the average are likely to be about 25 percent with three years of data and 35 percent with one year of data. Corresponding error rates for overall false positive and negative errors are 10 and 20 percent, respectively. Lower error rates can be achieved if schools are the performance unit. The results suggest that policymakers must carefully consider likely system error rates when using value-added estimates to make high-stakes decisions regarding educators….

Using rigorous statistical methods and realistic performance measurement schemes, this report presents evidence that value-added estimates are likely to be quite noisy using the amount of data that are typically used in practice for estimation. ….

If only three years of data are used for estimation (the amount of data typically used in practice), Type I and II errors for teacher-level analyses will be about 26 percent each. This means that in a typical performance measurement system, 1 in 4 teachers who are truly average in performance will be erroneously identified for special treatment, and 1 in 4 teachers who differ from average performance by 3 to 4 months of student learning will be overlooked. Corresponding error rates will be lower if the focus is on overall false positive and negative error rates for the full population of affected teachers. With three years of data, these misclassification rates will be about 10 percent.

These results strongly support the notion that policymakers must carefully consider system error rates in designing and implementing teacher performance measurement systems based on value-added models, especially when using these estimates to make high-stakes decisions regarding teachers (such as tenure and firing decisions)….

Studies have found only moderate year-to-year correlations—ranging from 0.2 to 0.6—in the value-added estimates of individual teachers (McCaffrey et al. 2009; Goldhaber and Hansen 2008) or small to medium-sized school grade-level teams (Kane and Staiger 2002b). As a result, there are significant annual changes in teacher rankings based on value-added estimates. Studies from a wide set of districts and states have found that one-half to two-thirds of teachers in the top quintile or quartile of performance from a particular year drop below that category in the subsequent year (Ballou 2005; Aaronson et al. 2008; Koedel and Betts 2007; Goldhaber and Hansen 2008; McCaffrey et al. 2009).
While previous work has documented instability in value-added estimates post hoc using several years of available data, the specific ways in which performance measurement systems should be designed ex ante to account for instability of the estimates have not been examined. This paper is the first to systematically examine this precision issue from a design perspective focused on the following question: “What are likely error rates in classifying teachers and schools in the upper elementary grades into performance categories using student test score gain data that are likely to be available in practice?” These error rates are critical for assessing appropriate sample sizes for a performance measurement system that aims to reliably identify low- and high-performing teachers and schools.

Leonie Haimson
Executive Director
Class Size Matters

Thursday, June 17, 2010

Pondering Legal Implications of Value-Added Teacher Evaluation

June 2, 2010

Pondering Legal Implications of Value-Added Teacher Evaluation

http://schoolfinance101.wordpress.com/category/race-to-the-top/value-added-teacher-evaluation/

I’m going out on a limb here. I’m a finance guy. Not a lawyer. But, I do have a reasonable background on school law thanks to colleagues in the field like Mickey Imber at U. of Kansas and my frequent coauthor Preston Green at Penn State. That said, any screw ups in my legal analysis below are my own and not attributable to either Preston or Mickey. In any case, I’ve been wondering about the validity of the claim that some pundits seem to be making that these new teacher evaluation policies are going to make it easier and less expensive to dismiss teachers.

=====

A handful of states have now adopted legislation which mandates that teacher evaluation be linked to student test data. Specifically, legislation adopted in states like Colorado, Louisiana and Kentucky and legislation vetoed in Florida follow a template of requiring that teacher evaluation for pay increase, for retaining tenure and ultimately for dismissal must be based 50% or 51% on student “value-added” or “growth” test scores alone. That is, student test score data could make or break a salary increase decision, but could also make or break a teacher’s ability to retain tenure. Pundits backing these policies often highlight provisions for multi-year data tracking on teachers so that a teacher would not lose tenure status until he/she shows poor student growth for 2 or 3 years running. These provisions are supposed to eliminate the possibility that random error or a “bad crop of students” alone could determine a teacher’s future.

Pundits are taking the position that these new evaluation criteria will make it easier to dismiss teachers and will reduce the costs of dismissing a teacher that result from litigation. Oh, how foolish!

The way I see it, this new crop of state statutes and regulations which include arbitrary use of questionable data, applied in a questionably appropriate way will most likely lead to a flood of litigation like none that has ever been witnessed.

Why would that be? How can a teacher possibly sue the school district for being fired because he/she was a bad teacher? Simply writing into state statute or department regulations that one’s “property interest” to tenure and continued employment must be primarily tied to student test scores does not by any stretch of the legal imagination guarantee that dismissal based on student test scores will stand up to legal challenges – good and legitimate legal challenges.

There are (at least) two very likely legal challenges that will occur once we start to experience our first rounds of teacher dismissal based on student assessment data.

Due Process Challenges

Removing a teacher’s tenure status is denial of a teacher’s property interest and doing so requires “due process.” That’s not an insurmountable barrier, even under typical teacher contracts that don’t require dismissal based on student test scores. Simply declaring that “a teacher will be fired if he/she shows 2 straight years of bad student test scores (growth or value-added)” and then firing a teacher for as much does not mean that the teacher necessarily was provided due process. Under a policy requiring that 51% of the employment decision be based on student value added test scores, a teacher could be wrongly terminated due to:

a) Temporal instability of the value-added measures

http://www.urban.org/UploadedPDF/1001266_stabilityofvalue.pdf

Ooooh…Temporal instability… what’s that supposed to mean? What it means is that teacher value-added ratings, which are averages of individual student gains, tend not to be that stable over time. The same teacher is highly likely to get a totally different value added rating from one year to the next. The above link points to a policy brief which explains that the year to year correlation for a teacher’s value added rating is only about .2 or .3. Further, most of the change or difference in the teacher’s value added rating from one year to the next is unexplainable – not by differences in observed student characteristics, peer characteristics or school characteristics. 87.5% (elementary math) to 70% (8th grade math) noise! While some statistical corrections and multi-year measures might help, it’s hard to guarantee or even be reasonably sure that a teacher wouldn’t be dismissed simply as a function of unexplainable low performance for 2 or 3 years in a row. That is, simply due to noise, and not the more troublesome issue of how students are clustered across schools, districts and classrooms.

b) Non-random assignment of students

The only fair way to compare teachers’ ability to produce student value-added is to randomly assign all students, statewide to all teachers… and then of course, to have all students live in exactly comparable settings with exactly comparable support structures outside of school, etc., etc. etc. That’s right. We’d have to send all of our teachers and all of our students to a single boarding school location somewhere in the state and make sure, absolutely sure that we randomly assigned students, the same number of students to each and every teacher in the system.

Obviously, that’s not going to happen. Students are not randomly sorted and the fact that they are not has serious consequences for comparing teachers’ ability to produce student value-added. See: http://gsppi.berkeley.edu/faculty/jrothstein/published/rothstein_vam2.pdf

c) Student manipulation of test results

As she travels the nation on her book tour, Diane Ravitch raises another possibility for how a teacher might find him/herself out of a job by no real fault of actual bad teaching. As she puts it, this approach to teacher evaluation puts the teacher’s job directly in the students’ hands. And the students can, if they wish, choose to consciously abuse that responsibility. That is, the students could actually choose to bomb the state assessments to get a teacher fired, whether it’s a good teacher or a bad one. This would most certainly raise due process concerns.

d) A whole bunch of other uncontrollable stuff

A recent National Academies report noted:

“A student’s scores may be affected by many factors other than a teacher — his or her motivation, for example, or the amount of parental support — and value-added techniques have not yet found a good way to account for these other elements.”

http://www8.nationalacademies.org/onpinews/newsitem.aspx?RecordID=1278

This report generally urged caution regarding overemphasis of student value-added test scores in teacher evaluation – especially in high stakes decisions. Surely, if I was an expert witness testifying on behalf of a teacher who had been wrongly dismissed, I’d be pointing out that the National Academies said that using the student assessment data in this way is not a good idea.

Title VII of the Civil Rights Act Challenges

The non-random assignment of students leads to the second likely legal claim that will flood the courts as student testing based teacher dismissals begin – Claims of racially disparate teacher dismissal under Title VII of the Civil Rights Act of 1964. Given that students are not randomly assigned and that poor and minority – specifically black – students are densely clustered in certain schools and districts and that black teachers are much more likely to be working in schools with classrooms of low-income black students, it is highly likely that teacher dismissals will occur in a racially disparate pattern. Black teachers of low-income black students will be several times more likely to be dismissed on the basis of poor value-added test scores. This is especially true where a statewide fixed, rigid requirement is adopted and where a teacher must be de-tenured and/or dismissed if he/she shows value-added below some fixed value-added threshold on state assessments.

So, here’s how this one plays out. For every 1 white teacher dismissed on value-added basis, 10 or more black teachers are dismissed - relative to the overall proportions of black and white teachers. This gives the black teachers the argument that the policy has racially disparate effect. No, it doesn’t end there. A policy doesn’t violate Title VII merely because it has racially disparate effect. That just starts the ball rolling – gets the argument into court.

The state gets to defend itself – by claiming that producing value-added test scores is a legitimate part of a teacher’s job and then explaining how the use of those scores is, in fact neutral with respect to race. It just happens to have the disparate effect. Right? But, as the state would argue, that’s a good thing because it ensures that we can put better teachers in front of these poor minority kids, and get rid of the bad ones.

But, the problem is that the significant body of research on non-random assignment of students and its effect of value added scores indicates that it’s not necessarily differences in the actual effectiveness of black versus white teachers, but that the black teachers are concentrated in the poor black schools and that student clustering and not teacher effectiveness is leading to the disparate rates of teacher dismissal. So they weren’t fired because they were precisely measurably ineffective, they were fired because they had classrooms of poor minority students year after year? At the very least, it is statistically problematic to distill one effect from the other! As a result, it’s statistically problematic to argue that the teacher should be dismissed! There is at least equal likelihood that the teacher is wrongly dismissed as there is that the teacher is rightly dismissed. I suspect a court might be concerned by this.

Reduction in Force

Note that many of these same concerns apply to all of the recent rhetoric over teacher layoffs and the need to base those layoffs on effectiveness rather than seniority. It all sounds good, until you actually try to go into a school district of any size and identify the 100 “least effective” teachers given the current state of data for teacher evaluation. Simply writing into a reduction in force (RIF) policy a requirement of dismissal based on “effectiveness” does not instantly validate the “effectiveness” measures. And even the best “effectiveness” measures, as discussed above, remain really problematic, providing tenured teachers reduced on grounds of ineffectiveness multiple options for legal action.

Additional Concerns

These two legal arguments ignore the fact that school districts and states will have to establish two separate types of contracts for teachers to begin with, since even in the best of statistical cases, only about 1/5 of teachers (those directly responsible for teaching math or reading in grades three through eight) might possibly be evaluated via student test scores (see: http://schoolfinance101.wordpress.com/2009/12/04/pondering-the-usefulness-of-value-added-assessment-of-teachers/)

I’ve written previously about the technical concerns over value-added assessment of teachers and my concern that pundits are seemingly completely ignorant of the statistical issues. I’m also baffled that few others in the current policy discussion seem even remotely aware of just how few teachers might – in the best possible case – be evaluated via student test scores, and the need for separate contracts. But, I am perhaps most perplexed that no-one seems to be acknowledging the massive legal mess likely to ensue when (or if) these poorly conceived policies are put into action.

I’ll save for another day the discussion of just who will be waiting in line to fill those teaching vacancies created by rigid use of test scores for disproportionately dismissing teachers in poor urban schools. Will they, on average, be better or perhaps worse than those displaced before them? Just who will wait in this line to be unfairly judged?

For a related article on the use of certification exams for credentialing teachers, see:

Green, P.C., Sireci, S.G. (2005) Legal and Psychometric Criteria for Evaluating Teacher Certification Tests. Educational Measurement: Issues and Practice. Volume 19 Issue 1, Pages 22 – 31