Showing posts with label testing. Show all posts
Showing posts with label testing. Show all posts

Thursday, September 10, 2009

Experts: Test focus driving education wrong way, not preparing college-bound NYC students

NYC Parent Steve Koss writes:

In her article in yesterday's NY Daily News, Meredith Kolodner wrote that, "In 2002, a student could pass the math Regents exam by getting about 61% of the questions right. This year, that number dropped to 42%." Actually, and sadly, that number of 42% is correct but also doesn't tell the whole story. In comparing to June 2002, she used the Math A exam where a 61% raw score (52 out of 85 possible points) gave you a 65 scaled score and compared that to January 2009's final Math A exam where a 42% (just 35 out of 84 possible raw score points) "earned" (more like gifted) you a 65.

The rest of the story, so to speak, is that the Integrated Algebra exam that has replaced Math A has lowered the bar even further, something that hardly seemed possible without being called out for academic malpractice. Be that as it may, that's what has happened, so that in June and August of 2009, the aspiring 9th Grade student needs score only 30 out of 87 raw score points (just 34.5%) to be given (yes, given, not earned) a 65.

I am quoted in the Daily News story (see full text below), but here's part of what I wrote to Ms. Kolodner by email (I'm in China right now) when she was preparing her story (she used part of the first sentence in her story):

My own view of the Regents exams (at least those in my field - Math) is that they have been so simplified and the raw (cut) score for passing has been so reduced that the exams are now virtually meaningless as measures of mathematical knowledge or preparedness. I would argue that a student who is granted a converted score of 65 on Integrated Algebra (having managed to earn 30 raw score points out of 87, or just 34.5% of the points available) is in no way, shape, or form prepared for the next level of high school mathematics and is, in fact, very likely to experience failure in his or her math course work and Regents exams throughout the rest of his/her high school career. NY State is now graduating students with an assertion that they have achieved some base level of mathematical competency (passing the Integrated Algebra Regents is their only math diploma requirement) based on an exam in which sheer, blind guessing on 30 multiple choice questions will, on average, get you halfway (15 points) to a passing score. A student who knows the correct answers to just 10 multiple choice questions can blind-guess the remaining 20 and, on average, will pass. This leaves out any consideration of the remainder of the exam (extended answer questions) as well as educated guessing (where the most obviously wrong answer(s) can be eliminated and the correct answer guessed from the reduced set of options, increasing the odds of guessing correctly).

This is hardly the type of foundation you would want to build for students' future math work in high school or for any sort of college preparation. As a result, we end up with a benign, almost paternalistic form of academic fraud in which students are told that they have "passed math" and have an acceptable level of competency for whatever lies before them. Far worse, their parents, most of whom are not aware of how the Regents exams are structured, graded, o r scaled, are consequently led to believe that their children have demonstrated that level of proficiency on a NYS-administered exam, so it must be legitimate. How many parents would react with shock or dismay if they wre told the truth, that Johnny or Janie's Integrated Math Regents exam score of 65% this last June was in reality a 34% (30 out of 87), or that his or her 70% was actually a 40% (35 out of 87)? Do anyone believe those parents would be relaxed about their chidren's education, believing they were being well-prepared for college or adulthood? What scaled scoring has done has been to rob parents of their understanding of their children's true educational achievement levels, and you see next to nothing done by NYS, NYC, or individual schools to educate their parent communities otherwise. You might call it a case of the emperor's new clothes, but in this instance, those clothes (lack thereof) are being worn by our own children who are actually parading forward unclothed.


Steve Koss


DAILY NEW STORY BELOW:

Experts: Test focus driving education wrong way, not preparing college-bound NYC students


Read more: http://www.nydailynews.com/ny_local/education/2009/09/09/2009-09-09_reading_writing__worry_experts_test_focus_is_driving_ed_wrong_way.html#ixzz0QhofjRSi

Some of the state-mandated Regents tests have been dumbed down in the past eight years, experts say, and many students' SAT scores leave them unprepared for college.
"Unless you were in an AP [Advanced Placement] course or an honors class, they didn't prepare you for college," said Rianna Moustapha, 18, who graduated from Leon M. Goldstein High School in Sheepshead Bay, Brooklyn, last year.
Moustapha, a sophomore at Brooklyn College, said she was taught only to pass the Regents.
Teachers complain the tests have become less comprehensive and rigorous.
"We could be doing a lot better," said Saul Cohen, a former Queens College president, who heads the state Regents committee charged with looking at state standards.
"The complaints we get from higher ed people over and over [are that] most youngsters are not well-prepared for college - unless, of course, they've taken APs or international baccalaureates."
In 2002, a student could pass the math Regents exam by getting about 61% of the questions right. This year, that number dropped to 42%.
"The exams are now virtually meaningless as measures of mathematical knowledge or preparedness," said Steve Koss, a former New York City math teacher who is completing a study of state math exams.
"The Regents have committed to doing whatever it takes to meet the President's standard for college readiness," said state Education Department spokesman Tom Dunn.
"This will include a thorough review of the learning standards, the core curricula and the state assessments."
The national average on each section of the SAT tests has hovered at about 500 - of a perfect 800 points - for several years.
Last year, only 10% of city schools averaged above 500 in math and 7% did so in reading and writing.
At the same time, more than half of city schools' average SAT scores were below 400. A score of 200 is awarded just for showing up.
The city Education Department points to increases among those scoring higher than 600 on the SATs and the drop in the percentage of students entering the City University of New York system who must take remedial classes - 51% last year, down from 59% in 2002 - as evidence of achievement.
"The data suggests that more kids are graduating," spokesman Andrew Jacob said. "But a higher percentage are ready for college-level work."
Nonetheless, of the 56% of city students who graduate from high school, many report trouble at college.
"It's like we had to memorize, but not learn," said Jordan Woodward, 19, a Brooklyn College sophomore who graduated last year from Bedford Academy with an advanced Regents diploma and an A-minus average.
"It kind of hit me in October of my first semester when I was getting my exams back, and the grades weren't very good.
"In high school, I didn't really have to study," said Woodward, who grew up in Bedford-Stuyvesant, Brooklyn. "I've gotten a lot of help from my academic adviser, and I'm doing better. It's a work in progress."

Read more: http://www.nydailynews.com/ny_local/education/2009/09/09/2009-09-09_reading_writing__worry_experts_test_focus_is_driving_ed_wrong_way.html#ixzz0QhosUdS9



__._,_.___

Monday, June 23, 2008

Mayor Sees a Test Scores Triumph

Or is it a case of inflation of results?

By ELIZABETH GREEN, Staff Reporter of the Sun
June 23, 2008
http://www.nysun.com/new-york/mayor-sees-a-test-scores-triumph/80476/

Mayor Bloomberg will announce an education victory today: Test scores are up across the city, by double digits at some schools. But a cloud is already gathering, as education experts are raising the possibility that these gains and others across the country could suggest score inflation and not real learning gains.

The scores being released today show a nine-percentage-point gain in math citywide versus last year, and a seven-point gain on the reading test. The gains are even more remarkable when viewed over the six-year timeline since Mr. Bloomberg took office: Three-quarters of city students now score proficient at math, up from 37% in 2002.

This year's gains were larger than the increases statewide, though smaller than in other cities, such as Buffalo, Rochester, and Yonkers.

Mr. Bloomberg is scheduled to announce the results at a press conference at P.S. 175 in Harlem this afternoon.

The mayor has often greeted test-score increases as evidence that he is fulfilling his promise to improve public schools. "I'm happy, thrilled, ecstatic," he said last year, announcing gains on the state math test.

Since then, concerns have grown that rises in state test scores in New York and elsewhere do not reflect real improvement, but rather "inflations" — either due to easier tests, deliberate cheating, or more subtle "gaming" that helps students perform better without actually having to learn more material.

Local education experts last September called for an audit of the state test after a respected national test showed the city posting significant gains in only one academic area, despite reports from the state test showing larger gains across the board.

A report grading state tests recently delivered New York State a C+. And a study by the teachers union found that a reading test dropped in difficulty by as many as six grade levels between 2004 and 2005.

The concern in New York follows a pattern around the country.

"Measuring Up: What Educational Testing Really Tells Us," a new book by a testing expert and Harvard Graduate School of Education professor, Daniel Koretz, calls score inflation the "dirty secret" of high-stakes testing.

Although some test-score gains represent the real, hard work of teachers and students, others "are entirely illusory," Mr. Koretz writes.

A recent study of Texas schools is the latest in a string of academic observations on the effects of high-stakes testing. The study concludes that educators are "gaming" their state tests by preventing low-achieving students from taking them and teaching only material they expect to appear on the test, rather than the wider span of material tests are supposed to represent.

A study by a pscyhometrician and professor at the University of Iowa, Andrew Ho, found that two-thirds of state tests are publishing higher gains than a national test.

New York is not immune from the phenomenon, according to another researcher studying state tests, Bruce Fuller, a professor of education and public policy at the University of California at Berkeley.

"We've got great rhetoric and great pressure on teachers to teach to the test, but at the end of the day when Albany reports out the share of kids that are proficient, we can't really trust their claims, especially when put up against the federal definition of proficiency," Mr. Fuller said.

The State Education Department defended its tests.

"All of New York's tests are checked many times to be sure that a score this year means the same next year," a spokesman, Tom Dunn, said in a statement. "The only way for a student to improve performance is by learning the curriculum — reading, writing, and math."

At the city Department of Education, where scores are used to help determine school closures, teacher and principal salaries, and promotion decisions from one grade to another, a senior official who oversees testing, Jennifer Bell-Ellwanger, said she has confidence that state tests are reliable.

Ms. Bell-Ellwanger said the department takes into account the potential errors of testing by looking at trends rather than individual data points.

Shown the results, the education historian Diane Ravitch noted that scores were up statewide, and that some cities had even larger gains than New York City.

"What this suggests to me is that the state lowered the bar and it is easier to pass the exams," Ms. Ravitch said.

She said the lesson of the results should be that the state "needs an independent agency to conduct the state tests and report on results."

The results will have an effect on schools. Teachers who signed up for a pilot project on merit-based performance bonuses could become eligible to receive them. Schools teetering on the edge of a failing report card grade could be pushed to another side.

Students will also be affected.

Yesterday, the principal at I.S. 349 in Bushwick section of Brooklyn, Roy Parris, said that his school's high test scores — he said they shot up about 10 points in both math and reading — mean that only one seventh-grader of 155 total is eligible to be held back this year, down from about 15 seventh-graders last year.

Sunday, January 27, 2008

Accountability Tests' Instructional Insensitivity: The Time Bomb Ticketh

COMMENTARY
By W. James Popham

Education Week [American Education's Newspaper of Record], Wednesday, November 14, 2007, Volume 27, Issue 12, pp. 30-31. http://www.edweek.org/ew/articles/2007/11/14/12popham.h27.html?print=1

Would you ever want your temperature to be taken with a thermometer that was unaffected by heat? Of course not; that would be dumb. Or would you ever want to weigh yourself with bathroom scales that weren't influenced by the weight of the person using them? Of course not again; that would be equally dumb. But today's educators are allowing their instructional success to be judged by students' scores on accountability tests that are essentially incapable of distinguishing between effective and ineffective instruction. Talk about dumb.

What's worse is that we are now racing toward the 2014 deadline of the federal No Child Left Behind Act, the point at which all students are supposed to have attained test-based "proficiency." But the 2002-2014 schedules that most states devised when establishing their goals for annual required numbers of proficient students will soon demand some staggering increases in how many students must earn proficient scores on state NCLB tests each year. These balloon-payment improvement schedules were, in most instances, adopted as a way of deferring the pain stemming from having too many state schools and districts flop in reaching their goals for adequate yearly progress, or AYP.

Such cunningly crafted, soft-to-start improvement schedules will lead in a very few years to altogether unrealistic requirements for improved test scores. Without such improvements, huge numbers of U.S. schools and districts will be seen as AYP failures. If the American public is skeptical now about the quality of public schools, how do you think citizens will react when, in the next several years, test-based AYP failure becomes the rule rather than the exception? Can you hear the ticking of this nontrivial time bomb?

How could American educators let themselves get into a situation in which the tests being used to evaluate their instruction are unable to distinguish between effective and ineffective teaching? The answer, though simple, is nonetheless disquieting. Most American educators simply don't know that their state's NCLB tests are instructionally insensitive. Educators, and the public in general, assume that because such tests are "achievement tests," they accurately measure how much students have learned in schools. That's just not true.

Two types of accountability tests are currently being used to satisfy the No Child Left Behind law's assessment requirements. About half of the nation's NCLB tests consist of traditional, off-the-shelf, standardized achievement tests, usually supplemented by a sprinkling of new items, so that the slightly expanded tests will supposedly be better aligned with a particular state's content standards. Other NCLB tests are made-from-scratch, customized standards-based accountability tests, built specifically for a given state. Let's see, briefly, why both these types of tests are instructionally insensitive.

Traditional standardized achievement tests, such as the Stanford Achievement Test-10th Edition, are intended to provide comparative information about test-takers. So the performance of a student who scores at, for instance, the 96th percentile can be contrasted with that of students who score at lower percentiles. To accomplish this comparative-measurement mission, these tests must produce a substantial degree of "score spread," so there are ample numbers of high scores, middle scores, and low scores. Most items on such tests are of middle-difficulty levels because such items, statistically, maximize score spread.

Over the years, however, many of these middle-difficulty items turn out to be closely linked to students' socioeconomic status. More-affluent kids tend to answer these socioeconomically linked items correctly, while less-affluent kids tend to miss them. This occurs because socioeconomic status, or SES, is a nicely distributed variable, and one that doesn't change rapidly; SES-linked items help generate the score spread required by traditional standardized achievement tests. When such tests are used as accountability assessments, however, they tend to measure the socioeconomic composition of a school's student body, rather than the effectiveness with which those students have been taught. The more SES-linked items there are on a traditional standardized achievement test, the more instructionally insensitive that test is bound to be.
-----------------------
SIDEBAR: How could American educators let themselves get into a situation in which the tests being used to evaluate their instruction are unable to distinguish between effective and ineffective teaching?
-----------------------
The other type of NCLB accountability test used in the United States is usually described as a "standards-based test," because such tests are deliberately built to assess students' mastery of a given state's content standards, that is, its curricular aims. In all but a few states, though, the number of content standards to be assessed is so large that there is no way to accurately assess-via an annual accountability test-students' mastery of this immense array of skills and knowledge. Instead, each year's accountability test must sample from the profusion of the state's curricular aims. Such a sampling-based approach to annual assessment means that teachers end up guessing about which curricular aims will be assessed each year. And, given the huge numbers of potentially assessable curricular targets, most teachers guess wrong.

After a few years of incorrect guessing, many teachers simply give up on trying to mesh their teaching with what's to be assessed on each year's accountability tests. And when this happens, it turns out that the major determinant of how well a school's students perform on accountability tests is the very same factor that governed students' performances on traditional standardized achievement tests: socioeconomic status. Thus, even on customized standards-based tests, a school's scores are influenced less by what students are taught than by what the students brought to that school. Most standards-based accountability tests are every bit as instructionally insensitive as traditional standardized achievement tests.

The instructional insensitivity of accountability tests does not represent an insuperable problem, however. Remember when, several decades ago, we began to recognize that there was considerable test bias in our high-stakes educational assessments? Once this difficulty had been identified, it was attacked with both empirical and judgmental bias-detection procedures. As a consequence, today's educational tests are markedly less biased than were their predecessors. Once the test-bias problem had been identified, we set out to fix it-and in less than a decade, we did.

That's precisely what we need to do now. Using a mildly technical definition, a test's instructional sensitivity represents the degree to which students' performances on that test accurately reflect the quality of instruction specifically provided to promote students' mastery of what is being assessed. We need to discover how to build accountability tests that will be instructionally sensitive and, therefore, can provide valid inferences about effective and ineffective instruction. It may take several years to get the required procedures in place, but we need to get started right now.

In the short term, though, we must make citizens, and especially educational policy makers, understand that almost all of today's accountability tests yield an invalid picture of how well students are being taught. Accountability systems based on the use of such instructionally insensitive tests are flat-out senseless. We need accountability tests capable of distinguishing between students who have been properly taught and those who have not. Until such tests are at hand, we might as well re-label our accountability systems as what they are-elaborate and costly socioeconomic-status identifiers.
------------------------------
W. James Popham is a professor emeritus in the graduate school of education and information studies of the University of California, Los Angeles. He now lives in Wilsonville, Ore.

Wednesday, November 21, 2007

Tweedle Dum vs Tweedle Dee. A plague on both their data pools!

by Sean Ahern

Critics of the Bloomklein 'reform' use the NAEP results to refute the 'success' of Mayoral control but I think this sort of 'critique' leaves us chasing our tails and leads to no positive alternative model of accountability.

Those of us in the schools would do well to change the channel and create new bottom up systems and ethics of accountability. I think that would be "important".

This conversation is going on in many parts of the country but I fear the incessant chatter and noise from the pundits, union leadership and professional advocates that proliferate here in NYC and Washington DC in particular may be a distraction.

Diane Ravitch and the AFT/UFT are part of the testocracy regime. They harp on the discrepancy between federal and state scores but to what end? To advocate for additional Federal controls over curriculum, certification and accountability through testing? To secure their own special place as gatekeepers in a second Clinton administration? Would the tests be more valid if they designed and administered them?

There is a shibboleth much esteemed by the Shankerites, that public schools create the national identity and should be uniform, ignoring the fact that our "nation" is an internal empire under a white male supremacist oligarchy. Education here is hierarchical, segregated by race and class, unavoidably a battleground and rightfully so. Uniformity and standards is the language of social control, vastly overrated as beneficial to all by the managerial classes.

Top down accountability systems, be they Fed or SED, Tweedle Dee or Tweedle Dum, need to be replaced by a bottom up version of accountability where managers serve the people not vice versa.

Sean Ahern

Wednesday, September 26, 2007

Calls Intensify For State Test Audit

By ELIZABETH GREEN
Staff Reporter of the Sun
September 26, 2007

http://www.nysun.com/article/63411

Differences between state tests and a national benchmark test whose results were released yesterday are ramping up calls for an independent audit of the
New York State test.
Results from the test known as the nation's report card, the National Assessment of Educational Progress, showed the scores of New York State students remaining essentially flat in most categories but one, fourth-grade math, where students showed improvements. State tests, meanwhile, have shown declining fourth-grade scores — but improvements in eighth-grade scores.
A
Manhattan Institute scholar, Sol Stern, called the discrepancies a "spanking" for the New York State Education Department. "Its claims of fabulous improvements in eighth-grade reading and math scores for 2007 have proven to be just more hype," he said.
A spokesman for the state education department,
Tom Dunn, said concerns are flat-out wrong. The NAEP test is not given to all students but to a sample, about 2% of students in New York per grade, and each student takes only a portion of the test, he said. The U.S. Department of Education, which administers the test, advises that results be treated as "estimates," he added.
The president of the teachers union,
Randi Weingarten, called the discrepancies a "cloud" on returns showing achievement overall went up nationwide.
"What these tests suggest is New York State has a very serious problem with its testing program," an education historian,
Diane Ravitch, said, calling for an audit of the program.
The state education commissioner, Richard Mills, praised gains made by minority students. Since 1998, the report showed, the percentage of black and Hispanic students who are proficient in reading and math increased significantly, with increases hovering around 20 points.

Friday, September 07, 2007

State Guts Its Test of Reading: Union Study Sees Inflated Scores

BY ELIZABETH GREEN - Staff Reporter of the Sun

September 7, 2007
URL: http://www.nysun.com/article/62089

The difficulty of a reading test used to judge students across New York State dropped by as many as six grade levels between 2004 and 2005, according to an internal study by the New York City teachers union obtained by The New York Sun.


The study, written in March 2006, found that passages in the 2005 test hovered around third- and fourth-grade reading levels, down from a ninth-grade level in 2004. It also found that the 2004 test was characterized by longer passages, smaller print, crammed text, and more complex questions, such as asking a student to make an inference versus asking the main idea. Despite this apparent drop in difficulty, however, the number of correct answers needed to pass — known as the "cut score" — was just slightly higher in 2005 than in 2004.


Schools across the state had reported fourth-grade reading gains in some cases of more than 40 points in 2005, and New York City pupils registered a 10-point gain on average.

In New York City, low scores on state tests can prevent students from advancing to the next grade level or lead to a school being shut down.


Coming during Mayor Bloomberg's re-election campaign, the touted surge raised many eyebrows, including those of the then-chairwoman of the City Council's Education Committee, Eva Moskowitz of Manhattan, who held a six-hour hearing on the test scores, and the city's public advocate, Betsy Gotbaum, who handed copies of the 2004 and 2005 tests to researchers at the United Federation of Teachers for examination.


Shown the UFT study yesterday afternoon, a state education department spokesman, Tom Dunn, called its conclusion "completely wrong."


"The passages combined with the questions are at the same grade level for both years," he said.

The UFT study was done in consultation with at least three researchers, according to a March 1, 2006, memo to the union president, Randi Weingarten.


In an interview yesterday, Ms. Weingarten said she chose not to publicize the study out of concerns that doing so would make her appear "anti-test." She also said the study could not be considered comprehensive because her researchers are not psychometricians and lack access to some specific data about the test.


She said the study did concern her.


"It's part of why I keep saying, be careful about data. Standardized test scores can't be used for these high-stakes measures for kids or for teachers," she said.


The director of New York University's Center for Research on Teaching and Learning, Robert Tobias, said the UFT study added empirical confirmation to the concerns he had raised in 2005. But he said the study's grade-level measure, called the Fry Readability test, which uses two indicators to determine "readability" — the average number of syllables per 100 words and the average number of sentences — was too simplistic. A test tabulating more nuanced characteristics — "vocabulary, the complexity of the sentence structure, the number of subordinate clauses, that type of thing" — would be more conclusive, he said.


Mr. Tobias praised the part of the UFT study that describes more subtle differences between the tests, such as differences in the length of passages and question types.

"To me, that's some pretty powerful stuff," he said.


A freelance data analyst who has worked for the Department of Education, Frederick Smith, said a readability study he has conducted using the Fry method — but testing every word on the tests rather than a sample — is more precise than the UFT study. His results, which he showed to The New York Sun, found a drop of five reading levels, to third-grade in 2005 from eighth-grade in 2004.



Mr. Smith said the results confirm his longstanding calls for an independent review board to oversee state testing. Since as long ago as a 1982 Newsday article, Mr. Smith has called for a testing ombudsman, citing the importance of an independent reviewer to temper politicians' inclination to boast of gains under their watch.


Michael Cohen, a former education adviser to President Clinton who now works with 30 states to improve their annual assessments through his nonprofit group, Achieve, said he believes that some states have "dumbed down" tests in response to pressures to get better academic results — with deleterious results.


"The tests, generally speaking, are not all that rigorous to begin with. So almost any dumbing-down is moving the expectations in the wrong direction," he said.


Mr. Cohen said that he would be very surprised, however, if New York was one of those states. "I do know Rick Mills," he said, referring to the state education commissioner. "This hardly sounds like what he would do."


A vice president of a Washington, D.C.-based think tank, the Thomas Fordham Foundation, Michael Petrilli, said the federal No Child Left Behind law has caused several states to make their tests easier — some publicly and some without announcement, a conclusion that will be issued in a Fordham report due later this month, Mr. Petrilli said.


Responding to a Daily News story questioning the reliability of state math tests published on the first day of the new school year Tuesday, Mr. Bloomberg pointed out that New York City can use state tests to compare its schools to schools in other districts across the state. Between 1999 and 2007, the portion of New York City students who passed the state reading test has risen by more than 15 percentage points, to 51% from 35%, while cities such as Buffalo have had single-digit increases.