kitchen table math, the sequel: death by data
Showing posts with label death by data. Show all posts
Showing posts with label death by data. Show all posts

Monday, March 24, 2014

Death by data, Part I've-lost-count

A compelling vision of the data-driven future of K-12 schooling? Or a chilling description of a brave new educational world in which even students' smallest actions are converted to digital data and used to build permanent "learner profiles"?

A new report from London- and New York City-based educational publishing powerhouse Pearson is likely to generate both reactions, depending on whom you ask.

Released this week, "Impacts of the Digital Ocean on Education" is intended as "an aspirational vision of what success might look like" in the rapidly changing world of "big" educational data and personalized learning.

Report authors Kristen DiCerbo and John Behrens of Pearson's Center for Digital Data, Analytics, and Adaptive Learning sketch out a vision in which end-of-year, summative tests of narrowly defined skills and content knowledge are replaced by a constant stream of digital data generated by "in vivo naturalistic tasks" that thoroughly blur the line between assessment and instruction.

Such data—generated from a variety of sources and activities, and focusing on students' social connections and interactions, rather than just their isolated individual experiences—would be constantly tracked and used to update profiles that follow each student across classrooms, grades, and schools, helping facilitate more customized learning experiences for each.

"The devices and digital environments with which we interact are designed to record and store experiences, thereby creating a slowly rising ocean of digital data," DiCerbo and Behrens write in the report. "We believe the ability to capture data from everyday formal and informal learning activity should fundamentally change how we think about education."

Just as big data and analytics have transformed finance, insurance, retail, and professional sports, the report says, they will change education. Until recently, the authors write, data collection and storage was expensive, limited, and isolated, and students' educational records were not portable, easy to share, or able to be quickly analyzed.

But the digital revolution has changed that reality, DiCerbo and Behrens contend. They argue that the abundance of increasingly fine-grained data available to educators "can help pinpoint the moments when learning occurs or a learner's approach to a problem changes" and then be used to help tailor suggestions or recommendations to help each student's learning continue.

'Ocean' of Digital Data to Reshape Education, Pearson Report Predicts
By Benjamin Herold on March 19, 2014 11:24 AM
A tidal wave of data could personalize learning

Thursday, July 25, 2013

Big data strikes again

The New York Times is reliably fun to read on the subject of technology:
SAN FRANCISCO — Although certain kinds of engineers are in short supply in the United States, plenty of potential candidates exist for thousands of positions for which companies want to import guest workers, according to an analysis of three million résumés of job seekers in the United States.

[snip]

[T]he technology industry argues there are not enough qualified Americans [to fill tech positions]. Its critics, including labor groups, say bringing in guest workers is a way to depress wages in the industry.

Many economists take issue with the industry’s argument, too. One side points out that wages have not gone up across the board for engineers, suggesting that there is no stark labor shortage.

[snip]

“I didn’t expect this result,” said Steve Goodman, Bright’s chief executive.

[snip]

“We’re Silicon Valley people, we just assumed the shortage was true,” Mr. Goodman said. “It turns out there is a little Silicon Valley groupthink going on about this, though it’s not comfortable to say that.”

[snip]

The Senate immigration bill, passed last month, nearly doubles the number of H-1B visas that companies can seek every year. Industry lobbied heavily for it, bulldozing efforts to add language that would force companies to try to hire an equally qualified American first.

[snip]

The age of workers, which the study did not look at, may also play a role....[A]mong 32 technology companies surveyed, only six had a work force with a median age over 35. At Monster, the job search portal, the median age was 30; at Google, 29; and at Facebook, 28. The median age of American workers over all is 42.3 years old, according to the Bureau of Labor Statistics.

As if to underline the study’s findings, Mr. Goodman spoke from a conference room that looked out on decorated ping-pong tables, a liquor bar and tiki-themed snacks. Later that day, Bright was having a party, partly to attract new talent, he said, including foreign programmers here on H-1B visas.

Big Data Analysis Adds to Guest Worker Woes
By QUENTIN HARDY and SOMINI SENGUPTA
JULY 23, 2013, 10:23 AM
Groupthink in Silicon Valley, groupthink amongst the punditry, groupthink in the White House....

[pause]

This part was interesting:
For a few job categories, like computer systems analysts, there are relatively few “good fits” among American applicants, Bright found. Computer systems analyst jobs, considered relatively low-skilled in the tech world, had four openings for every American candidate. For others, like high-skilled computer programmers, there were more than enough potential candidates in the United States, the company found.
As I recall, there's a section in the Steve Jobs book where Jobs explains to President Obama that the workers they really can't find are skilled workers with Associate degrees.

I'll have to find that and post.

For the record, I had absolutely no idea there wasn't a shortage of engineers until Kitchen Table Math readers explained the world to me. I never questioned the narrative; I just took all the Silicon lamentations at face value.

Sunday, March 13, 2011

Mark Roulo on baseball statistics

re: the 32-parameter value-added model in NYC
One thing I like doing for discussions like this is to try and find a sports analogy. People tend no to get so hung up on non-PC conclusion in sports, but also often care a lot. This can lead to enlightenment.

So ... baseball:

(1) You can get a *VERY* good handle on how valuable a batter is with just two values, which can be combined into one number. You need on-base percentage (OBP) , which is, for every 100 times he comes to the plate, how often does he get on base? And you need slugging percentage (SPG) , which says how many bases he gets each time he has an at-bat. In both cases, more is better. And you can combine them with this: (OBP*3 + SPG)/2 to get a number that works the way most people who follow baseball can understand.

There ARE more sophisticated models, but they don't improve on this one by much. So ... two parameters, both of which are pretty easily understood.

(2) For pitchers it is a bit more complicated, but you can basically track strikeouts, walks and home runs and then put them together to get a single number. Again, one can improve (for starting pitchers, you also care about how "efficient" they are), but basically you'll get the right answer for ranking pitchers with just these three.

I get that teaching is more complicated. But 32 parameters is nuts.

Monday, November 15, 2010

21st century skills

One of the goals of No Child Left Behind is to increase the availability of data. Part of the implicit model underlying No Child Left Behind is that with improved information, parents will recognize good and bad schools. Principals will identify good and bad teachers. District administrators will identify weak and strong principals, and state  administrators will recognize struggling school districts. Armed with this information, parents will choose with their feet, and the other actors will undertake the necessary reforms to improve education.

As an empirical economist I am, of course, sympathetic to the use of data, and as a school board member I pushed for more thorough evaluation of our programs. But the gap between the rhetoric and the ability to use education data effectively is large.

Few school districts have the resources to analyze statistical data in even remotely sophisticated ways. In the early days of the Massachusetts Comprehensive Assessment System (MCAS) tests, I visited the Assistant Superintendent for Curriculum and Instruction who was anxious to use the testing data to help Brookline address its achievement gap. The state Department of Education had provided each district with a CD with the complete results of each student’s MCAS test. In principle, it would be possible to pinpoint the exact questions on which the gap was greatest. The problem was that no one in the central administrative offices could figure out how to read the CD. I loaded the CD onto my laptop and quickly ascertained that the file could be read with Excel. Shortly thereafter, our Assistant Superintendent attended a meeting of her counterparts from the western (generally affluent) suburbs of Boston and discovered that Brookline was the only system that had succeeded in reading the CD. Districts have become somewhat more savvy about using data. A younger generation of administrators has more experience with computers, but relatively few would be able to link student report cards generated by the school district with SAT scores and the state tests.

Principals, district administrators, and even state-level administrators generally begin their careers as teachers, and relatively few teachers have strong backgrounds in statistical reasoning. In my experience, the people who rise to senior administrative positions in public education are smart. They understand in a general sense that estimates come with standard errors attached, but faced with a report that last year 43 percent and this year 56 percent of black students in fourth grade were profifi cient in math, few could tell you whether with 75 students each year, the change was statistically signifificant.

When I stepped down from the school board, one of my colleagues joked that they could all go back to treating correlation as causality. In education policy settings, one repeatedly hears statements like: “Students who take Algebra II in eighth grade meet the profifi ciency standard in grade ten. We must require all students to take Algebra II in eighth grade.” “Students taking math curriculum A and curriculum B get similar math SAT scores. The curricula are equally good.” “Students who are retained in grade continue to fall further behind. Retention is a bad policy.”2

School administrators may understand at some level that they are only looking at  correlations, but almost none have the training to address the issue of causality, and faced with a correlation, they will often interpret it causally in the absence of evidence to the contrary. The capacity to address causality, weaknesses of various measures, and other strengths and weaknesses of statistics is very limited. The Public Schools of Brookline recently recruited for a Director of Data Management and Evaluation. Although school board members generally are not (and should not be) involved in personnel decisions other than those involving the Superintendent, in this specific case the Superintendent asked me to participate in the candidate interviews. Many of the candidates held or had held similar positions in other districts. I asked each candidate how we could decide whether a math curriculum used by some, but not all, of our students was effective. Many of the candidates did not think of this question in statistical terms at all. Only one addressed the issue of selection—and we hired him.

Measurement Matters: Perspectives on Education Policy from an Economist and School Board Member (pdf file)
by Kevin Lang

Apparently they don't cover Excel in ed school.

Friday, August 20, 2010

sorry, Johnny

Sorry, Johnny. You have to stay in the crumbling, failing, dangerous public school in your neighborhood because your freedom of choice is not statistically justified.

commenter David reacts to freaknomics post on futility of school choice


Very droll.

I keep thinking the politics of choice will be affected by the economy. Charter schools (and vouchers) are cheaper than public schools, and parents are happier (or at least spend fewer years of their lives being unhappy - pdf file).

Same academic outcomes for less money, with less stress on the family: put it that way, some of us are going to take that deal.

Tuesday, February 3, 2009

voodoo correlations in social neuroscience

I've always been skeptical of big behavioral claims based on brain scan data.

Turns out I was right.

Here's Andrew Gelman (and here, too).

Brain imaging studies under fire (naturenews)

interview: Have the Results of Some Brain Scanning Experiments Been Overstated? (Scientific American)
LEHRER: What is a "voodoo correlation"?

VUL: We use that term as a humorous way to describe mysteriously high correlations produced by complicated statistical methods (which usually were never clearly described in the scientific papers we examined)—and which turn out unfortunately to yield some very misleading results. The specific issue we focus on, which is responsible for a great many mysterious correlations, is something we call “non-independent” testing and measurement of correlations. Basically, this involves inadvertently cherry-picking data and it results in inflated estimates of correlations.

To go into a bit more detail:

An fMRI scan produces lots of data: a 3-D picture of the head, which is divided into many little regions, called voxels. In a high-resolution fMRI scan, there will be hundreds of thousands of these voxels in the 3-D picture.

When researchers want to determine which parts of the brain are correlated with a certain aspect of behavior, they must somehow choose a subset of these thousands of voxels. One tempting strategy is to choose voxels that show a high correlation with this behavior. So far this strategy is fine.

The problem arises when researchers then go on to provide their readers with a quantative measure of the correlation magnitude measured just within the voxels they have pre-selected for having a high correlation. This two-step procedure is circular: it chooses voxels that have a high correlation, and then estimates a high average correlation. This practice inflates the correlation measurement because it selects those voxels that have benefited from chance, as well as any real underlying correlation, pushing up the numbers.
One can see closely analogous phenomena in many areas of life. Suppose we pick out the investment analysts whose stock picks for April 2005 did best for that month. These people will probably tend to have talent going for them, but they will also have had unusual luck (and some finance experts, such as Nassim Taleb, actually say the luck will probably be the bigger element). But even assuming they are more talented than average—as we suspect they would be—if we ask them to predict again, for some later month, we will invariably find that as a group, they cannot duplicate the performance they showed in April. The reason is that next time, luck will help some of them and hurt some of them—whereas in April, they all had luck on their side or they wouldn’t have gotten into the top group. So their average performance in April is an overestimate of their true ability—the performance they can be expected to duplicate on the average month.

It is exactly the same with fMRI data and voxels. If researchers select only highly correlated voxels, they select voxels that "got lucky," as well as having some underlying correlation. So if you take the correlations you used to pick out the voxels as a measure of the true correlation for these voxels, you will get a very misleading overestimate.

This, then, is what we think is at the root of the voodoo correlations: the analysis inadvertently capitalized on chance, resulting in inflated measurements of correlation. The tricky part, which I can’t go into here, was that investigators were actually trying to take account of the fact they were checking so many different brain areas—but their precautions made the problem that I am describing worse, not better!
Of course, now I'm wondering whether there is anything I think I know about brain & behavior that is not based on non-independent analysis.

Sunday, June 1, 2008

data loops and other strange beasts

from Education Week:

Whether schools will know how to make use of data collected through value-added statistical techniques is an open question, however.

Daniel F. McCaffrey, a senior statistician in the Pittsburgh office of the Santa Monica, Calif.-based RAND Corp., studied 32 Pennsylvania school districts taking part in the first wave of a state pilot program aimed at providing districts with value-added student-achievement data in mathematics.

He and his research colleagues surveyed principals, other administrators, teachers, and parents in the districts involved in the program and compared their responses with those from other districts having similar demographic characteristics.

“We found it was really having no effect relative to the comparison districts,” Mr. McCaffrey said.

Even though educators, for instance, seemed to like the data they were getting and viewed the information as useful, few were doing anything with the results, he said. Twenty percent of the principals didn’t know they were participating in the study, Mr. McCaffrey said, noting also that the program was still young at that point in the evaluation process.

Despite such challenges, other speakers at the conference argued that the use of value-added methodology should become more widespread. Said Robert Gordon, a senior fellow at the Center for American Progress, a Washington think tank: “The way we will learn about implementation problems, I think, is to implement.”

New Uses Explored for ‘Value Added’ Data
by Debra Viadero
May 28, 2008


Oh yes, I agree. There is much to be learned from implementations of all types. One after another. Implementation upon implementation upon implementation.

Implementation is the path to enlightenment.


data-driven loops & noise

Tuesday, April 22, 2008

data-driven loops & noise

more from anonymous on data-driven instruction:

There are two kinds of data loops in play. The first, I'll call loop 1. The second I'll call loop 2.

Loop 1 is the "We've measured your school's performance and found it lacking" loop. This is aggregated data that has been massaged to produce your AYP (measure of Adequate Yearly Progress, Mass. MCAS). From this measure, schools are to produce a School Improvement Plan (SIP) which is basically goal setting sans concomitant resources to actually effect a change. Since the SIP is a dead end, that particular loop looks more like a croquet wicket. It's not a loop at all.

Loop 2 is the "Your last year's students failed MCAS" loop which is given to teachers about 5 months after the students in question have left your embrace. Usually this is aggregated also, at least in my school it was handed to us on a printout. I had granularity down to the question but not down to the student.

Somewhere in this measurement system is precise, standard by standard knowledge of each student's current (actually 5 month old) ability. We don't get that.

The real problem is that loop 1 should be used for "oh my, I think we need a remediation here" and it is used for "oh my, time to reset the goal posts." If loop 1 is not used to drive some kind of structural change, then anything you learn (and you can't learn much) in loop 2 becomes a 'nice to know' kind of thing but it doesn't fix Johnny's inability to add.

Data not acted upon is noise and old data is rancid.

Call me crazy, but I don't see the problem here.

This school should just tell all its struggling students to Seek Extra Help.

Then, when the struggling students don't Seek Extra Help, or do Seek Extra Help but Extra Help doesn't Help, they should tell the parents, "If your child doesn't come in for Extra Help, there's nothing I can do."*

That's what my school does.

It works, too.


* direct quote

more fun with numbers
data-driven instruction redux
data-driven loops & noise

data-driven instruction redux

a comment on data-driven instruction left by Anonymous:

Don't even get me started on this one! I'm in a district that places heavy emphasis on being "data driven". In spite of this emphasis, teachers don't have access to much of the data. They are either lacking hardware, privileges, or timeliness to make it accessible and relevant. This is all before we get to the training issues.

Worse, let's say your data tells you that Johnny is in the sixth grade and can't add, there is no system in place to do anything about it. Sure, you can try to get him to stay late for help or you can differentiate in class (at the expense of what he's supposed to be current with). But, there is no way (especially with 50% of your kids in this condition) to get Johnny remediated.

As long as curriculum fills every inch of available space, teachers aren't going to use objective data to replace subjective data when neither can be used as a force for change.


Interesting.

Tell us more if you get a chance.

And thanks!


more fun with numbers
data-driven instruction redux
data-driven loops & noise

Monday, April 21, 2008

more fun with numbers

In our non-degree professional development programs at Harvard University, I have taken to routinely asking the assembled administrators and teachers how many of them have taken a basic course on educational measurement. In an audience of 50 to 100 participants, the usual count is two or three. These people are usually “ringers”—they are typically assistant superintendents for measurement and evaluation. That is, they run the testing operation in their school systems. Now, imagine what the state of health care would be if practicing physicians didn’t know how to read EKGs, EEGs or chest x-rays, didn’t know how to interpret a basic blood analyses, or didn’t know anything about the test-retest reliability of these simple diagnostic measures. Imagine what it would be like if your basic family practitioner in a health maintenance organization didn’t know how to interpret a piece of current medical research questioning the validity of the standard test for colo-rectal cancer. Imagine what it would be like to be a practitioner in a health care organization in which every piece of evidence required for patient care came from a standard test of morbidity and mortality administered once a year in the organization. The organization you are imagining is a school system.

Leadership as the Practice of Improvement (pdf fie)
Richard F. Elmore

This probably accounts for the existence of books with titles like: Getting Excited about Data Second Edition: Combining People, Passion, and Proof to Maximize Student Achievement by Edie L. Holcomb.

This text is actually written in language that educators can understand, even if they aren't especially data-savy. The guidelines for incorporating data for school improvement are actually practical and easy to replicate. A great guide for joining the data-driven school improvement movement!
5 stars

In theory, my district is using data to inform instruction. We have a data warehouse.

Thus far, however, the data is apparently showing that all learning problems in the gen-ed population can be attributed to student failure to Seek Extra Help. Either that, or Weak Inferential Thinking.

Which is pretty much what all learning problems were attributed to before we had a data warehouse as far as I know.





more fun with numbers
data-driven instruction redux
data-driven loops & noise

Monday, February 4, 2008

Steve H on schools and assumptions

. . . . This relates to a main theme I have been pushing for years; education based on individuals rather than statistics. I can understand that the government wants to improve averages (statistics), but this comes at the expense of individual educational opportunities. As long as education is based on statistics, the affluent will provide the needed opportunities and the poor will get the baseline.... It's nice that psychologists will add some real science to the debate, but it's not the solution.

Years ago, I had a discussion with a member of our school committee who really liked the idea of IEP's for all students. I thought it odd at the time because she was a major proponent of mixed-ability, child-centered learning. I guess she thought schools could have it both ways. Differentiated Learning sounds nice, but most schools use it as cover for their fundamental belief in mixed-ability learning. As I mentioned long ago, our school started calling it Differentiated Learning instead of Differentiated Instruction because the teachers don't instruct and they want the kids to take responsibility for their learning.

How can psychologists set a baseline for instruction when schools do not believe in instruction? Will it be a baseline for instruction just to meet NCLB? [ed.: the answer is yes]

What's missing from this discussion are parents and their opinions of what constitutes a good education for their individual children. I emphasize the word opinions because this is not the domain of ed school graduates or psychologists, who now seem to be playing the statistical baseline game.

I would rather see psychologists stick with the individual. I'd rather see psychologists define what is an expected learning level for individual children. But what I would really like to see is schools assuming that all kids can get into Harvard (no matter what a psychologist says), not that all kids can get over the minimal NCLB requirements, or what they call "all kids can learn."

I may have to tattoo this onto my forehead.

My own district, which includes parents who attended Harvard themselves, scorns the very idea that a parent might wish his child to follow in his footsteps.

"Don't push your child."

"Every child has his place."

"Parents need to let go."

"Let your child self-advocate."

etc.

Richard Elmore posts coming right up.

The rich really are different from you and me. Rich school districts, too.

Thursday, October 11, 2007

case studies versus "data"

This explanation of the case study approach, as opposed to the "data" approach, is going to be extremely useful here in Irvington:

Schools just don't do (and appreciate) what parents do at home. Most parents make sure learning happens. Schools don't do that. That's why I believe a lot can be learned by looking at individual kids and what their parents do, not at statistics. Schools want more parental involvement because they see a correlation with student results (statistics). They just don't realize what that connection really involves (study individual cases). This involves things that the school should be doing. Nobody can argue against parental involvement, but what are the details - practicing math facts at home? In other words, doing their job?

I'm going to send this to our administrators and board. People are beginning to grapple with this problem here, I think, and this is perfect.

Previously the administration has had two ways of looking at tutoring:

  • Westchester parents hire tutors when they don't need them.
  • Tutoring is really just a form of providing your child with very reduced class size.

We had a terrific meeting with the assistant superintendent and the assistant principal of the middle school today. It seems clear that the administration is beginning to understand the enormity of the parent reteaching that goes on here. It's not just "tutoring"; it's not just "help with homework." Many students here are experiencing almost a separate school at home, at least in the subjects in which they need a separate school at home.

Today, in the meeting, we gave two examples of what is taking place:


Last week Ed spent 2 hours walking a 9th grader we know through a h.s. writing assignment no 9th grader could possibly do. No graduate student would have been able to do it, either. Essentially, the assignment asked students to write a book, or perhaps a series of books, in two paragraphs. Not possible.

Fortunately, this student just so happened to have a professional historian as a familiy friend. So he and his mom came over, and Ed figured out a way for the student and his Ph.D.-bearing mother to do the assignment. The final product wouldn't be good -- it couldn't be -- but it would be as good as it could be under the circumstances.

What happens to students who don't happen to have a historian friend who can "help with homework"?

The other story we told is heartbreaking.

One of Christopher's sweetest friends -- this is such a great kid -- has some family problems (divorce), parents aren't professionals, etc.

Last year Chris' ELA class was given an assignment that was way over the kids' heads.

(I'm going to add my standard disclaimer here: we think the world of this teacher -- I might even be able to document the amount C. learned in her class. Here, I'm talking about a particular problem with writing instruction that we're seeing in many, many classes.)

Anyway, the writing assignment was far too advanced.

Ed spent hours breaking it down, teaching each part, selecting an appropriate text for C. to write about, and so on.

C. ended up with an A.

His friend, who was on his own, got a D or perhaps even an F. He was sad and demoralized; Chris was proud and happy. Ed said, afterwards, "It's like emotional blackmail. If you don't help your kid he's going to be miserable."

What conclusion does C's friend draw?

He told us this summer. One of C's other friends was telling the others that his mom was going to get him a game system if he made honor roll. C's average-student friend said, in his sweet, direct voice, without a trace of envy, "That would be difficult for me."

Then he looked at Ed and me and said, "I'm an average student."

It breaks your heart.

We brought these stories to the meeting, and for the first time, I think, the stories were heard.

Steve's comment about the need for case studies will explain part of what is needed here.

But how does one level the playing field?

I have some thoughts about that, courtesy of Susan J.

Does anyone else?

Sunday, September 30, 2007

"statistics hides problems"

Keeper Comment from Ken's post on ambiguity in teaching:

I've said in the past that many teachers think that the problem of education is defined by what walks into their classroom. As kids get older, it's very easy to blame the kids and external causes.

"As Engelmann suggests, let's save the excuse making until we clean up our instructional act."


Unfortunately, this is done by trying to bring more kids up to very low cut-off levels. I'll call this the guess-and-check approach to educational improvement.

This reminds me of students trying to fix a computer program that has many different internal errors. The errors interact and produce all sorts of odd results. There are no clear cause and effect relationships. Invariably, students try to fix the program by changing something and looking at the results. The program might be fixed in one area, but the fix might cause a problem in another location. This happens because they don't take the time to really understand what is going on in the code line by line. Things change, but they don't really know why.

A lot of research abhors individual anecdotes. It's almost a dirty word. However, by performing a detailed analysis of individual cases, one really understands what's going on. You see a direct connection between cause and effect. Fixing this problem won't fix all of your problems, but it is a necessary step in the process.

Statistics hides problems. It's a process of reducing large amounts of data (many errors) into a more manageable amount. In doing so, information (problems) can be lost or confused. You might think you know what's going on, but you don't. If you want to understand why some kids are successful and some are not, you have to analyze a lot of individual cases. You aren't looking for one error and one solution. You're looking for many.

Why do so many educators try to find the "one thing", like better teacher preparation, that will solve the problem?

Guess and check.

..............................

Statistics hides problems.

Ed calls this death by data.

Monday, September 10, 2007

the fourth goal

The fourth goal is data.

Data analysis, data warehousing, data-driven decisionmaking.

................................

I am in favor of data.

However, it's obvious that data is going to be misconstrued and misused, intentionally or not. Especially in the world of education, where there are few scientific standards for research, even fewer qualms about flinging around the words "research shows," we're going to need independent audits of data and citizen's oversight committees.

I don't trust Michael Bloomberg on NYC school data for a minute.


Matthew K: youth wants to know!
2005 math scores in NY state - test was easier than others

2005 math scores in NY state

re: Brett's post about the easier 2005 test

As Brett mentioned, the reading test wasn't the only problem. The math test was easier, too.

The News obtained technical details on high-stakes math tests given to fourth-graders across the state over the past six years and found that in every year when scores went up, testmakers had identified the questions as easier during pretest trials.

In years when scores were lower, pretest trials showed the questions were harder.

"That's pretty strong evidence that something is just not right with the test," said New York University Prof. Robert Tobias, who ran the Board of Education's testing department for 13 years.

"If this were a single year's data or two years' data, I would say it would be inappropriate to make conclusions," Tobias said. "But with the pattern over time ...that's prima facie evidence that something's not right."

In 2005, for example, when a record-breaking 85% of New York State's fourth-graders passed the test, the questions had the highest average easy score in years. The easy score was .73 - meaning the average question was answered correctly by 73% of the kids who participated in pretest trials.

In contrast, when 68% of kids passed the state test in 2002, the easy score was .61.


here's more:

Before any high-stakes test is given to kids in New York, testmakers subject every possible question to an experimental trial called a field test.

[snip]

Every question gets an easy score - called a P-value - that comes from the percent of field-testers who correctly answered a question.

If 61% of kids get a question right - as field-testers did for the average question on the 2002 fourth-grade exam - the question has a P-value of .61.

Kids in New York get the same number of points for correct answers regardless of whether a question is rated easy or difficult. One way testmakers equalize exams is by requiring more correct answers on easier tests.

If the 2005 test was easier than the 2002 test, that wasn't done. Kids needed 40 points to pass the 2002 test but only 39 points to pass in 2005.

"Wow!" said NYU testing expert Robert Tobias. "This is really good evidence that the test was easier, substantially easier."

State officials deny that the 2005 test was easier. Testmaker CTB/McGraw-Hill used the highly regarded Item Response Theory to ensure equivalent exams.

Item Response Theory considers the difficulty of a question, its ability to separate smart kids from struggling ones and the odds that a kid can guess the answer correctly.

Test experts who reviewed technical reports from state exams for the Daily News said McGraw-Hill used state-of-the-art equating methods, but they said they couldn't know if the equating was done properly without a thorough audit.

Some experts were troubled by the fact that test scores have gone up and down over the past six years in the same pattern as the easiness ratings.

"It is worrisome that the average P-values for those items on the field test do tend to track the overall passing rate," said Columbia University testing expert James Corter.


This report is significant in terms of Irvington's public use of data. Last year the administration presented data on our ELA scores to the board. In that meeting we learned that the 8th grade class had scored poorly on the ELA exam. Where 43% of the class had earned 4s in 4th grade, only 16.7% of them had earned 4s in 8th grade.

The administration offered 3 explanations:

  • a couple of ELA teachers took sudden leaves, so many 8th graders were taught by substitutes

  • 18 new students moved into the district, 14 of whom were "receiving services" (mostly 504C or "building support"); these low-scoring students depressed the scores of the rest of the class (total class size approximately 150)

  • you really can't compare one year's kids to any other year's kids anyway because "the scaling might be different"

And there it was left.

It is a mathematical impossibility for 18 new students entering a class of 150 to cause a decline from from 43% earning 4s to 16.7%.

That said, the percentages are beside the point.

What matters is that, in 4th grade, 68 students in this class scored 4s on ELA; in 8th grade, only 25 students scored 4s. You can jigger the figures in a couple of ways (fewer kids were tested in 8th grade than in 4th, for instance), but any way you slice it, a bunch of 4s turned into 3s. The district offered excuses, then told us that the 8th grade test was "unnecessarily difficult" and the scores would "bounce back."

What the administration meant was that the scores on the 8th grade ELA exam would bounce back on the high school Regents, which has low cut scores. The tests aren't equivalent.*

Moreover, the 8th grade ELA test, which is in fact more difficult than any of the other state ELA tests, is the only test that comes within calling distance of matching NAEP. (Look at the column for New York.)

This use of data by district administrators is unsound.


a district email to the community

Not long after that, things heated up in the district for a number of reasons, one of them being the fact that I had begun to write up and distribute to the community data parents had not seen. This post in particular, on the recentering of SAT scores in 1995, was forwarded widely around the community I'm told.

A couple of other things occurred, including a public threat by the teacher's union to sue me (or to look into suing me - that's the threat that gets made around here)......

One thing led to another and presently the superintendent issued an email filled with all good data, including the news that:

This year’s 5th grade, the first to use Trailblazers in both 3rd and 4th grades, scored 96% of students at proficient or mastery level. Sixty-one percent of students scored at mastery level.

There was no mention made of the "bad" data discussed at the board meeting. As far as I'm aware, the administration has never mentioned the 8th grade ELA scores outside of an untelevised school board meeting attended by only a handful of parents.

Now we learn that the 2005 math test, the test that yielded Irvington's 96% 3s and 4s, was easier than the test taken by older students using the old curriculum (SRA Math).

It strikes me as unlikely that the administration will mention this development to the wider community. Nor do I know whether the administration has plans to discuss these reports with the board.


auditing the data

I believe that school districts need, at a minimum, routine independent audits of data, data analysis, and the use of data to make decisions. This is simply good practice. When I was on the board of NAAR, we were required to undergo an independent audit each and every year of operation. Businesses must do the same.

For a school, standardized test scores are money; scores directly support real estate value.

I would also like to see citizen's oversight committees set up to give data and the district's use of data a second look, and to and offer guidance on the way in which data is used by schools. A number of statisticians and researchers live in Irvington; we need these people looking at our data and offering the administration their expertise.


2005 math scores in NY state - test far easier
All Quizzes Not Created Equal (Daily News)
4th and 5th graders subjected to comparison study
2002 NY state math test
2005 NY state math test
answer keys
if you can't improve the results, make the test easier (2005 reading scores in NY)


* No links--sorry. I'm not going to spend hours of my life running down data on the NY edu-web site, which has become even more impossible to navigate than it was last year. Scores don't bounce; this is a core psychometric truth. The burden is on my district to prove that high school Regents is equivalent to the 8th grade ELA in difficulty, not on me to prove that it's not. They're the ones asserting an anomaly as a reality.

If you can't improve results, make the test easier

This is just galling.

The difficulty of a reading test used to judge students across New York State dropped by as many as six grade levels between 2004 and 2005, according to an internal study by the New York City teachers union obtained by The New York Sun.

The study, written in March 2006, found that passages in the 2005 test hovered around third- and fourth-grade reading levels, down from a ninth-grade level in 2004. It also found that the 2004 test was characterized by longer passages, smaller print, crammed text, and more complex questions, such as asking a student to make an inference versus asking the main idea. Despite this apparent drop in difficulty, however, the number of correct answers needed to pass -- known as the "cut score" -- was just slightly higher in 2005 than in 2004.

Many states have reduced the difficulty of their assessments in order to post better results without actually improving - see here for more on that. But the boldness of what was done in New York - whether it's reducing reading difficulty by six grade levels or five - is simply unbelievable.

Update: It's not just reading, either.

Cross-posted at The DeHavilland Blog.


2005 math scores in NY state - test far easier
if you can't improve the results, make the test easier (2005 reading scores in NY)

Sunday, March 25, 2007

help desk - statistics





I desperately need a course in statistics.

My question concerns this passage from a terrific article: The structure of human intelligence: It is verbal perceptual and image rotation (VPR), not fluid and crystallized by Wendy Johnson & Thomas J. Bouchard Jr.* (pdf file)

Interestingly, though the correlations between the verbal and perceptual and perceptual and image rotation factors were high (0.80 and 0.85), the correlation between the verbal and image rotation factors was much lower, 0.41.

This study sets out to determine the "relative statistical performance of three major psychoetric models of huam intelligence," those being:


The fluid-crystallized model, which has been dominant for some time now, didn't work.

Good.

I'm glad.

I'm glad because according to the fluid-crystallized model people get dumber as they age and their "fluid" intelligence gets less fluid. Or something.

This is why we're always hearing that for us old folk "experience and wisdom" have to make up for "ability to solve novel problems" or what-have-you.

Turns out that "experience and wisdom" and the "ability to solve novel problems" are the same thing.

Or so I gather. (If anyone who actually researches intelligence stumbles across this entry, I yearn to be fact-checked on this. Please. Chime in.)

At any rate: Bouchard's new study is good news for geezers because the fluid-crystallized model did not work out.

Nor did the "three-strata model." (Don't know what the three-strata model is; not going to find out any time soon.)

What did work is the verbal-perceptual model, which is pretty much the common-sense understanding of human intelligence most of us non-experts have always believed in.

I still don't really understand the distinction between verbal intelligence and perceptual intelligence. Generally speaking, however, it breaks down this way:

  • "verbal: verbal fluency and divergent thinking [ed.: what is divergent thinking?] as well as verbal scholastic knowledge and numerical abilities"
  • "perceptual speed, and psychomotor and physical abilities such as proprioception in addition to spatial and mechanical abilities"

Clear as mud!

What's interesting about Johnson's and Bouchard's study is that they discovered that one needs to add a third category, which is the ability to mentally rotate, manipulate, and twist two-and three-dimensional objects.

As we all know, this is a guy thing:

Interestingly, it is the image rotation abilities that have repeatedly shown the most robust sex differences among cognitive abilities (favoring males, Voyer, Voyer, & Bryden, 1995).

I'll get back to that.

In another post.

Here's my question.

I'm not understanding how verbal intelligence can correlate highly with perceptual intelligence, and perceptual intelligence can correlate highly with image rotation intelligence, but image rotation intelligence does not correlate highly with verbal intelligence.

How does that work?

Am I reading the passage incorrectly?

Is this an expression of a sex difference?

Or what?

Thanks in advance!



* Bouchard directs the Minnesota Twin Study.

good schools raise IQ, bad schools lower IQ, part 1
good schools raise IQ, bad schools lower IQ, part 2
good schools raise IQ, bad schools lower IQ, part 3
Seth Roberts on IQ

fuzzy math makes you smarter
IQ quiz
school raises IQ
intelligence is verbal, perceptual, and image rotation
math isn't English


Sunday, March 18, 2007

why we need statistics

from rightwingprof, one of the most succinct & clear statements I've seen:

Before I go on, let me quickly address why we must analyze the data statistically, and cannot just report means. If we gave the same kids the same proficiency exams on two different days, say only a week apart, their scores would be different. Anytime we see a difference between scores, without statistics, we do not know if those differences are due to random variation or not. We cannot without statistics point to two different scores or means and say, "See? The scores increased!"

Also, let me mention a few crucial points.

  • The more data we have, the more reliable our statistical analysis will be (this will become an issue later on).
  • Means (averages) alone do not give us a complete picture, particularly when they are means of aggregated data, as these are (this is why I look at other descriptive statistics).
  • Statistics always deals with probability (uncertainty), and we calculate our statistics to a specific probability, 95% here (sometimes statistics are calculated to a 99% probability). This is the level of sensitivity (alpha), here, 0.05.
  • We are assuming here either that the proficiency exam standards did not change between the two years or that the proficiency reports for the two years are comparable (if they are not, then Wisconsin cannot make any statement about their proficiency levels over time).

Saturday, March 10, 2007

this article is the schnizzell

D-Ed Reckoning: Schemo gets pwned

D-Ed, performs a smackdown on a NYT article about Madison, Wi schools.

To complicated to explain, except it involves deception, statistics, nation standards, more deception, lying, and deception.

Curious aren't you?

Wednesday, March 7, 2007

First Grade Probability


My first grader brought home this worksheet today that she completed in class.

"If we spin the arrow 10 times, I think:

green red yellow

will come up the most."

My daughter circled red. I asked why she picked red. "I like red."

As far as she was concerned that was the end of the matter. Since she still had no idea that she was wrong, or that liking a color wasn't the best basis for making a prediction, I have to assume that they didn't spend a lot of time being instructed on probability. But they did spend a good portion of time spinning the spinner and noting their results.

So I have two issues: 1 -- a lot of time spent, without instruction, resulting in kids reaching incorrect conclusions; 2 -- she will be doing this exact same problem for the next five years. My fifth grader is still spinning spinners, they just have more colors.

I believe Catherine has noted before the arrogance with which the schools waste our children's time. They have spent almost no time on the fundamental building blocks that lead to success in math. In first grade, almost 3/4 of the way through the year, they are adding numbers up through 3, and not regularly.