05 February 2011
Rubric and Grading Issues
19 June 2009
Grading and Reporting at CHS
This got a little longer than I anticipated. The introduction explains some of my thinking and the research section on page two includes quotes and summaries from leading researchers on assessment and grading.
Introduction
- This weekend I tried to think back to why I began doing 1-5 grading in the first place. It actually had nothing to do with percent grades. I knew from the outset that the conversion would cause problems. My purpose was to clearly communicate to students exactly where they were on a learning continuum.
- I was convinced by Bob Marzano that 5-9 categories were about the most anyone could reliably use to judge students.
- Rick Stiggins and Anne Davies convinced me that assessment should be FOR learning. That it should not communicate and end point but that it should let a student know where they are and what they can do to go to the next level.
- Tom Guskey convinced me that I really couldn’t reliably sort students into the 101 categories available on the 0-100 scale. (And I certainly couldn’t sort them into the 1001 categories available on the 0.0-100 scale.)
- Grant Wiggins made me reconsider averages. Why, he asked, would we give a student the average when they could do it at the end of a course?
- Dylan Wiliam gave me the example of the driver’s license—no matter how many times you fail—we all get the same driver’s license once we finally pass.
So putting this all together I began to use a system where I scored each part of each assignment using the 1-5 scale. I stressed early and often that even if you had a 1 you were still a good person—we are all at the beginning at some point in our lives. I stopped giving zeros and gave incompletes. I would say to students that I couldn’t give them a score until I had that piece of work. I stopped talking about percent grades at all—the only thing I would talk about with students was 1, 2, 3, 4 or 5.
It seemed to work. Conversations with students transitioned from being about their grade to being about what they can to improve their understanding of a specific topic. The only time it didn’t work was for any of the 8 reporting periods when I was limited to giving students only a 2 digit grade to summarize their performance. Which is where I was 4 years ago and where I am today…
So…
Here is an attempt to compile the thoughts of leading researchers on grading. First of all I think it is important to remember that all experts recommend a systematic approach to rethinking grading. They are all pretty similar so I will use the steps that Tom Guskey and Jane Bailey advise. (The first 5 would go in order. Six through 9 give information about steps 1-5.)
- Purpose—what is the purpose of grades?
- Define the impetus for change—why are we re-examining our grading and reporting system?
- Exploring the history of grading and reporting—what is the history of grading?
- Laying a foundation for change—what is the research that surrounds grading?
- Building a Grading and Reporting System
- Grading and Reporting Methods I: Letter grades, percentage grades, and other categorical grading
- Grading and Reporting Methods II: Standards-based, pass/fail, mastery grading, and narratives
- Grading and reporting for students with special needs
- Special problems in grading and reporting
Marzano recommends the following order:
Phase 1—have a vanguard team of teachers experiment with competency based assessment and record keeping.
Phase 2—identify the competencies and the software that will be used.
Phase 3—implement the system in stages.
Research
Anne Davies
Dr. Davies does not mention grading scales in her works or presentations. She does strongly and continuously stress that a change in grading should be a change from “assessment of learning,” to “assessment FOR learning.” Her take is that assessment is changing from something that happens to students to something that is done to help students.
“Research shows that when students are involved in the learning process—learning to articulate what they have learned and what they still need to work on—achievement improves.”[1]
Ken O’Connor
“Impreciseness is the main point of those who argue for letter grades rather than percentage grades; they believe that dividing student achievement into a limited number of categories is all that we can ever hope to do with any pretense of real meaning. According to this argument, using a 101 point scale gives a false sense of precision and, therefore, detracts from the main purpose of grades—meaningful communication of student achievement.
This argument has a great deal of merit, especially for elementary and middle schools, where grades are not involved in high stakes decisions, except pass/fail. However, where grades are involved in high stakes decisions about students’ educational future—such as college entrance, graduate school acceptance, and employment opportunities, numbers may be preferable to letters because there are more scale points available.”[2]
O’Connor then illustrates an example of a student who gets all 89s and gets a B and one who gets all 90s gets an A. The 5 point scale of A-F would amplify differences between students that really aren’t that great. BUT O’Connor is not against a 5 point system. He goes on to say,
“If all the guidelines and principles described in chapters 1-8 [Nearly the entire book] are applied then letter grades based on teachers’ professional judgments using a detailed descriptive scale will produce the best grades. But if teachers crunch numbers to arrive at grades, especially in high school and college, then percentage grades are probably fairer, and therefore better, than letter grades.”[3]
Again, the second sentence here might seem to be a rejection of a 5 point system. But, go back and read the first sentence again. What O’Connor advocates throughout the book, and when he presents, is for teachers to use grades, professional judgment, conversations with students, and multiple measures to determine grades. He is the man who told me about the following simple way of explaining the 5 point scale.
5=Wow!
4=Great!
3=Got it!
2=Nearly there!
1=Oops!
O’Connor would look at the 2 students described above and using professional judgment determine whether all 89s should merit the designation of A.
Guskey
“Letter grades offer a brief description of students’ achievement and level of performance, along with some idea of the adequacy of that performance (Payne, 1974). Because most parents experienced letter grades during their school years, they also have a general sense of what letter grades mean. For this reason, parents often prefer letter grades to newer, less traditional reporting methods (Libit, 1999).
Despite their simplicity, however, letter grades also have their shortcomings. First and probably most important, their use requires the combination of lots of different forms of evidence into a single symbol (Stiggins, 2001). As described in Chapter Three, many teachers combine product, process, and progress evidence in a single grade. This makes the grade a confusing hodgepodge that’s impossible to interpret, rather than a meaningful summary of students’ achievement and performance (Brookhart, 1991; Cross and Frary, 1996).
Second, despite educators’ best efforts; many parents interpret letter grades in strictly norm-referenced terms. Probably because the letter grades they received as students reflected their standing in comparison to classmates, parents frequently assume the same is true for their children. To them, a C doesn’t represent achievement at the third level of a five point scale, similar to a middle level belt in a karate class. Instead, a C means “average” or “in the middle of the class.”
A third shortcoming of letter grades is that the cutoffs between grade categories are always arbitrary and difficult to justify. If the teacher decides that the scored for a grade of B will range from 80 to 89, for example, the student with a score of 80 will receive the same grade as the student with a score of 89, even though there is a nine-point difference in their scores. But the student with a score of 79—a one point difference—receives a grade of C. Why? Because the teacher set the cutoff for a B grade at 80. Although cutoffs are absolutely necessary in any multilevel grading method, where they are set is always arbitrary.
Finally, letter grades lack the richness of other, more detailed reporting methods, such as standards-based grading or narratives. Although they offer a brief description of adequacy of students’ achievement and performance, letter grades provide no information that can be used to identify students’ unique accomplishments, their particular learning strengths, or their specific areas of weakness.”[4]
Guskey goes on to say, “letter grades should always be based on clearly stated learning criteria, not on norm-referenced criteria.”
Guskey and Bailey
The seriousness of arguments over plus and minus grades contrasts sharply with the simplicity of the issue involved. Basically, the issue comes down to whether is is better to have a 5-category grade system (A, B, C, D, and F), or a 12 category system (A, A-, B+, B, B-, and so on0. But if more categories are better, one might ask, “Why stop at 12?” There’s nothing sacred or particularly special about using 12 categories. Instead, we might consider a scale similar to the one used to express grade point average: 0.0-4.0.”[5]
This would equal 41 categories if you stick to just tenths place. Or you could go on to percents and get to 101 categories. Or even more if you go to percents and decimals. 92.34 for example.
Guskey and Bailey go on to say, “Research on rating scales shows that increasing the number of rating categories from 4 to just 6 generally lowers both the reliability and validity of the measures (Chang, 1993, 1994). Other studies indicate that scaled of 5 to possibly 9 categories are about as many as any qualified judge can reliable distinguish (Hargis, 1990, p. 14). Moreover, as the number of potential grades or grade categories increases, especially beyond 5 or 6, the reliability of grade assignments decreases. This means that the chance of two equally competent judges looking at the same collection of evidence and coming up with exactly the same grade is drastically reduced.”
Guskey and Bailey’s Recommendation
“Although to our knowledge no research evidence to date confirms that more affirming grade-category labels reduce stigma attached to low grades, we remain optimistic that this may be true. Certainly the connotation of Novice or Beginning is far less negative than that of Failing….
At the more advanced grade levels, we also believe that it is much more advantageous to assign a grade of I or Incomplete to students’ work and expect additional effort than it is to assign a letter grade of F (see the discussion of “Grades as Punishments” in Chapter 3).”[6]
Marzano
Marzano first writes a book on grading theory and only reluctantly gets to conversions about 7/8s of the way in. From having seen him talk it is clear that in his mind converting scores to anything is of little use. But I have also seen him say that grades in their traditional sense will probably always be necessary in grades 10-12 at least.
In essence Marzano explains that giving a score for each of the measurement topics (competencies) in a class would be preferable. But if a “district or school…wishes to use the traditional A, B, C, D, and F grading protocol,” it would need “a translation, such as the following:
3.00-4.00=A
2.50-2.99=B
2.00-2.49=C
1.50-1.99=D
Below 1.50=F”[7]
“Summary and Conclusions
Various Techniques can be use for computing final scores for topics and translating these scores to grades. Computer software that is suited to the system described in this book has three characteristics. First, the software should allow teachers to easily enter multiple scores for an assessment. Second, it should provide for the most accurate estimate of a student’s final score for each topic. Third, it should provide graphs depicting student progress.”[8]
Stiggins
Rick Stiggins and Anne Davies go hand in hand when they talk about Assessment FOR Learning. He wrote the original paper where he began always writing the for in all caps. They both focus most of their work on the idea that assessment (to sit beside) should be something that helps a student learn. In terms of conversions Stiggins suggests a “Decision Rule.”
“[If} at least 50% of the ratings are 5s and the rest are 4s, the grade is an A, [if] at least 75% of the ratings are 4s or better and the other 25% are not lower than 3, then the grade is a B, and [if] 40% of the ratings are 3s or better and the other 60% are not lower than two then the [grade] is a C.”[9]
“What about the situation in which a student receives a B, but it’s a high B or a low B? Over the course of an entire year, the difference will not be significant in terms of mastery, and mastery is what grades are based on, not averages. This isn’t being dismissive, but the reality is that the difference in learning (mastery) between the high and low versions of one particular grade is not that much. In larger grading scales, for example, the difference between a B and a B+ is just a few points. How exact can we be when identifying a student’s true mastery of something? Does a 0.01 (1 percent) difference in a grade-point average really mean a discernable, significant difference in mastery? No. It’s splitting hairs.
There are some teachers who disagree with this. They claim that there are a large number of mastery points wrapped into each percentage point due to multiple and influential assessments over a long period of time, and that the difference of one percentage point can describe mastery or lack of mastery of a significant amount of material. If this is the case, then whittling grades down to their exact and relative values (offering 2.75s for example) may be necessary. Each time we are tempted to do this however, let’s remember how elusive declarative mastery is, as well as how subjective we are in the micro-moment of grading each product from each student, and how we make it even more subjective when we aggregate a variety of data for a summative grade. And let’s wonder whether having done this, even justifiably, will have any lasting impact ten years down the road.”[11]
“One caution: If we primarily use a 4-point scale, many students and their parents will equate the highest numerical value (4.0) [or 5 in our case] with an A,…They will wonder why we just don’t write A, B, C, D and F if that’s what they really are.”[12]
Thanks for reading—if you made it this far.
Tom Crumrine
[1] Research of Black and Wiliam 1998 and Stiggins 2001 reported in Conferencing and Reporting by Gregory, Cameron and Davies.
[2] How to Grade for Learning—Ken O’Connor. Page 200.
[3] How to Grade for Learning—Ken O’Connor. Page 200.
[4] How’s My Kid Doing?-Tom Guskey. Pages 45-46.
[5] Developing Grading and Reporting Systems for Student Learning. Tom Guskey and Jane Bailey. Page 70.
[6] Developing Grading and Reporting Systems for Student Learning. Tom Guskey and Jane Bailey. Page 77.
[7] Classroom Assessment and Grading that Work—Robert J. Marzano. Page 122.
[8] Classroom Assessment and Grading that Work—Robert J. Marzano. Page 124.
[9] Fair Isn’t Always Equal—Rick Wormeli. Page 154. Wormeli quotes Stiggins here.
[10] Fair Isn’t Always Equal—Rick Wormeli. Page 154.
[11] Fair Isn’t Always Equal—Rick Wormeli. Page 154-155.
[12] Fair Isn’t Always Equal—Rick Wormeli. Page 157.
13 February 2009
Mid Year Experiment
Middle of the night draft
By Tom Crumrine
In November I attended a conference and Douglas Reeves suggested giving midyears early, providing corrective instruction and giving the midyear again. A sysnopsis is found here:
So, I decided to try the experiment myself…
Introduction
The idea was to give the midyear at an earlier point so I decided to go with the logical pre-Christmas pre-exam. I have always wondered why we give first semester exams after a 2 week holiday so I just gave them early.
Upon our return to class in January I did the following things:
- Over the holidays I “graded” the exams by highlighting areas where students could add more information.
- On the first day back I gave students the exams and asked them to provide more information in the highlighted areas.
- When I got the test back I marked the scored versus the standards.
- With the scores v. the standards I was able to develop a plan to educate groups of students.
- In the two weeks after the holidays and before exams I provided corrective instruction
When it came time to take the exam I created a test that tested the same standards but with different questions. If students had already scored a 5 on a particular measurement topic they did not have to respond to the new question. So some students had to do all 9 questions and some had to do only 1. (All students completed a common part of the exam that assessed basic science skills—calculation of density, metric conversions, lab safety, etc.)
The Results
Figure 1: Items 1-9 are the measurement topics. The numbers at the bottom are the averages of all student scores. The test dates were exactly 1 month apart with a 2 week holiday break at the beginning of the month and four 90 minute classes of corrective instruction prior to exam week.
Reflection
Obviously I was excited about the graph and the fact that the average for all measurement topics went up. But there are both positives and concerns with this approach.
Concerns
We spent four 90 minute periods on the corrective part. I will argue later that this is not wasted time but we did not go on to new material.
The students who did well the first time around actually did go on to new material but they were self directed as I spent most of the time working with the students who needed more help. The students worked on an extension of the atomic structure unit where they investigated isotopes by looking at the poisoning death of Russian agent Alexandre Litvenienko. The students that were good at being active self directed learners told me that they enjoyed this project and they were glad that they got to do it. But those students that were not good at self directed learning did not get much out of this extension project.
I scored both tests. I tried very hard to eliminate any bias that would come from doing this by not looking at the December scores when scoring the January test and by creating specific rubrics for each question—but the fact remains…
Positives
The corrective instruction took some time but all student scores went up. The scores were based on understanding of the topic so the understanding of my students increased with the extra two weeks of instruction. The research backs up depth over breadth but as a teacher who learned science in a different era—the one of breadth over depth with the memorization of tons of factoids—it feels like I’m doing something wrong.
A couple of students asked if they could forgo the extra project and help other students. I granted this request and the results were great. Seeing one student teach another how to describe the model of the atom is the kind of scene that makes you nearly tear up.
Students loved the fact that each test was essentially customized to them. Once they got over the newness of the fact that they only had to answer certain questions they really liked the approach. This also was the “reward” for those students who did well the first time around.
Conclusion
While I have concerns the positives do outweigh them in my mind. In order for high school to change we must change some things about high school. This is one experiment that will be worth repeating.
08 November 2008
Averaging and Zeroes
This is a clip of Dr. Douglas Reeves speaking to a Canadian audience about what he calls "toxic" grading practices. Reeves is the author of more than 25 books and countless articles on education.
In the clip Reeves talks about zeroes and averaging. We showed this clip to our high school faculty and there were wide ranging responses:
- Students deserve the average because this helps differentiate the steady performing student from the student that does poorly all the time but well at the end.
- Students are a sum of all of their performances so the average is the correct score for them.
- Zeroes are an essential part of grading.
- I want students to see zeroes and I want them to be calculated in the grade.
Here I want to clarify and expand upon some of what Dr. Reeves was getting at. First in the case of the average.
While in the clip he states provocatively that all averaging must go, what he and other researchers have been preaching for the past decade is to end a mindless devotion to the average as the only way of evaluating students. In fact the average might be the right score for a given student. I, along with Reeves and other researchers, am arguing that the average is not the best evaluation for all students.
It comes back as always to a conversation about standards. Is our goal to get them to the standard (in our parlance competency)? If we work hard as teachers and students work hard at learning and understanding and they make it to the standard what should the grade represent? Why should it be the average in this case? If everyone can meet the standard at the end but then we average scores it not only hurts students but it hurts us. The scores are a poor representation of how we operated as teachers. Grading is always a subjective process. As professionals we strive to minimize the subjectivity but we cannot eliminate it. As professionally trained practitioners we should allow ourselves to award students the score, the evaluation, that most appropriately matches their ability.
As far as zeroes I have posted on this many times but take one more example.
- 100-40=A
- 39-30=B
- 29-20=C
- 19-10=D
- 9-0=F
When you present this example people say that is ridiculous! But this system is as mathematically unsound as the system where the top four categories are 10 or 11 points and the last category is 59 points. So this is the first argument against zeroes--it is simply mathematically unsound.
The second argument is that it does not increase motivation. And has been shown by Dr. Reeves to have a role in whether struggling students stay in school or leave. Zeroes motivate only one type of student--good ones, ones like teachers used to be when they were in the classroom. The students that we worry about the most are not motivated to do work by receiving a zero. To the contrary they are encouraged to give up because when zeros mount the combination of their extra mathematical weight and the increase in a feeling of hopelessness cause students to shut down.
I feel strongly that zeroes should not be used and the average should not be used in all cases. That said, if a teacher still wants to use zeroes and averages as the only way to go I would ask them to continue that practice only after reflecting on exactly why they want to do it that way. As professionals we will always evaluate in different ways--the question is: Is the way you reach the evaluation of a student the best representation of what they can do?
