Tuesday, September 4, 2012

Item Analysis


ITEM ANALYSIS
            Evaluation is indispensable part of and different types of tests are used for assessment and consequently evaluation. Tests play important role in giving feedback stakeholders in education on various aspects therefore quality of tests has always been a hot issue since long; consequently literature is full of comprehensive discussion on validity, reliability and the characteristics of quality assessment programs (Stephen & Polly, 2006) so that to bring improvement in feedback. For this reasons item analysis is widely used to improve test quality through knowing about item statistics. Item statistics are not only used for improvement of test but also in item revision (Lange, 1967). Item analysis allows us to observe the characteristics of a particular item and can be used to ensure that questions are of an appropriate standard for inclusion in a test. A comprehensive knowledge of the factors leading to construct a good test item can enable us to create more effective test besides standardizing the existing tests. Improvement in the test through item analysis can save a lot of time and energy on the part of teachers and test developer. Typically, in analysis of a test, two values are computed, a difficulty level and a discrimination index.
            One of the most important tasks confronting by the faculty members is the student performance evaluation. Once designed, the evaluative procedure must be administered and then scored, interpreted and graded. After that, feedback must be presented to students. Doing these tasks demands a broad range of cognitive, technical and interpersonal resources on the part of faculty. But there is an even more critical task remains which is investigating the quality of the evaluative procedure.
            What constitutes a good exam item? The students seem to know or at least believe they know but are they correct when they claim that an item was too difficult, too tricky or too unfair. According to Lewis Aiken, item analysis is a group of procedures for assessing the quality of exam items. The purpose of an item analysis is to improve the quality of an exam by identifying items that are candidates for retention, revision or removal. It can also clarify what concepts the examinees have and have not mastered.
PURPOSE:
1.      Assemble or write a relatively large number of items of the type you want on the test.
2.      Analyze the items carefully using item format analysis to make sure they are well-written and clear.
3.      Pilot the items using a group of students similar to the group that will ultimately be taking the test.
4.      4. Analyze the results statistically using item analysis techniques.
5.      Select the most effective items and make a shorter, more effective revised version of the test.
2 BROAD CATEGORIES:
Qualitative Item Analysis
            Qualitative item analysis procedures include careful proofreading of the exam prior to its administration for typographical errors, for grammatical cues that might inadvertently tip off examinees to the correct answer, and for the appropriateness of the reading level of the material. This procedure can also include small group discussion of the quality of the exam and its items with the examinees who have already taken the test or with departmental student assistants or even experts in the field. Some teachers use “think-aloud test administration” in which examinees are asked to verbally express their opinions or what they are thinking as they respond to each of the items on the exam. It can assist the teachers in determining whether certain students misinterpreted particular items and it can also determine why students may have misinterpreted a particular item.
Quantitative Item Analysis
            Specifically, three numerical indicators are often derived during an item analysis: item difficulty, item discrimination and distractor power.
1. Item Difficulty Index (p)
            The item difficulty statistic is an appropriate choice for achievement or aptitude tests when the items are scored dividedly (correct and incorrect). It can be derived for true-false, multiple choice, matching items and essay items where the instructor can convert the range of possible point values into categories “passing” and “failing.”
            The item diificulty index or p can be computed by:
         p =         Number of test takers who answered the item correctly
                        Total number of students who answered the item
            p can range between 0.00 (no examinees answered the item correctly) and 1.00 (all examinees answered the item correctly. It can also range from 0% to 100%, the higher the value, the easier the item. Pvalues above 0.90 are very easy items and might be a concept not worth testing. P-values below 0.20 indicate difficult items and should be reviewed for possible confusing language or the contents needs re-instruction. Optimum difficulty level is 0.50 for maximum discrimination between high and low achievers.
 No test item need have only one p value. Not only may the p value vary with each class group that takes the test, the teacher may gain insight by computing the item difficulty level for a number of different subgroups within a class, those who did well on the exam overall and those who performed more poorly.
For example, the difficulty level is 0.20 so it means 20% of the examinees answered the item correctly. Does this item mean that the item was challenging for all? Does it mean that the teacher failed in his or her attempt to teach the particular topic assessed by the item? Does it mean that the students failed to learn the material? Does it mean that the item was poorly written? Teachers must also rely on other item analysis procedure, qualitative and quantitative.
2. Item Discrimination Index (D)
            Item discrimination analysis deals with the fact that often different test takers will answer a test item in different ways. It addresses the validity of the items on a test, the extent to which the items tap the attributes they were intended to assess.
            It is the point-bacterial relationship between students’ performance on individual item and total test score. This value ranges between 0.0 and 1.00. The higher the value, more discriminating the item. A highly discriminating item indicates that the students who had high tests scores got the item correct whereas students who had low test scores got the item incorrect.
            Teachers test because they want to find out if students know the material but they learn is how they did on the exam we gave them. Item discrimination index tests the test in the hope of keeping the correlation between knowledge and exam performance as close as it can be in an admittedly imperfect system.
            It is calculated by:
A.    Divide the group of test takers into 2 groups. (high scoring and low scoring).
B.     Compute the item difficulty levels separately for the upper (Pupper) and lower (Plower) scoring groups.
C.     D = pupper-  plower
How can this be interpreted?
            Example: Half of the examinees answered a particular item correctly and that all of the examinees who scored above the median on the exam answered the item correctly and all of the examinees who scored below the median answered incorrectly. Pupper = 1.00 and plower = 0.00. Then, D = 1.00 and the item is somehow a perfect positive discriminator. This suggests that those examinees who knew the material and were well-prepared passed the item while the others failed it.
            Difficulty and discrimination are not independent. If all students in both the upper and lower levels either pass or fail an item, there is nothing in the data to indicate whether the item itself was good or not. The value of the item discrimination index will be maximized when only half of the test takers overall answer an item correctly. The ideal situation is one in which the half who passed the item were students who all did well on the exam overall. There are many reasons to include at least some such items. Very easy items can reflect the fact that some relatively straightforward concepts were taught well and mastered by all students. The teacher may choose to include some very difficult items on the exam to challenge even the best-prepared students. The teacher should be aware that neither of these types of items functions well to make discriminations among those taking the test.
3. Item Distractor Analysis
            This mainly applies particularly to multiple-choice item. The incorrect alternatives are called distractors. Item distractor analysis examines the percentage of examinees who select each incorrect alternative, to determine whether the distractors are functioning as intended. On a well-designed multiple choice item, those who know the material and are well-prepared for the exam should select the correct alternative. Those who are not well-prepared should guess or select almost randomly from among the available distractors. Such an item would be a very good discriminator and would very likely be a candidate for retention for use in future exams.
            It can also provide useful diagnostic information in other situations. Candidate for removal from the exam is the item that was passed by more of those who did poorly on the exam overall that those who were well-prepared and knew the material
Caution in Item Analysis
• Item analysis data is just a reflection of internal consistency and therefore should not be treated as item validity which requires an external criteria e.g. experts’ opinions to accurately judge the validity of test items.
• A low discrimination index does not make an item to be dropped from a test because extremely difficult or easy items may have low ability to discriminate. But such items can be included in a test to sample course content adequately. Similarly an item may have low discrimination because multidimensionality of a test.
• Item analysis data are tentative due to influences by factors like sample of students, quality of instruction, and chance errors.
(“Teste Item Analysis”, Online,
www.utexas.edu/academic/mec/research/.../itemanalysishandout.pdf, accessed on 15-08-2010)
Conclusion:
            Item difficulty and discrimination analysis programs are often included in the software used in processing exams answered on Scantron or other optically scannable forms. These analyses can often be performed for students by personnel in the computer services office. Item analysis can certainly help determine whether or not the items on the exams were good ones and to determine which items to retain, revise or replace.
References:
(Zurawski, R. Making the Most of Exams: Procedures for Item Analysis.)


item analysis


THE PURPOSE OF ITEM ANALYSIS
There must be a match between what is taught and what is assessed. However, there must also be an effort to test for more complex levels of understanding, with care taken to avoid over-sampling items that assess only basic levels of knowledge. Tests that are too difficult (and have an insufficient floor) tend to lead to frustration and lead to deflated scores, whereas tests that are too easy (and have an insufficient ceiling) facilitate a decline in motivation and lead to inflated scores. Tests can be improved by maintaining and developing a pool of valid items from which future tests can be drawn and that cover a reasonable span of difficulty levels.
Item analysis helps improve test items and identify unfair or biased items. Results should be used to refine test item wording. In addition, closer examination of items will also reveal which questions were most difficult, perhaps indicating a concept that needs to be taught more thoroughly. If a particular distracter (that is, an incorrect answer choice) is the most often chosen answer, and especially if that distracter positively correlates with a high total score, the item must be examined more closely for correctness. This situation also provides an opportunity to identify and examine common misconceptions among students about a particular concept.
In general, once test items have been created, the value of these items can be systematically assessed using several methods representative of item analysis: a) a test item's level of difficulty, b) an item's capacity to discriminate, and c) the item characteristic curve. Difficulty is assessed by examining the number of persons correctly endorsing the answer. Discrimination can be examined by comparing the number of persons getting a particular item correct with the total test score. Finally, the item characteristic curve can be used to plot the likelihood of answering correctly with the level of success on the test.

ITEM DIFFICULTY

In test construction, item difficulty is determined by the number of people who answer a particular test item correctly. For example, if the first question on a test was answered correctly by 76% of the class, then the difficulty level (p or percentage passing) for that question isp = .76. If the second question on a test was answered correctly by only 48% of the class, then the difficulty level for that question is p = .48. The higher the percentage of people who answer correctly, the easier the item, so that a difficulty level of .48 indicates that question two was more difficult than question one, which had a difficulty level of .76.
Many educators find themselves wondering how difficult a good test item should be. Several things must be taken into consideration in order to determine appropriate difficulty level. The first task of any test maker should be to determine the probability of answering an item correctly by chance alone, also referred to as guessing or luck. For example, a true-false item, because it has only two choices, could be answered correctly by chance half of the time. Therefore, a true-false item with a demonstrated difficulty level of only p = .50 would not be a good test item because that level of success could be achieved through guessing alone and would not be an actual indication of knowledge or ability level. Similarly, a multiple-choice item with five alternatives could be answered correctly by chance 20% of the time. Therefore, an item difficulty greater than .20 would be necessary in order to discriminate between respondents' ability to guess correctly and respondents' level of knowledge. Desirable difficulty levels usually can be estimated as halfway between 100 percent and the percentage of success expected by guessing. So, the desirable difficulty level for a true-false item, for example, should be around p = .75, which is halfway between 100% and 50% correct.
In most instances, it is desirable for a test to contain items of various difficulty levels in order to distinguish between students who are not prepared at all, students who are fairly prepared, and students who are well prepared. In other words, educators do not want the same level of success for those students who did not study as for those who studied a fair amount, or for those who studied a fair amount and those who studied exceptionally hard. Therefore, it is necessary for a test to be composed of items of varying levels of difficulty. As a general rule for norm-referenced tests, items in the difficulty range of .30 to .70 yield important differences between individuals' level of knowledge, ability, and preparedness. There are a few exceptions to this, however, with regard to the purpose of the test and the characteristics of the test takers. For instance, if the test is to help determine entrance into graduate school, the items should be more difficult to be able to make finer distinctions between test takers. For a criterion-referenced test, most of the item difficulties should be clustered around the criterion cut-off score or higher. For example, if a passing score is 70%, the vast majority of items should have percentage passing values of
Figure 1ILLUSTRATION BY GGS INFORMATION SERVICES. CENGAGE LEARNING, GALE.
p = .60 or higher, with a number of items in the p > .90 range to enhance motivation and test for mastery of certain essential concepts.

DISCRIMINATION INDEX

According to Wilson (2005), item difficulty is the most essential component of item analysis. However, it is not the only way to evaluate test items. Discrimination goes beyond determining the proportion of people who answer correctly and looks more specifically at who answers correctly. In other words, item discrimination determines whether those who did well on the entire test did well on a particular item. An item should in fact be able to discriminate between upper and lower scoring groups. Membership in these groups is usually determined based on their total test score, and it is expected that those scoring higher on the overall test will also be more likely to endorse the correct response on a particular item. Sometimes an item will discriminate negatively, that is, a larger proportion of the lower group select the correct response, as compared to those in the higher scoring group. Such an item should be revised or discarded.
One way to determine an item's power to discriminate is to compare those who have done very well with those who have done very poorly, known as the extreme group method. First, identify the students who scored in the top one-third as well as those in the bottom one-third of the class. Next, calculate the proportion of each group that answered a particular test item correctly (i.e., percentage passing for the high and low groups on each item). Finally, subtract the p of the bottom performing group from the p for the top performing group to yield an item discrimination index (D). Item discriminations of D = .50 or higher are considered excellent. D = 0 means the item has no discrimination ability, while D = 1.00 means the item has perfect discrimination ability.
In Figure 1, it can be seen that Item 1 discriminates well with those in the top performing group obtaining the correct response far more often (p = .92) than those in the
Figure 2ILLUSTRATION BY GGS INFORMATION SERVICES. CENGAGE LEARNING, GALE.
low performing group (p = .40), thus resulting in an index of .52 (i.e., .92 - .40 = .52). Next, Item 2 is not difficult enough with a discriminability index of only .04, meaning this particular item was not useful in discriminating between the high and low scoring individuals. Finally, Item 3 is in need of revision or discarding as it discriminates negatively, meaning low performing group members actually obtained the correct keyed answer more often than high performing group members.
Another way to determine the discriminability of an item is to determine the correlation coefficient between performance on an item and performance on a test, or the tendency of students selecting the correct answer to have high overall scores. This coefficient is reported as the item discrimination coefficient, or the point-biserial correlation between item score (usually scored right or wrong) and total test score. This coefficient should be positive, indicating that students answering correctly tend to have higher overall scores or that students answering incorrectly tend to have lower overall scores. Also, the higher the magnitude, the better the item discriminates. The point-biserial correlation can be computed with procedures outlined in Figure 2.
In Figure 2, the point-biserial correlation between item score and total score is evaluated similarly to the extreme group discrimination index. If the resulting value is negative or low, the item should be revised or discarded. The closer the value is to 1.0, the stronger the item's discrimination power; the closer the value is to 0,
Figure 3ILLUSTRATION BY GGS INFORMATION SERVICES. CENGAGE LEARNING, GALE.
the weaker the power. Items that are very easy and answered correctly by the majority of respondents will have poor point-biserial correlations.

CHARACTERISTIC CURVE

A third parameter used to conduct item analysis is known as the item characteristic curve (ICC). This is a graphical or pictorial depiction of the characteristics of a particular item, or taken collectively, can be representative of the entire test. In the item characteristic curve the total test score is represented on the horizontal axis and the proportion of test takers passing the item within that range of test scores is scaled along the vertical axis.
For Figure 3, three separate item characteristic curves are shown. Line A is considered a flat curve and indicates that test takers at all score levels were equally likely to get the item correct. This item was therefore not a useful discriminating item. Line B demonstrates a troublesome item as it gradually rises and then drops for those scoring highest on the overall test. Though this is unusual, it can sometimes result from those who studied most having ruled out the answer that was keyed as correct. Finally, Line C shows the item characteristic curve for a good test item. The gradual and consistent positive slope shows that the proportion of people passing the item gradually increases as test scores increase. Though it is not depicted here, if an ICC was seen in the shape of a backward S, negative item discrimination would be evident, meaning that those who scored lowest were most likely to endorse a correct response on the item.

BIBLIOGRAPHY

Anastasi, A., & Urbina, S. (1997). Psychological testing (7th ed.). Upper Saddle River, NJ: Prentice Hall.
Brown, F. (1983). Principles of education and psychological testing(3rd ed.). New York: Holt, Rinehart, & Winston.
DeVellis, R. (2003). Scale development: Theory and applications (2nd ed.). Thousand Oaks, CA: Sage.
Grunlund, N. (1993). How to make achievement tests and assessments (5th ed.). Boston: Allyn and Bacon.
Kaplan, R., & Saccuzzo, D. (2004). Psychological testing: Principles, applications, and issues (6th ed.) Pacific Grove, CA: Brooks/Cole.
Kehoe, J. (1995). Basic item analysis for multiple-choice tests.Practical Assessment, Research & Evaluation, 4(10), retrieved April 1, 2008, from http://pareonline.net/getvn.asp?v=4&n=10.
Patten, M. (2001). Questionnaire research: A practical guide (2nd ed.). Los Angeles: Pyrczak.
Wilson, M. (2005). Constructing measures: An item response modeling approach. Mahwah, NJ: Lawrence Erlbaum.

Article: Item Analysis

FACT SHEET 24A

Title: Item Analysis Assumptions (Difficulty & Discrimination Indexes)           
 Date: June 2008

Details: Mr. Imran Zafar, Database Administrator, Assessment Unit, Dept of Medical Education.  Ext. 47142

Introduction:

It is widely believed that “Assessment drives the curriculum”. Hence it can be argued that if the quality of teaching, training, and learning is to be upgraded, assessment is the obvious starting point. However, upgrading assessment is continuous process. The cycle of planning, and constructing assessment tools, followed by testing, validating, and reviewing has to be repeated continuously

When tests are developed for instructional purposes, to assess the effects of educational programs, or for educational research purposes, it is very important to conduct item and test analyses. These analyses evaluate the quality of the items and of the test as a whole. Such analyses can also be employed to revise and improve both items and the test as a whole.

Quantitative item analysis is a technique that enables us to assess the quality or utility of an item. It does so by identifying distractors or response options that are underperforming.

Item-analysis procedures are intended to maximize test reliability. Because maximization of test reliability is accomplished by determining the relationship between individual items and the test a whole, it is important to insure that the overall test is measuring what it is supposed to measure. It this not the case, the total score will be a poor criterion for evaluating each item.

The use of a multiple-choice format for hour exams at many institutions leads to a deluge of statistical data, which are often neglected or completely ignored. This paper will introduce some of the terms encountered in the analysis of test results, so that these data may become more meaningful and therefore more useful

Need for Item Analysis

1)                  Provision of information about how the quality of test items compare. The comparison is necessary if subsequent tests of the same material are going to be better.

2)                  Provision of diagnostic information about the types of items that students most often get incorrect. This information can be used as a basis for making instructional decision.

3)                  Provision of a rational basis for discussing test results with students.

4)                  Communication to the test developer which items needs to be improved or eliminated, to be replaced with better items.

What is the Output of Item Analysis?


Item analysis could yield the following outputs:

  • Distribution of responses for each distractor of each item or frequencies of responses (histogram).
  • Difficulty index for each item of the test
  • Discrimination Index for each item of the test.
  • Measure of exam internal consistency reliability

Total Score Frequencies


Issues to consider when interpreting the distribution of students’ total scores:

Distribution:

  • Is this the distribution you expected?
  • Was the test easier, more difficult than you anticipated?
  • How does the mean score of this year’s class compare to scores from previous classes?
  • Is there a ceiling effect – that is, are all scores close to the top?
  • Is there a floor effect – that is, are all scores close to the lower possible?

Spread of Scores:

  • Is the spread of scores large?
  • Are there students who are scoring low marks compared to the majority of the students?
  • Can you determine why they are not doing as well as most other students?
  • Can you provide any extra assistance? Is there a group of students who are well ahead of the other students?

Difficulty Index


It actually tells us how easy the item was for the students in that particular group. The higher the difficulty index the easier the question; the lower the difficulty index, the more difficult the question. The difficulty index, in fact, equals to “Easiness Index”.

Issues to consider in relation to the Difficulty Index

  • Are the easiest items in your test, i.e. those with the lowest difficulty ranking, the first items in the test?
  • If the more difficult items occur at the start of the test, students can become upset because they feel, early on, that they can not succeed.





Literature quotes following (generalized) interpretation of Difficulty Index.

Indexed Range
Inference to Question
0.85 – 1.00
Very Easy
0.70 – 0.84
Easy
0.30 – 0.69
Optimum
0.15 – 0.29
Hard
0.00 – 0.14
Very Hard

Item Discrimination Index (DI)


This is calculated by subtracting the proportion of students correct in the lower group from the proportion correct in the upper group. It is assumed that persons in the top third on total scores should have a greater proportion with the item correct than the lower third.

The calculation of the index is an approximation of a correlation between the scores on an item and the total score. Therefore, the DI is a measure of how successfully an item discriminates between students of different abilities on the test as a whole. Any item which did not discriminates between the lower and upper group of students would have a DI=0. An item where the lower group performed better than the upper group would have a negative DI.

The discrimination index is affected by the difficulty of an item, because by definition, if an item is very easy everyone tends to get it right and it does not discriminate. Likewise, if it is very difficult everyone tends to get it wrong. Such items can be important to have in a test because they help define the range of difficulty of concepts assessed. Items should not be discarded just because they do not discriminate.

Issues to consider in relation to the Item Discrimination Index

  • Are there any items with a negative discrimination index (DI)? That is, terms where students in the lower third of the group did better than students in the upper third of the group?
  • Was this a deceptively easy item?
  • Was the correct answer key used?
  • Are there any items that do not discriminate between the students i.e. where the DI is 0.00 or very close to 0.0?
  • Are these items which are either very hard or very easy and therefore where you could have a DI of 0?

Literature quotes following (generalized) interpretation of Discrimination Index.

Indexed Range
Inference to Question
Below 0.19
Poor
0.20 – 0.29
Dubious
0.30 – 1.00
Okay





FACT SHEET 24B

Title: Item Analysis Assumptions

Measure of exam internal consistency (reliability) Kuder-Richardson 20 (KR20)

Date: June 2008

Details: Mr. Imran Zafar, Database Administrator, Assessment Unit, Dept of Medical Education.  Ext. 47142

Test Validity and Reliability

Test reliability measures the accuracy, stability, and consistency of the test scores.  Reliability is affected by the characteristics of the students, characteristics of the test, and conditions affecting test administration and scoring.

The two factors that determine overall test quality are test validity and test reliability.

Test validity is the appropriateness of the test for the subject area and students being tested.  Validity cannot be measured by a computer.  It is up to the instructor to design valid test items that best measure the intended subject area.  By definition, valid tests are reliable.  However, a reliable test is not necessarily valid.  For example, a math test comprised entirely of word problems may be measuring as much verbal skills as math ability.

The reliability of a test refers to the extent to which the test is likely to produce consistent scores.  The KR-20 index is the appropriate index of test reliability for multiple-choice examinations.

What does the KR-20 measure?

The KR-20 is a measure of internal consistency reliability or how well your exam measures a single cognitive factor.  If you administer an Embryology exam, you hope all test items relate to this broad construct. Similarly, an Obstetrics/Gynecology test is designed to measure this medical specialty.

Reliability coefficients theoretically range in value from zero (no reliability) to 1.00 (perfect reliability). In practice, their approximate range is from .50 to .90 for about 95% of the classroom tests.

High reliability means that the questions of a test tended to "pull together." Students who answered a given question correctly were more likely to answer other questions correctly. If a parallel test were developed by using similar items, the relative scores of students would show little change.

Low reliability means that the questions tended to be unrelated to each other in terms of who answered them correctly. The resulting test scores reflect peculiarities of the items or the testing situation more than students' knowledge of the subject matter.




The KR-20 formula includes

(1)        the number of test items on the exam,
(2)        student performance on every test item, and
(3)        the variance (standard deviation squared) for the set of student test scores. 

The index ranges from 0.00 to 1.00. A value close to 0.00 means you are measuring many unknown factors but not what you intended to measure. You are close to measuring a single factor when your KR-20 is near 1.00.  Most importantly, we can be confident that an exam with a high KR-20 has yielded student scores that are reliable (i.e., reproducible or consistent; or as psychometricians say, the true score). A medical school test should have a KR-20 of 0.60 or better to be acceptable.

How do you interpret the KR-20 value?

Reliability                               Interpretation

.90 and above                          Excellent reliability; at the level of the best standardized tests
.80 - .90                                   Very good for a classroom test
.70 - .80                                   Good for a classroom test; in the range of most. There are probably a few items which could be improved.
.60 - .70                                   Somewhat low. This test needs to be supplemented by other measures (e.g., more tests) to determine grades. There are probably some items which could be improved.
.50 - .60                                   Suggests need for revision of test, unless it is quite short (ten or fewer items). The test definitely needs to be supplemented by other measures (e.g., more tests) for grading.
.50 or below                            Questionable reliability. This test should not contribute heavily to the course grade, and it needs revision.

Standard Error of Measurement

The standard error of measurement is directly related to the reliability of the test. It is an index of the amount of variability in an individual student's performance due to random measurement error. If it were possible to administer an infinite number of parallel tests, a student's score would be expected to change from one administration to the next due to a number of factors. For each student, the scores would form a "normal" (bellshaped) distribution. The mean of the distribution is assumed to be the student's "true score," and reflects what he or she "really" knows about the subject. The standard deviation of the distribution is called the standard error of measurement and reflects the amount of change in the student's score which could be expected from one test administration to another.

Whereas the reliability of a test always varies between 0.00 and 1.00, the standard error of measurement is expressed in the same scale as the test scores. For example, multiplying all test scores by a constant will multiply the standard error of measurement by that same constant, but will leave the reliability coefficient unchanged.

A general rule of thumb to predict the amount of change which can be expected in individual test scores is to multiply the standard error of measurement by 1.5. Only rarely would one expect a student's score to increase or decrease by more than that amount between two such similar tests. The smaller the standard error of measurement, the more accurate the measurement provided by the test.


A CAUTION in Interpreting Item Analysis Results

Each of the various item statistics provides information which can be used to improve individual test items and to increase the quality of the test as a whole. Such statistics must always be interpreted in the context of the type of test given and the individuals being tested. W. A. Mehrens and I. J. Lehmann provide the following set of cautions in using item analysis results (Measurement and Evaluation in Education and Psychology. New York: Holt, Rinehart and Winston, 1973, 333-334):

1.                  Item analysis data are not synonymous with item validity. An external criterion is required to accurately judge the validity of test items. By using the internal criterion of total test score, item analyses reflect internal consistency of items rather than validity.

2.                  The discrimination index is not always a measure of item quality. There is a variety of reasons an item may have low discriminating power:

a)      extremely difficult or easy items will have low ability to discriminate but such items are often needed to adequately sample course content and objectives;

b)      an item may show low discrimination if the test measures many different content areas and cognitive skills. For example, if the majority of the test measures "knowledge of facts," then an item assessing "ability to apply principles" may have a low correlation with total test score, yet both types of items are needed to measure attainment of course objectives.

3.                  Item analysis data are tentative. Such data are influenced by the type and number of students being tested; instructional procedures employed, and chance errors. If repeated use of items is possible, statistics should be recorded for each administration of each item.

Conclusions

Developing the perfect test is the unattainable goal for anyone in an evaluative position. Even when guidelines for constructing fair and systematic tests are followed, a plethora of factors may enter into a student's perception of the test items. Looking at an item's difficulty and discrimination will assist the test developer in determining what is wrong with individual items. Item and test analysis provide empirical data about how individual items and whole tests are performing in real test situations.

One of the principal advantages of having MCQ tests scored and analyzed by computer at College of Medicine, King Saud bin Abdulaziz University for Health Sciences is the feedback available on how well the test has performed. Careful consideration of the results of item analysis can lead to significant improvements in the quality of exams. 

Reference:  com.ksau-hs.edu.sa/eng/images/DME_Fact_Sheets/fs_24.doc


Item Analysis
Item Analysis allows us to observe the characteristics of a particular question (item) and can be used to ensure that questions are of an appropriate standard and select items for test inclusion.
IntroductionItem Analysis describes the statistical analyses which allow measurement of the effectiveness of individual test items. An understanding of the factors which govern effectiveness (and a means of measuring them) can enable us to create more effective test questions and also regulate and standardise existing tests.
There are three main types of Item Analysis: Item Response Theory, Rasch Measurement and Classical Test Theory. Although Classical Test Theory and Rasch Measurement will be discussed, this document will concentrate primarily on Item Response Theory.
The Models
Classical Test Theory

Classical Test Theory (traditionally the main method used in the United Kingdom) utilises two main statistics - Facility and Discrimination.
  • Facility is essentially a measure of the difficulty of an item, arrived at by dividing the mean mark obtained by a sample of candidates and the maximum mark available. As a whole, a test should aim to have an overall facility of around 0.5, however it is acceptable for individual items to have higher or lower facility (ranging from 0.2 to 0.8).
  • Discrimination measures how performance on one item correlates to performance in the test as a whole. There should always be some correlation between item and test performance, however it is expected that discrimination will fall in a range between 0.2 and 1.0.
The main problems with Classical Test Theory are that the conclusions drawn depend very much on the sample used to collect information. There is an inter-dependence of item and candidate.
Item Response Theory
Item Response Theory (IRT) assumes that there is a correlation between the score gained by a candidate for one item/test (measurable) and their overall ability on the latent trait which underlies test performance (which we want to discover). Critically, the 'characteristics' of an item are said to be independent of the ability of the candidates who were sampled.
Item Response Theory comes in three forms: IRT1, IRT2, and IRT3 reflecting the number of parameters considered in each case.
  • For IRT1, only the difficulty of an item is considered,
    (difficulty is the level of ability required to be more likely to correctly answer the question than answer it wrongly).
  • For IRT2, difficulty and discrimination are considered,
    (discrimination is how well the question is at separating out candidates of similar abilities).
  • For IRT3, difficulty, discrimination and chance are considered,
    (chance is the random factor which enhances a candidates probability of success through guessing.
Rasch Measurement
Rasch measurement is very similar to IRT1 - in that it considers only one parameter (difficulty) and the ICC is calculated in the same way. When it comes to utilising these theories to categorise items however, there is a significant difference. If you have a set of data, and analyse it with IRT1, then you arrive at an ICC that fits the data observed. If you use Rasch measurement, extreme data (e.g. questions which are consistently well or poorly answered) is discarded and the model is fitted to the remaining data.

Why Item Analysis Important?
Item Analysis is an important (probably the most important) tool to increase test effectiveness. Each items contribution is analyzed and assessed.
To write effective items, it is necessary to examine whether they are measuring the fact, idea, or concept for which they were intended. This is done by studying the student’s responses to each item. When formalized, the procedure is called “item analysis”. It is a scientific way of improving the quality of tests and test items in an item bank.
An item analysis provides three kinds of important information about the quality of test items.
  • Item difficulty: A measure of whether an item was too easy or too hard.
  • Item discrimination: A measure of whether an item discriminated between students who knew the material well and students who did not.
  • Effectiveness of alternatives: Determination of whether distractors (incorrect but plausible answers) tend to be marked by the less able students and not by the more able students.
2. Article on system analysis please refer to the URL below:
a. www1.carleton.ca/edc/ccms/wp-content/ccms.../Item-Analysis.pdf
b. http://www.education.com/reference/article/item-analysis/#A
c. www.washington.edu/oea/pdfs/resources/item_analysis.pdf