Creating reliable knowledge tests goes beyond simply writing questions and recording answers. Test developers need to ensure each question contributes meaningfully to the assessment’s overall purpose. Item difficulty and discrimination indices provide the quantitative data needed to evaluate how well each test question performs. These statistical measures help identify which items effectively measure knowledge and which ones need improvement or removal, transforming test development from guesswork into a systematic, evidence-based process.

Table of Contents

What is item difficulty?

Item difficulty measures how challenging a test question is for the people taking it. Despite its name, the difficulty index actually indicates ease rather than difficulty. It’s calculated as the proportion of test-takers who answer an item correctly, typically expressed as a p-value ranging from 0.0 to 1.0.

For example, if 80 out of 100 students answer a question correctly, the item difficulty index is 0.80 or 80%. This would be considered a relatively easy item. Higher p-values indicate easier items, while lower values indicate more difficult ones. The calculation is straightforward for single-correct-answer items: simply divide the number of correct responses by the total number of responses.

Interpreting difficulty levels

Test developers typically categorize items into three difficulty levels:

Easy items: Difficulty index above 0.75 (more than 75% answer correctly)

Moderate items: Difficulty index between 0.30 and 0.75

Difficult items: Difficulty index below 0.30 (less than 30% answer correctly)

The ideal difficulty level depends on the test’s purpose. Mastery tests designed to verify minimum competency may intentionally include easier items. Competitive examinations might feature more difficult items to distinguish top performers. However, most measurement experts recommend targeting moderate difficulty with indices between 0.30 and 0.70 for discriminating questions.

Understanding item discrimination

While difficulty tells us how many people got an item correct, discrimination reveals whether the right people got it correct. A well-designed test item should be answered correctly more often by high-performing students than by low-performing students. The discrimination index measures this differentiating power.

Calculating the discrimination index

The standard method identifies the upper and lower groups by selecting the top 27% and bottom 27% of test-takers based on total test scores. This percentage provides an optimal balance between group size and discriminating power. The discrimination index is then calculated by subtracting the proportion of the lower group who answered correctly from the proportion of the upper group who answered correctly.

For instance, if 90% of the upper group answers an item correctly and only 50% of the lower group does, the discrimination index would be 0.40 (90% – 50% = 40%). The discrimination index ranges from -1.0 to +1.0, with higher positive values indicating better discrimination.

Interpreting discrimination values

Discrimination indices are typically classified as follows:

Excellent discrimination: 0.40 or higher

Good discrimination: 0.30 to 0.39

Acceptable discrimination: 0.20 to 0.29

Poor discrimination: Below 0.20

Problematic: Negative values

A negative discrimination index signals a serious problem. It means more low-performing students answered correctly than high-performing students, suggesting the item may be confusing, mis-keyed, or poorly constructed. Items with negative discrimination should be examined immediately and typically revised or removed.

The relationship between difficulty and discrimination

Item difficulty and discrimination are interconnected. Very easy or very difficult items tend to have low discrimination because they don’t provide enough variation in responses. When almost everyone gets an item right or wrong, it becomes difficult to distinguish between high and low performers.

The maximum potential for discrimination occurs when an item has a difficulty index around 0.50. At this moderate difficulty level, there’s the greatest opportunity for the item to differentiate between students with different levels of knowledge. This creates an inverted U-shaped relationship: discrimination increases as difficulty approaches 0.50 from either extreme, then decreases as difficulty moves away from this optimal point.

However, not every item needs high discrimination. Essential knowledge items that all students should know might appropriately have low discrimination indices if most students answer correctly. The key is matching the item’s statistical properties to its intended purpose within the test.

Conducting item analysis

Item analysis is the systematic evaluation of test items using difficulty and discrimination indices. This process helps test developers make evidence-based decisions about which items to retain, revise, or discard.

Steps in item analysis

First, administer the test to a representative sample of examinees. Larger samples provide more stable estimates of item statistics. For classical difficulty values, samples of 100-200 test-takers typically yield reliable results.

Second, calculate the difficulty index for each item by determining the proportion of examinees who answered correctly. This provides initial insight into whether items are too easy or too difficult.

Third, calculate the discrimination index by comparing performance between upper and lower scoring groups. This reveals which items effectively distinguish between knowledgeable and less knowledgeable test-takers.

Fourth, analyze the results together. An item’s difficulty and discrimination should be considered simultaneously rather than in isolation. A difficult item with good discrimination may be valuable for identifying top performers, while an easy item with poor discrimination might need revision.

Making decisions based on indices

Test developers typically follow these guidelines when reviewing item analysis results:

Keep: Items with moderate difficulty (0.30-0.70) and good discrimination (0.30 or higher) should be retained as they’re performing well.

Revise: Items with appropriate difficulty but marginal discrimination (0.10-0.29) may benefit from refinement. Focus on clarifying ambiguous wording, improving distractors, or eliminating unintended clues.

Review closely: Very easy or difficult items should be examined even if discrimination is acceptable. Determine whether they serve a specific purpose or simply aren’t contributing to the test’s effectiveness.

Discard: Items with negative or very low discrimination (below 0.10) should typically be removed, regardless of difficulty level. These items aren’t helping measure what the test intends to measure.

Enhancing test validity and reliability

Systematic item analysis directly improves two critical test qualities: validity and reliability. Validity refers to whether a test measures what it claims to measure. By identifying and removing items with poor discrimination, test developers ensure the final assessment accurately measures the intended knowledge or skills rather than random guessing or test-taking strategies.

Reliability reflects consistency-whether a test produces stable results across different administrations or with similar groups. Items with good discrimination contribute to higher reliability because they measure a coherent underlying construct. When items effectively distinguish between those who know the material and those who don’t, the test as a whole becomes more reliable.

The item analysis process creates a feedback loop for continuous improvement. After each test administration, developers can examine item statistics, make informed revisions, and gradually build a high-quality item bank. This iterative approach transforms tests into increasingly effective measurement tools.

What do you think? How might regularly analyzing item difficulty and discrimination change the quality of assessments in your field? Have you ever encountered test questions that seemed to measure something other than knowledge of the subject matter?

How useful was this post?

Click on a star to rate it!

Average rating 4 / 5. Vote count: 1

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://www.questionmark.com/resources/blog/item-analysis-report-item-difficulty-index/
  2. https://www.washington.edu/assessment/scanning-scoring/scoring/reports/item-analysis/
  3. https://phoenixmed.arizona.edu/assessment/item-analysis
  4. https://maxinity.co.uk/blog/item-discrimination-index/

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Research Methodology

1 Selection of Research Problem

  1. Science and Characteristics of Scientific Knowledge
  2. Need for Scientific Methodology
  3. Identification of Research Problem
  4. Statement of the Problem and Objectives

2 Review of Literature

  1. Review of Literature: Sources and Classification
  2. Uses of Review of Literature
  3. Steps in Review of Literature
  4. Writing Review of Literature and Theoretical Orientation
  5. Citation
  6. Writing Bibliographical Details of a Reference

3 Concept and Variables, Formulation and Testing of Hypothesis

  1. Concept, Construct and Variables
  2. Types of Variables
  3. Hypothesis
  4. Types and Forms of Hypothesis
  5. Characteristics, Function and Testing of Hypothesis

4 Research Design

  1. Characteristics of Research Design
  2. Criteria of a Research Design
  3. Max-Min-Con Principle
  4. Classification of Research Design
  5. Experimental Research Design
  6. Descriptive Research Design

5 Descriptive and Survey Research Design

  1. Characteristics of Descriptive Research Design
  2. Steps in Descriptive Research
  3. Aims of Descriptive Research Design
  4. Types of Descriptive Research Design
  5. Case Studies
  6. Observational Studies
  7. Historical Studies
  8. Field Studies
  9. Diagnostic Studies
  10. Explorative Studies
  11. Longitudinal Studies
  12. Correlational Studies
  13. Cross-Sectional Studies
  14. Action Research
  15. Evaluation Research
  16. Survey Research

6 Experimental Research

  1. Testing of hypothesis
  2. t-test
  3. ฯ‡2-test
  4. F-test
  5. Principles of Experimental Designs
  6. Completely Randomised Designs
  7. Randomized Complete Block Design
  8. Latin Square Design
  9. Factorial Experiments
  10. 2n factorial experiment
  11. 3n factorial experiment

7 Levels of Measurement

  1. Concept of Measurement
  2. Postulates of Measurement
  3. Nominal Scale
  4. Ordinal Scale
  5. Interval Scale
  6. Ratio Scale

8 Knowledge Test Constructions

  1. Knowledge Test
  2. Characteristics of a Good Test
  3. Steps in Standardised Test Construction
  4. Item Analysis
  5. Writing Test Items
  6. Preliminary Administration
  7. Reliability of the Final Test
  8. Validity of the Final Test
  9. Norms of the Final Test
  10. Item Difficulty and Discrimination

9 Data Collection

  1. Secondary Data Sources
  2. Instruments Used for Collecting Primary Data
  3. Validity, Data Editing, and Coding
  4. Data Tabulation and Presentation

10 Sampling Technique

  1. Importance of Sampling
  2. Types of Sampling Techniques
  3. Probability based Sampling Techniques
  4. Non-Probability based Sampling Techniques
  5. Sample Size Determination
  6. Sampling and Non-Sampling Errors

11 Quantitative Techniques

  1. Frequency Distribution
  2. Measures of Central Tendency
  3. Measures of Dispersion
  4. Correlation
  5. Regression
  6. Multiple Regressions
  7. Dummy Variable Analysis
  8. Discriminant Function Analysis
  9. Factor Analysis
  10. Principal Component Analysis

12 Qualitative Techniques

  1. Observation Method
  2. Interview Method
  3. Questionnaire Method
  4. Case Study Method
  5. Projective Techniques

13 Statistical Analysis and Packages

  1. ฯ‡2- test
  2. t-test
  3. F-test
  4. Basic Experimental Designs
  5. Factorial Experiments
  6. Non-Parametric Tests
  7. Run Test
  8. Sign Test
  9. Wilcoxon Signed Rank Test
  10. Mann-Whitney U-Test
  11. Kruskal-Wallis One-way Analysis of Variance
  12. Friedman Two-way Analysis of Variance

14 Report Writing

  1. Research Report
  2. Steps in Preparing the Report: Preliminary Considerations
  3. Main Components of a Research Report
  4. Diagrammatic Presentation
  5. Common Weaknesses in Research Report Writing