Scoring Model Calibration Using External Proficiency Measures
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional automated scoring models for constructed responses are biased towards response length and may disadvantage certain populations, as they assume reliable human-assigned scores and are susceptible to being 'gamed' by lengthening responses without adding substance, leading to unfair scoring outcomes.
Innovation Solution
A computer-implemented method for calibrating a scoring model that analyzes training responses to derive feature values and uses external measures of proficiency not derived from the training responses to determine weights for the scoring model, reducing the influence of response length and improving fairness across demographic groups.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional automated scoring models use human-assigned scores for training, then the model can be trained with available data, but the model becomes biased towards response length and unreliable
Solution Approach 1:
The patent introduces an intermediary variable (response quality metrics such as vocabulary diversity, grammatical complexity, and content coherence) that mediates between the raw response text and the final score. This intermediary layer filters out the spurious correlation with response length while preserving the meaningful relationship between response quality and proficiency, thereby resolving the contradiction between reliability and measurement precision.
Solution Approach 2:
The patent transforms the scoring parameters by replacing the direct use of response length as a feature with normalized quality metrics. Specifically, it changes the parameter representation from raw text characteristics (including length) to standardized quality indicators that are independent of response length, thus eliminating the bias while maintaining scoring accuracy.
2Device complexity
If the scoring model weights response length heavily, then the model can be simple to implement, but the model becomes susceptible to being 'gamed' and produces unfair scores
Solution Approach 1:
The patent segments the response evaluation into multiple independent dimensions (vocabulary diversity, grammatical complexity, content coherence, and response length) rather than relying on a single aggregate metric. By segmenting the scoring process, the model can independently evaluate each dimension and combine them weighted towards quality metrics, preventing examinees from gaming the system by simply lengthening responses while maintaining manageable model complexity.
3Ease of manufacture
If the scoring model is trained using traditional methods, then the training process is straightforward, but the model unfairly disadvantages certain populations
Solution Approach 1:
The patent applies preliminary normalization and standardization to the training data before model training, adjusting for demographic variables and response length effects in advance. This preliminary action ensures that the model learns from pre-processed data that already accounts for demographic differences, making the training process remain straightforward while significantly improving demographic fairness and model adaptability.
Data Source
AI summary
Systems and methods are described for generating a scoring model for responses. A computer-implemented method of calibrating a scoring model using a processing system for scoring examinee responses includes accessing a plurality of training responses for training the scoring model. The plurality of training responses are analyzed to derive values of multiple features (variables) of the training responses. The scoring model is trained based on the values of the multiple features of the training responses and one or more external measures of proficiency for each individual associated with a training response utilized in the training. The one or more external measures are not derived from the training responses. Based on the training, a weight for each of the multiple features is determined. The scoring model is calibrated to include the weights for at least some of the features for scoring examinee responses.


