Automated Scoring Engine for Integrated Reasoning Test Items
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Standardized tests face challenges in accurately measuring new academic skills due to evolving educational needs, requiring continuous research and feedback loops in test development, and scoring processes that can be influenced by various factors, leading to inconsistencies in assessing examinees' proficiency.
Innovation Solution
Automated methods and systems for presenting and scoring items in different question formats, including graphics interpretation, two-part analysis, table analysis, and multi-source reasoning, where each response opportunity is independent, and only full credit is assigned if all responses are correct, using a processor-controlled user interface and scoring engine.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If automated scoring is implemented for multiple response opportunities, then scoring consistency and reliability are improved, but system complexity increases
Solution Approach 1:
The test item is divided into multiple independent response opportunities, each scored separately. This segmentation allows the scoring system to evaluate each component independently, ensuring consistency while managing complexity through modular processing of discrete response elements rather than holistic evaluation
Solution Approach 2:
The scoring system uses electronic copying and digital representation of responses and scoring rules, replacing manual scoring processes. This digital replication ensures identical scoring criteria are applied uniformly across all examinees, improving reliability while the automation of this copying process manages the inherent system complexity
2Measurement precision
If all response opportunities must be correct to assign credit, then measurement precision is improved, but the difficulty of the test increases
Solution Approach 1:
Different scoring weights or credit assignments are applied to different response opportunities based on their local importance. Critical components that must be correct receive full weight, while supplementary components may have different weighting, allowing precise measurement of skill mastery without uniformly increasing test difficulty across all elements
Solution Approach 2:
The scoring system allows flexibility in adjusting the credit assignment parameters. Test administrators can modify whether all responses must be correct or if partial credit is awarded based on the number of correct responses, enabling adaptation of test difficulty while maintaining measurement precision through configurable scoring thresholds
3Ease of operation
If independent response opportunities are used, then ease of operation is improved, but loss of information occurs
Solution Approach 1:
The test interface serves as an intermediary that presents response opportunities in independent, manageable units while maintaining the contextual framework of the overall question. Each response opportunity is clearly delineated with instructions that provide necessary context, allowing examinees to operate independently on each component without losing understanding of the broader task requirements
4Reliability
If extensive field tests are conducted, then reliability of test development is improved, but loss of time increases
Solution Approach 1:
Extensive field testing and validation are conducted during the preliminary test development phase before the test is deployed for actual use. This preliminary action ensures that the automated scoring system and question formats are thoroughly validated in advance, improving reliability while minimizing time loss during actual test administration and scoring operations
Data Source
AI summary
Automated methods and systems are provided for presenting and scoring items which are presented in different question formats that include graphics interpretation, two-part analysis, table analysis, and multi-source reasoning. There are a plurality of response opportunities for each item, and each response opportunity is independent of each other. In operation, an item is presented in the respective question format on a processor-controlled user interface display screen. A response engine receives the response for each response opportunity. A scoring engine in communication with the response engine receives the plurality of responses and scores the plurality of responses by determining whether the response selected for each response opportunity is correct, and assigning credit for the item only if each of the responses for each of the response opportunities is correct, and assigning no credit for the item if at least one of the responses for each of the response opportunities is incorrect.


