NLP Quality Assurance for Situational Judgement Test Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing assessment systems for constructed-response tests, such as situational judgement tests (SJTs), face challenges in ensuring the validity of scores due to the reliance on human judgment for rating and scoring, which can be subjective and inconsistent.

Innovation Solution

A distributed computer assessment system that employs natural language processing (NLP) for quality assurance, using generative models and large language models to generate predicted rating data and feedback, thereby automating the assessment and quality assurance process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If human judgment is used for rating and scoring constructed-response tests, then the assessment can be performed with current technology, but the scoring becomes subjective and inconsistent

Engineering Contradiction:
Improvescoring consistencyVSAvoidautomation level
Core Design Contradiction:
ReliabilityVSExtent of automation

Solution Approach 1:

The patent replaces the mechanical human judgment system with an automated NLP-based system. The NLP engine automatically processes constructed-response test answers, applies consistency rules, and generates scores without human intervention, thereby eliminating subjectivity and improving scoring consistency while increasing automation extent.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-service by automatically rating its own outputs through the NLP engine. The consistency rules enable the system to self-evaluate and self-correct, ensuring uniform scoring standards are applied without external human influence, thus improving reliability while maintaining high automation.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If NLP and generative models are used to automate assessment, then scoring consistency and accuracy improve, but system complexity increases

Engineering Contradiction:
Improvescoring accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the assessment system into distinct functional modules: an NLP engine for processing responses, a consistency rules module for quality assurance, and a scoring module for generating results. This segmentation manages system complexity by organizing complex NLP and generative model components into separate, manageable units that can be independently optimized and maintained.

Inventive Principle:
Principle #1Segmentation

3Device complexity

If human raters are used for constructed-response tests, then the system is simpler to implement, but quality assurance and validation of scores become difficult

Engineering Contradiction:
Improvesystem simplicityVSAvoidquality assurance
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent implements feedback mechanisms where the NLP engine provides real-time quality assurance feedback on scored responses. The consistency rules generate validation feedback that can be reviewed and adjusted, creating a feedback loop that improves reliability while keeping the system manageable through automated monitoring and control.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250036880A1System and method for natural language processing for quality assurance for constructed-response tests
Publication Date: 2025.01.30 ACUITY INSIGHTS INC
  • US20250036880A1 patent drawing
  • US20250036880A1 patent drawing
  • US20250036880A1 patent drawing

AI summary

Embodiments described herein provide systems and processes for natural language processing for quality assurance and response assessment of a situational judgement test. For example, system can use natural language processing engine for sentiment analysis and unsupervised text classification to automatically score or rate response data. The system can provide quality assurance for rating data by generating predicted scorings or ratings that can be compared to human rating data.