Staged LLM Examination Question Generation for Quality Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing large language model (LLM)-enabled systems for generating examination questions face challenges such as overfitting to patterns, accuracy issues, ambiguity, relevance and alignment with examination objectives, lack of depth, difficulty in creating higher-order thinking questions, context errors, and loss of pedagogical intent, leading to inefficient use of computing resources.
Innovation Solution
A computer-implemented method involving multiple LLM prompts and user feedback to generate, refine, and store structured question units, ensuring adherence to examination criteria and standards, using a system that integrates LLMs with human expertise to enhance question generation efficiency and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If LLMs are used extensively to generate examination questions, then the quantity of questions increases, but the quality deteriorates due to overfitting to specific patterns and repetitive content
Solution Approach 1:
The question generation process is segmented into multiple stages: initial question generation, fact-checking verification, ambiguity detection, and human review. Each stage independently processes specific aspects of question quality, allowing the system to maintain high question quantity while filtering out low-quality outputs through sequential validation layers.
Solution Approach 2:
The system implements feedback mechanisms where generated questions are evaluated by fact-checking modules and human reviewers, with results fed back into the generation process. This iterative feedback loop enables continuous improvement of question quality while maintaining high output volume through automated filtering and refinement.
2Productivity
If LLMs generate examination questions automatically, then productivity increases, but accuracy decreases due to potential incorrect information and lack of fact-checking
Solution Approach 1:
Fact-checking modules and verification systems serve as intermediary components between the LLM generation process and final question output. These intermediaries independently verify the accuracy of generated questions, allowing high-speed automated generation while maintaining information accuracy through dedicated verification layers.
Solution Approach 2:
The system performs preliminary fact-checking and verification actions on generated questions before they are finalized. This preliminary verification step prevents incorrect information from entering the final question bank, enabling high productivity while ensuring accuracy through advance validation.
3Productivity
If LLMs generate examination questions without human review, then efficiency improves, but quality deteriorates due to ambiguity and poor phrasing
Solution Approach 1:
The system implements self-service quality control where the LLM generates questions that are then automatically evaluated by detection modules for ambiguity and poor phrasing. The system serves its own quality assurance needs through automated self-evaluation, maintaining high efficiency while filtering out problematic questions through built-in detection mechanisms.
Solution Approach 2:
Ambiguity detection modules provide feedback on question clarity and phrasing quality, identifying problematic questions and sending them for human review or regeneration. This feedback loop maintains high generation efficiency while ensuring question quality through automated monitoring and selective human intervention.
4Device complexity
If LLMs generate higher-order thinking questions, then the complexity of questions increases, but the capability to create such questions decreases
Solution Approach 1:
The question generation capability is segmented into different modules: basic question generation, fact-checking, ambiguity detection, and higher-order thinking verification. This segmentation allows the system to handle complex higher-order questions through specialized verification modules while maintaining the core generation capability through automated processing.
Solution Approach 2:
Specialized verification modules act as intermediaries that specifically evaluate higher-order thinking aspects of questions. These intermediaries provide targeted feedback on complex reasoning, analysis, and synthesis capabilities, enabling the system to generate and verify high-complexity questions through dedicated verification layers.
5Reliability
If human experts review all LLM-generated questions, then quality improves, but time consumption increases
Solution Approach 1:
Human review is applied locally only to questions that fail automated quality checks or are flagged for potential issues. Most questions pass automated fact-checking and ambiguity detection without human intervention, maintaining high quality through selective human review only where necessary, thus reducing overall time consumption while preserving quality.
Solution Approach 2:
Automated quality control systems provide continuous feedback on question quality, allowing human reviewers to focus only on problematic cases. This feedback mechanism optimizes the balance between automated processing speed and human quality assurance, minimizing time loss by directing human review only to questions that require it.
Data Source
AI summary
Methods, devices, and processor-readable media for automated generation of examination questions. A first large language model (LLM) prompt is generated by inserting a first set of one or more parameters into a first LLM prompt template, the first LLM prompt including instructions to provide a second set of one or more parameters pertaining to a question scenario. A first prompt response for the first LLM prompt includes the second set of one or more parameters. A second LLM prompt is generated by inserting the second set of one or more parameters into a second LLM prompt template, the second LLM prompt including instructions to provide structured question unit content based on the second set of one or more parameters. A second prompt response is received for the second LLM prompt, the second prompt response including the structured question unit content.


