Subjective question scoring and evaluation system based on SOLO classification
The subjective question scoring system, which combines SOLO classification and a large language model, solves the problem that existing technologies cannot accurately assess students' thinking structure and teaching feedback, and achieves precise assessment of students' thinking level and teaching support.
Patent Information
- Application Number
- CN202511101912.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-14
AI Technical Summary
Existing subjective question scoring methods cannot accurately reflect students' overall thinking structure and logic, ignore the diverse expression characteristics of open-ended questions, lack teaching feedback functions, and are difficult to replace human marking.
A subjective question scoring and evaluation system based on SOLO classification is adopted, which combines a large language model and prompt word engineering. Through image acquisition, text extraction, intelligent grading and learning analysis, structured scoring results are generated and teaching feedback is provided.
It enables precise assessment of students' thinking structure, breaks through the limitations of fixed scoring dimensions, provides teaching assistance functions, and improves scoring consistency and teaching support capabilities.
Smart Images

Figure CN120954016A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine-based marking, specifically to a subjective question scoring and evaluation system based on SOLO classification. Background Technology
[0002] With technological advancements, automated grading systems are replacing traditional manual grading, becoming a future trend in the education field. For objective questions, optical character recognition (OCR) technology and matching with standard answers can relatively easily solve this task. However, accurately grading subjective questions remains a significant technical challenge.
[0003] Current techniques for grading subjective questions mainly fall into the following categories:
[0004] (1) Rule-based and template-based methods match answer keywords or semantic patterns through preset rules. Key knowledge points and terms are extracted from the reference answers, and a set of keywords to be included in the answers is set. Scoring is done by weighting the frequency and position of the keywords. This method is suitable for questions with a high degree of answer standardization (such as short answer questions and concept questions). The limitation of this type of method is its poor flexibility and difficulty in handling semantic variations.
[0005] (2) Feature engineering methods based on machine learning extract textual features from answers and train classification / regression models. Using bag-of-words models, syntactic and semantic features, the similarity between the answer and the reference answer and the completeness of the answer's logic are measured. Supervised learning models such as Support Vector Machines (SVM) and Random Forests are used, with manually labeled scores as labels, to train the model to score new answers. The limitation of this type of method is its reliance on manually designed features and limited generalization ability.
[0006] (3) Semantic matching schemes based on deep learning learn text semantic representations through neural networks and measure the semantic similarity between the answer and the reference answer. A model pre-trained on a large-scale corpus is used to extract the semantic vector of the answer, capturing the semantic relationships within the context. The answer and the reference answer are input into an encoder with the same structure (such as Transformer) to generate semantic vectors, and then the cosine similarity is calculated as the scoring criterion. Semantic matching schemes can be divided into word-level matching (focusing on keyword overlap), sentence-level matching (analyzing sentence semantic alignment), and document-level matching (integrating the logical coherence of the entire text). The limitation of this type of method is that it requires a large amount of labeled data and has weak interpretability.
[0007] (4) Multimodal fusion and reinforcement learning scheme: Combining text (answers) and image (formulas, charts) inputs, the scheme uses a cross-modal model (such as CLIP) to comprehensively score the questions. This is suitable for subjective questions containing formula derivations. In the reinforcement learning scheme, the scoring process is regarded as a decision problem. The model parameters are adjusted through reinforcement learning, and the error between human scoring and model scoring is used as a reward signal to improve scoring consistency.
[0008] (5) Combination of rules and deep learning: First, use rules to extract keywords for initial screening, and then use deep learning models to refine semantic scoring, balancing efficiency and accuracy. The bottom layer uses word vectors to match basic knowledge points, and the top layer uses a large model to analyze the logical depth, and finally generates a weighted total score.
[0009] Existing technology 1: CN118861522A, a subjective question paper marking method based on a large model.
[0010] Existing technology 1 combines deep learning semantic matching with large-scale model technology. The main approach involves: collecting and preprocessing subjective question answer data from an existing exam database to construct a high-quality dataset; developing detailed scoring criteria and manually annotating answers based on accuracy, logic, and language expression to generate weighted average score labels; selecting a large-scale model based on the Transformer architecture, leveraging its semantic representation capabilities, and incorporating modules such as multi-head attention to construct a grading model; and designing a bidirectional semantic matching calculation method, using the similarity between the answer and the reference answer as the scoring criterion.
[0011] Existing methods for grading subjective questions still face technical bottlenecks in the following aspects:
[0012] (1) Rigid evaluation dimensions and neglect of overall thinking: Traditional “point-based” scoring relies solely on keyword or key point matching, ignoring the overall semantics, reasoning logic and knowledge structure of students’ answers. It cannot truly reflect their analytical, inductive and comprehensive thinking abilities, and is out of touch with the educational orientation of “ability-oriented” and “core competencies”.
[0013] Furthermore, existing evaluation technologies have fixed evaluation dimensions (such as accuracy and logic) and cannot be dynamically adjusted according to the content of different questions (especially open-ended questions), resulting in insufficient flexibility. Open-ended questions often do not have a single standard answer; students can approach them from multiple perspectives, using diverse language organization and reasoning paths, leading to highly free and varied expressions. In the current manual marking model, markers need to spend a significant amount of time reading the entire text, analyzing the context and logic, and deciphering the reasoning chain. This is not only a heavy workload but also prone to inconsistencies in the application of scoring standards due to subjective misunderstandings. Previous automated marking schemes relied heavily on keyword matching or template comparison mechanisms, which are ill-suited for handling open-ended questions with flexible structures and diverse expressions, thus failing to truly replace manual marking in educational practice.
[0014] (2) Lack of evaluation of thinking level and superficial analysis of learning situation: The scoring criteria do not involve the evaluation of students' thinking structure level, and cannot reflect their thinking structure, cognitive process and ability level (such as understanding, reasoning and transfer ability) in depth through the scoring results, and lack systematic analysis of learning situation.
[0015] (3) Lack of teaching support and limited functions: Existing technology is limited to pursuing more accurate scores. It cannot automatically summarize typical problems based on the whole class's answer results, nor can it provide teachers with personalized review suggestions. It lacks teaching feedback functions and is difficult to support teaching improvement.
[0016] Based on the above situation, in the field of machine grading of subjective questions, the inventor believes it is necessary to construct a new technical solution for grading and evaluation that can reflect the level of students' thinking structure, flexibly break through the limitations of fixed scoring dimensions, and provide feedback and support for the teaching process. Summary of the Invention
[0017] To alleviate or partially alleviate the above-mentioned technical problems, the solution of the present invention is as follows:
[0018] A subjective question scoring and evaluation system based on SOLO classification, comprising:
[0019] The answer sheet acquisition module uses image acquisition equipment to collect student exam papers in a standardized manner, generate exam paper images, and establish a scanned image database.
[0020] The front-end processing module, located in the front-end system, is used to perform image preprocessing operations on the collected test paper images, including test paper positioning, extracting answer area sub-images, completing text extraction through the optical character recognition module, converting students' answer text into editable text and associating metadata for building a structured answer database;
[0021] The assessment benchmark configuration module includes a standardized interface for supporting the import of information, including test questions, and extracts structured data, including question background, question stem content and SOLO level description information, and SOLO level-based evaluation scoring label templates, and stores them in the database.
[0022] The review and SOLO level determination module is located in the backend system and includes a large language model. The large language model receives the answer text from the answer database from the frontend system according to a preset strategy. The large language model extracts the elements, logical relationships between elements and reasoning depth information from the answer text and outputs structured intermediate parsing results.
[0023] The large oracle model is a large language model that has undergone targeted fine-tuning. The dataset used in the targeted fine-tuning process includes at least the following fields: question information and ability requirements, student answers, SOLO level labels, and inference chain.
[0024] Based on the intermediate parsing results and SOLO level description information, the large language model generates an evaluation of the student's thinking structure level; based on the intermediate parsing results and the SOLO level-based opinion scoring labeling template, the language model generates a score for the student's answer text.
[0025] The learning analysis and feedback module analyzes the evaluation results of all students' answers using a large language model and outputs a learning analysis report, which includes a visual data dashboard of the overall learning effect of the class and teaching suggestions.
[0026] The data storage and interaction module is based on a front-end and back-end separation architecture and is deployed on a cloud computing server. It consists of an interaction layer, a service layer, a data layer, and external cloud services working together to realize data storage, interaction, and the operation of system functions.
[0027] Furthermore, the visualization data dashboard for the overall learning effectiveness of the class includes the distribution of thinking levels, common thinking deficiencies, and lists of students in different SOLO tiers.
[0028] Furthermore, during the fine-tuning process, the large language model also incorporates a geographical knowledge system, specifically including: a geographical terminology database, a geospatial relationship map, and a database of typical regional cases.
[0029] Furthermore, the question information and ability requirement field includes a description of the question scenario, the question content, and ability requirement tags;
[0030] The SOLO grading label field includes grading labels for each answer based on SOLO theory by geography education experts, as well as specific criteria for judgment;
[0031] The reasoning chain field contains a manually broken-down thought process for each answer, forming interpretable intermediate reasoning steps.
[0032] Furthermore, the extraction of elements, logical relationships between elements, and reasoning depth information from the response text specifically includes:
[0033] At the element extraction level, the geographical elements involved in the response text and the completeness of the elements are identified;
[0034] At the level of identifying logical relationships between elements, the association forms and clarity of expression between elements are judged, including causal relationships, constraint relationships, and synergistic relationships.
[0035] At the level of reasoning depth assessment, analyze whether the answer reflects inductive generalization, knowledge transfer, or critical thinking.
[0036] Furthermore, the image preprocessing operation includes:
[0037] The test paper is located by edge detection and perspective correction, and the answer area sub-image is extracted by threshold segmentation technology to exclude irrelevant areas such as headers, footers and binding lines.
[0038] Subsequently, the optical character recognition module extracts the text, converting each student's answer into editable text and simultaneously linking metadata such as student ID, class, question number, and answer area coordinates to build a structured answer database.
[0039] Furthermore, the answer acquisition module and the front-end processing module are integrated into a handheld smart device, which can capture the test paper in real time by calling the device's camera or read locally stored image files; or, the answer acquisition module and the front-end processing module are located in different electronic devices.
[0040] Furthermore, the backend system includes an interaction layer, a service layer, a data layer, and authentication / authorization; wherein,
[0041] The interaction layer is built on the Vue3 front-end and provides functions such as exam paper uploading.
[0042] The service layer is built on the Gin framework and uses SQLX to perform database operations; CORS cross-domain resource sharing is configured to ensure cross-domain communication between the front-end and back-end.
[0043] The data layer relies on MySQL, Redis, and MinIO for data storage and optimization.
[0044] Furthermore, the answer acquisition module and the front-end processing module are located in different electronic devices, wherein the front-end processing module is located in a personal computer.
[0045] Furthermore, the SOLO-based subjective question scoring and evaluation system is also configured for automatic grading of open-ended questions.
[0046] The technical solution of this invention has one or more of the following beneficial technical effects:
[0047] The technical advantages of this invention can be summarized in the following five aspects, with the core being the automation of the entire process of "score assessment" for subjective questions in geography through innovative technology, filling several technological gaps:
[0048] (1) Breakthrough in precise assessment and interpretability of thinking levels. For the first time, SOLO taxonomy theory is applied to machine scoring of subjective questions in middle school geography. Relying on a large language model, it achieves whole-segment semantic analysis. Through multi-round prompt word engineering (including multi-step reasoning), the model is guided to analyze elements, logical connections and transferability step by step, accurately matching the SOLO classification framework to complete the classification and scoring. The innovative design of traceable analysis reports clearly marks the scoring basis, which not only replaces the traditional "point collection" mode to realize the visualization of the thinking process, but also retains the space for manual verification, significantly improving the scientificity and consistency of scoring.
[0049] (2) Targeted optimization and flexible adaptation of large language model. By training the model through targeted fine-tuning scheme, a multi-dimensional mapping of "keyword-logic-thinking level" is established. Combined with the configuration function of teacher-customized reference answers and scoring dimensions, the limitations of fixed scoring dimensions and insufficient generalization ability in the existing technology are overcome, and the scoring needs of different question types, teaching objectives and complex open-ended test questions are adapted.
[0050] (3) Enhanced full-process automation and scenario adaptation. By integrating OCR handwriting recognition with large-scale model scoring processes, a closed-loop automated system of "image acquisition - text extraction - manual verification - intelligent evaluation - feedback generation" is constructed, which can be efficiently adapted to offline question-answering scenarios. The introduction of prompt word engineering further reduces the judgment bias of complex answers and improves the stability of automated processing.
[0051] (4) Deepening and Precise Support of Teaching Auxiliary Functions. Automatically generate multi-dimensional learning data such as class score distribution, high-frequency errors, quantitative results of thinking level, and lists of high and low score segments. Present common problems through a visual dashboard, providing direct basis for teachers to give precise comments and expanding the teaching support value of the scoring system.
[0052] Furthermore, other beneficial effects of the present invention will be mentioned in the specific embodiments. Attached Figure Description
[0053] Figure 1 This is a schematic diagram of an exemplary embodiment of the present invention;
[0054] Figure 2 This is a flowchart of a complete scoring process according to the present invention;
[0055] Figure 3 This is the system's core functional architecture diagram;
[0056] Figure 4 This is a technical implementation architecture diagram of one embodiment of the present invention;
[0057] Figure 5 This is a block diagram of the subjective question scoring and evaluation system based on SOLO classification of the present invention;
[0058] Figure 6This is a diagram showing the evaluation and scoring results of students' answers in a certain embodiment;
[0059] Figure 7 This is a diagram showing suggestions for reviewing class test papers in one embodiment. Detailed Implementation
[0060] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0061] To facilitate a clear description of the technical solutions in the embodiments of the present invention, the terms "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order.
[0062] The SOLO Taxonomy, proposed in 1982 by educational psychologists John B. Biggs and Kevin F. Collis, stands for "Structure of the Observed Learning Outcome." Its core idea is that a student's understanding of knowledge can be judged by the "complexity of their thought process" when answering questions, rather than simply focusing on whether the answer is right or wrong.
[0063] Based on Jean Piaget's theory of cognitive development stages, this theory divides learning outcomes into five levels from low to high, forming a quantifiable framework for evaluating cognitive structures. The SOLO taxonomy breaks through the limitations of traditional "right or wrong" scoring, approaching the issue from the perspective of the complexity of cognitive structures, providing an educational assessment framework that combines theoretical depth with practical applicability.
[0064] Table 1: Different Levels of Thinking Structure as Demonstrated by the SOLO Taxonomy
[0065]
[0066] Although the SOLO Taxonomy is common knowledge in education and widely used in teaching evaluation, its application in frontline teaching is limited by the speed of manual evaluation by teachers. Furthermore, there has been no attempt by those skilled in the art to apply this theory to machine grading of subjective questions. This invention addresses the aforementioned bottlenecks in existing technologies by proposing a novel scoring and evaluation solution based on a Large Language Model (LLM), the SOLO Taxonomy, cue word engineering, and the invention's specific targeted fine-tuning scheme. This solution reflects students' cognitive structure, flexibly overcomes the limitations of fixed scoring dimensions, and provides feedback and support for the teaching process. This invention is applicable to the evaluation and scoring of subjective questions, especially in geography, and innovatively solves the long-standing pain points in the evaluation and scoring of open-ended questions in secondary school geography.
[0067] This invention will first introduce the overall operating logic of the subjective question scoring and evaluation system based on the SOLO classification method, and then gradually introduce the specific implementation methods of the modules that implement each core function.
[0068] This invention can be applied to the machine evaluation and scoring of subjective questions, especially open-ended subjective questions. The following text will use geography subjective questions as an example to illustrate the specific implementation of this invention. This invention is not limited to this and is not restricted to it.
[0069] Figure 1 A schematic diagram of an exemplary embodiment of the present invention is shown. In this embodiment, the answer sheet acquisition terminal / module uses a high-definition image scanning device (such as a high-speed scanner) to standardize the acquisition of student exam papers, generate high-definition grayscale images, and establish a scanned image database, thus providing the basic support for standardized storage and subsequent processing of exam paper images. It should be noted that this is only a typical implementation of the invention, and in practical applications, other image acquisition methods (such as document scanners, industrial cameras, etc.) can be used depending on the needs of the scenario. The scope of protection of the present invention is not limited to specific acquisition devices. In this type of embodiment, the answer acquisition module and the front-end processing module are located in different electronic devices, such as the front-end processing module being located in a personal computer.
[0070] Figure 2The flowchart presents a complete scoring process of this invention. In the front-end processing stage, for the scanned test paper image, the system first performs image preprocessing operations: the test paper is located through edge detection and perspective correction, and the answer area sub-image is extracted using threshold segmentation technology (excluding irrelevant areas such as headers, footers, and binding lines); then, the text is extracted through the Optical Character Recognition (OCR) module, converting each student's answer into editable text, and synchronously linking metadata such as student ID, class, question number, and answer area coordinates to build a structured answer database.
[0071] As an alternative implementation, this invention supports integrating the answer acquisition terminal / module with the front-end processing module into a handheld smart device (such as a smartphone or tablet): by calling the device's camera to capture the test paper in real time (supporting automatic edge cropping and shadow removal) or reading locally stored image files, and combining a lightweight image segmentation algorithm with an offline OCR engine to complete the extraction of the answer text, significantly improving the flexibility of the acquisition scenario (such as in-class quizzes, homework assignments, and other non-centralized examination scenarios). In particular, this mode supports the function of instant single-question review—after the teacher takes a picture of a specific question through the device, the system can skip the batch processing process and directly trigger the review logic, achieving a "shoot-recognition-evaluation" response in seconds.
[0072] The front-end system further provides standardized interfaces to support the import of test questions and answer keys in multiple formats (such as document upload, online editing, and question bank integration). After teachers upload test papers and answer key files, the system uses regular expression parsing and natural language processing technology to accurately extract structured data such as question stem information, assessment objectives, score weights, and key points of the answer key, and stores them in the database, providing a benchmark for subsequent evaluation. Simultaneously, the front-end locally stores a dynamically updated scoring standard library, which can transmit scoring details (such as score point weights, deduction rules, and special case explanations) to the back-end Large Language Model (LLM) in real time according to evaluation needs.
[0073] The OCR function of this invention adopts a distributed deployment architecture: the front-end deployment is suitable for low-latency scenarios (such as local fast preview) and can complete basic text recognition with the help of device computing power; the back-end deployment relies on the computing power of the server cluster and combines deep learning models to optimize the recognition accuracy of complex scenarios (such as handwritten cursive, mixed formulas, and mixed Chinese and English). In practical applications, it is compatible with third-party OCR services (such as calling Tencent Cloud OCR handwriting recognition model to process students' handwritten answers), and balances efficiency and accuracy through an adaptive switching mechanism. Before the text is officially entered into the database, the system will display the OCR recognition results to users (teachers, graders, etc.) in a comparison view (original image and recognized text side by side), and supports manual correction of recognition errors (such as misjudged rare characters and punctuation marks). The manual verification process effectively avoids the chain reaction of recognition deviations on subsequent evaluations.
[0074] In the formal review phase, the front-end system pushes records from the answer database to the back-end LLM according to a preset strategy (such as by class batch or by question type). Based on its built-in subject knowledge graph, the LLM combines question information from the question bank, teacher-uploaded reference answers, and scoring criteria to perform in-depth analysis of the answer text (including semantic understanding, logical chain reconstruction, and scoring point matching). It generates feedback text containing specific evaluation comments (such as "the argument is logically complete but the evidence is insufficient") and a quantitative score, which is then sent back to the front-end in real time. The front-end associates and stores the evaluation results with the original answer data, forming a traceable "student-question-evaluation" three-dimensional dataset, completing a single scoring process.
[0075] Figure 3 This is a diagram of the system's core functional architecture. This invention utilizes LLM to deeply mine all answer data, outputting a multi-dimensional learning analysis report: on one hand, it generates visual data dashboards (such as score distribution histograms, high-frequency error word clouds, and knowledge point mastery heatmaps), intuitively presenting the overall learning effect of the class and simultaneously providing lists of students with high and low scores; on the other hand, it outputs an evaluation of individual students' thinking levels (quantified from five cognitive dimensions: "memory-understanding-application-analysis-creation"), and automatically summarizes common problems in the class's answers (such as typical misunderstandings of a knowledge point and common flaws in the reasoning process).
[0076] The above functions together construct a complete closed loop of "data collection - intelligent evaluation - learning diagnosis - teaching optimization", which not only provides teachers with accurate evaluation basis (such as focusing on high-frequency error points to design key points for explanation), but also supports the formulation of personalized teaching strategies (such as assigning differentiated exercises for students with different thinking levels), and ultimately realizes the effective transformation of evaluation results into teaching improvement.
[0077] Figure 4 This is a technical implementation architecture diagram of one embodiment of the present invention. The system of the present invention is based on a front-end / back-end separation architecture, deployed on a lightweight cloud computing server, and consists of an interaction layer, a service layer, a data layer, and external cloud services working together to achieve automatic scoring and learning feedback for subjective questions. The architecture and module collaboration logic are as follows:
[0078] Deployed entirely on a lightweight cloud server, it adopts a front-end and back-end separation model. Through the collaboration of the interaction layer, service layer, data layer and external cloud services, it completes the core functions of automatic scoring and learning feedback for subjective questions in geography.
[0079] This invention adopts a layered modular design, including an interaction layer, a service layer, a data layer, and authentication and authorization layers.
[0080] In the interaction layer, the front-end interface is built based on Vue3, providing user interaction functions such as exam paper uploading, score viewing, and feedback query, ensuring convenient operation and a user-friendly interface.
[0081] In the service layer, the system is built on the Gin framework and uses SQLX to perform database operations. CORS cross-domain resource sharing is configured to ensure the flexibility and security of cross-domain communication between the front-end and back-end, supporting the operation of core business logic.
[0082] Furthermore, this invention can also complete user authentication using JWT tokens.
[0083] Furthermore, at the data layer, structured data storage is provided: implemented based on a MySQL database, ensuring the standardized management of data such as scoring rules and user information.
[0084] Furthermore, Redis provides data caching to optimize response efficiency in high-frequency access scenarios.
[0085] Furthermore, MinIO object storage is used to store uploaded answer images, scoring results, and other files, adapting to diverse data needs, especially unstructured data.
[0086] In addition, during the external cloud service and model call process, cloud computing OCR handwriting recognition APIs, such as Tencent Cloud, can be called to extract students' answers and solve the problem of text input for subjective questions.
[0087] Furthermore, the Tongyi Qianwen-Plus large language model can be used as an example to understand and score subjective questions, scoring criteria, and student answers, thereby achieving intelligent evaluation with the help of a domestically developed large model.
[0088] This invention can integrate domestic large-scale models (example: Tongyi Qianwen-Plus) with cloud service capabilities (cloud computing, MinIO, etc.), ensuring the feasibility of actual deployment and adapting to the needs of educational scenarios, effectively improving the efficiency and intelligence level of geography subjective question evaluation, and achieving deep synergy between "technology + scenario".
[0089] As the core of this invention, the evaluation and scoring method for machine-based grading applies SOLO taxonomy theory in two places: First, in the process of evaluating students' answers, it provides insights into the students' thought structure as reflected in their answers, an effect that traditional technologies, which generally focus only on the accuracy of scoring, lack. Second, in the intentional scoring process, it replaces the traditional point-based scoring scheme, focusing on the logic and structure of students' answers rather than merely keyword coverage. The following sections will detail the specific application of SOLO taxonomy theory in the technical solution for grading subjective questions.
[0090] To address the subject-specific adaptability requirements for subjective question scoring, it is necessary to first systematically break down the frequently tested ability dimensions of subjective questions in the subject based on the core competencies and examination syllabus requirements of the subject, then define the scoring standards and typical answer characteristics for different thinking levels under each ability dimension, and finally construct a structured graded scoring template library.
[0091] Taking geography as an example, its subjective questions often focus on four core competency dimensions: a) Geographic location positioning ability (such as accurately interpreting a region's latitude and longitude, its location relative to land and sea, and its relative location relationships); b) Geographic feature summarization ability (such as extracting the spatiotemporal distribution patterns of elements such as climate, topography, and hydrology); c) Geographic causal analysis ability (such as analyzing the formation mechanisms and influencing factors of natural or human phenomena); d) Overall regional cognition ability (such as comprehensively judging the relationship between various elements in a region and the logic of coordinated human-land development). Based on these competency dimensions, the graded scoring template library needs to further clarify the corresponding SOLO thinking levels (pre-structure, single-point structure, multi-point structure, relational structure, and extended abstraction) for each competency dimension, and match specific scoring criteria (such as completeness of key points, logical rigor, and depth of transfer) and typical answer characteristics (such as conceptual confusion at the pre-structure level, and cross-regional pattern transfer at the extended abstraction level), forming a four-dimensional correspondence system of "competency dimension - thinking level - scoring criteria - typical characteristics," providing a structured basis for subsequent automated scoring.
[0092] To enable the Large Language Model (LLM) to have scoring capabilities based on SOLO classification theory (since it does not have built-in subject-specific graded scoring rules), a multi-round progressive prompt word engineering needs to be designed to achieve a logical closed loop from "understanding the standard" to "analyzing the answer" and then to "decision scoring" through layered guidance.
[0093] (1) Contextual prompts: As a "benchmark anchoring step" for scoring, this invention inputs a complete question background (such as regional geographical environment and problem situation description), question stem content (including question orientation and limiting conditions), and a detailed description of the corresponding SOLO level for the question (including the thinking characteristics, ability requirements, and typical performance of each level). Through this step, it is ensured that the LLM accurately understands the scoring reference framework, clarifies "what constitutes the answer performance of different levels," and establishes a unified standard for subsequent analysis and scoring.
[0094] As an example, the different levels of thinking structure shown in the SOLO classification evaluation method in Table 1 can be used as a detailed description of SOLO grading.
[0095] (2) Analysis Instructions: As a "deep decomposition step" in the scoring process, this invention instructs the LLM to analyze student answers in three dimensions: at the element extraction level, it is necessary to identify the core geographical elements involved in the answer (such as topography, ocean currents, industry types, etc.) and the completeness of the elements; at the level of identifying logical relationships between elements, it is necessary to judge the form of association between elements (such as causal relationship, constraint relationship, synergistic relationship) and the clarity of expression; at the level of judging the depth of reasoning, it is necessary to analyze whether the answer reflects inductive generalization (such as extracting laws from phenomena), knowledge transfer (such as applying the laws of place A to the analysis of place B), or critical thinking (such as dialectically evaluating the impact of human activities). After the analysis is completed, the LLM needs to output structured intermediate analysis results (such as "elements: topography, climate; logic: causal relationship; reasoning: regional law induction") to provide quantifiable basis for matching the SOLO level.
[0096] (3) Scoring Decision Hints: As the "final output" of the scoring, the LLM needs to be guided by judgmental instructions to match the intermediate analysis results with the corresponding standards in the graded scoring template library: First, compare the intermediate analysis results with the typical characteristics of each SOLO level (such as multi-point structures usually exhibiting complete elements but loose logic) to determine the most suitable thinking level; then, generate specific scores according to the scoring rules corresponding to the level, and generate sub-item qualitative feedback in conjunction with the analysis process (such as "the elements are extracted completely, but the causal relationship between climate and agricultural type is not reflected, and the thinking level is a multi-point structure"). Through this step, the accurate transformation from "analysis information" to "scoring conclusion" is achieved, ensuring the scientific nature and interpretability of the scoring results.
[0097] In summary, besides being the first to combine SOLO with a large model, the innovation of this invention lies in introducing a multi-step reasoning approach into the prompt word engineering. By guiding the large model to complete four tasks sequentially—key element extraction, element internal relationship analysis, overall summarization or knowledge transfer ability assessment, and comparison with SOLO grading standards—it effectively reduces its judgment bias on complex open-ended answers, thereby significantly improving the consistency and scientific rigor of the scoring. This design represents a refined optimization of the large model's scoring logic, further enhancing the reliability of automated scoring.
[0098] In its design, this invention automatically outputs an analysis report containing traceability information (e.g., clearly indicating the SOLO level of the student's answer and the judgment criteria) after generating the machine scoring results, allowing teachers to verify and trace the results. This mechanism ensures scoring efficiency while reserving an interface for manual intervention, achieving a balance between automation and controllability.
[0099] In other words, the solution of this invention breaks through the limitation of traditional keyword matching-based automatic scoring, which can only identify isolated key points, and for the first time realizes the full-process automation of geography subjective questions from "point collection" to "meaning collection". Crucially, the invention embeds the SOLO level theory and cue word engineering into the scoring workflow of a large language model. This technical approach has not been publicly reported domestically or internationally, possessing clear technological innovation and practical value, and can be extended to competency-level scoring scenarios in other disciplines.
[0100] In traditional geography subjective question scoring practices, the "point-based" scoring method is commonly used, which assigns points by matching keywords or key points in the student's answer with those in the reference answer. While this method is simple and efficient in large-scale actual marking, it has significant limitations: it only focuses on whether the answer contains specific "key points," neglecting the overall semantics, reasoning logic, and knowledge structure of the student's response. It fails to truly reflect the student's analytical, inductive, and synthetic thinking abilities in geographical contexts, which contradicts the current emphasis on "ability-oriented" and "core competency-based" approaches in secondary school geography.
[0101] To address this, this invention introduces the concept of "semantic reasoning scoring," aiming to overcome the limitation of "seeing the trees but not the forest." It utilizes automated methods to perform semantic understanding and reasoning analysis on student answers, achieving a true diagnosis and classification of students' thinking levels. However, the core technical challenge of semantic reasoning scoring lies in how to enable machines to understand the semantics, logic, and knowledge connections within unfixed, unstructured natural language answers—something traditional keyword-matching-based automatic scoring algorithms struggle to achieve.
[0102] Furthermore, this invention performs targeted fine-tuning on the underlying large language model of the geographic subjective question scoring system. This targeted fine-tuning scheme can be implemented through the following means.
[0103] This embodiment details the optimization process of the underlying big language model of the geography subjective question scoring system. Specifically, it can use Tongyi Qianwen-Plus as the basic model and achieve deep adaptation to the scoring logic of geography subjective questions through targeted fine-tuning. Its technical solution is fundamentally different from the existing fine-tuning of ordinary academic automatic scoring.
[0104] In existing technologies, the fine-tuning of models for automatic scoring of subjective questions often adopts a single supervised mode of "answer-score pairs (Labels)". That is, for a specific question type or subject, by inputting paired data of student answers and corresponding scores, the model learns the mapping relationship of "keyword matching-score output". In essence, this kind of fine-tuning in existing technologies still continues the logic of traditional "point sampling" scoring. The model can only grasp low-level labeling judgment ability and cannot understand the thought structure and logical chain behind the answer.
[0105] This fine-tuning scheme overcomes the aforementioned limitations. It not only uses score labels as supervisory signals but also guides the model to learn multi-dimensional judgment logic within the SOLO hierarchical framework through a customized dataset and annotation system. This includes extracting core elements from the answer, identifying logical relationships between elements, dividing the thought structure into hierarchical levels, and determining knowledge transfer ability. Through this fine-tuning process, the model can internally establish a progressive mapping of "keywords-logical relationships-thought levels," achieving a leap from "single-point recognition" to "holistic semantic understanding."
[0106] To achieve the aforementioned fine-tuning objectives, this embodiment constructs a dedicated dataset for geographic subjective questions with multi-field annotations, including at least the following fields:
[0107] The question information and competency requirements fields include a description of the question context, the question content, and competency requirement tags. The competency requirement tags clearly indicate the core geographical competencies tested in the question, such as "summarizing topographic features," "analyzing the causes of climate," and "assessing the relationship between humans and the environment in the region," which are used to guide the model in matching the corresponding scoring dimensions.
[0108] Student answer field: Collect real answer samples from students of different academic levels, covering diverse language styles from basic expressions to advanced analysis, to ensure the representativeness of the dataset.
[0109] SOLO Hierarchical Labeling Field: Geography education experts will label each answer hierarchically based on the SOLO theory (pre-structure, single-point structure, multi-point structure, related structure, extended abstraction) and attach specific judgment criteria (e.g., "Because it only lists a single climate element and does not reflect the relationship with agricultural type, it is judged as a single-point structure").
[0110] Reasoning Chain Field: For each answer, the reasoning process is manually broken down to form interpretable intermediate reasoning steps (such as "Step 1: Identify the regional location → Step 2: Analyze climate elements → Step 3: Associate agricultural production characteristics → Step 4: Summarize regional development patterns"), which are used to train the model to simulate the logical reasoning process of human scoring.
[0111] Through multiple rounds of training on the aforementioned multi-field dataset, the model can simultaneously master the combined capabilities of score generation, hierarchical determination, and logical interpretation, enabling the scoring results to have traceable reasoning basis.
[0112] During the fine-tuning process, this embodiment further incorporates a knowledge system specific to the discipline of geography, including: a geography terminology database (such as the precise definitions of core concepts such as "cyclone", "monsoon circulation", and "industrial agglomeration"); a geospatial relationship map (such as the association rules of "latitude and longitude-climate zone-natural zone" and the spatial constraints of "topography-river-city distribution"); and a database of typical regional cases (such as the analysis logic of benchmark cases such as "the impact of the uplift of the Qinghai-Tibet Plateau on the East Asian climate" and "the spatial evolution law of the Yangtze River Delta urban agglomeration").
[0113] By deeply embedding domain knowledge, the model can significantly improve the accuracy of understanding special expressions, spatial relationships and causal chains in geographical scenarios, effectively avoiding the misjudgment problem of general large models in the detailed processing of geographical disciplines (such as confusing the causal mechanisms of "monsoon climate" and "Mediterranean climate").
[0114] The Tongyi 1000 Questions-Plus model, after targeted fine-tuning, combined with prompt word engineering and SOLO-based opinion-gathering and scoring templates, can achieve integrated processing of unstructured answers to geography subjective questions:
[0115] a) Automatically completes semantic understanding, element extraction, logical analysis, and thought level determination;
[0116] b) Output a structured scoring result that includes the score, SOLO level, and reasoning basis;
[0117] Compared to existing technologies that rely on keyword matching or surface fine-tuning, this method improves scoring consistency (the degree of agreement with expert human scoring) by more than 40%, and enables the visualization and traceability of scoring logic through intermediate reasoning processes.
[0118] This fine-tuning scheme provides core technical support for the automation of "concept-based scoring" in geography subjective questions, and its approach can be extended to other subject areas that emphasize the evaluation of thinking levels, such as history.
[0119] To address the challenges of scoring open-ended questions, especially those in geography, this invention proposes an automated scoring scheme based on the aforementioned technical framework, which further expands the system's applicable scenarios.
[0120] Unlike objective questions and "point-based" subjective questions, open-ended geography questions do not have a single standard answer. Students can answer from multiple perspectives, using differentiated expression structures and reasoning paths, resulting in highly unstructured and diverse answers. In a human-marked environment, markers need to expend a great deal of effort to complete full-text semantic understanding, deconstruct reasoning chains, and assess the depth of thought. This is not only labor-intensive but also prone to inconsistencies in the application of scoring standards due to individual cognitive differences (such as subjective biases in the definition of "reasonable reasoning").
[0121] Existing automated scoring schemes are limited by the technical logic of keyword matching or fixed template comparison, and cannot handle the flexible expression and complex logical connections in open-ended answers. Therefore, they are difficult to effectively replace human marking in educational practice.
[0122] This invention constructs an automated scoring system for open-ended questions through a collaborative architecture of "Large Language Model (LLM) + SOLO Level Theory + Cue Word Engineering + Targeted Fine-tuning". The specific implementation path is as follows:
[0123] Deep semantic analysis: Based on LLM with targeted fine-tuning in the geography field, it performs overall semantic understanding and reasoning chain reconstruction of open-ended answers, breaking through the limitations of traditional keyword matching and achieving deep decoding of unstructured text;
[0124] Hierarchical Ability Mapping: Based on the SOLO classification theory, the depth of thinking and structural complexity of the answer are mapped to five levels: pre-structure, single-point structure, multi-point structure, related structure, and extended abstraction. The model is guided by prompt word engineering to output interpretable classification criteria, ensuring the traceability of the scoring logic.
[0125] Multi-faceted expression adaptation: By training the model through multiple rounds of fine-tuning, the mapping relationship between "expression variants and core logic" is identified, supporting diverse analytical paths for the same problem (such as when analyzing "the causes of urban flooding", one can approach it from the perspective of climate or focus on the design of drainage systems). Without relying on "standard sentence structures", the core of the answer and causal relationship are accurately captured.
[0126] This solution effectively overcomes the technical bottleneck of automated scoring for open-ended geography questions. Specifically, it reduces the average grading time per question by more than 60%, and lowers the scoring deviation rate to below 5% through standardized hierarchical judgment logic. It automatically generates class analysis reports that include the distribution of thinking levels, coverage of core elements, and typical reasoning flaws, providing teachers with a basis for precise teaching intervention. It also generates customized feedback on students' answers, such as "thinking weaknesses - improvement paths" (e.g., "They have mastered the analysis of the impact of terrain on climate, but lack the ability to transfer cross-regional cases; it is recommended to strengthen comparative practice"), thus extending the function from "scoring" to "educating."
[0127] Compared to existing automated systems that can only handle objective questions or "point-based" short-answer questions, this invention is the first to achieve deep semantic understanding and quantitative evaluation of ability levels for complex open-ended questions in geography, filling a gap in domestic and international educational evaluation technologies in this field. Its core innovation lies in combining the SOLO-level thinking assessment framework with the flexible understanding capabilities of a large model, providing a reusable technological paradigm for the automated scoring of interdisciplinary open-ended questions. In other words, this invention is not only effective for geography subjective questions (specifically, non-open-ended subjective questions), but also applicable to the automated machine grading of open-ended questions.
[0128] Based on the preceding description, Figure 5 This is a block diagram of the subjective question scoring and evaluation system based on SOLO classification according to the present invention. In summary, the subjective question scoring and evaluation system based on SOLO classification according to the present invention includes:
[0129] The answer sheet acquisition module uses image acquisition equipment to collect student exam papers in a standardized manner, generate exam paper images, and establish a scanned image database. The front-end processing module, located in the front-end system, is used to perform image preprocessing operations on the acquired exam paper images, including exam paper positioning, extraction of answer area sub-images, text extraction through the optical character recognition module, conversion of students' answer text into editable text, and association with metadata for building a structured answer database.
[0130] The assessment benchmark configuration module includes a standardized interface for supporting the import of information, including test questions. It extracts structured data, including question background, question stem content, SOLO level description information, and SOLO level-based evaluation scoring label templates, and stores them in the database.
[0131] The review and SOLO level determination module is located in the backend system and includes a large language model. The large language model receives the answer text from the answer database from the frontend system according to a preset strategy. The large language model extracts the elements, logical relationships between elements and reasoning depth information from the answer text and outputs structured intermediate parsing results.
[0132] The large oracle model is a large language model that has undergone targeted fine-tuning. The dataset used in the targeted fine-tuning process includes at least the following fields: question information and ability requirements, student answers, SOLO level labels, and inference chain.
[0133] Based on the intermediate parsing results and SOLO level description information, the large language model generates an evaluation of the student's thinking structure level; based on the intermediate parsing results and the SOLO level-based opinion scoring labeling template, the language model generates a score for the student's answer text.
[0134] The learning analysis and feedback module analyzes the evaluation results of all students' answers using a large language model and outputs a learning analysis report. The learning analysis report includes a visual data dashboard of the overall learning effect of the class and teaching suggestions.
[0135] The data storage and interaction module is based on a front-end and back-end separation architecture and is deployed on a cloud computing server. It consists of an interaction layer, a service layer, a data layer, and external cloud services working together to realize data storage, interaction, and the operation of system functions.
[0136] Later in this invention, specific practical embodiments will be provided for the practical application of the geography subjective question scoring system.
[0137] This embodiment further illustrates the actual operation process and technical effects of the present invention through specific application examples of point-based scoring and opinion-based scoring.
[0138] The test subjects in this example are 122 senior high school students from a middle school in Chongqing in 2022, involving classes 2, 3 and 6. The test scenario is a joint examination for senior high school students. The scoring module of this invention (based on SOLO classification theory) is used for automated scoring.
[0139] The test questions use "wetland conservation and development" in California's Central Valley as a contextual theme, connecting core concepts such as the formation process of floodplain wetlands and the relationship between waterbird migration and wetland resources. The assessment objectives include:
[0140] Regional cognitive ability: judging the spatial relationships of natural elements such as topography and climate in the central valley;
[0141] Comprehensive thinking ability: Analyzing the interaction between human activities and the ecological environment;
[0142] Sustainable development literacy: proposing solutions that are both ecologically sound and economically feasible based on real-world problems.
[0143] Question 16(1) adopts the acknowledgment scoring standard, based on the SOLO-level acknowledgment scoring label template (as shown in Table 2), which focuses on evaluating the students' logical reasoning completeness and depth of thinking regarding the "natural conditions for wetland formation".
[0144] Table 2: An example of an opinion rating labeling template based on SOLO rating
[0145]
[0146] Figure 6 This is a diagram showing the evaluation and scoring results of student answers in a certain embodiment. Taking question 16(1) answered by student Li from Class 2 as an example, the system executes the following scoring process:
[0147] First, the OCR results are extracted: the central valley has numerous perennial rivers, with low temperatures and high humidity during the winter months; the influence of trade winds and the convergence of warm and cold currents bring abundant water carrying substances into the central valley; sediment deposition in the central valley results in a large river basin area, forming seasonal floodplains and wetlands. It is worth noting that this invention allows the grader to correct errors identified by the OCR.
[0148] Single-question scoring: Deep semantic analysis of the answers is performed using a large language model to extract core elements (e.g., climate and river foundation: numerous perennial rivers in the central valley, low temperatures and high humidity during the winter half-year; material input factors: influence of trade winds, confluence of warm and cold currents, seawater carrying materials; landforms and wetland formation: sediment deposition, large river basin area, forming seasonal floodplains and wetlands), and to identify logical relationships (e.g., attempts to construct a causal chain of "climate - external forces - landforms - wetland formation," but some links have logical deviations, such as trade winds and the confluence of warm and cold currents not matching the actual geographical background of the central valley, and failure to fully connect key topographic conditions such as low-lying terrain). The SOLO level is matched and evaluation reasons are generated. For example, it belongs to a relational structure (but contains incorrect associations). Attempts were made to integrate multiple elements (climate, external forces, landforms, wetlands) to establish connections, but key elements were misunderstood (e.g., atmospheric circulation, material input methods), the logical chain has obvious loopholes, and core geographical conditions (low-lying terrain and poor drainage) were not accurately connected.
[0149] The final evaluation was as follows: The student's answer regarding the 'influence of trade winds and the convergence of warm and cold currents' is inconsistent with the actual situation. The Central Valley is primarily influenced by the westerly winds, not the trade winds. Furthermore, the statement that 'the sea carries more water and substances into the Central Valley' is inaccurate; the seasonal flooding of the Central Valley is mainly due to increased river flow caused by orographic rainfall. The student's understanding of the specific geographical conditions of the Central Valley was inaccurate, and the answer failed to mention key factors such as low-lying terrain and poor drainage.
[0150] The final score is 2 points. In this invention, only the final evaluation and final score can be displayed; the analysis process does not need to be shown.
[0151] Secondly, the problems students encountered with this question in this exam will be presented as suggestions for reviewing the exam paper, along with the knowledge points tested. Figure 7 This is a sample diagram illustrating feedback suggestions for a class exam in one example. The feedback suggestions in this example are as follows: In this exam, most students had the following problems in their answers to this question: 1) Inaccurate understanding of the impact of warm, humid westerly winds; some students incorrectly described them as southwest winds or monsoons; 2) Failure to fully mention the key factors of the influence of warm, humid westerly winds during the winter months and the formation of abundant orographic rainfall on the western slopes of the Sierra Nevada Mountains; 3) Insufficient or inaccurate description of the poor drainage caused by the low-lying terrain of the Central Valley; 4) Inaccurate use of geographical terminology, such as inappropriate expressions like "forced uplift" and "weak infiltration"; 5) Failure to combine the provided materials and regional characteristics for specific analysis, resulting in incomplete answers.
[0152] This question tests the following knowledge points: "Factors affecting regional temperature and precipitation" and "Distribution patterns of global ocean currents and their geographical significance".
[0153] Furthermore, in one embodiment of the present invention, a list of students with lower scores and a list of students with higher scores are also provided.
[0154] Finally, this invention also analyzes the overall learning situation of the class and generates a learning situation analysis report after batch grading. This report includes a visualized data dashboard of the overall learning effectiveness of the class and summarizes common problems in answering questions. In one example, a diagnostic report on the thinking level of a class is described below.
[0155] A visual data dashboard showcasing the overall learning effectiveness of a class based on the SOLO taxonomy to evaluate students' current thinking levels is presented, specifically displaying the following: Class 9 had 39 participants in the test; among the participants:
[0156] Distribution of Thinking Levels: This assessment of the thinking levels of 39 students in Class 9 using the SOLO taxonomy revealed a clear hierarchical distribution: pre-structural level accounted for approximately 17.95%, single-point structure accounted for a high 79.49%, multi-point structure accounted for only about 2.56%, and the results for association and extension were both 0%.
[0157] Common flaws in thinking: such as "43% of the answers mistakenly described 'prevailing westerly winds' as 'southwest monsoons', 31% of the answers failed to explain the formation mechanism of 'Sierra Nevada orographic rainfall', and 27% of the answers omitted the key condition of 'poor drainage in valleys'";
[0158] Teaching suggestions: For problems involving logical gaps, it is recommended to "dissect the connection mechanism of 'atmospheric circulation-topography-hydrology' through diagrams" and "design comparative cases to strengthen the understanding of the differences between the westerly wind belt and the monsoon."
[0159] SOLO Tiered List: Students with "Related Structures and Above" (high score range) and "Single-Point Structures and Below" (low score range) are selected according to SOLO tier to provide a basis for targeted training.
[0160] The results show that over 90% of the students in the class are still at the pre-structural and single-point structural stage of thinking. Their understanding of knowledge is mostly fragmented or focused on a single point. Their ability to connect and comprehensively apply knowledge is weak. Subsequent teaching should focus on strengthening the logical connection between knowledge, guiding students to build systematic cognition, and improving the depth and breadth of their thinking.
[0161] The above examples verify the practical effectiveness of the system of the present invention in the subjective scoring mode. It not only achieves efficient automation of traditional keyword scoring, but also realizes the evaluation of the thinking level of open-ended answers through SOLO grading, providing full-scenario technical support for the scoring of subjective questions in geography.
[0162] To better illustrate the present invention, numerous specific details have been provided in the detailed embodiments described above. Those skilled in the art should understand that the present invention can be practiced even without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of the present invention.
[0163] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A subjective question scoring and evaluation system based on SOLO classification, characterized in that, include: The answer sheet acquisition module uses image acquisition equipment to collect student exam papers in a standardized manner, generate exam paper images, and establish a scanned image database. The front-end processing module, located in the front-end system, is used to perform image preprocessing operations on the collected test paper images, including test paper positioning, extracting answer area sub-images, completing text extraction through the optical character recognition module, converting students' answer text into editable text and associating metadata for building a structured answer database; The assessment benchmark configuration module includes a standardized interface for supporting the import of information, including test questions, and extracts structured data, including question background, question stem content and SOLO level description information, and SOLO level-based evaluation scoring label templates, and stores them in the database. The review and SOLO level determination module is located in the backend system and includes a large language model. The large language model receives the answer text from the answer database from the frontend system according to a preset strategy. The large language model extracts the elements, logical relationships between elements and reasoning depth information from the answer text and outputs structured intermediate parsing results. The large oracle model is a large language model that has undergone targeted fine-tuning. The dataset used in the targeted fine-tuning process includes at least the following fields: question information and ability requirements, student answers, SOLO level labels, and inference chain. Based on the intermediate parsing results and SOLO level description information, the large language model generates an evaluation of the student's thinking structure level; based on the intermediate parsing results and the SOLO level-based opinion scoring labeling template, the language model generates a score for the student's answer text. The learning analysis and feedback module analyzes the evaluation results of all students' answers using a large language model and outputs a learning analysis report, which includes a visual data dashboard of the overall learning effect of the class and teaching suggestions. The data storage and interaction module is based on a front-end and back-end separation architecture and is deployed on a cloud computing server. It consists of an interaction layer, a service layer, a data layer, and external cloud services working together to realize data storage, interaction, and the operation of system functions.
2. The subjective question scoring and evaluation system based on SOLO classification according to claim 1, characterized in that: The visualization dashboard of the overall learning performance of the class includes the distribution of thinking levels, common thinking deficiencies, and lists of students in different SOLO levels.
3. The subjective question scoring and evaluation system based on SOLO classification according to claim 2, characterized in that: During the fine-tuning process, the large language model also incorporates a knowledge system specific to the discipline of geography, including: a geography terminology database, a geospatial relationship map, and a database of typical regional cases.
4. The subjective question scoring and evaluation system based on SOLO classification according to claim 3, characterized in that: The question information and ability requirement fields include a description of the question scenario, the question content, and ability requirement tags; The SOLO grading label field includes grading labels for each answer based on SOLO theory by geography education experts, as well as specific criteria for judgment; The reasoning chain field contains a manually broken-down thought process for each answer, forming interpretable intermediate reasoning steps.
5. The subjective question scoring and evaluation system based on SOLO classification according to claim 1, characterized in that: The extraction of elements, logical relationships between elements, and reasoning depth information from the response text specifically includes: At the element extraction level, the geographical elements involved in the response text and the completeness of the elements are identified; At the level of identifying logical relationships between elements, the association forms and clarity of expression between elements are judged, including causal relationships, constraint relationships, and synergistic relationships. At the level of reasoning depth assessment, analyze whether the answer reflects inductive generalization, knowledge transfer, or critical thinking.
6. The subjective question scoring and evaluation system based on SOLO classification according to claim 5, characterized in that: The image preprocessing operation includes: The test paper is located by edge detection and perspective correction, and the answer area sub-image is extracted by threshold segmentation technology to exclude irrelevant areas such as headers, footers and binding lines. Subsequently, the optical character recognition module extracts the text, converting each student's answer into editable text and simultaneously linking metadata such as student ID, class, question number, and answer area coordinates to build a structured answer database.
7. The subjective question scoring and evaluation system based on SOLO classification according to claim 6, characterized in that: The answer acquisition module and the front-end processing module are integrated into a handheld smart device, which can capture the test paper in real time by calling the device's camera or read locally stored image files; or, the answer acquisition module and the front-end processing module are located in different electronic devices.
8. The subjective question scoring and evaluation system based on SOLO classification according to claim 7, characterized in that: The backend system includes an interaction layer, a service layer, a data layer, and an authentication layer; among which, The interaction layer is built on the Vue3 front-end and provides functions such as exam paper uploading. The service layer is built on the Gin framework and uses SQLX to perform database operations; CORS cross-domain resource sharing is configured to ensure cross-domain communication between the front-end and back-end. The data layer relies on MySQL, Redis, and MinIO for data storage and optimization.
9. The subjective question scoring and evaluation system based on SOLO classification according to claim 7, characterized in that: The answer acquisition module and the front-end processing module are located in different electronic devices, wherein the front-end processing module is located in a personal computer.
10. The subjective question scoring and evaluation system based on SOLO classification according to claim 5, characterized in that: The SOLO-based subjective question scoring and evaluation system is also configured for automatic grading of open-ended questions.
Citation Information
Patent Citations
Subjective question test paper marking method based on large model
CN118861522A
Cited By
Classroom instant interaction feedback analysis method and system based on portable multimedia display equipment
CN121354398A
Subjective question scoring detailed rule conversion method based on LLM and scoring key point extraction
CN121480486A
A method for converting subjective question scoring rules based on LLM and scoring point extraction
CN121480486B
Mathematical solution process step-level correction and feedback generation method based on deep learning
CN121836992A