Chain scoring method and device based on attention mechanism and multi-modal semantic understanding
Patent Information
- Application Number
- CN202610952585.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-29
AI Technical Summary
[0008]为此,本发明目的在于至少一定程度上解决现有技术中的复杂教育场景中主观题自动批改面临的作答形式多样、逻辑路径复杂、多模态信息融合困难以及跨页分布等技术难题,从而提出一种基于注意力机制与多模态语义理解的链式判分方法及装置
[0023]本发明提供了一种基于注意力机制与多模态语义理解的链式判分方法及装置,所述方法包括:获取题目内容、标准答案以及学生答案并进行处理,得到相对应的多模态语义特征序列,通过预设大语言模型将多个所述多模态语义特征序列分别拆分为多个最小推理步骤,并按照推理顺序构建对应的思维链推理序列;分别对多个所述思维链推理序列中的最小推理步骤进行上下文语义编码,通过多头自注意力机制提取多个最小推理步骤之间的逻辑关系,生成所述标准答案和所述学生答案的文本语义特征序列;通过多头自注意力机制计算所述标准答案的文本语义特征序列和所述学生答案的文本语义特征序列的语义匹配度;提取所述学生答案中的非文本信息得到视觉特征,将所述视觉特征与所述学生答案的文本语义特征序列进行联合映射得到多模态联合语义表示,其中,所述学生答案包括所述非文本信息;基于所述多模态联合语义表示和所述语义匹配度从多个维度分别对所述学生答案的每个最小推理步骤进行评分,得到对应的评估结果以及综合得分,同时根据各个维度评分一致性计算置信度,并根据所述置信度输出批改结果。通过本发明提供的方法,利用预设大语言模型与思维链技术,将题目与学生答案解构为最小推理步骤,结合多头注意力机制,实现学生答案与标准答案在推理路径上的语义匹配与逻辑等价性判断;同时,通过跨模态特征融合,从推理步骤覆盖度、逻辑连贯性及结论正确性等维度进行分步骤评分,实现主观题自动批改的高准确性、鲁棒性与可解释性。
Smart Images

Figure CN122840236A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart education technology, and in particular to a chain-based scoring method and apparatus based on attention mechanism and multimodal semantic understanding. Background Technology
[0002] Currently, with the rapid development of online education, intelligent marking, and educational big data analysis, how to conduct high-precision and intelligent evaluation of subjective questions in complex educational scenarios has become a key technological bottleneck in the application of digital and intelligent education. Traditional subjective question marking methods heavily rely on manual marking or automatic scoring technology based on keyword matching and rule templates, resulting in low efficiency and high costs. Furthermore, they struggle to accurately understand the deep semantic and logical relationships in students' answers and cannot adapt to open-ended and diverse answer formats (such as essays, mathematical or physics derivation questions). Especially in subjects like Chinese, mathematics, physics, and chemistry, subjective question answers often contain multimodal information such as text, formulas, charts, and layout structure. Their problem-solving processes exhibit clear logic and hierarchy. Existing methods still suffer from performance bottlenecks in key areas such as complex handwriting recognition, step-by-step analysis, and essay theme comprehension. They struggle to achieve step-by-step scoring, logical equivalence judgment of multiple solution methods, and cannot perform error attribution, student profiling, or personalized guidance, thus hindering the construction of an intelligent education ecosystem.
[0003] Furthermore, current automatic scoring technology for subjective questions faces the following multiple challenges:
[0004] 1. The answer formats are diverse and the logical paths are complex, making it difficult for traditional literal matching methods to cover all situations;
[0005] 2. The answers contain multimodal information such as text, formulas, charts, and layout structure, making unified modeling and deep semantic understanding difficult;
[0006] 3. The methods lack step-by-step explanations and analyses, making it difficult to provide teachers and students with effective feedback on answering strategies and error identification;
[0007] 4. Robustness and generalization ability in scenarios involving multiple pages, unstructured responses, and multiple solutions are still insufficient. Summary of the Invention
[0008] Therefore, the purpose of this invention is to at least partially solve the technical challenges faced by the automatic grading of subjective questions in complex educational scenarios in the prior art, such as diverse answer formats, complex logical paths, difficulties in multimodal information fusion, and cross-page distribution. In this way, a chain-based scoring method and device based on attention mechanism and multimodal semantic understanding is proposed.
[0009] In a first aspect, the present invention provides a chain-based scoring method based on attention mechanisms and multimodal semantic understanding, the method comprising:
[0010] The question content, standard answer, and student answers are obtained and processed to obtain the corresponding multimodal semantic feature sequence. The multiple multimodal semantic feature sequences are then split into multiple minimum reasoning steps using a preset large language model, and the corresponding thought chain reasoning sequence is constructed according to the reasoning order.
[0011] The context semantic encoding is performed on the minimum reasoning steps in the multiple thought chain reasoning sequences respectively, and the logical relationship between the multiple minimum reasoning steps is extracted through a multi-head self-attention mechanism to generate the text semantic feature sequences of the standard answer and the student answer;
[0012] The semantic matching degree between the text semantic feature sequence of the standard answer and the text semantic feature sequence of the student answer is calculated using a multi-head self-attention mechanism.
[0013] Visual features are obtained by extracting non-textual information from the student's answer. The visual features are then jointly mapped with the textual semantic feature sequence of the student's answer to obtain a multimodal joint semantic representation, wherein the student's answer includes the non-textual information.
[0014] Based on the multimodal joint semantic representation and the semantic matching degree, each minimum reasoning step of the student's answer is scored from multiple dimensions to obtain the corresponding evaluation results and comprehensive scores. At the same time, the confidence level is calculated based on the consistency of the scores in each dimension, and the grading results are output based on the confidence level.
[0015] Secondly, the present invention provides a chain-based scoring device based on attention mechanism and multimodal semantic understanding, the device comprising:
[0016] Deconstruction module: used to acquire and process the question content, standard answer and student answer to obtain the corresponding multimodal semantic feature sequence. The multimodal semantic feature sequence is then broken down into multiple minimum reasoning steps by a preset large language model, and the corresponding thought chain reasoning sequence is constructed according to the reasoning order.
[0017] Encoding module: used to encode the contextual semantics of the minimum reasoning steps in the multiple thought chain reasoning sequences respectively, extract the logical relationship between the multiple minimum reasoning steps through a multi-head self-attention mechanism, and generate the textual semantic feature sequences of the standard answer and the student answer;
[0018] Calculation module: used to calculate the semantic matching degree between the text semantic feature sequence of the standard answer and the text semantic feature sequence of the student answer through a multi-head self-attention mechanism;
[0019] Fusion module: used to extract non-textual information from the student's answer to obtain visual features, and to jointly map the visual features with the textual semantic feature sequence of the student's answer to obtain a multimodal joint semantic representation, wherein the student's answer includes the non-textual information;
[0020] Evaluation module: It is used to score each minimum reasoning step of the student's answer from multiple dimensions based on the multimodal joint semantic representation and the semantic matching degree, to obtain the corresponding evaluation results and comprehensive scores. At the same time, it calculates the confidence level based on the consistency of the scores of each dimension, and outputs the grading results based on the confidence level.
[0021] Thirdly, the present invention also provides a chain-based scoring device based on attention mechanism and multimodal semantic understanding, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the various steps of the chain-based scoring method based on attention mechanism and multimodal semantic understanding as described in the first aspect.
[0022] Fourthly, the present invention also provides a storage medium storing a computer program thereon, which, when executed, implements the various steps of the chain-based scoring method based on attention mechanism and multimodal semantic understanding as described in the first aspect.
[0023] This invention provides a chain-based scoring method and apparatus based on attention mechanisms and multimodal semantic understanding. The method includes: acquiring and processing question content, standard answers, and student answers to obtain corresponding multimodal semantic feature sequences; decomposing the multiple multimodal semantic feature sequences into multiple minimal inference steps using a preset large language model, and constructing corresponding thought chain inference sequences according to the inference order; performing contextual semantic encoding on the minimal inference steps in the multiple thought chain inference sequences; extracting the logical relationships between the multiple minimal inference steps using a multi-head self-attention mechanism to generate textual semantic feature sequences of the standard answer and the student answers; and using a multi-head... The self-attention mechanism calculates the semantic matching degree between the text semantic feature sequence of the standard answer and the text semantic feature sequence of the student answer; non-textual information in the student answer is extracted to obtain visual features, and the visual features are jointly mapped with the text semantic feature sequence of the student answer to obtain a multimodal joint semantic representation, wherein the student answer includes the non-textual information; based on the multimodal joint semantic representation and the semantic matching degree, each minimum reasoning step of the student answer is scored from multiple dimensions to obtain the corresponding evaluation result and comprehensive score, and confidence is calculated based on the consistency of scores in each dimension, and the grading result is output based on the confidence. The method provided by this invention utilizes a pre-set large language model and thought chain technology to deconstruct the question and student answer into minimum reasoning steps, and combines a multi-head attention mechanism to achieve semantic matching and logical equivalence judgment between the student answer and the standard answer on the reasoning path; simultaneously, through cross-modal feature fusion, step-by-step scoring is performed from dimensions such as reasoning step coverage, logical coherence, and conclusion correctness, achieving high accuracy, robustness, and interpretability of automatic grading of subjective questions. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0025] Figure 1 This is a flowchart illustrating the chain-based scoring method based on attention mechanism and multimodal semantic understanding of the present invention.
[0026] Figure 2 This is a schematic diagram of a sub-process of the chain-based scoring method based on attention mechanism and multimodal semantic understanding of the present invention;
[0027] Figure 3This is a schematic diagram of the structure generated by the chain-based scoring method based on attention mechanism and multimodal semantic understanding of the present invention, which generates multimodal semantic feature sequences and minimum inference steps.
[0028] Figure 4 This is a schematic diagram of another sub-process of the chain-based scoring method based on attention mechanism and multimodal semantic understanding of the present invention;
[0029] Figure 5 This is a schematic diagram of another sub-process of the chain-based scoring method based on attention mechanism and multimodal semantic understanding of the present invention;
[0030] Figure 6 This is a schematic diagram of another sub-process of the chain-based scoring method based on attention mechanism and multimodal semantic understanding of the present invention;
[0031] Figure 7 This is a schematic diagram of another sub-process of the chain-based scoring method based on attention mechanism and multimodal semantic understanding of the present invention;
[0032] Figure 8 This is a schematic diagram of another sub-process of the chain-based scoring method based on attention mechanism and multimodal semantic understanding of the present invention;
[0033] Figure 9 This is a schematic diagram of the step-by-step scoring and iterative optimization process of the chain-based scoring method based on attention mechanism and multimodal semantic understanding of the present invention;
[0034] Figure 10 This is a schematic diagram of the program modules of the chain-based scoring device based on attention mechanism and multimodal semantic understanding of the present invention. Detailed Implementation
[0035] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0036] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating the chain-based scoring method based on attention mechanism and multimodal semantic understanding in this embodiment. In this embodiment, the chain-based scoring method based on attention mechanism and multimodal semantic understanding includes:
[0037] Step 101: Obtain the question content, standard answer, and student answers and process them to obtain the corresponding multimodal semantic feature sequence. Then, use a preset large language model to split the multiple multimodal semantic feature sequences into multiple minimum reasoning steps and construct the corresponding thought chain reasoning sequence according to the reasoning order.
[0038] In this embodiment, the question content, standard answer, and student answer are first obtained, and then the obtained question content, standard answer, and student answer are processed to obtain a multimodal semantic feature sequence corresponding to the question content, standard answer, and student answer.
[0039] Specifically, please refer to Figure 2 and Figure 3 , Figure 2 This is a schematic diagram of a sub-process of the chain-based scoring method based on attention mechanism and multimodal semantic understanding in an embodiment of this application. Figure 3 This is a schematic diagram illustrating the structure of the multimodal semantic feature sequence and minimum inference steps generated by the chain-based scoring method based on attention mechanism and multimodal semantic understanding in this embodiment. In this embodiment, the step of obtaining and processing the question content, standard answer, and student answer to obtain the corresponding multimodal semantic feature sequence includes:
[0040] Step 201: After preprocessing and uniformly encoding the various modal information within the question content, standard answer, and student answer, cross-modal fusion is then performed.
[0041] Step 202: Obtain the multimodal semantic feature sequence corresponding to the question content, the standard answer, and the student's answer.
[0042] In this embodiment, after preprocessing and uniformly encoding various modal information within the question, the standard answer, and the student answer, cross-modal fusion is performed to generate multimodal semantic feature sequences corresponding to the question content, the standard answer, and the student answer, thus providing a foundation for subsequent semantic deconstruction and alignment.
[0043] Specifically, the question content, standard answer, and student answers all include multimodal information such as text, formulas, charts, and layout structure. These elements are input into a unified multimodal semantic encoding module. The text of the question content, standard answer, and student answers undergoes normalization processing: a pre-trained language model generates contextual semantic vectors; formulas undergo symbolic parsing and structural encoding to obtain logical features; and image regions are extracted using a visual Transformer to extract local and global visual features. A layout structure analysis model is then used to analyze the question stem and answer areas based on the layout structure. The processed multimodal information from the question content, standard answer, and student answers is then fused and serialized through a multi-layer cross-modal fusion network to generate corresponding multimodal semantic feature sequences for each element, providing a complete foundation for subsequent semantic structure and reasoning. The pre-trained language model is pre-trained with the goal of optimizing the semantic relationships between multimodal information through self-supervised learning, enabling it to jointly understand the multimodal semantics of the question content, standard answer, and student answers.
[0044] After obtaining the multimodal semantic feature sequences corresponding to the question content, standard answer, and student answer, the three multimodal semantic feature sequences are input into the preset large language model, which is then broken down into multiple minimum reasoning steps. These minimum reasoning steps are then constructed into corresponding thought chain reasoning sequences in sequence.
[0045] Specifically, please refer to Figure 3 and Figure 4 , Figure 3 This is a schematic diagram illustrating the structure generated by the multimodal semantic feature sequence and minimum inference steps of the chain-based scoring method based on attention mechanism and multimodal semantic understanding in the embodiments of this application. Figure 4 This is another sub-process diagram of the chain-based scoring method based on attention mechanism and multimodal semantic understanding in this application embodiment. In this embodiment, the step of splitting the multiple multimodal semantic feature sequences into multiple minimum reasoning steps by using a preset large language model, and constructing the corresponding thought chain reasoning sequence according to the reasoning order, includes:
[0046] Step 301: Input the multimodal semantic feature sequences corresponding to the question content, standard answer, and student answer into the preset large language model;
[0047] Step 302: Using the multimodal semantic feature sequence corresponding to the content of the question as the context constraint, the multimodal semantic feature sequence corresponding to the standard answer as the reasoning benchmark, and the multimodal semantic feature sequence corresponding to the student's answer as the parsing object, perform semantic deconstruction and decomposition to obtain the corresponding multiple minimum reasoning steps;
[0048] Step 303: After breaking down the question content, standard answer, and student answer, generate the corresponding thought chain reasoning sequence according to the logical order of the multiple smallest reasoning steps obtained.
[0049] In this embodiment, a multimodal educational dataset is constructed, containing text, formulas, charts, and layout information of question content, standard answers, and student answers. This dataset includes training and testing sets. The training set data is manually annotated, with annotations including the location of questions within documents and corresponding structured question-answer pairs. The pre-trained language model, based on the Transformer architecture, is pre-trained using the training set data. The specific training process is as follows:
[0050] 1. Data Processing: Using a large language model pre-trained based on Transformer, semantic parsing is performed on the text of the question content, standard answer, and student answers. Based on the reasoning process of the thought chain, the problem-solving process is broken down into steps to obtain multiple corresponding minimum reasoning steps.
[0051] 2. Construction of thought chain: Construct corresponding thought chain sequences by matching multiple sets of minimum reasoning steps according to the problem-solving logic, so that each minimum reasoning step corresponds to an independent and interpretable reasoning step.
[0052] 3. Training objective: Through comparative learning or supervised signal optimization, the generated sequence of minimum reasoning steps (i.e., the thought chain reasoning sequence) can accurately preserve the logical structure of the answer, providing structured input for semantic alignment and scoring.
[0053] The specific processing steps of the pre-trained pre-defined large language model include: inputting the multimodal semantic feature sequences corresponding to the question content, standard answer, and student answer into the pre-defined large language model; using the multimodal semantic feature sequence corresponding to the question content as contextual constraints, the multimodal semantic feature sequence corresponding to the standard answer as the reasoning benchmark, and the multimodal semantic feature sequence corresponding to the student answer as the parsing object; then semantically deconstructing the multimodal semantic feature sequences corresponding to the question content, standard answer, and student answer; and breaking down the question content, standard answer, and student answer into multiple minimum reasoning steps through sequential logical decomposition. Subsequently, according to the logical order, the multiple minimum reasoning steps corresponding to the question content, standard answer, and student answer are constructed into corresponding thought chain reasoning sequences to achieve a structured representation of each minimum reasoning step in the student answer, that is, representing each minimum reasoning step in the student answer as an independent and interpretable reasoning node. Each reasoning node retains the text semantics and potential logical relationships. This process ensures that each step of reasoning in the student answer is traceable and interpretable, and facilitates subsequent multimodal matching.
[0054] Step 102: Contextual semantic encoding is performed on the minimum reasoning steps in the multiple thought chain reasoning sequences, and the logical relationship between the multiple minimum reasoning steps is extracted through a multi-head self-attention mechanism to generate the text semantic feature sequences of the standard answer and the student answer.
[0055] In this embodiment, the minimum reasoning steps in multiple thought chain reasoning sequences are each subject to contextual semantic encoding. A multi-head attention mechanism is then used to extract the contextual dependencies between different minimum reasoning steps to generate textual semantic feature sequences for the standard answer and student answers. Contextual semantic encoding not only preserves the semantic information of each minimum reasoning step but also captures the logical connections between them.
[0056] Specifically, please refer to Figure 5 , Figure 5 This is another sub-process diagram of the chain-based scoring method based on attention mechanism and multimodal semantic understanding in the embodiments of this application. In this embodiment, the step of performing contextual semantic encoding on the minimum reasoning steps in the multiple thought chain reasoning sequences, extracting the logical relationship between the multiple minimum reasoning steps through a multi-head self-attention mechanism, and generating the textual semantic feature sequences of the standard answer and the student answer includes:
[0057] Step 401: Sequential modeling of multiple minimum reasoning steps in each of the thought chain reasoning sequences is performed using position encoding;
[0058] Step 402: Extract the contextual dependencies between multiple minimal inference steps using a multi-head self-attention mechanism to obtain the text semantic feature sequence of the standard answer and the text semantic feature sequence of the student answer.
[0059] In this embodiment, firstly, the minimal inference steps are sequentially modeled using positional encoding; secondly, a multi-head self-attention mechanism is used to extract the contextual dependencies between different minimal inference steps, and the association strength between different minimal inference steps is calculated using attention weights, supporting synonym substitution, step transformation, and logical equivalence matching; finally, the text semantic feature sequences of the standard answer and the student answer are generated, providing a basic representation for subsequent semantic matching and scoring. For example, the attention calculation between the semantic sequence of the student answer (S_t) and the semantic sequence of the standard answer (S_s) can be represented as: This formula measures the semantic similarity between student answers and standard answers in terms of reasoning steps, where... This is the transpose matrix of the semantic feature sequence of the standard answer. This is the scaling factor.
[0060] Step 103: Calculate the semantic matching degree between the text semantic feature sequence of the standard answer and the text semantic feature sequence of the student answer using a multi-head self-attention mechanism.
[0061] In this embodiment, the degree of matching between the text semantic feature sequence of the standard answer and the text semantic feature sequence of the student's answer on the reasoning path is calculated through a multi-head self-attention mechanism to achieve logical equivalence judgment, thereby ensuring that even if the student uses synonym substitution, step adjustment or different expression, the logical structure of the standard answer can still be matched.
[0062] Specifically, please refer to Figure 6 , Figure 6 This is another sub-process diagram of the chain-based scoring method based on attention mechanism and multimodal semantic understanding in the embodiments of this application. In this embodiment, the step of calculating the semantic matching degree of the text semantic feature sequence of the standard answer and the text semantic feature sequence of the student answer through a multi-head self-attention mechanism includes:
[0063] Step 501: Calculate the association weights between the text semantic feature sequence of the standard answer and the text semantic feature sequence of the student answer at different minimum inference steps using a parallel attention head;
[0064] Step 502: Based on the association weight, and combined with semantic distance and logical relationship strength, generate the semantic matching degree between the two at different minimum inference steps.
[0065] In this embodiment, a dynamic semantic alignment network is first constructed. The dynamic semantic alignment network is based on a multi-head attention mechanism. The dynamic semantic alignment network is pre-trained, and its training objective is to optimize the dynamic semantic alignment network so that it can accurately judge the logical equivalence between student answers and standard answers under multiple expression methods.
[0066] The textual semantic feature sequences of the standard answer and the student answer are input into a pre-trained dynamic semantic alignment network. Attention weights are used to calculate the association weights between every two minimum inference steps in the student and standard answer textual semantic feature sequences. Based on these association weights, and combined with the semantic distance and logical relationship strength of the minimum inference steps, a fine-grained semantic matching degree between the student and standard answers at each minimum inference step is generated. For example, the minimum inference steps in the standard answer textual semantic feature sequence include A, B, and C, while the minimum inference steps in the student answer textual semantic feature sequence include 1, 2, and 3. Parallel attention heads are used to calculate the association weights between A and 1, B and 2, and C and 3. Based on these association weights, and combined with the semantic distance and logical relationship strength of the relationships between A and 1, B and 2, and C and 3, the semantic matching degree between A and 1, B and 2, and C and 3 is obtained.
[0067] Step 104: Extract non-textual information from the student's answer to obtain visual features, and perform joint mapping between the visual features and the textual semantic feature sequence of the student's answer to obtain a multimodal joint semantic representation, wherein the student's answer includes the non-textual information.
[0068] In this embodiment, for non-textual information in student answers, a cross-modal feature extraction module is used to obtain visual features, and the visual features are jointly mapped with the text semantic feature sequence to generate a multimodal joint semantic representation.
[0069] Specifically, please refer to Figure 7 , Figure 7 This is another sub-process diagram of the chain-based scoring method based on attention mechanism and multimodal semantic understanding in this application embodiment. In this embodiment, the step of extracting non-textual information from the student's answer to obtain visual features, and jointly mapping the visual features with the textual semantic feature sequence of the student's answer to obtain a multimodal joint semantic representation, includes:
[0070] Step 601: Extract the visual features of non-textual information from the student's answer, wherein the non-textual information includes images and formulas;
[0071] Step 602: Map the visual features and the textual semantic feature sequence of the student's answer to a unified semantic space, and fuse them through a fusion attention mechanism to obtain the multimodal joint semantic representation.
[0072] In this embodiment, visual features are extracted from images, formulas, and other information in student answers. Then, a cross-modal mapping module projects these visual features and textual semantic feature sequences onto a unified semantic space. A cross-modal attention mechanism is used to uniformly map and fuse the visual features and textual semantic feature sequences, generating a multimodal joint semantic representation to achieve collaborative semantic understanding of multimodal information. Specifically, the cross-modal fusion attention mechanism integrates the complementarity of information from different modalities, preserving the textual logical hierarchy while incorporating visual and layout structural information to form a global semantic feature representation. This feature representation provides reliable input for step-by-step scoring, logical equivalence judgment, and interpretability feedback, ensuring scoring accuracy and multimodal consistency. The cross-modal mapping module is pre-trained; its training aims to enhance the understanding of the overall semantics of student answers through cross-modal fusion, providing high-quality multimodal feature support for step-by-step scoring.
[0073] Step 105: Based on the multimodal joint semantic representation and the semantic matching degree, score each minimum reasoning step of the student's answer from multiple dimensions to obtain the corresponding evaluation results and comprehensive scores. At the same time, calculate the confidence level based on the consistency of the scores in each dimension, and output the grading results based on the confidence level.
[0074] In this embodiment, each minimum reasoning step of the student's answer is scored from multiple dimensions based on multimodal joint semantic representation and semantic matching degree. Each dimension is scored using a weighted mechanism to obtain the corresponding evaluation result. The comprehensive score is obtained by combining the evaluation results of each minimum reasoning step. At the same time, the confidence level is calculated based on the consistency of the scores of each dimension, and the grading result is output based on the confidence level.
[0075] Specifically, please refer to Figure 8 and Figure 9 , Figure 8 This is another sub-process diagram of the chain-based scoring method based on attention mechanism and multimodal semantic understanding in the embodiments of this application. Figure 9 This is a flowchart illustrating the step-by-step scoring and iterative optimization process of the chain-based scoring method based on attention mechanism and multimodal semantic understanding in this embodiment. In this embodiment, the multiple dimensions include reasoning step coverage, logical coherence, and conclusion correctness. The student's answer is scored from multiple dimensions based on the multimodal joint semantic representation and the semantic matching degree to obtain the corresponding evaluation results and a comprehensive score. Simultaneously, confidence is calculated based on the consistency of scores across all dimensions, and the grading results are output based on the confidence scores, including:
[0076] Step 701: Based on the multimodal joint semantic representation and the semantic matching degree, score each of the minimum inference steps from three dimensions: inference step coverage, logical coherence, and conclusion correctness, to obtain the evaluation result of each of the minimum inference steps;
[0077] Step 702: The multiple evaluation results are processed using a weighted scoring mechanism to obtain the comprehensive score, and an interpretable feedback result is obtained by combining the thought chain reasoning sequence corresponding to the student's answer.
[0078] Step 703: Simultaneously, the confidence level is calculated based on the consistency of the scores for each minimum reasoning step according to the three dimensions of reasoning step coverage, logical coherence, and conclusion correctness.
[0079] Step 704: Compare the confidence level with the preset threshold. If the confidence level is greater than the preset threshold, the comprehensive score and the interpretable feedback result are used as the correction result and output. If the confidence level is less than the preset threshold, iterative optimization is performed until the confidence level reaches the preset threshold or the maximum number of iterations is reached.
[0080] In this embodiment, based on multimodal joint semantic representation, scoring is performed from three dimensions: reasoning step coverage, logical coherence, and conclusion correctness. Each dimension scores each minimum reasoning step of the student's answer, resulting in an evaluation result for each minimum reasoning step. A weighted scoring mechanism is then used to process the evaluation results of each minimum reasoning step to obtain a comprehensive score. This comprehensive score is further combined with the matching degree of the thought chain reasoning sequence corresponding to the student's answer to output interpretability feedback, providing a comprehensive scoring basis for both teachers and students. For example, the comprehensive score can be expressed as: .in, For coverage, For logical coherence, For the correctness of the conclusion, weights It can be adjusted according to the characteristics of the subject.
[0081] A comprehensive score is generated based on the step-by-step scoring results. Confidence is calculated based on the consistency of scores across three dimensions: reasoning step coverage, logical coherence, and conclusion correctness. Confidence = 1 - (max(C, L, R) - min(C, L, R)). The closer the scores are across the three dimensions, the more consistent the evaluation results from different perspectives, and the more credible the score. Conversely, if the three dimensions differ significantly, it indicates uncertainty in the score, resulting in low confidence. When the confidence falls below a preset threshold τ (e.g., 0.7), the iterative semantic matching optimization module is invoked. Based on the differences in the three-dimensional scores, the reasons for the low confidence are diagnosed, and the reasoning path and attention weights are adjusted accordingly. The semantic alignment process is iteratively optimized until the score confidence reaches the preset threshold or the maximum number of iterations is reached, thereby improving the scoring accuracy and interpretability of complex, multimodal, and logically complex answers.
[0082] For low-confidence results with confidence levels below a preset threshold, the system uses the student's multimodal joint semantic features, initial comprehensive score, and score confidence as states, and inference path adjustment and attention weight update as actions. It diagnoses the reasons for low confidence based on the differences in the three-dimensional scores and generates corresponding optimization strategies. It iteratively adjusts the semantic alignment process until the score confidence reaches the set threshold or the maximum number of iterations is reached.
[0083] This application provides a chain-based scoring method based on attention mechanisms and multimodal semantic understanding. The method includes: acquiring and processing the question content, standard answer, and student answer to obtain corresponding multimodal semantic feature sequences; using a pre-set large language model to decompose the multiple multimodal semantic feature sequences into multiple minimal inference steps, and constructing corresponding thought chain inference sequences according to the inference order; performing contextual semantic encoding on the minimal inference steps in the multiple thought chain inference sequences; extracting the logical relationships between the multiple minimal inference steps through a multi-head self-attention mechanism to generate textual semantic feature sequences of the standard answer and the student answer; and using multi-head self-attention... The intention mechanism calculates the semantic matching degree between the text semantic feature sequence of the standard answer and the text semantic feature sequence of the student answer; extracts non-textual information from the student answer to obtain visual features, and jointly maps the visual features with the text semantic feature sequence of the student answer to obtain a multimodal joint semantic representation, wherein the student answer includes the non-textual information; based on the multimodal joint semantic representation and the semantic matching degree, each minimum reasoning step of the student answer is scored from multiple dimensions to obtain the corresponding evaluation result and comprehensive score, and the confidence level is calculated based on the consistency of the scores in each dimension, and the grading result is output based on the confidence level. Through the method provided by this invention, the question and student answer are deconstructed into minimum reasoning steps using a preset large language model and thought chain technology, and combined with a multi-head attention mechanism, semantic matching and logical equivalence judgment of the student answer and the standard answer on the reasoning path are achieved; at the same time, through cross-modal feature fusion, step-by-step scoring is performed from dimensions such as reasoning step coverage, logical coherence, and conclusion correctness, achieving high accuracy, robustness, and interpretability of automatic grading of subjective questions.
[0084] Furthermore, the specific implementation steps of this application embodiment are as follows:
[0085] Step 1: Input multimodal educational data containing question content, standard answer, and student answers. Preprocess and uniformly encode the text, formulas, images, and layout structure respectively to generate multimodal semantic feature sequences corresponding to the questions, standard answers, and student answers.
[0086] Step 2: Using a pre-defined large language model based on the Transformer architecture, semantic deconstruction is performed with the question content as contextual constraints, the standard answer as the reasoning benchmark, and the student's answer as the parsing object. The model is broken down into multiple minimum reasoning steps, and the corresponding thought chain reasoning sequence is constructed.
[0087] Step 3: Encode the contextual semantics of the ordered minimum reasoning steps in the thought chain reasoning sequence, extract the logical relationship between the minimum reasoning steps through a multi-head self-attention mechanism, and generate the textual semantic feature sequence of the standard answer and the student's answer.
[0088] Step 4: Construct a dynamic semantic alignment network. Calculate the semantic matching degree of the textual semantic feature sequences of student answers and standard answers on the reasoning path through an attention mechanism. This will support subsequent step-by-step scoring and iterative optimization, enabling logical equivalence judgment under multiple expression methods.
[0089] Step 5: For non-text content such as images and formulas in student answers, use the cross-modal feature extraction module to obtain visual features, and perform joint mapping with the text semantic feature sequence of student answers generated in Step 3 to generate a multimodal joint semantic representation;
[0090] Step 6: Based on multimodal joint semantic representation and semantic matching degree, score the steps in three dimensions: reasoning step coverage, logical coherence and conclusion correctness. Generate a comprehensive score through a weighted mechanism and output interpretability feedback results in combination with the thought chain reasoning sequence.
[0091] Step 7: Simultaneously calculate the confidence score based on the consistency of scores across the three dimensions of reasoning step coverage, logical coherence, and conclusion correctness. For low-confidence results with confidence scores below a preset threshold, further perform iterative semantic matching optimization.
[0092] In summary, this application's embodiments construct a multimodal semantic understanding framework that integrates text, formulas, images, and layout structure information. Utilizing a large language model and thought chain reasoning technology, it deconstructs questions and student answers into minimal reasoning steps. Through cross-modal semantic alignment and hierarchical evaluation mechanisms, it achieves step-by-step scoring, logical equivalence judgment under multiple expressions, and error attribution. This provides high-quality and reliable technical support for intelligent education, personalized teaching, and digital educational applications. The beneficial effects also include:
[0093] 1. Achieving Multimodal Deep Semantic Understanding and Logical Equivalence Determination. This invention constructs a multimodal semantic understanding framework that integrates text, formulas, images, and layout structure information. Utilizing a large language model and thought chain technology, it deconstructs the question content and student answers into minimal reasoning steps. Through cross-modal semantic alignment and hierarchical evaluation mechanisms, it achieves step-by-step scoring of student answers and logical equivalence determination of multiple solution methods. This fundamentally overcomes the limitations of traditional literal matching methods in complex answer scenarios, thereby significantly improving the accuracy and reliability of automatic grading of subjective questions.
[0094] 2. Possesses high interpretability and step-by-step feedback capabilities. This invention generates a traceable, minimal-step reasoning sequence through thought chain reasoning, and performs weighted scoring based on the coverage of reasoning steps, logical coherence, and correctness of the conclusion. This enables interpretability analysis and feedback for each step of the answer, providing teachers and students with suggestions for optimizing answer strategies and attributing errors, thereby enhancing the interactive effect and teaching guidance value in intelligent education scenarios.
[0095] 3. Strong cross-modal information fusion capability, adaptable to diverse answer formats. This invention employs a cross-modal feature extraction and joint semantic mapping mechanism to map text, formulas, charts, and visual information to a unified semantic space. It utilizes an attention mechanism for joint semantic representation, enabling the system to effectively understand different expressions and multiple solutions. This achieves accurate information matching for complex handwritten text and multi-page questions, significantly enhancing input robustness and scene adaptability.
[0096] 4. Iterative Optimization and High-Confidence Scoring Guarantee. This invention introduces a confidence assessment mechanism based on three-dimensional score consistency and a feedback iterative optimization strategy. It can automatically diagnose the causes of low confidence and adjust the semantic alignment process accordingly until the scoring confidence reaches a preset threshold. This achieves dual guarantees of scoring accuracy and reliability, ensuring that high-quality and stable scoring results can still be output in complex, multimodal, and logically complex answer evaluation scenarios.
[0097] 5. Promoting the scalability of intelligent education and digital applications. The structured multimodal semantic data generated by this invention can be directly used for the construction of automated question banks, learning analysis, and personalized learning guidance, providing data support for AI teaching assistants, automatic grading, and educational research, improving the efficiency of educational data utilization, promoting the digitalization, intelligentization, and scalable application of educational resources, and possessing both academic innovation and industrial application value.
[0098] Furthermore, this application also provides a chain-based scoring device based on attention mechanism and multimodal semantic understanding, please refer to... Figure 10 , Figure 10 This is a schematic diagram of the program modules of the chain-based scoring device based on attention mechanism and multimodal semantic understanding in an embodiment of this application. In this embodiment, the chain-based scoring device based on attention mechanism and multimodal semantic understanding includes:
[0099] Deconstruction module 801: is used to acquire and process the question content, standard answer and student answer to obtain the corresponding multimodal semantic feature sequence. The multimodal semantic feature sequence is then split into multiple minimum reasoning steps by a preset large language model, and the corresponding thought chain reasoning sequence is constructed according to the reasoning order.
[0100] Encoding module 802: is used to perform contextual semantic encoding on the minimum reasoning steps in the multiple thought chain reasoning sequences respectively, extract the logical relationship between the multiple minimum reasoning steps through a multi-head self-attention mechanism, and generate the text semantic feature sequence of the standard answer and the student answer;
[0101] Calculation module 803: used to calculate the semantic matching degree between the text semantic feature sequence of the standard answer and the text semantic feature sequence of the student answer through a multi-head self-attention mechanism;
[0102] Fusion module 804: used to extract non-textual information from the student's answer to obtain visual features, and to jointly map the visual features with the textual semantic feature sequence of the student's answer to obtain a multimodal joint semantic representation, wherein the student's answer includes the non-textual information;
[0103] Evaluation module 805: is used to score each minimum reasoning step of the student's answer from multiple dimensions based on the multimodal joint semantic representation and the semantic matching degree, to obtain the corresponding evaluation results and comprehensive scores, and to calculate the confidence level based on the consistency of the scores of each dimension, and output the grading results based on the confidence level.
[0104] This application provides a chain-based scoring device based on attention mechanism and multimodal semantic understanding, which can: acquire and process the question content, standard answer, and student answer to obtain corresponding multimodal semantic feature sequences; decompose the multiple multimodal semantic feature sequences into multiple minimum reasoning steps using a preset large language model, and construct corresponding thought chain reasoning sequences according to the reasoning order; perform contextual semantic encoding on the minimum reasoning steps in the multiple thought chain reasoning sequences; extract the logical relationship between the multiple minimum reasoning steps using a multi-head self-attention mechanism to generate the text semantic feature sequences of the standard answer and the student answer; and use a multi-head self-attention mechanism to... An attention mechanism calculates the semantic matching degree between the text semantic feature sequence of the standard answer and the text semantic feature sequence of the student answer; non-textual information in the student answer is extracted to obtain visual features, and the visual features are jointly mapped with the text semantic feature sequence of the student answer to obtain a multimodal joint semantic representation, wherein the student answer includes the non-textual information; based on the multimodal joint semantic representation and the semantic matching degree, each minimum reasoning step of the student answer is scored from multiple dimensions to obtain the corresponding evaluation result and comprehensive score, and confidence is calculated based on the consistency of scores in each dimension, and the grading result is output based on the confidence. The method provided by this invention utilizes a pre-set large language model and thought chain technology to deconstruct the question and student answer into minimum reasoning steps, and combines a multi-head attention mechanism to achieve semantic matching and logical equivalence judgment between the student answer and the standard answer on the reasoning path; simultaneously, through cross-modal feature fusion, step-by-step scoring is performed from dimensions such as reasoning step coverage, logical coherence, and conclusion correctness, achieving high accuracy, robustness, and interpretability of automatic grading of subjective questions.
[0105] Furthermore, this application also provides a chain-based scoring device based on attention mechanism and multimodal semantic understanding, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the various steps of the chain-based scoring method based on attention mechanism and multimodal semantic understanding as described above.
[0106] Furthermore, this application also provides a storage medium storing a computer program thereon, which, when executed, implements the various steps of the chain-based scoring method based on attention mechanism and multimodal semantic understanding as described above.
[0107] In the various embodiments of this invention, the functional modules can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0108] Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0109] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, as some steps can be performed in other orders or simultaneously according to the present invention. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention. In the above embodiments, the descriptions of each embodiment have their own emphasis; for parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0110] For those skilled in the art, based on the ideas of the embodiments of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A chain-based scoring method based on attention mechanism and multimodal semantic understanding, characterized in that, The method includes: The question content, standard answer, and student answers are obtained and processed to obtain the corresponding multimodal semantic feature sequence. The multiple multimodal semantic feature sequences are then split into multiple minimum reasoning steps using a preset large language model, and the corresponding thought chain reasoning sequence is constructed according to the reasoning order. The context semantic encoding is performed on the minimum reasoning steps in the multiple thought chain reasoning sequences respectively, and the logical relationship between the multiple minimum reasoning steps is extracted through a multi-head self-attention mechanism to generate the text semantic feature sequences of the standard answer and the student answer; The semantic matching degree between the text semantic feature sequence of the standard answer and the text semantic feature sequence of the student answer is calculated using a multi-head self-attention mechanism. Visual features are obtained by extracting non-textual information from the student's answer. The visual features are then jointly mapped with the textual semantic feature sequence of the student's answer to obtain a multimodal joint semantic representation, wherein the student's answer includes the non-textual information. Based on the multimodal joint semantic representation and the semantic matching degree, each minimum reasoning step of the student's answer is scored from multiple dimensions to obtain the corresponding evaluation results and comprehensive scores. At the same time, the confidence level is calculated based on the consistency of the scores in each dimension, and the grading results are output based on the confidence level.
2. The method according to claim 1, characterized in that, The process of acquiring and processing the question content, standard answer, and student answers to obtain the corresponding multimodal semantic feature sequence includes: After preprocessing and uniformly encoding the various modal information within the question content, standard answer, and student answer, cross-modal fusion is then performed. The multimodal semantic feature sequence corresponding to the question content, the standard answer, and the student's answer is obtained.
3. The method according to claim 1, characterized in that, The step of splitting multiple multimodal semantic feature sequences into multiple minimum reasoning steps using a preset large language model, and constructing corresponding thought chain reasoning sequences according to the reasoning order, includes: The multimodal semantic feature sequences corresponding to the question content, standard answer, and student answer are input into the preset large language model; Using the multimodal semantic feature sequence corresponding to the content of the question as the context constraint, the multimodal semantic feature sequence corresponding to the standard answer as the reasoning benchmark, and the multimodal semantic feature sequence corresponding to the student's answer as the parsing object, semantic deconstruction and decomposition are performed respectively to obtain the corresponding multiple minimum reasoning steps; The question content, standard answer, and student answer are broken down into multiple minimum reasoning steps, which are then used to generate a corresponding thought chain reasoning sequence in logical order.
4. The method according to claim 3, characterized in that, The process of encoding the contextual semantics of the minimum reasoning steps in the multiple thought chain reasoning sequences, extracting the logical relationships between the multiple minimum reasoning steps through a multi-head self-attention mechanism, and generating textual semantic feature sequences of the standard answer and the student answer includes: Sequential modeling is performed on multiple minimum reasoning steps in each of the aforementioned thought chain reasoning sequences using position encoding; The contextual dependencies between multiple minimal inference steps are extracted using a multi-head self-attention mechanism to obtain the text semantic feature sequence of the standard answer and the text semantic feature sequence of the student answer.
5. The method according to claim 1, characterized in that, The step of calculating the semantic matching degree between the text semantic feature sequence of the standard answer and the text semantic feature sequence of the student answer using a multi-head self-attention mechanism includes: The association weights between the text semantic feature sequence of the standard answer and the text semantic feature sequence of the student answer at different minimum inference steps are calculated using a parallel attention head. Based on the association weights, and combined with semantic distance and logical relationship strength, the semantic matching degree between the two is generated at different minimum inference steps.
6. The method according to claim 1, characterized in that, The step of extracting non-textual information from the student's answer to obtain visual features, and then jointly mapping the visual features with the textual semantic feature sequence of the student's answer to obtain a multimodal joint semantic representation, includes: Extract visual features of non-textual information from the student's answers, wherein the non-textual information includes images and formulas; The visual features and the textual semantic feature sequences of the student's answers are mapped to a unified semantic space and fused through a fusion attention mechanism to obtain the multimodal joint semantic representation.
7. The method according to claim 3, characterized in that, The multiple dimensions include reasoning step coverage, logical coherence, and conclusion correctness. Based on the multimodal joint semantic representation and the semantic matching degree, each minimum reasoning step of the student's answer is scored from multiple dimensions to obtain corresponding evaluation results and a comprehensive score. Simultaneously, confidence levels are calculated based on the consistency of scores across all dimensions, and grading results are output based on these confidence levels, including: Based on the multimodal joint semantic representation and the semantic matching degree, each minimum reasoning step is scored from three dimensions: reasoning step coverage, logical coherence, and conclusion correctness, to obtain the evaluation result of each minimum reasoning step; A weighted scoring mechanism is used to process multiple evaluation results to obtain a comprehensive score, and an interpretable feedback result is obtained by combining the thought chain reasoning sequence corresponding to the student's answer. Meanwhile, the confidence level is calculated based on the consistency of the scores for each minimum reasoning step across three dimensions: coverage of reasoning steps, logical coherence, and correctness of conclusion. The confidence level is compared with a preset threshold. If the confidence level is greater than the preset threshold, the comprehensive score and the interpretable feedback result are used as the correction result and output. If the confidence level is less than the preset threshold, iterative optimization is performed until the confidence level reaches the preset threshold or the maximum number of iterations is reached.
8. A chain-based scoring device based on attention mechanism and multimodal semantic understanding, characterized in that, The device includes: Deconstruction module: used to acquire and process the question content, standard answer and student answer to obtain the corresponding multimodal semantic feature sequence. The multimodal semantic feature sequence is then broken down into multiple minimum reasoning steps by a preset large language model, and the corresponding thought chain reasoning sequence is constructed according to the reasoning order. Encoding module: used to encode the contextual semantics of the minimum reasoning steps in the multiple thought chain reasoning sequences respectively, extract the logical relationship between the multiple minimum reasoning steps through a multi-head self-attention mechanism, and generate the textual semantic feature sequences of the standard answer and the student answer; The calculation module is used to calculate the semantic matching degree between the text semantic feature sequence of the standard answer and the text semantic feature sequence of the student answer through a multi-head self-attention mechanism. Fusion module: used to extract non-textual information from the student's answer to obtain visual features, and to jointly map the visual features with the textual semantic feature sequence of the student's answer to obtain a multimodal joint semantic representation, wherein the student's answer includes the non-textual information; Evaluation module: It is used to score each minimum reasoning step of the student's answer from multiple dimensions based on the multimodal joint semantic representation and the semantic matching degree, to obtain the corresponding evaluation results and comprehensive scores. At the same time, it calculates the confidence level based on the consistency of the scores of each dimension, and outputs the grading results based on the confidence level.
9. A chain-based scoring device based on attention mechanism and multimodal semantic understanding, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements each step of the chain-based scoring method based on attention mechanism and multimodal semantic understanding as described in any one of claims 1-7.
10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the chain-based scoring method based on attention mechanism and multimodal semantic understanding as described in any one of claims 1-7.