Student problem-solving process analysis method and system based on large model
Patent Information
- Application Number
- CN202610858773.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-09-01
AI Technical Summary
[0002]在教育领域,对学生的解题过程进行分析与诊断是提升教学效果的核心环节,传统的教学场景中,教师通过批改学生的作业和试卷,判断解题结果的正确性,但无法很难快速有效地深入洞察导致每一错误节点的错因;
本发明通过获取题目题干、标准解题步骤序列和学生作答步骤序列,利用大语言模型提取各步骤的核心论据,包括已知条件、中间推导结果和概念定义;基于核心论据及最后步骤的计算参量与结果,分别构建作答逻辑链和标准逻辑链;将作答逻辑链与标准逻辑链比对,获取数值错误节点、概念定义错误指向边、节点错误指向边以及缺失节点和缺失指向边;根据比对结果判断各数值错误节点的错误原因,实现精准识别学生的作答步骤过程中所存在的错误节点以及对应的错误成因;
Smart Images

Figure CN122674701A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and system for analyzing student problem-solving processes based on large models. Background Technology
[0002] In the field of education, analyzing and diagnosing students' problem-solving process is a core element in improving teaching effectiveness. In traditional teaching scenarios, teachers judge the correctness of students' solutions by grading their homework and test papers, but it is difficult to quickly and effectively gain a deep understanding of the causes of each error. In existing intelligent grading technologies, such as the invention patent with application number "CN201910529878.7", a method, system, and intelligent terminal for grading homework based on image recognition are disclosed. This solution acquires image data of students' answers, performs image recognition to obtain the question stem, question type, and answer content of the question to be graded. If the question type is objective, the corresponding standard answer is obtained from the question stem, and the answer content of the answer area of the question to be graded is compared with the standard answer to obtain the corresponding objective question grading result. If the question type is subjective, the standard answers and corresponding scoring criteria for all solution methods are obtained from the question stem, and the answer content of the answer area of the question to be graded is compared with the standard answer to obtain the corresponding subjective question grading result. However, this solution is limited to comparing the answer content with the standard answer to provide a grading result; it cannot analyze the causes of errors at each error point in the student's answering process. Summary of the Invention
[0003] To address the technical problems existing in the prior art, this invention provides a method for analyzing student problem-solving processes based on a large model, comprising the following steps: S1. Obtain the question stem, standard solution step sequence, and student's answer step sequence for the target question. Use a large language model to perform semantic parsing on the question stem, standard solution step sequence, and answer step sequence, and extract the core arguments for each step: extract the first core argument corresponding to each answer step in the answer step sequence and extract the second core argument corresponding to each solution step in the standard solution step sequence; the core arguments include known conditions, intermediate derivation results, and specific concept definitions; S2. Construct the answer logic chain based on the first core argument, the required calculation parameters of the final answer step, and the corresponding calculation results; construct the standard logic chain based on the second core argument, the required calculation parameters of the final problem-solving step, and the corresponding calculation results. S3. Compare the answer logic chain with the standard logic chain to obtain the numerical error nodes, concept definition error pointing edges and node error pointing edges in the answer logic chain, as well as the missing nodes and missing pointing edges that are missing from the answer logic chain compared to the standard logic chain. S4. Based on the definitions of erroneous pointing edges, node erroneous pointing edges, missing nodes, and missing pointing edges, determine the cause of error for each numerically erroneous node.
[0004] Furthermore, the extraction of the first core argument corresponding to each answer step in the sequence of answer steps specifically includes: Obtain the answer formula in the current answering step, and identify the corresponding concept definition from the relevant preset subject knowledge base based on the answer formula; Identify the known parameters and their corresponding values from the question stem to form the first set; identify the intermediate derived parameters and their corresponding calculated values from all the answer steps before the current answer step to form the second set. Identify the parameters required for calculation from the answer formula, and match them from the first set and the second set respectively to obtain the corresponding known conditions and intermediate derivation results; The identification of the corresponding concept definition, known conditions, and intermediate derivation results will serve as the first core argument for the current answer step.
[0005] Furthermore, the standard problem-solving step sequence is obtained by querying the target question from the answer database used to store the answers to each question. The sequence of answering steps is obtained by receiving the answer images uploaded by students for the target question, and then using optical character recognition technology to extract each answering step from the answer images to form the sequence of answering steps.
[0006] Furthermore, the step of identifying the corresponding concept definition from the relevant preset subject knowledge base based on the answer formula specifically involves: Extract the formula features from the current answer formula and encode them to form a formula structure feature vector. The formula features include: the number of each parameter, the type of each operator and the corresponding number of times it appears, the value of each numerical constant and the corresponding number of times it appears, and the hierarchical structure of operator operation priority. The similarity is calculated between the formula structure feature vector corresponding to the current answer formula and the formula structure feature vector corresponding to each concept formula in the corresponding preset subject knowledge base. Concept formulas with similarity greater than or equal to the preset similarity threshold are used as candidate matching concept formulas for the current answer formula. From each candidate matching concept formula, the candidate matching concept formula with the highest similarity is selected as the matching concept formula for the current answer formula. The concept definition corresponding to the matching concept formula in the preset subject knowledge base is used as the concept definition corresponding to the current answer formula.
[0007] Furthermore, the answering logic chain uses the required calculation parameters and corresponding calculation results of the final answering step, as well as the known conditions and intermediate step derivation results in the first core argument, as each first node, and the specific concept definition in the first core argument as the first pointing edge corresponding to the two first nodes. The standard logic chain uses the required computational parameters and corresponding computational results of the final problem-solving step, as well as the known conditions and intermediate derivation results in the second core argument, as each second node, and the specific concept definitions in the second core argument as the second pointing edges corresponding to the two second nodes.
[0008] Furthermore, the acquisition of the missing nodes and missing pointing edges is specifically as follows: Match the current second node with each of the first nodes one by one, and query the first nodes that have the same parameters as the current second node to form a matching node pair with the current second node. If the query result is empty, then the current second node is a missing node in the answer logic chain. Match the current second pointing edge with each of the first pointing edges one by one. Query the first pointing edges that form a matching node pair with the starting second node of the current second pointing edge as the starting first node and a matching node pair with the ending second node of the current second pointing edge as the ending first node. Form a matching pointing edge pair with the current second pointing edge. If the query result is empty, then the current second pointing edge is the missing pointing edge of the answer logic chain.
[0009] Furthermore, the acquisition of the numerical error nodes, concept definition error pointing edges, and node error pointing edges is specifically as follows: In the current matching node pair, compare whether the value corresponding to the parameter in the first node is the same as the value corresponding to the parameter in the second node. If they are not the same, then the first node is regarded as the numerical error node in the answer logic chain. In the current matching edge pair, compare whether the concept definition corresponding to the first edge is consistent with the concept definition corresponding to the second edge. If they are inconsistent, the first edge is regarded as the concept definition error edge of the answer logic chain. The first pointing edge that does not form a matching pointing edge pair with any second pointing edge is regarded as the node error pointing edge of the answer logic chain.
[0010] Furthermore, based on the conceptual definitions of erroneous pointing edges, erroneous pointing edges of nodes, missing nodes, and missing pointing edges, the reasons for errors in each numerically erroneous node are determined, specifically as follows: S41. Using the current numerical error node as the first terminal node, query all corresponding first pointing edges, which are the third pointing edges. At the same time, using the current numerical error node as the second terminal node, query the missing node and missing pointing edge corresponding to the answering logic chain in the standard logic chain. S42. Determine whether there is a concept definition error pointing edge from the third pointing edge. If it exists, the error reason of the current numerical error node includes concept definition error. Determine whether there is an incorrectly pointing edge from the third pointing edge. If so, the error of the current numerically incorrect node may be due to an incorrect parameter type. Determine from the third pointing edge whether there is a third pointing edge that starts with the node with the numerical error. If it exists, the error reason of the current numerical error node includes the value of the substituted parameter. If a missing node and / or missing pointing edge are found based on the current numerical error node, the error reason for the current numerical error node includes missing input parameters; S43. If the error cause of the current numerical error node does not include any of the situations in step S42, then the error cause of the current numerical error node is determined to be a simple calculation error.
[0011] Furthermore, the following steps are included after step S4: S5. Based on the frequency of each error reason, analyze the severity of the various error reasons the student made on the target questions:
[0012] The severity of the i-th error cause, Let N be the number of occurrences of the i-th error cause, and N be the number of error cause types.
[0013] This invention also provides a student problem-solving process analysis system based on a large model, which applies any of the student problem-solving process analysis methods based on large models described above, including: The large language model is used to perform semantic parsing on the question stem, standard problem-solving step sequence, and student's answer sequence for the target question, and extract the core arguments for each step: extracting the first core argument corresponding to each answer step in the answer sequence and extracting the second core argument corresponding to each problem-solving step in the standard problem-solving step sequence; the core arguments include known conditions, intermediate derivation results, and specific concept definitions; The logic chain construction module is used to construct the answer logic chain based on the first core argument, the required calculation parameters of the final answer step, and the corresponding calculation results, and to construct the standard logic chain based on the second core argument, the required calculation parameters of the final problem-solving step, and the corresponding calculation results. The comparison module is used to compare the answer logic chain with the standard logic chain, obtain the numerical error nodes, concept definition error pointing edges and node error pointing edges in the answer logic chain, and obtain the missing nodes and missing pointing edges that the answer logic chain lacks compared to the standard logic chain. The error analysis module is used to determine the error cause of each numerical error node based on the definition of error pointing edge, node error pointing edge, missing node, and missing pointing edge.
[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention obtains the question stem, standard problem-solving steps sequence, and student answer steps sequence, and uses a large language model to extract the core arguments for each step, including known conditions, intermediate derivation results, and concept definitions. Based on the core arguments and the calculation parameters and results of the final step, it constructs an answer logic chain and a standard logic chain respectively. The answer logic chain is compared with the standard logic chain to obtain numerical error nodes, concept definition error pointing edges, node error pointing edges, and missing nodes and missing pointing edges. Based on the comparison results, the cause of each numerical error node is determined, achieving accurate identification of error nodes and their corresponding causes in the student's answer process. By constructing a logical chain for answering questions based on core arguments and a standard logical chain, the basis for each step of the answer is accurately analyzed. Based on the logical chain for answering questions and the standard logical chain, the second node is matched with the first node, and the second pointing edge is matched with the first pointing edge. Missing nodes and missing pointing edges are identified. At the same time, by comparing whether the values in the nodes are consistent, whether the conceptual definitions of the pointing edges are consistent, and unmatched pointing edges, nodes with incorrect values, pointing edges with incorrect conceptual definitions, and pointing edges with incorrect nodes are identified. This effectively detects key steps that students may miss in the problem-solving process, improves the accuracy of error location, and improves the accuracy of the causes of errors in each answering step. Based on numerical error nodes, we analyze their associated pointing edges, missing nodes, and missing pointing edges to comprehensively determine the cause of error for each error node, thereby achieving multidimensional cause diagnosis and improving the effectiveness of error cause identification. Attached Figure Description
[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart of a student problem-solving process analysis method based on a large model according to the present invention; Figure 2 This is an example diagram of the response logic chain of the present invention; Figure 3 This is a structural block diagram of a student problem-solving process analysis system based on a large model according to the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0019] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indication will also change accordingly.
[0020] Furthermore, the use of terms such as "first" and "second" in this invention is for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" and "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of the various embodiments can be combined with each other, but only on the basis of being achievable by those skilled in the art. When the combination of technical solutions is contradictory or impossible to implement, such a combination of technical solutions should be considered non-existent and not within the scope of protection claimed by this invention.
[0021] Example 1 See Figure 1 As shown, the present invention provides a method for analyzing student problem-solving processes based on a large model, which specifically includes the following steps: S1. Obtain the question stem, standard solution step sequence, and student answer step sequence of the target question. Use a large language model to perform semantic parsing on the question stem, standard solution step sequence, and answer step sequence and extract the core arguments of the steps: extract the first core argument corresponding to each answer step in the answer step sequence and extract the second core argument corresponding to each solution step in the standard solution step sequence. S2. Construct the answer logic chain based on the first core argument, the required calculation parameters of the final answer step, and the corresponding calculation results; construct the standard logic chain based on the second core argument, the required calculation parameters of the final problem-solving step, and the corresponding calculation results. S3. Compare the answer logic chain with the standard logic chain to obtain the numerical error nodes, concept definition error pointing edges and node error pointing edges in the answer logic chain, as well as the missing nodes and missing pointing edges that are missing from the answer logic chain compared to the standard logic chain. S4. Based on the definitions of erroneous pointing edges, node erroneous pointing edges, missing nodes, and missing pointing edges, determine the cause of error for each numerically erroneous node.
[0022] Each step is explained in detail: S1. Obtain the question stem, standard solution step sequence, and student's submitted answer step sequence for the target question. Use a large language model to perform semantic parsing on the question stem, standard solution step sequence, and answer step sequence, and extract the core arguments for each step: extract the first core argument corresponding to each answer step in the answer step sequence and extract the second core argument corresponding to each solution step in the standard solution step sequence. Among them, the Large Language Model (LLM) is a model that simulates the human language understanding and generation ability through deep learning algorithms. It can recognize complex textual contexts and logical connections, and has built-in knowledge bases of multiple disciplines. In this solution, any existing large language model is selected.
[0023] The core arguments refer to the factual elements such as conceptual definitions and numerical values that support the corresponding steps, including known conditions, intermediate derivation results, and specific conceptual definitions.
[0024] In step S1, the standard problem-solving step sequence is obtained by querying the target problem from the answer database used to store the answers to each problem.
[0025] In step S1, the sequence of answering steps is obtained by receiving the answer image uploaded by the student for the target question, and then using optical character recognition technology to extract each answering step from the answer image to form the sequence of answering steps.
[0026] In step S1, the extraction of the first core argument corresponding to each answer step in the answer step sequence specifically includes: S11. Obtain the answer formula in the current answering step, and identify the corresponding concept definition from the corresponding preset subject knowledge base according to the answer formula; S12. Identify the known parameters and their corresponding values from the question stem to form the first set. Identify the intermediate parameters and their corresponding calculated values from all the answer steps before the current answer step to form the second set. S13. Identify the parameters required for calculation from the answer formula, and match them from the first set and the second set respectively to obtain the corresponding known conditions and intermediate derivation results. S14. Identify the corresponding concept definition, known conditions, and intermediate derivation results as the first core argument for the current answer step.
[0027] In step S11, the corresponding concept definition is identified from the relevant preset subject knowledge base according to the answer formula, specifically as follows: S111. Extract the formula features from the current answer formula and encode them to form a formula structure feature vector. The formula features include: the number of each parameter, the type of each operator and the corresponding number of times it appears, the value of each numerical constant and the corresponding number of times it appears, and the operator operation priority hierarchy. S112. Calculate the similarity between the formula structure feature vector corresponding to the current answer formula and the formula structure feature vector corresponding to each concept formula in the corresponding preset subject knowledge base. Select the concept formula with a similarity greater than or equal to the preset similarity threshold as the candidate matching concept formula of the current answer formula. Select the candidate matching concept formula with the highest similarity from each candidate matching concept formula as the matching concept formula of the current answer formula. S113. The concept definition corresponding to the matching concept formula in the preset subject knowledge base is used as the concept definition corresponding to the current answer formula.
[0028] The preset subject knowledge base includes the definitions of various concepts in the corresponding subject, the corresponding concept formulas for each concept definition, and the formula structure feature vectors corresponding to each concept formula. The method of obtaining the formula structure feature vectors corresponding to the concept formulas is the same as the method of obtaining the formula structure feature vectors of the answer formulas in step S111.
[0029] The operator precedence hierarchy is the order in which each operator is arranged according to its precedence in the corresponding formula.
[0030] The specific implementation method of the second core argument corresponding to each problem-solving step in the sequence of extracted standard problem-solving steps corresponds to the specific implementation method of the first core argument corresponding to each answer step in the sequence of extracted answer steps, that is: Obtain the problem-solving formula in the current problem-solving step, and identify the corresponding concept definition from the relevant preset subject knowledge base based on the problem-solving formula; The first set is formed by identifying the known parameters and their corresponding values from the problem stem, and the third set is formed by identifying the intermediate parameters and their corresponding calculated values from all the problem-solving steps before the current solution step. Identify the parameters required for calculation from the solution formula, and match them with the first and third sets respectively to obtain the corresponding known conditions and intermediate derivation results; The corresponding concept definition, known conditions, and intermediate derivation results are identified as the second core argument for the current problem-solving step.
[0031] To better understand, the following is a simple example: Question: A 2kg object is subjected to a horizontal pulling force of 10N on a horizontal surface. The coefficient of friction is 0.2. Find the acceleration of the object.
[0032] The student's answer sequence: 1. Frictional force f = μmg = 0.2 × 2 × 9.8 = 3.92 N; Core arguments: The formula for friction (i.e., the specific conceptual definition), and given conditions μ=0.2, m=2kg, g=9.8. ; 2. The resultant force F_resultant = F_f = 10 - 3.92 = 6.08 N; Core argument: Composition of forces (i.e., specific conceptual definition), given F = 10 N, intermediate steps derive result f = 3.92 N; 3. a = F_total / m = 6.08 / 2 = 3.04 ; Core argument: Newton's second law (net force F=ma), given m=2kg, intermediate steps derive result F_net=6.08N.
[0033] S2. Construct the answer logic chain based on the first core argument, the required calculation parameters for the final answer step, and the corresponding calculation results. Construct the standard logic chain based on the second core argument, the required calculation parameters for the final problem-solving step, and the corresponding calculation results. In step S2, the answering logic chain uses the required calculation parameters and corresponding calculation results of the final answering step, as well as the known conditions and intermediate step derivation results in the first core argument, as each first node, and the specific concept definition in the first core argument as the first pointing edge corresponding to the two first nodes. The standard logic chain uses the required computational parameters and corresponding computational results of the final problem-solving step, as well as the known conditions and intermediate derivation results in the second core argument, as each second node, and the specific concept definitions in the second core argument as the second pointing edges corresponding to the two second nodes.
[0034] See Figure 2 The diagram shown is an example of the answer logic chain obtained from the above example of the student's answer sequence.
[0035] S3. Compare the answer logic chain with the standard logic chain to obtain the numerical error nodes, concept definition error pointing edges, and node error pointing edges in the answer logic chain, as well as the missing nodes and missing pointing edges that are missing from the answer logic chain compared to the standard logic chain: In step S3, the acquisition of the missing node and the missing pointing edge specifically involves: Match the current second node with each of the first nodes one by one, and query the first nodes that have the same parameters as the current second node to form a matching node pair with the current second node. If the query result is empty, then the current second node is a missing node in the answer logic chain. Match the current second pointing edge with each of the first pointing edges one by one. Query the first pointing edges that form a matching node pair with the starting second node of the current second pointing edge as the starting first node and a matching node pair with the ending second node of the current second pointing edge as the ending first node. Form a matching pointing edge pair with the current second pointing edge. If the query result is empty, then the current second pointing edge is the missing pointing edge of the answer logic chain.
[0036] The first starting node is the starting node of the first pointing edge in the answer logic chain (i.e., the node from which the first pointing edge originates), and the first ending node is the ending node of the first pointing edge in the answer logic chain (i.e., the node pointed to by the first pointing edge). The second starting node and the second ending node are similar.
[0037] In step S3, the acquisition of the numerical error node, the concept definition error pointing edge, and the node error pointing edge is specifically as follows: In the current matching node pair, compare whether the value corresponding to the parameter in the first node is the same as the value corresponding to the parameter in the second node. If they are not the same, then the first node is regarded as the numerical error node in the answer logic chain. In the current matching edge pair, compare whether the concept definition corresponding to the first edge is consistent with the concept definition corresponding to the second edge. If they are inconsistent, the first edge is regarded as the concept definition error edge of the answer logic chain. The first pointing edge that does not form a matching pointing edge pair with any second pointing edge is regarded as the node error pointing edge of the answer logic chain.
[0038] S4. Based on the definitions of erroneous pointing edges, erroneous pointing edges of nodes, missing nodes, and missing pointing edges, determine the cause of error for each numerically erroneous node: In step S4, the error cause of each numerical error node is determined based on the conceptual definitions of erroneous pointing edges, node erroneous pointing edges, missing nodes, and missing pointing edges. Specifically: S41. Using the current numerical error node as the first terminal node, query all corresponding first pointing edges, which are the third pointing edges. At the same time, using the current numerical error node as the second terminal node, query the missing node and missing pointing edge corresponding to the answering logic chain in the standard logic chain. S42. Determine whether there is a concept definition error pointing edge from the third pointing edge. If it exists, the error reason of the current numerical error node includes concept definition error. Determine whether there is an incorrectly pointing edge from the third pointing edge. If so, the error of the current numerically incorrect node may be due to an incorrect parameter type. Determine from the third pointing edge whether there is a third pointing edge that starts with the node with the numerical error. If it exists, the error reason of the current numerical error node includes the value of the substituted parameter. If a missing node and / or missing pointing edge are found based on the current numerical error node, the error reason for the current numerical error node includes missing input parameters; S43. If the error cause of the current numerical error node does not include any of the situations in step S42, then the error cause of the current numerical error node is determined to be a simple calculation error.
[0039] In some embodiments, the student problem-solving process analysis method based on a large model further includes the following steps: S5. Based on the frequency of each error reason, analyze the severity of the various error reasons the student made on the target questions:
[0040] The severity of the i-th error cause, Let N be the number of occurrences of the i-th error cause, and N be the number of error cause types.
[0041] Example 2 See Figure 3 As shown, the present invention also provides a student problem-solving process analysis system based on a large model, specifically including: The large language model is used to perform semantic parsing on the question stem, standard problem-solving step sequence, and student's answer sequence for the target question, and extract the core arguments for each step: extracting the first core argument corresponding to each answer step in the answer sequence and extracting the second core argument corresponding to each problem-solving step in the standard problem-solving step sequence; the core arguments include known conditions, intermediate derivation results, and specific concept definitions; The logic chain construction module is used to construct the answer logic chain based on the first core argument, the required calculation parameters of the final answer step, and the corresponding calculation results, and to construct the standard logic chain based on the second core argument, the required calculation parameters of the final problem-solving step, and the corresponding calculation results. The comparison module is used to compare the answer logic chain with the standard logic chain, obtain the numerical error nodes, concept definition error pointing edges and node error pointing edges in the answer logic chain, and obtain the missing nodes and missing pointing edges that the answer logic chain lacks compared to the standard logic chain. The error analysis module is used to determine the error cause of each numerical error node based on the definition of error pointing edge, node error pointing edge, missing node, and missing pointing edge.
[0042] In some embodiments, the system further includes an error severity analysis module, used to analyze the severity of various error causes made by the student on the target question based on the frequency of occurrence of each error cause:
[0043] The severity of the i-th error cause, Let N be the number of occurrences of the i-th error cause, and N be the number of error cause types.
[0044] Example 3 The present invention also provides an electronic device, including: a processor, a transmitting device, an input device, an output device, and a memory. The processor may be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit, or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory may be implemented using a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), and is used to store computer program code, which includes computer instructions. When the processor executes the computer instructions, the electronic device executes a method as described in any of the above possible implementation methods.
[0045] Example 4 The present invention also provides a computer-readable storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor of an electronic device, cause the processor to perform a method as described in any of the above possible implementations.
[0046] The beneficial effects of this invention are as follows: This invention obtains the question stem, standard problem-solving steps sequence, and student answer steps sequence, and uses a large language model to extract the core arguments for each step, including known conditions, intermediate derivation results, and concept definitions. Based on the core arguments and the calculation parameters and results of the final step, it constructs an answer logic chain and a standard logic chain respectively. The answer logic chain is compared with the standard logic chain to obtain numerical error nodes, concept definition error pointing edges, node error pointing edges, and missing nodes and missing pointing edges. Based on the comparison results, the cause of each numerical error node is determined, achieving accurate identification of error nodes and their corresponding causes in the student's answer process. By constructing a logical chain for answering questions based on core arguments and a standard logical chain, the basis for each step of the answer is accurately analyzed. Based on the logical chain for answering questions and the standard logical chain, the second node is matched with the first node, and the second pointing edge is matched with the first pointing edge. Missing nodes and missing pointing edges are identified. At the same time, by comparing whether the values in the nodes are consistent, whether the conceptual definitions of the pointing edges are consistent, and unmatched pointing edges, nodes with incorrect values, pointing edges with incorrect conceptual definitions, and pointing edges with incorrect nodes are identified. This effectively detects key steps that students may miss in the problem-solving process, improves the accuracy of error location, and improves the accuracy of the causes of errors in each answering step. Based on numerical error nodes, we analyze their associated pointing edges, missing nodes, and missing pointing edges to comprehensively determine the cause of error for each error node, thereby achieving multidimensional cause diagnosis and improving the effectiveness of error cause identification.
[0047] In the description of this specification, the references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0048] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes: USB flash drive, mobile hard drive, read-only memory (ROM). ROM (ROM-only memory), RAM (random access memory), magnetic disks, optical disks, and other media that can store programs.
[0049] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for analyzing student problem-solving processes based on a large model, characterized in that, Includes the following steps: S1. Obtain the question stem, standard solution step sequence, and student's answer step sequence for the target question. Use a large language model to perform semantic parsing on the question stem, standard solution step sequence, and answer step sequence, and extract the core arguments for each step: extract the first core argument corresponding to each answer step in the answer step sequence and extract the second core argument corresponding to each solution step in the standard solution step sequence; the core arguments include known conditions, intermediate derivation results, and specific concept definitions; S2. Construct the answer logic chain based on the first core argument, the required calculation parameters of the final answer step, and the corresponding calculation results; construct the standard logic chain based on the second core argument, the required calculation parameters of the final problem-solving step, and the corresponding calculation results. S3. Compare the answer logic chain with the standard logic chain to obtain the numerical error nodes, concept definition error pointing edges and node error pointing edges in the answer logic chain, as well as the missing nodes and missing pointing edges that are missing from the answer logic chain compared to the standard logic chain. S4. Based on the definitions of erroneous pointing edges, node erroneous pointing edges, missing nodes, and missing pointing edges, determine the cause of error for each numerically erroneous node.
2. The student problem-solving process analysis method based on a large model according to claim 1, characterized in that, The first core argument corresponding to each answer step in the sequence of answer steps is specifically as follows: Obtain the answer formula in the current answering step, and identify the corresponding concept definition from the relevant preset subject knowledge base based on the answer formula; Identify the known parameters and their corresponding values from the question stem to form the first set; identify the intermediate derived parameters and their corresponding calculated values from all the answer steps before the current answer step to form the second set. Identify the parameters required for calculation from the answer formula, and match them from the first set and the second set respectively to obtain the corresponding known conditions and intermediate derivation results; The identification of the corresponding concept definition, known conditions, and intermediate derivation results will serve as the first core argument for the current answer step.
3. The student problem-solving process analysis method based on a large model according to claim 2, characterized in that, The standard problem-solving step sequence is obtained by querying the target problem from the answer database used to store the answers to each problem. The sequence of answering steps is obtained by receiving the answer images uploaded by students for the target question, and then using optical character recognition technology to extract each answering step from the answer images to form the sequence of answering steps.
4. The student problem-solving process analysis method based on a large model according to claim 2, characterized in that, The step of identifying the corresponding concept definition from the relevant preset subject knowledge base based on the answer formula is as follows: Extract the formula features from the current answer formula and encode them to form a formula structure feature vector. The formula features include: the number of each parameter, the type of each operator and the corresponding number of times it appears, the value of each numerical constant and the corresponding number of times it appears, and the hierarchical structure of operator operation priority. The similarity is calculated between the formula structure feature vector corresponding to the current answer formula and the formula structure feature vector corresponding to each concept formula in the corresponding preset subject knowledge base. Concept formulas with similarity greater than or equal to the preset similarity threshold are used as candidate matching concept formulas for the current answer formula. From each candidate matching concept formula, the candidate matching concept formula with the highest similarity is selected as the matching concept formula for the current answer formula. The concept definition corresponding to the matching concept formula in the preset subject knowledge base is used as the concept definition corresponding to the current answer formula.
5. The student problem-solving process analysis method based on a large model according to claim 2, characterized in that, The answering logic chain uses the required calculation parameters and corresponding calculation results of the final answering step, as well as the known conditions and intermediate step derivation results in the first core argument, as each first node, and the specific concept definition in the first core argument as the first pointing edge of the corresponding two first nodes. The standard logic chain uses the required computational parameters and corresponding computational results of the final problem-solving step, as well as the known conditions and intermediate derivation results in the second core argument, as each second node, and the specific concept definitions in the second core argument as the second pointing edges corresponding to the two second nodes.
6. The student problem-solving process analysis method based on a large model according to claim 5, characterized in that, The acquisition of the missing nodes and missing pointing edges is specifically as follows: Match the current second node with each of the first nodes one by one, and query the first nodes that have the same parameters as the current second node to form a matching node pair with the current second node. If the query result is empty, then the current second node is a missing node in the answer logic chain. Match the current second pointing edge with each of the first pointing edges one by one. Query the first pointing edges that form a matching node pair with the starting second node of the current second pointing edge as the starting first node and a matching node pair with the ending second node of the current second pointing edge as the ending first node. Form a matching pointing edge pair with the current second pointing edge. If the query result is empty, then the current second pointing edge is the missing pointing edge of the answer logic chain.
7. The student problem-solving process analysis method based on a large model according to claim 6, characterized in that, The acquisition of numerical error nodes, concept definition error pointing edges, and node error pointing edges is specifically as follows: In the current matching node pair, compare whether the value corresponding to the parameter in the first node is the same as the value corresponding to the parameter in the second node. If they are not the same, then the first node is regarded as the numerical error node in the answer logic chain. In the current matching edge pair, compare whether the concept definition corresponding to the first edge is consistent with the concept definition corresponding to the second edge. If they are inconsistent, the first edge is regarded as the concept definition error edge of the answer logic chain. The first pointing edge that does not form a matching pointing edge pair with any second pointing edge is regarded as the node error pointing edge of the answer logic chain.
8. The student problem-solving process analysis method based on a large model according to claim 2, characterized in that, Based on the definitions of erroneous pointing edges, erroneous pointing edges of nodes, missing nodes, and missing pointing edges, determine the cause of error for each numerically erroneous node. Specifically: S41. Using the current numerical error node as the first terminal node, query all corresponding first pointing edges, which are the third pointing edges. At the same time, using the current numerical error node as the second terminal node, query the missing node and missing pointing edge corresponding to the answering logic chain in the standard logic chain. S42. Determine whether there is a concept definition error pointing edge from the third pointing edge. If it exists, the error reason of the current numerical error node includes concept definition error. Determine whether there is an incorrectly pointing edge from the third pointing edge. If so, the error of the current numerically incorrect node may be due to an incorrect parameter type. Determine from the third pointing edge whether there is a third pointing edge that starts with the node with the numerical error. If it exists, the error reason of the current numerical error node includes the value of the substituted parameter. If a missing node and / or missing pointing edge are found based on the current numerical error node, the error reason for the current numerical error node includes missing input parameters; S43. If the error cause of the current numerical error node does not include any of the situations in step S42, then the error cause of the current numerical error node is determined to be a simple calculation error.
9. The student problem-solving process analysis method based on a large model according to claim 2, characterized in that, The following steps are included after step S4: S5. Based on the frequency of each error reason, analyze the severity of the various error reasons the student made on the target questions: ; The severity of the i-th error cause, Let N be the number of occurrences of the i-th error cause, and N be the number of error cause types.
10. A student problem-solving process analysis system based on a large model, employing the student problem-solving process analysis method based on a large model as described in any one of claims 1 to 9, characterized in that, include: The large language model is used to perform semantic parsing on the question stem, standard problem-solving step sequence, and student's answer sequence for the target question, and extract the core arguments for each step: extracting the first core argument corresponding to each answer step in the answer sequence and extracting the second core argument corresponding to each problem-solving step in the standard problem-solving step sequence; the core arguments include known conditions, intermediate derivation results, and specific concept definitions; The logic chain construction module is used to construct the answer logic chain based on the first core argument, the required calculation parameters of the final answer step, and the corresponding calculation results, and to construct the standard logic chain based on the second core argument, the required calculation parameters of the final problem-solving step, and the corresponding calculation results. The comparison module is used to compare the answer logic chain with the standard logic chain, obtain the numerical error nodes, concept definition error pointing edges and node error pointing edges in the answer logic chain, and obtain the missing nodes and missing pointing edges that the answer logic chain lacks compared to the standard logic chain. The error analysis module is used to determine the error cause of each numerical error node based on the definition of error pointing edge, node error pointing edge, missing node, and missing pointing edge.
Citation Information
Patent Citations
A method, system, and smart terminal for grading assignments based on image recognition.
CN112116840B