Iterative problem solving framework construction method and system based on multi-modal large model
By using an iterative problem-solving framework based on a multimodal large model, and leveraging a dual verification mechanism and progressive interactive analysis, the shortcomings of the traditional Socratic method in terms of personalization and scalability are addressed, achieving efficient coverage and precise adaptation of personalized teaching.
Patent Information
- Application Number
- CN202511832270.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-08
- Publication Date
- 2026-03-17
- Estimated Expiration
- 2045-12-08
AI Technical Summary
The traditional Socratic method is insufficient in terms of personalization and scalability, making it difficult to meet the personalized teaching needs of a large number of students. Furthermore, assessment and question generation cannot be adapted in real time and with precision.
An iterative problem-solving framework based on a multimodal large model is adopted. Through two rounds of progressive questioning and interactive content recognition, combined with interactive content association verification and progressive interaction time analysis, a dual verification mechanism is constructed to optimize the solution framework and achieve personalized teaching.
It accurately captures users' knowledge mastery, generates corresponding interaction scores, constructs an optimal solution framework adapted to users, enhances the scalability and personalization of teaching, breaks through resource limitations, and achieves efficient coverage of differentiated learning needs.
Smart Images

Figure CN121279462B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and more specifically to a method and system for constructing an iterative problem-solving framework based on a multimodal large model. Background Technology
[0002] With the rapid development of artificial intelligence technology, especially large language models, and the growing demand for personalized and interactive learning in the education field, the traditional Socratic teaching method, which guides students to think deeply through questioning, aligns with the educational goal of cultivating students' critical thinking. The emergence of large language models provides technical support for its intelligent upgrade.
[0003] However, the traditional Socratic method relies on human teachers for implementation. While it is effective in guiding deep thinking, it has many practical limitations: on the one hand, teacher resources are limited, making it difficult to meet the personalized teaching needs of a large number of students at the same time; on the other hand, the assessment of students' knowledge levels, the relevance of questions generated, and the dynamism of dialogue adjustments are limited by human resources and energy, making it difficult to achieve real-time and accurate adaptation, resulting in insufficient scalability and personalization of teaching. Summary of the Invention
[0004] This application provides a method and system for constructing an iterative problem-solving framework based on a multimodal large model, aiming to solve the technical problems of insufficient scalability and personalization in existing teaching technologies.
[0005] In view of the above problems, this application provides a method and system for constructing an iterative problem-solving framework based on a multimodal large model.
[0006] Firstly, this application provides a method for constructing an iterative problem-solving framework based on a multimodal large model, including:
[0007] Output the first question, and obtain the user's first interaction content in response to the first question, identify it, and obtain the first interaction score;
[0008] Output the second question, obtain the second interaction content of the user's interaction with the second question, identify it, obtain the second interaction score, and obtain the progressive interaction time;
[0009] The first and second interactive content are associated and verified to obtain a first verification score, and a second verification score is obtained based on the interaction time.
[0010] Based on the first verification score and the second verification score, a fusion verification score is calculated. The first interaction score and the second interaction score are corrected to obtain a fusion interaction score. The solution framework is optimized to obtain the optimal solution framework, and the solution is taught to the user.
[0011] Secondly, this application provides a system for constructing an iterative problem-solving framework based on a multimodal large model, including:
[0012] The initial interaction scoring module is used to output the first question, obtain the user's first interaction content in response to the first question, identify it, and obtain the first interaction score;
[0013] The progressive interaction timing module is used to output the second question, obtain the second interaction content of the user's interaction with the second question, identify it, obtain the second interaction score, and obtain the progressive interaction time.
[0014] The association verification analysis module is used to verify and identify the association between the first interaction content and the second interaction content, obtain a first verification score, and analyze and obtain a second verification score based on the interaction time.
[0015] The fusion correction and optimization module is used to calculate the fusion verification score based on the first verification score and the second verification score, correct the first interaction score and the second interaction score to obtain the fusion interaction score, optimize the solution framework to obtain the optimal solution framework, and provide solution instruction to the user.
[0016] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0017] This application provides a method and system for constructing an iterative problem-solving framework based on a multimodal large model. Through two rounds of progressive questioning and interactive content recognition, it accurately captures users' knowledge mastery and generates corresponding interaction scores. Combining interactive content correlation verification and progressive interaction time analysis, a dual verification mechanism is constructed to ensure assessment accuracy. By integrating verification scores to correct the interaction scores, an optimal problem-solving framework suitable for users is formed, enabling targeted teaching. Ultimately, this effectively overcomes the resource limitations of traditional teaching models, significantly improves the scalability and personalization of teaching, and allows deep thinking-guided teaching to efficiently cover the diverse learning needs of more users. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating the method for constructing an iterative problem-solving framework based on a multimodal large model, as provided in this application embodiment;
[0020] Figure 2A schematic diagram of the structure of the iterative problem-solving framework construction system based on a multimodal large model provided in the embodiments of this application;
[0021] The components represented by each number in the attached diagram are explained below:
[0022] Initial interaction scoring module 11, progressive interaction timing module 12, correlation verification analysis module 13, and fusion correction and optimization module 14. Detailed Implementation
[0023] This application provides a method and system for constructing an iterative problem-solving framework based on a multimodal large model, which is used to address the technical problem of insufficient scalability and personalization in existing teaching technologies.
[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0025] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to these processes, methods, products, or devices.
[0026] Example 1, as Figure 1 As shown, this application provides a method for constructing an iterative problem-solving framework based on a multimodal large model, the method comprising:
[0027] S100: Output the first question, and obtain the first interaction content of the user's interaction with the first question, identify it, and obtain the first interaction score.
[0028] In this embodiment, a first question is output, and the user's first interaction content with the first question is acquired, identified, and a first interaction score is obtained. In the iterative problem-solving framework based on a multimodal large model, step S100 is the initial exploration stage of the framework. By identifying and scoring both the result and the cause, an accurate initial user profile can be provided for subsequent progressive questioning. This is an important foundation for achieving personalized teaching and framework scalability, and can prevent subsequent steps from deviating from the teaching direction due to initial assessment bias.
[0029] Step S100 in the method provided in this application embodiment includes:
[0030] Output the first question and obtain the user's first interaction content in response to the first question;
[0031] Based on natural language processing, extract the first interaction result content and the first interaction reason content of the first interaction content;
[0032] The first interaction result content and the first interaction reason content are input into the first interaction recognizer, and the first interaction result score and the first interaction reason score are output to obtain the first interaction score.
[0033] First, the system outputs the first question and obtains the user's initial interactive content based on that question. The first question is an initial exploratory question designed for the knowledge point the user needs to master. It should guide the user to both output the answer and explain the reasoning behind it, avoiding simply obtaining a single result. The first interactive content is the complete, unprocessed, raw response from the user to the first question, provided they input complete feedback via text, voice, or other means. For example, the first question might be: Given a right triangle with legs of length 3 cm and 4 cm, find the length of the hypotenuse and explain the underlying principles and calculation process. After thinking and calculating, the user might input the following: The hypotenuse is 5 cm. The Pythagorean theorem was used because it's a right triangle, and the sum of the squares of the two legs equals the square of the hypotenuse. Therefore, 3 squared plus 4 squared equals 9 + 16 = 25, which is then taken as the square root, resulting in 5.
[0034] Secondly, based on natural language processing (NLP), the first interaction result and the first interaction reason are extracted from the first interaction content. NLP is a technology that uses AI algorithms to analyze human natural language. Here, the function is information extraction, that is, separating result-type information and reason-type information from the user's spoken feedback, excluding irrelevant expressions such as interjections and repetitive content. The first interaction result refers to the information in the user feedback that directly corresponds to the answer to the first question, i.e., the final conclusion. The first interaction reason refers to the logical process information in the user feedback that explains how the result was obtained, including the principles and calculation steps. The NLP module is called to process the first interaction content input by the user. For example, the extracted result is as follows: First interaction result: The hypotenuse of the right triangle is 5 centimeters; First interaction reason: The solution is based on the Pythagorean theorem. The calculation steps are to add the square of 3 and the square of 4 to get a sum of 25, and take the square root of 25 to get the length of the hypotenuse.
[0035] Furthermore, the first interaction result content and the first interaction reason content are input into the first interaction recognizer, and the first interaction result score and the first interaction reason score are output to calculate the first interaction score.
[0036] Specifically, the first interaction result content and the first interaction reason content are input into the first interaction recognizer, and the first interaction result score and the first interaction reason score are output. The first interaction score is calculated, including:
[0037] Obtain a pre-trained first interaction recognizer, wherein the first interaction recognizer includes a first interaction result recognition branch and a first interaction reason recognition branch. The first interaction result recognition branch is trained using a sample first interaction result content set and a sample first interaction result score set for the first question. The first interaction reason recognition branch is trained using a sample first interaction reason content set and a sample first interaction reason score set for the first question.
[0038] Input the first interaction result content and the first interaction reason content into the first interaction result recognition branch and the first interaction reason recognition branch in the first interaction recognizer, and output to obtain the first interaction result score and the first interaction reason score;
[0039] The first interaction score is calculated based on the first interaction result score and the first interaction reason score.
[0040] First, a pre-trained first interaction recognizer is obtained. This first interaction recognizer includes a first interaction result recognition branch and a first interaction reason recognition branch. The first interaction result recognition branch is trained using a sample set of first interaction result content and a sample set of first interaction result scores for the first question. The first interaction reason recognition branch is trained using a sample set of first interaction reason content and a sample set of first interaction reason scores for the first question. The first interaction recognizer is an AI evaluation model using a dual independent sub-neural network model architecture, containing two independent branches for evaluating results and reasons respectively. Through pre-training, the model possesses the ability to score interaction content for specific knowledge points. The first interaction result recognition branch focuses on evaluating the correctness of the result content, with the user's feedback result content as input and the corresponding score as output. The first interaction reason recognition branch focuses on evaluating the logicality of the reason content, with the user's feedback reason content as input and the corresponding score as output. The sample set of first interaction result content refers to the set of sample data used to train the result recognition branch, containing a large number of possible user output results for the first question, such as a correct result of 5 cm, an incorrect result of 6 cm, 7 cm, etc. The first interaction result score set refers to the set of annotations corresponding one-to-one with the result sample, with only 0 or 1 points marked. Correct results are marked with 1 point, and incorrect results with 0 points. The first interaction cause content set refers to the sample set used to train the cause identification branch, containing various possible cause contents for the first question. The first interaction cause score set refers to the set of annotations corresponding one-to-one with the cause sample, with annotations being decimals in the range of 0-1. Completely conforming to standard logic is marked with 1 point, partially close to it is marked with 0.3-0.8 points, and completely irrelevant is marked with 0 points.
[0041] For example, taking the application of the Pythagorean theorem for right triangles as an example, the training process of the first interaction recognizer based on the dual independent sub-neural network model architecture is as follows: For the result recognition branch, prepare a sample first interaction result content set: 2000 result contents such as 5 cm, 6 cm, 7 cm, 4 cm, etc.; corresponding sample first interaction result score set annotation: 5 cm marked with 1 point, 6 cm marked with 0 points, 7 cm marked with 0 points, 4 cm marked with 0 points, etc., 2000 result scores. For the cause recognition branch, prepare a sample first interaction cause content set: [Based on the Pythagorean theorem, 3... 2 +4 2 =25, 25 square root = 5], [According to the Pythagorean theorem, 3 plus 4 squared equals 49], [According to the sum of both sides, 3+4=7], [No explanation of principle], etc., totaling 2000 reasons. The corresponding sample first interaction reason score set is labeled as follows: correct principle + complete steps = 1 point, correct principle but calculation error = 0.4 points, incorrect principle = 0 points, no principle = 0 points, etc., totaling 2000 reasons. Using the above sample set, two branches are trained respectively, so that the result recognition branch can accurately output 0 or 1 points, and the reason recognition branch can output a decimal of 0-1 for the matching closeness. When the result recognition branch reaches an accuracy of over 95% and the mean square error of the reason recognition branch is less than 0.05, the two sub-neural network branches are trained and completed. Combining the two forms a complete first interaction recognizer, which has the ability to accurately score the first question interaction content in the Pythagorean theorem application scenario in two dimensions.
[0042] Secondly, the content of the first interaction result and the content of the first interaction reason are input into the first interaction result recognition branch and the first interaction reason recognition branch within the first interaction recognizer, and the first interaction result score and the first interaction reason score are output. For example, a correct case: the user feedback result is 5 centimeters, the reason being based on the Pythagorean theorem, 3 2 +4 2 =25, 25 square root = 5. Input the first interaction result into the result recognition branch. If the branch judgment is correct, output the first interaction result score of 1 point. Input the first interaction reason into the reason recognition branch. If the branch judgment is completely consistent with the standard logic, output the first interaction reason score of 1 point. Deviation case: User feedback result is 5 cm, the reason being based on the Pythagorean theorem, 3 plus 4 equals 7, 7 minus 2 equals 5: Input the first interaction result into the result recognition branch, output the first interaction result score of 1 point. Input the first interaction reason into the reason recognition branch. The branch judgment principle is correct, but the calculation step is incorrect, close to the standard logic part, output the first interaction reason score of 0.4 points.
[0043] Finally, the first interaction score is calculated based on the first interaction result score and the first interaction reason score. According to preset rules, the arithmetic mean of the first interaction result score and the first interaction reason score is taken as the first interaction score, rounded to two decimal places. The calculation formula is: the first interaction score equals the sum of the first interaction result score and the first interaction reason score divided by 2. For example, in a correct case: first interaction result score 1 point, first interaction reason score 1 point, first interaction score = (1+1) / 2 = 1 point; in a flawed case: first interaction result score 1 point, first interaction reason score 0.4 points, first interaction score = (1+0.4) / 2 = 0.7 points.
[0044] In this embodiment, a closed-loop process of initial questioning, content extraction, and two-dimensional scoring is used to achieve accurate evaluation. This process can determine the user's mastery of knowledge points and the degree of logical understanding, avoiding evaluation bias caused by accidental correct results. The output of the first interaction score and the details of the results and reasons behind the score can provide accurate basis for subsequent steps. This is the key initial step for the entire iterative solution framework to achieve personalized teaching and large-scale expansion.
[0045] S200: Output the second question, obtain the second interaction content of the user's interaction with the second question, identify it, obtain the second interaction score, and obtain the progressive interaction time.
[0046] In this embodiment, a second question is output, and the user's second interaction content in response to the second question is acquired, identified, and a second interaction score is obtained. The progressive interaction time is also acquired. The interaction result of a single question may be accidental; for example, a user might arrive at the correct result based solely on guesswork, which cannot accurately reflect their actual mastery of the knowledge point. By designing a second question that is completely consistent with the core knowledge point of the first question, the user's understanding and application stability of the knowledge point can be repeatedly verified, eliminating interference from accidental factors. Simultaneously, the progressive interaction time can intuitively reflect the user's proficiency with the knowledge point; the shorter the interval, the more solid the user's mastery of the knowledge point and the more proficient their application, providing data support for subsequent verification and evaluation, and further improving the accuracy of personalized teaching.
[0047] Step S200 in the method provided in this application embodiment includes:
[0048] Output the second question and obtain the second interaction content of the user's interaction with the second question, wherein the first question and the second question are related questions;
[0049] Extract the second interaction result content and the second interaction reason content based on natural language processing;
[0050] The second interaction result content and the second interaction reason content are input into the second interaction recognizer, and the second interaction result score and the second interaction reason score are output. The second interaction score is calculated. The second interaction recognizer includes a second interaction result recognition branch and a second interaction reason recognition branch.
[0051] The time from the output of the first question to the collection of the second interactive content is used as the progressive interaction time.
[0052] First, output the second question and obtain the user's second interaction content regarding the second question. The first and second questions are related problems. Related problems refer to different specific scenario problems designed around the same knowledge point, used to verify the user's ability to transfer and apply knowledge. The second question is a verification question that completely matches the core knowledge point of the first question. By adjusting the given conditions and the angle of the question, different specific scenarios are created, ensuring the knowledge point remains unchanged but avoiding the user mechanically reusing the solution approach of the first question, thus achieving repeated verification of the user's mastery of the knowledge point. The second interaction content refers to the complete feedback information input by the user regarding the second question through text, voice, etc., and must include the solution result and the corresponding principle and calculation steps. For example, based on the Pythagorean theorem, adjust the given conditions and the direction of the question to design the second question: Given a right triangle with one leg of length 5 cm and a hypotenuse of length 13 cm, find the length of the other leg of the triangle and explain the principle and calculation process used to solve it. This question matches the knowledge point of the first question but has a different direction of questioning, examining the ability to apply the Pythagorean theorem in reverse, avoiding the user merely mechanically memorizing the solution steps of the first question. The user inputs the second interactive content: The length of the other right-angled side is 12 centimeters. The solution is based on the Pythagorean theorem, which states that the sum of the squares of the two right-angled sides of a right triangle is equal to the square of the hypotenuse. Therefore, the square of the other right-angled side is equal to the square of the hypotenuse minus the square of the known right-angled side, i.e., 13 squared minus 5 squared equals 169 minus 25 equals 144. The arithmetic square root of 144 is 12, so the length is 12 centimeters.
[0053] Secondly, the second interaction result content and the second interaction reason content are extracted from the second interaction content using natural language processing. The second interaction result content refers to the final solution conclusion in the user feedback that directly corresponds to the answer to the second question. The second interaction reason content refers to the logical information in the user feedback explaining the origin of the solution result, including the underlying principles and specific calculation steps. The natural language processing module is invoked to process the second interaction content input by the user, eliminating irrelevant expressions and separating the result and reason information. For example, the extracted result is as follows: Second interaction result content: The length of the other right-angled side is 12 centimeters; Second interaction reason content: The solution is based on the Pythagorean theorem. The calculation steps are: 13 squared minus 5 squared to get 144, and the arithmetic square root of 144 is 12 centimeters.
[0054] Furthermore, the second interaction result content and the second interaction reason content are input into the second interaction recognizer, which outputs the second interaction result score and the second interaction reason score, and calculates the second interaction score. The second interaction recognizer includes a second interaction result recognition branch and a second interaction reason recognition branch. The second interaction recognizer is an AI evaluation model with a structure completely consistent with the first interaction recognizer, employing a dual independent sub-neural network model architecture with two independent branches. Its scoring rules are consistent with S100, ensuring the comparability of the two interaction scores. The second interaction result recognition branch focuses on evaluating the correctness of the second interaction result content, outputting a score of 0 or 1. A score of 1 is awarded if the solution matches the standard answer, and 0 is awarded if it does not. The second interaction reason recognition branch focuses on evaluating the closeness of the second interaction reason content to the standard application logic of the Pythagorean theorem, outputting a decimal in the range of 0-1. A larger value indicates that the principle of the reason is correct and the steps are more complete than the standard logic. The second interaction score is calculated using the same method as the first interaction score, i.e., the arithmetic mean of the second interaction result score and the second interaction reason score. The formula is: the second interaction score equals the sum of the second interaction result score and the second interaction reason score divided by 2.
[0055] For example, the pre-training logic of the second interaction recognizer is consistent with that of the first interaction recognizer. The training sample set consists of various results and reasons related to the Pythagorean theorem, and the annotation rules are completely consistent with S100. Itemized score output: Input the second interaction result content into the second interaction result recognition branch. If the branch judgment result is consistent with the standard answer of 12 cm, output a second interaction result score of 1 point. Input the second interaction reason content into the second interaction reason recognition branch. If the branch judgment principle is correct and the steps are complete, and it is completely consistent with the standard logic, output a second interaction reason score of 1 point. Calculate the second interaction score: (1+1) / 2 = 1 point.
[0056] Finally, the time from outputting the first question to collecting the second interactive content is recorded as the progressive interaction time. The progressive interaction time, measured in seconds, is the time interval from the moment the first question is output to the moment the user's second interactive content for the second question is successfully collected. The progressive interaction time is positively correlated with the user's mastery of the knowledge point. A shorter interval generally indicates a more proficient application and better mastery of the knowledge point; a longer interval may reflect difficulties in understanding or a lack of proficiency in applying the knowledge point. For example, if the moment of outputting the first question is recorded as T1, and the moment the second interactive content is successfully received from the user is recorded as T2, the progressive interaction time is calculated as T2 minus T1, resulting in 20 seconds. This indicates that the user is relatively proficient in applying the Pythagorean theorem. If the user's grasp of the Pythagorean theorem is weak, the progressive interaction time may be longer. For example, a progressive interaction time of 75 seconds may reflect difficulties in recalling the principle or hesitation in calculation steps, indicating that the user's mastery needs improvement.
[0057] In this embodiment, by asking the same knowledge point twice in different scenarios, the evaluation bias caused by the user guessing the correct result is effectively avoided, random factors are eliminated, and the true knowledge mastery status is accurately captured. The second interaction recognizer continues the structure and scoring logic of the first interaction recognizer, so that the scores of the two interactions can be directly compared and analyzed, maintaining evaluation consistency and providing reliable data for subsequent correlation verification. The progressive interaction time intuitively reflects the user's proficiency in applying the knowledge point. The shorter the interval, the better the mastery effect, providing multi-dimensional support for subsequent verification analysis and framework optimization, and further enhancing the personalization and accuracy of the entire solution framework.
[0058] S300: Associate and verify the first and second interactive content to obtain a first verification degree, and analyze and obtain a second verification degree based on the interaction time.
[0059] In this embodiment, the first and second interactive content are correlated and verified to obtain a first verification score, and a second verification score is obtained based on the interaction time. A single score from two interactions only reflects the response to each question and cannot directly determine whether the user's understanding of the knowledge point is stable, nor can it be converted into a quantifiable mastery indicator. In actual teaching, a user may obtain the correct result in a single interaction but express the reason incorrectly; the score alone is insufficient to distinguish whether the user has thoroughly mastered the relevant knowledge point. By verifying the consistency of the reasons for two interactions, it is possible to determine whether the user's understanding of the knowledge point's principles is stable. Simultaneously, the progressive interaction time reflects application proficiency; shorter times indicate smoother application of the knowledge point. Converting this into a quantifiable second verification score, combined with the first, provides a dual basis of authenticity and proficiency for subsequent evaluation, avoiding the one-sidedness of a single score.
[0060] Step S300 in the method provided in this application embodiment includes:
[0061] Obtain the first interaction reason content and the second interaction reason content within the first interaction content and the second interaction content;
[0062] The consistency of the first interaction reason content and the second interaction reason content is verified to obtain the first verification degree;
[0063] The interaction time is input into the verification classification table, and the output is the second verification degree. The verification classification table includes a mapping relationship between the sample interaction time set and the sample second verification degree set, and the interaction time and the second verification degree are negatively correlated.
[0064] First, the first interaction reason content and the second interaction reason content within the first interaction content and the second interaction content are obtained. For example, the first interaction reason content is extracted from the processing result of S100: the solution is based on the Pythagorean theorem, the calculation step is to add the square of 3 and the square of 4 to get 25, and take the arithmetic square root to get 5 centimeters; the second interaction reason content is extracted from the processing result of S200: the solution is based on the Pythagorean theorem, the calculation step is to subtract the square of 5 from the square of 13 to get 144, and take the arithmetic square root to get 12 centimeters.
[0065] Secondly, the consistency of the first interaction reason content and the second interaction reason content is verified to obtain the first verification degree.
[0066] The process of verifying the consistency between the first interaction reason content and the second interaction reason content to obtain a first verification degree includes:
[0067] Obtain a pre-trained cause validator, wherein the cause validator is trained using a first set of sample interaction cause content, a second set of sample interaction cause content, and a first set of sample validation scores, and the first set of sample validation scores includes the cause similarity between the first set of sample interaction cause content and the second set of sample interaction cause content;
[0068] The first interaction reason content and the second interaction reason content are input into the reason verifier, and the first verification degree is obtained by outputting the result.
[0069] First, a pre-trained cause validator is obtained. This validator is trained using a first set of interactive cause content, a second set of interactive cause content, and a first set of validation scores. The first validation score includes the cause similarity between the first and second interactive cause content. The cause validator is a pre-trained model based on AI algorithms, specifically designed to determine the consistency between the first and second interactive cause content, outputting a quantified first validation score. A BERT-based dual-input fine-tuning model with deep semantic understanding capabilities is selected to construct the cause validator. Through bidirectional modeling of text semantics, the model accurately captures the semantic correlation between the two cause contents, transforming subjective consistency into an objective, quantified first validation score, thus improving validation accuracy. The first set of validation scores refers to a set of labels that corresponds one-to-one with the first and second sets of interactive cause content, containing only 1 or 0. A label of 1 is used when the first and second interactive cause content are completely correct in principle and have the same logical steps; otherwise, a label of 0 is used.
[0070] For example, the training process of the cause validator is as follows: The base model is selected from the BERT-base model, whose bidirectional semantic understanding capability accumulated during pre-training can accurately capture key semantic information such as principle descriptions and logical steps in the cause text. Sample data preparation: Sample first interaction cause content set: [Based on the Pythagorean theorem, 3 2 +4 2 =25, 25 square root = 5], [3+4-2=5, don't know the principle], [Pythagorean theorem, 3+4=7] and other 2000 reasons; Sample second interactive reason content set: [According to the Pythagorean theorem, 13 2 -5 2 =144, 144 square root = 12], [13-5+4=12, arbitrarily calculated], [13-5=8, Pythagorean theorem], etc., totaling 2000 reasons; the first validation set of the samples includes 2000 labeled values: 1 (all using the Pythagorean theorem and the steps are logically consistent), 0 (the principle is correct but the steps are incorrect), 0 (all guesses without a principle), 0 (no principle and the expression is irrelevant). The BERT-base model is fine-tuned using the above 2000 sets of samples, and training stops when the model's consistency judgment accuracy on the independent validation set reaches 95%. Finally, a pre-trained reason validator is obtained, which can output a first validation score of 0 or 1 when the first interaction reason content and the second interaction reason content are input.
[0071] Furthermore, the first interaction reason content and the second interaction reason content are input into the reason validator, and the first verification score is obtained by outputting the first verification score. The first interaction reason content and the second interaction reason content extracted in stage 1 of S300 are input into the pre-trained reason validator, and semantic analysis is used to determine whether the two are completely correct in principle and consistent in logical steps, and the output is 1 or 0 as the first verification score. For example, scenario 1: Input content: first interaction reason content [based on the Pythagorean theorem, 3 2 +4 2 =25, 25 square root = 5], the second interactive reason content [according to the Pythagorean theorem, 13 2 -5 2 =144, 144 square root = 12]; Reason verifier analysis: Both apply the Pythagorean theorem, the principle is completely correct, the logic is applied in the forward and reverse directions respectively, both conform to the theorem derivation rules, and the steps are consistent; Output first verification degree: 1. Scenario 2: Input content: First interaction reason content [Based on the Pythagorean theorem, 3 2 +4 2 =25, 25 square root = 5], Second interactive reason content [13-5=8, based on intuition]; Reason verifier analysis: The first reason principle is correct and the steps are consistent, the second reason principle is wrong; Output first verification degree: 0.
[0072] Furthermore, the interaction time is input into a verification classification table, and a second verification score is output. The verification classification table includes a mapping relationship between a set of sample interaction times and a set of sample second verification scores, with interaction time and second verification scores being negatively correlated. By using a pre-defined verification classification table, discrete time data is transformed into a 0-1 quantified second verification score, and the negative correlation between the two is clearly defined. This unifies the evaluation criteria, avoids subjective interpretation biases in the time dimension, and provides quantitative basis for subsequent integrated evaluation. The verification classification table includes a set of sample interaction times and a set of sample second verification scores. Interaction time and second verification score are negatively correlated: the longer the interaction time, the lower the second verification score; the shorter the interaction time, the higher the second verification score. The set of sample interaction times refers to typical time interval samples selected for users answering similar questions, covering scenarios such as rapid response, normal thinking, and hesitation / stuttering. The set of sample second verification scores refers to 0-1 quantified values corresponding one-to-one with the time intervals, labeled according to the negative correlation rule; the larger the value, the higher the proficiency. For example, using user response data from a Pythagorean theorem teaching scenario, a validation classification table is constructed as follows: Sample interaction time set: [Quick response: 10-30 seconds; Normal thinking: 31-60 seconds; Hesitation / Stuttering: 61-129 seconds; Severe stuttering: 121 seconds or more]; Sample second validation score set: [Quick response: 0.9; Normal thinking: 0.6; Hesitation / Stuttering: 0.3; Severe stuttering: 0.1]. The progressive interaction time is extracted from S200, and the validation classification table is queried to match the sample second validation score for the corresponding time interval, which is the final second validation score. For example, progressive interaction time: 20 seconds; querying the validation classification table, 20 seconds belongs to the 10-30 second interval, which is considered a quick response; output second validation score: 0.9.
[0073] In this embodiment, the first verification score effectively determines the stability of a user's understanding of the principles of a knowledge point by comparing the logical consistency of the two sets of reasons, thus avoiding accidental deviations from a single interaction. The second verification score converts the interaction time into a 0-1 metric, with a larger value for shorter time, intuitively reflecting the fluency of the user's access to the knowledge point. The combination of these two scores provides a dual basis for subsequent fusion evaluation, further improving the accuracy of judging the user's knowledge status.
[0074] S400: Calculate the fusion verification score based on the first verification score and the second verification score, correct the first interaction score and the second interaction score to obtain the fusion interaction score, optimize the solution framework to obtain the optimal solution framework, and provide solution instruction to the user.
[0075] In this embodiment, a fusion verification score is calculated based on the first verification score and the second verification score. A fusion interaction score is obtained by correcting the first interaction score and the second interaction score. The solution framework is then optimized to obtain the optimal solution framework for user instruction. The key to optimizing the solution framework is finding a personalized solution that is both effective and time-efficient. By evaluating both learning weight and efficiency weight, and iteratively filtering the framework using historical teaching data, the optimization process can both align with the user's actual learning ability and consider time costs, ultimately converging to the optimal framework with the highest teaching adaptability.
[0076] Step S400 in the method provided in this application embodiment includes:
[0077] The fusion verification score is calculated based on the first verification score and the second verification score.
[0078] Based on the first interaction score and the second interaction score, a basic interaction score is calculated, and a correction calculation is performed based on the fusion verification degree to obtain a fusion interaction score.
[0079] Based on the fusion interaction score and fusion verification degree, the user's solution framework is optimized to obtain the optimal solution framework.
[0080] First, the fusion verification score is calculated based on the first and second verification scores. The fusion verification score is a quantitative indicator that combines the first and second verification scores, with a value ranging from 0 to 1. By integrating the information from both scores through a weighted average, a comprehensive picture of the user's true mastery of the knowledge points and their fluency in application can be obtained, providing a comprehensive basis for subsequent evaluation. For example, if the first verification score is 1, the second verification score is 0.9, and the weights are set to 0.5 respectively, the weighted average yields a fusion verification score of 0.95; if the first verification score is 0, the second verification score is 0.3, and the weights are set to 0.5 respectively, the weighted average yields a fusion verification score of 0.15.
[0081] Secondly, based on the first and second interaction scores, a basic interaction score is calculated, and a fusion interaction score is obtained by correcting the result based on the fusion verification score. The basic interaction score is calculated by taking the arithmetic mean of the first and second interaction scores, reflecting the user's original performance in both responses. The fusion verification score is then used to weight the basic interaction score: Fusion Interaction Score = Basic Interaction Score × Fusion Verification Score. A high fusion verification score indicates high reliability of the original score, and the corrected score is close to the basic score. A low fusion verification score results in a lower corrected score, better reflecting the user's true performance. For example, if the first interaction score is 1, the second interaction score is 1, the basic interaction score is (1+1) / 2 = 1, and the fusion interaction score is 1 × 0.95 = 0.95, the original score has high reliability and remains high after correction. If the first interaction score = 1, the second interaction score = 0.05; the basic interaction score = (1 + 0.05) / 2 = 0.525; the fusion interaction score = 0.525 × 0.15 ≈ 0.08. The original score is significantly corrected due to its low authenticity, exposing the problem of false mastery.
[0082] Furthermore, based on the fusion interaction score and fusion verification degree, the user's solution framework is optimized to obtain the optimal solution framework.
[0083] Specifically, based on the fusion interaction score and fusion verification degree, the user's solution framework is optimized to obtain the optimal solution framework, including:
[0084] The fusion verification degree is used as the learning weight, and the efficiency weight is calculated.
[0085] Randomly select the first solution framework from the solution framework library, wherein the first solution framework includes a teaching plan that includes teaching content and the order of teaching content;
[0086] Obtain the historical teaching scores and historical teaching times of users with the fusion interaction scores taught using the first solution framework within a historical period, and use them as the first teaching score and the first teaching time. Calculate the ratio of the average time and the first teaching time to obtain the first efficiency score.
[0087] Based on the learning weight and efficiency weight, the first teaching fitness is obtained by weighting the first teaching score and the first efficiency score to obtain the first teaching fitness.
[0088] Continue iteratively optimizing the solution framework until convergence, obtaining the optimal solution framework with the greatest teaching adaptability.
[0089] First, the fusion verification score is used as the learning weight, and an efficiency weight is calculated. The learning weight, or fusion verification score, reflects the authenticity and proficiency of the user's knowledge acquisition. It serves as the primary weight for evaluating teaching effectiveness; a higher value indicates a greater potential for the user to absorb effective teaching. The sum of the efficiency weight and the learning weight is 1. The efficiency weight is used to balance teaching time efficiency, avoiding the pursuit of results at the expense of time cost. For example, if the fusion verification score is known to be 0.6, then the learning weight = 0.6, and the efficiency weight = 1 - 0.6 = 0.4.
[0090] Secondly, a first solution framework is randomly selected from the solution framework library. This first solution framework includes a teaching plan with both the teaching content and its order. The solution framework library is a collection of pre-stored teaching plans for relevant knowledge points, with each framework containing both teaching content and its order. The first solution framework is the initial teaching plan randomly selected from the framework library. For example, the Pythagorean theorem solution framework library contains three candidate plans: Framework A includes explanations of the Pythagorean theorem formula, examples of forward and reverse applications, and basic exercises, with the teaching order being principle first, then exercises; Framework B includes examples of forward applications, derivation of the Pythagorean theorem formula, exercises of reverse applications, and reinforcement of the principle, with the teaching order being examples first, then the principle; Framework C includes animated demonstrations of the Pythagorean theorem, formula decomposition, comparison of positive and negative examples, and comprehensive exercises, with the teaching order prioritizing visualization. Framework A is randomly selected as the first solution framework.
[0091] Furthermore, the historical teaching scores and historical teaching times for users with the aforementioned fusion interaction scores, taught using the first solution framework within a historical period, are obtained as the first teaching score and the first teaching time. The ratio of the average time to the first teaching time is calculated to obtain the first efficiency score. The average time is derived from historical big data statistics and refers to the average time spent teaching similar users with the same fusion interaction score. The higher the value of the first efficiency score, the shorter the teaching time of the first solution framework. For example, historical data shows that when using framework A to teach users with a fusion interaction score of 0.45, the first teaching score is 0.7, and the first teaching time is 20 minutes; the average teaching time for similar users is 25 minutes; the first efficiency score = 25 / 20 = 1.25, which is higher than 1, indicating that framework A saves 20% of the average teaching time compared to similar users.
[0092] Then, based on the learning weight and efficiency weight, the first teaching score and the first efficiency score are weighted and calculated to obtain the first teaching fitness. Teaching fitness is an indicator that measures the overall performance of the framework, taking into account both teaching effectiveness and efficiency. For example, given a learning weight of 0.6, an efficiency weight of 0.4, a first teaching score of 0.7, and a first efficiency score of 1.25, the first teaching fitness = (0.6 × 0.7) + (0.4 × 1.25) = 0.42 + 0.5 = 0.92.
[0093] Finally, the solution framework is iteratively optimized until convergence, yielding the optimal solution framework with the highest teaching fitness. The iterative logic involves repeatedly selecting a new framework, calculating its teaching score and efficiency score, and calculating its corresponding teaching fitness, until the difference in teaching fitness over N consecutive iterations does not exceed 0.01, indicating convergence. At this point, the framework with the highest teaching fitness is the optimal solution framework. For example, still targeting a user with a fusion interaction score of 0.45, the second iteration randomly selects frame B. In the corresponding historical data, this user's second teaching score is 0.65, and the second teaching time is 18 minutes; the second efficiency score = 25 / 18 ≈ 1.39; the second teaching fitness = (0.6 × 0.65) + (0.4 × 1.39) = 0.39 + 0.556 = 0.946; the third iteration randomly selects frame C, and similarly, the third teaching fitness is 0.936; the fourth iteration randomly selects frame B, and the fourth teaching fitness is calculated to be 0.945, which is 0.009 different from the third teaching fitness, not exceeding 0.01, satisfying the convergence condition; finally, frame B is determined to be the optimal solution frame, and its teaching fitness of 0.946 is the maximum value among the three iterations.
[0094] In this embodiment, the optimal solution framework is a personalized teaching plan that is iteratively selected based on user fusion interaction score and fusion verification degree combined with historical teaching data of similar users. It can accurately adapt to the user's current knowledge status and ensure stable and reliable teaching results. At the same time, it can significantly save unnecessary teaching time while ensuring the effect. It also provides a clear basis for subsequent teaching adjustments, effectively solving the problems of content mismatch, uncertain effect and wasted time in traditional teaching, and forming a complete closed loop of evaluation-optimization-teaching.
[0095] The embodiments of this application, through the specific implementation methods described above, achieve the following technical effects:
[0096] This application provides a method and system for constructing an iterative problem-solving framework based on a multimodal large model. It integrates authenticity and proficiency by calculating the fusion verification degree, then corrects the original performance with the fusion interaction score, and subsequently iteratively optimizes the solution framework by combining historical data. Finally, it outputs the optimal solution framework and implements a complete process for personalized teaching. This not only achieves accurate assessment of the user's knowledge status and personalized adaptation of teaching plans, but also stably ensures teaching effectiveness and saves teaching time efficiently. It effectively solves the problems of content misalignment, uncertain effects, and time-consuming redundancy in traditional teaching. At the same time, the closed-loop process provides a clear basis for subsequent teaching adjustments, forming a teaching support system that is scientifically evaluated, accurately adapted, highly efficient, and iteratively sustainable.
[0097] Example 2, as Figure 2 As shown, this application provides a system for constructing an iterative problem-solving framework based on a multimodal large model, the system comprising:
[0098] The initial interaction scoring module 11 is used to output the first question, obtain the first interaction content of the user's interaction with the first question, identify it, and obtain the first interaction score;
[0099] The progressive interaction timing module 12 is used to output the second question, obtain the second interaction content of the user's interaction with the second question, identify it, obtain the second interaction score, and obtain the progressive interaction time.
[0100] The association verification analysis module 13 is used to verify and identify the association between the first interaction content and the second interaction content, obtain a first verification degree, and analyze and obtain a second verification degree based on the interaction time.
[0101] The fusion correction and optimization module 14 is used to calculate the fusion verification score based on the first verification score and the second verification score, correct the first interaction score and the second interaction score to obtain the fusion interaction score, optimize the solution framework to obtain the optimal solution framework, and provide solution instruction to the user.
[0102] In one embodiment, the initial interaction scoring module 11 is further configured to:
[0103] Output the first question and obtain the user's first interaction content in response to the first question;
[0104] Based on natural language processing, extract the first interaction result content and the first interaction reason content of the first interaction content;
[0105] The first interaction result content and the first interaction reason content are input into the first interaction recognizer, and the first interaction result score and the first interaction reason score are output to obtain the first interaction score.
[0106] Specifically, the first interaction result content and the first interaction reason content are input into the first interaction recognizer, and the first interaction result score and the first interaction reason score are output. The first interaction score is calculated, including:
[0107] Obtain a pre-trained first interaction recognizer, wherein the first interaction recognizer includes a first interaction result recognition branch and a first interaction reason recognition branch. The first interaction result recognition branch is trained using a sample first interaction result content set and a sample first interaction result score set for the first question. The first interaction reason recognition branch is trained using a sample first interaction reason content set and a sample first interaction reason score set for the first question.
[0108] Input the first interaction result content and the first interaction reason content into the first interaction result recognition branch and the first interaction reason recognition branch in the first interaction recognizer, and output to obtain the first interaction result score and the first interaction reason score;
[0109] The first interaction score is calculated based on the first interaction result score and the first interaction reason score.
[0110] In one embodiment, the progressive interactive timing module 12 is further configured to:
[0111] Output the second question and obtain the second interaction content of the user's interaction with the second question, wherein the first question and the second question are related questions;
[0112] Extract the second interaction result content and the second interaction reason content based on natural language processing;
[0113] The second interaction result content and the second interaction reason content are input into the second interaction recognizer, and the second interaction result score and the second interaction reason score are output. The second interaction score is calculated. The second interaction recognizer includes a second interaction result recognition branch and a second interaction reason recognition branch.
[0114] The time from the output of the first question to the collection of the second interactive content is used as the progressive interaction time.
[0115] In one embodiment, the association verification analysis module 13 is further configured to:
[0116] Obtain the first interaction reason content and the second interaction reason content within the first interaction content and the second interaction content;
[0117] The consistency of the first interaction reason content and the second interaction reason content is verified to obtain the first verification degree;
[0118] The interaction time is input into the verification classification table, and the output is the second verification degree. The verification classification table includes a mapping relationship between the sample interaction time set and the sample second verification degree set, and the interaction time and the second verification degree are negatively correlated.
[0119] The process of verifying the consistency between the first interaction reason content and the second interaction reason content to obtain a first verification degree includes:
[0120] Obtain a pre-trained cause validator, wherein the cause validator is trained using a first set of sample interaction cause content, a second set of sample interaction cause content, and a first set of sample validation scores, and the first set of sample validation scores includes the cause similarity between the first set of sample interaction cause content and the second set of sample interaction cause content;
[0121] The first interaction reason content and the second interaction reason content are input into the reason verifier, and the first verification degree is obtained by outputting the result.
[0122] In one embodiment, the fusion correction optimization module 14 is further configured to:
[0123] The fusion verification score is calculated based on the first verification score and the second verification score.
[0124] Based on the first interaction score and the second interaction score, a basic interaction score is calculated, and a correction calculation is performed based on the fusion verification degree to obtain a fusion interaction score.
[0125] Based on the fusion interaction score and fusion verification degree, the user's solution framework is optimized to obtain the optimal solution framework.
[0126] Specifically, based on the fusion interaction score and fusion verification degree, the user's solution framework is optimized to obtain the optimal solution framework, including:
[0127] The fusion verification degree is used as the learning weight, and the efficiency weight is calculated.
[0128] Randomly select the first solution framework from the solution framework library, wherein the first solution framework includes a teaching plan that includes teaching content and the order of teaching content;
[0129] Obtain the historical teaching scores and historical teaching times of users with the fusion interaction scores taught using the first solution framework within a historical period, and use them as the first teaching score and the first teaching time. Calculate the ratio of the average time and the first teaching time to obtain the first efficiency score.
[0130] Based on the learning weight and efficiency weight, the first teaching fitness is obtained by weighting the first teaching score and the first efficiency score to obtain the first teaching fitness.
[0131] Continue iteratively optimizing the solution framework until convergence, obtaining the optimal solution framework with the greatest teaching adaptability.
[0132] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0133] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0134] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.
Claims
1. A method for constructing an iterative problem solving framework based on a multi-modal large model, characterized in that, The method comprises: output the first question, and obtain the first interaction content of the user interacting with the first question, identify, and obtain the first interaction score; output the second question, and obtain the second interaction content of the user interacting with the second question, identify, and obtain the second interaction score, and obtain the progressive interaction time; correlation verification and identification of the first interaction content and the second interaction content are obtained, and the second verification degree is analyzed according to the interaction time, including: obtain the first interaction reason content and the second interaction reason content in the first interaction content and the second interaction content; the first verification degree is obtained by performing consistency verification on the first interaction reason content and the second interaction reason content, including: obtain the pre-trained reason verifier, wherein the reason verifier is trained by using a sample first interaction reason content set, a sample second interaction reason content set, and a sample first verification degree set, and the sample first verification degree includes the reason similarity of the sample first interaction reason content and the sample second interaction reason content; the first interaction reason content and the second interaction reason content are input into the reason verifier, and the first verification degree is output; the interaction time is input into the verification classification table, and the second verification degree is output, wherein the verification classification table includes the mapping relationship between the sample interaction time set and the sample second verification degree set, and the interaction time and the second verification degree are negatively correlated; the fusion verification degree is calculated according to the first verification degree and the second verification degree, the first interaction score and the second interaction score are corrected to obtain the fusion interaction score, the solving framework is optimized, and the optimal solving framework is obtained, and the user is taught to solve, including: the fusion verification degree is calculated according to the first verification degree and the second verification degree; the base interaction score is calculated according to the first interaction score and the second interaction score, and the fusion interaction score is obtained by correction calculation based on the fusion verification degree; the optimal solving framework is obtained by optimizing the solving framework of the user according to the fusion interaction score and the fusion verification degree, including: the fusion verification degree is used as a learning weight, and an efficiency weight is calculated, wherein the efficiency weight is 1 minus the learning weight; a first solving framework is randomly selected in a solving framework library, wherein the first solving framework includes a teaching scheme of teaching content and teaching content order; the historical teaching score and the historical teaching time of the user with the fusion interaction score in the historical time are obtained by using the first solving framework for teaching, and are used as the first teaching score and the first teaching time, the ratio of the average time to the first teaching time is calculated, and the first efficiency score is obtained; the first teaching fitness is obtained by weighted calculation of the first teaching score and the first efficiency score according to the learning weight and the efficiency weight; the optimal solving framework with the maximum teaching fitness is obtained by continuing to perform solving framework iteration optimization until convergence, wherein the iteration is a process of repeatedly selecting a new framework, calculating a teaching score, an efficiency score, and a teaching fitness, and the convergence condition is that the difference between the teaching fitnesses of consecutive N iterations is not more than 0.
01. 2.The method of claim 1, wherein the method further comprises: receiving a first query from the user; and determining a first answer to the first query based on the first query and the first answer. Output the first question, and obtain the first interaction content of the user interacting with the first question, identify, and obtain the first interaction score, comprising: Output the first question, and obtain the first interaction content of the user interacting with the first question; Based on natural language processing, extract the first interaction result content and the first interaction reason content of the first interaction content; Input the first interaction result content and the first interaction reason content into the first interaction identifier, output the first interaction result score and the first interaction reason score, and calculate the first interaction score. 3.The method of claim 2, wherein, Input the first interaction result content and the first interaction reason content into the first interaction identifier, output the first interaction result score and the first interaction reason score, and calculate the first interaction score, comprising: Obtain a pre-trained first interaction identifier, wherein the first interaction identifier includes a first interaction result identification branch and a first interaction reason identification branch, the first interaction result identification branch is trained using a sample first interaction result content set and a sample first interaction result score set of the first question, and the first interaction reason identification branch is trained using a sample first interaction reason content set and a sample first interaction reason score set of the first question; Input the first interaction result content and the first interaction reason content into the first interaction result identification branch and the first interaction reason identification branch in the first interaction identifier, output the first interaction result score and the first interaction reason score; According to the first interaction result score and the first interaction reason score, calculate the first interaction score. 4.The method of claim 1, wherein, Output the second question, and obtain the second interaction content of the user interacting with the second question, identify, and obtain the second interaction score, and obtain the progressive interaction time, comprising: Output the second question, and obtain the second interaction content of the user interacting with the second question, wherein the first question and the second question are the same family question; Based on natural language processing, extract the second interaction result content and the second interaction reason content of the second interaction content; Input the second interaction result content and the second interaction reason content into the second interaction identifier, output the second interaction result score and the second interaction reason score, and calculate the second interaction score, wherein the second interaction identifier includes a second interaction result identification branch and a second interaction reason identification branch; Obtain the time from outputting the first question to collecting the second interaction content as the progressive interaction time.
5. The system for constructing an iterative problem solving framework based on a multi-modal large model, characterized in that, The system for implementing the method for constructing the iterative question solving framework based on the multi-modal large model of any one of claims 1-4, comprising: An initial interaction scoring module for outputting a first question, and obtaining the first interaction content of the user interacting with the first question, identifying, and obtaining the first interaction score; A progressive interaction timing module for outputting a second question, and obtaining the second interaction content of the user interacting with the second question, identifying, and obtaining the second interaction score, and obtaining the progressive interaction time; An association verification analysis module for association verification identification of the first interaction content and the second interaction content, obtaining the first verification degree, and analyzing the second verification degree according to the interaction time, comprising: Obtaining first interaction reason content and second interaction reason content in the first interaction content and the second interaction content; Consistency verification is carried out on the first interaction reason content and the second interaction reason content, and a first verification degree is obtained, including: Obtaining a pre-trained reason verifier, wherein the reason verifier is trained by a sample first interaction reason content set, a sample second interaction reason content set and a sample first verification degree set, and the sample first verification degree includes the reason similarity of the sample first interaction reason content and the sample second interaction reason content; The first interaction reason content and the second interaction reason content are input into the reason verifier, and the first verification degree is output; The interaction time is input into a verification classification table, and the second verification degree is output, wherein the verification classification table includes a mapping relationship between a sample interaction time set and a sample second verification degree set, and the interaction time and the second verification degree are negatively correlated; A fusion correction optimization module is used to calculate a fusion verification degree according to the first verification degree and the second verification degree, correct the first interaction score and the second interaction score to obtain a fusion interaction score, optimize the solving framework, and obtain an optimal solving framework for solving teaching of the user, including: According to the first verification degree and the second verification degree, a fusion verification degree is calculated; According to the first interaction score and the second interaction score, a basic interaction score is calculated, and a fusion interaction score is obtained by correction calculation based on the fusion verification degree; According to the fusion interaction score and the fusion verification degree, the solving framework of the user is optimized to obtain an optimal solving framework, including: The fusion verification degree is used as a learning weight, and an efficiency weight is calculated, wherein the efficiency weight is 1 minus the learning weight; A first solving framework is randomly selected in a solving framework library, wherein the first solving framework includes a teaching scheme of teaching content and teaching content order; The historical teaching score and the historical teaching time of the user with the fusion interaction score using the first solving framework in the historical time are obtained as the first teaching score and the first teaching time, the ratio of the average time and the first teaching time is calculated to obtain the first efficiency score; According to the learning weight and the efficiency weight, the first teaching score and the first efficiency score are weighted to obtain a first teaching fitness; The solving framework iteration optimization is continued until convergence, and the optimal solving framework with the maximum teaching fitness is obtained, wherein the iteration is a process of repeatedly selecting a new framework, calculating a teaching score, an efficiency score and a teaching fitness, and the convergence condition is that the difference of the teaching fitness of continuous N iterations is not more than 0.01.
Citation Information
Patent Citations
Interactive teaching method, device and equipment based on multi-modal large model and medium
CN120219122A
Self-adaptive intelligent teaching content recommendation system
CN120429482A