Training method of question explanation model and question explanation method
By using a multimodal large model training method, problem images and explanation framework samples are obtained to generate whiteboard text and multi-turn dialogues. This solves the problem of single-modal systems handling complex problems, realizes automated and accurate intelligent explanation, and improves the system's adaptability and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING YUANLI WEILAI SCI & TECH CO LTD
- Filing Date
- 2026-01-15
- Publication Date
- 2026-04-28
AI Technical Summary
Existing single-modal teaching systems struggle to handle problems with images, especially those with complex structures or multi-element relationships. This leads to distorted understanding of visual information or omission of key problem-solving clues, limiting the coverage and effectiveness of AI-powered explanations.
By acquiring training samples containing question images and explanation frameworks, a multimodal large model is used to train and generate whiteboard text and multi-turn dialogues, realizing an automated process from image understanding to explanation. This includes hierarchical construction of training samples and model fine-tuning, reducing reliance on manual intervention and improving the model's adaptability and accuracy.
It enables automatic and accurate understanding and personalized explanation of complex problems, significantly improving the adaptability and accuracy of the intelligent explanation system, reducing labor costs, and covering a wider range of question types, including math, physics and other subjects with diagrams.
Smart Images

Figure CN121938243A_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of computer technology, and in particular to a training method and a problem explanation method for a problem explanation model. Background Technology
[0002] In the field of intelligent education, model-based instructional guidance methods are gradually becoming important tools to assist learning. However, current mainstream methods are typically monomodal text input modes, meaning they can only handle questions and explanation requests in plain text format. In actual teaching scenarios, many exercises (especially in science and engineering subjects) are accompanied by diagrams, illustrations, or geometric figures. This visual information is crucial for understanding the question and solving it. Therefore, existing monomodal systems are difficult to directly apply to such illustrated questions, severely limiting the coverage and effectiveness of AI-powered explanations.
[0003] To address this challenge, existing technologies typically involve manually converting image content into text descriptions or structured codes (such as SVG, geometric relationship descriptions, etc.), and then embedding the converted text into the input for model processing. While this method can handle some simple, regular graphics, it often fails to fully and accurately convey visual information for images containing complex structures, multi-element relationships, or natural scenes, leading to distorted model understanding or missed key problem-solving clues. Therefore, there is an urgent need for a multimodal teaching guidance method that can automatically and accurately understand complex image content in problems and deeply integrate it with text information to achieve high-quality intelligent explanations for a wider range of problem types. Summary of the Invention
[0004] In view of the above, embodiments of this specification provide a training method for a problem explanation model and a problem explanation method. One or more embodiments of this specification also relate to a training device for a problem explanation model and a problem explanation device, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a method for training a problem explanation model is provided, comprising: Obtain a first training sample and a second training sample. The first training sample includes a sample question image, a sample explanation framework, and corresponding whiteboard text labels. The second training sample includes a sample question image, a sample explanation framework, sample whiteboard text, and a sample multi-turn dialogue, as well as corresponding explanation text labels. The sample explanation framework is determined based on the sample question image and sample question analysis. Based on the first and second training samples, the initial question explanation model is trained to obtain the target question explanation model that meets the training stopping condition.
[0006] According to a second aspect of the embodiments of this specification, a method for explaining problems is provided, including: Obtain the target question image and the target explanation framework, wherein the target explanation framework is determined based on the target question image and the target question analysis; Based on the target question image and the target explanation framework, the target whiteboard text is generated through the target question explanation model. The target question explanation model is trained using the training method of the above question explanation model. Obtain the target multi-turn dialogue content, which includes the historical dialogue content before the current turn and the user's response content in the current turn; Based on the target question image, target explanation framework, target whiteboard text, and target multi-turn dialogue content, the target explanation text for the current round is generated through the target question explanation model. Based on the target whiteboard text and the target explanation text, generate the synchronous explanation data for the current round.
[0007] According to a third aspect of the embodiments of this specification, a training apparatus for a problem explanation model is provided, comprising: The acquisition module is configured to acquire a first training sample and a second training sample. The first training sample includes a sample question image and a sample explanation framework, as well as corresponding whiteboard text labels. The second training sample includes a sample question image, a sample explanation framework, sample whiteboard text, and a sample multi-turn dialogue, as well as corresponding explanation text labels. The sample explanation framework is determined based on the sample question image and sample question parsing. The training module is configured to train the initial question explanation model based on the first training sample and the second training sample to obtain the target question explanation model that meets the training stopping condition.
[0008] According to a fourth aspect of the embodiments of this specification, a problem-solving apparatus is provided, comprising: The first acquisition module is configured to acquire the target question image and the target explanation framework, wherein the target explanation framework is determined based on the target question image and the target question analysis; The first generation module is configured to generate target whiteboard text based on the target question image and the target explanation framework, using the target question explanation model. The target question explanation model is trained using the training method of the above-mentioned question explanation model. The second acquisition module is configured to acquire the target multi-turn dialogue content, which includes the historical dialogue content before the current turn and the user response content of the current turn; The second generation module is configured to generate the target explanation text for the current round based on the target question image, the target explanation framework, the target whiteboard text, and the target multi-turn dialogue content, using the target question explanation model. The third generation module is configured to generate synchronized explanation data for the current round based on the target whiteboard text and the target explanation text.
[0009] According to a fifth aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the training method and the problem explanation method of the above-mentioned problem explanation model are implemented.
[0010] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer-executable instructions, which, when executed by a processor, implement the steps of the training method and the problem explanation method of the above-described problem explanation model.
[0011] According to a seventh aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the training method and the problem explanation method of the above-described problem explanation model.
[0012] One embodiment of this specification provides a training method for a question explanation model, comprising: acquiring a first training sample and a second training sample, wherein the first training sample includes a sample question image and a sample explanation framework, as well as corresponding whiteboard text labels; the second training sample includes a sample question image, a sample explanation framework, sample whiteboard text, and a sample multi-turn dialogue, as well as corresponding explanation text labels, wherein the sample explanation framework is determined based on the sample question image and sample question parsing; and training an initial question explanation model based on the first training sample and the second training sample to obtain a target question explanation model that meets the training stopping condition. First, the model is trained using a first training sample containing sample question images, sample explanation frames, and corresponding whiteboard text labels to learn the ability to extract key visual information from images and generate whiteboard text. Then, a second training sample containing sample question images, sample explanation frames, sample whiteboard text, sample multi-turn dialogues, and corresponding explanation text labels enables the model to further master the ability to perform contextual reasoning and personalized guidance based on dialogue history. This achieves full automation of the process, from automatically and accurately understanding visual information from complex question images to generating logically coherent whiteboard text, and then to providing strategic and personalized explanations based on dialogue history and real-time responses. This significantly improves the adaptability of the target question explanation model to a wide range of question types (especially those containing complex graphics) and reduces manual costs. Attached Figure Description
[0013] Figure 1This is a flowchart illustrating a training method for a problem explanation model provided in one embodiment of this specification; Figure 2 This is a flowchart illustrating a problem-solving method provided in one embodiment of this specification; Figure 3 This is a flowchart illustrating the processing steps of a training method for a problem explanation model provided in one embodiment of this specification. Figure 4 This is a schematic diagram of the structure of a training device for a problem explanation model provided in one embodiment of this specification; Figure 5 This is a schematic diagram of the structure of a problem-solving device provided in one embodiment of this specification; Figure 6 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0014] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0015] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0016] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0017] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0018] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0019] Large Language Models (LLMs), also known as large models, are artificial intelligence models that use machine learning to understand and generate human language. They can perform a wide range of tasks, including text summarization, translation, and sentiment analysis. LLMs are characterized by their massive scale, containing billions of parameters that help them learn complex patterns in language data.
[0020] Multimodal large models refer to large-scale deep learning models that can simultaneously understand, process, and generate multiple types of information such as text, images, audio, and video. Their core lies in mapping information from different modalities into a unified semantic representation space through pre-training with massive amounts of cross-modal data (such as text-image pairs and video-subtitle pairs). This enables the model to deeply grasp the intrinsic connections and alignments between language, visual, and other content, thereby giving rise to cross-modal deep understanding, reasoning, and creative capabilities.
[0021] Supervised Fine-Tuning of Large Models (SFT) is a machine learning strategy primarily used to handle large-scale pre-trained models. These models are typically trained to understand or learn from large amounts of unlabeled data, such as web pages, books, and other types of text. However, despite being trained, they may not directly solve specific tasks, such as text classification or sentiment analysis. This is where supervised fine-tuning comes in. This process mainly involves further training using labeled datasets. The parameters of some (or all) layers of the model are updated based on this new task; they are fine-tuned to better perform this specific task.
[0022] This specification provides a training method for a problem explanation model and a problem explanation method. This specification also relates to a training device for a problem explanation model and a problem explanation device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0023] See Figure 1 , Figure 1 A flowchart illustrating a training method for a problem-solving model according to an embodiment of this specification is shown. From a hardware perspective, the execution entity of this method can be a server. From a programming perspective, the execution entity of this method can be a model training program mounted on the server.
[0024] like Figure 1 As shown, the method specifically includes the following steps 102-104.
[0025] Step 102: Obtain the first training sample and the second training sample. The first training sample includes the sample question image and the sample explanation framework, as well as the corresponding whiteboard text labels. The second training sample includes the sample question image, the sample explanation framework, the sample whiteboard text, and the sample multi-turn dialogue, as well as the corresponding explanation text labels. The sample explanation framework is determined based on the sample question image and the sample question analysis.
[0026] The sample question image refers to the original image containing the question content used for model training. It is one of the model's multimodal inputs, designed to allow the model to directly learn to understand the question from the image, rather than relying on manually converted text descriptions.
[0027] The questions can be from various subjects such as mathematics, Chinese, physics, and chemistry. This application does not limit the specific subjects involved in the questions.
[0028] Among them, the question analysis refers to the standard answer and solution process of the question, which is the basis for generating the "explanation framework".
[0029] The explanation framework is an outline determined based on the analytical steps of the corresponding problem.
[0030] In practical applications, the explanation framework can provide a basis for determining the blackboard text in the first training sample. Correspondingly, the explanation framework can also provide a basis for determining the multi-turn dialogue and explanation text labels in the second training sample.
[0031] In practical applications, the explanation framework can include multiple steps, each with a corresponding correct answer. These correct answers serve as the basis for the model to determine whether the user's response is correct. Specifically, if the model asks the user question 1 (question 1 is about step 1 in the explanation framework), and the user's answer matches the correct answer for step 1 in the question guidance outline, the model can then ask the user question 2 (question 2 is about step 2 in the explanation framework). If the user's answer does not match the correct answer for step 1 in the explanation framework, the model can continue to ask more detailed questions about step 1 in the explanation framework.
[0032] For example, the problem image is a rectangle, labeled as 8cm long and 5cm wide. The problem text is: "Calculate the area of this rectangle." Corresponding explanation framework: Step 1: Identify the image and extract information.
[0033] Guiding question: "Observe the picture in the question. What shape is this? What information is labeled on the picture?" Correct answer / key information: This is a rectangle, length = 8 cm, width = 5 cm.
[0034] Step 2: Review the area formula.
[0035] Leading question: "Which formula do we need to use to calculate the area of a rectangle?" Correct answer / key information: The area of a rectangle = length × width.
[0036] Step 3: Substitute the numerical values for calculation.
[0037] Leading question: "Please substitute the length and width values into the formula and calculate the result." Correct answer / key information: Area = 8 × 5 = 40.
[0038] Step 4: Add units and final answer.
[0039] Guiding question: "The calculated number represents the area, and we need to add an appropriate unit to it. Think about it, what should the unit of area be?" Correct answer / key information: The area is 40 square centimeters (cm²).
[0040] Step 5: Quick Verification (Optional Extension) Guiding question: "How can we quickly verify the reasonableness of an answer? For example, what should the area be larger than, and what should it be smaller than?" Correct answer / key information: The area (40) should be greater than half of the perimeter ((8+5)*2 / 2=13) and less than the product of the length and width (8*5=40 itself).
[0041] Blackboard text labels: These refer to the reference blackboard content corresponding to the "sample question image" and "sample explanation framework." It typically includes the problem-solving approach, the necessary formulas, theorems, key steps, and conclusions, serving as the "standard answer" for the blackboard content generated by the model's learning process.
[0042] In practical applications, the written text can provide a basis for subsequently determining the labels of multi-turn dialogues and explanation texts in the second training sample.
[0043] Following the previous example, the text written on the blackboard could be: Question: Calculate the area of a rectangle. Shape: Rectangle Given: Length = 8 cm, Width = 5 cm Formula: Area = Length × Width Calculate: 8 × 5 = 40 Answer: 40 cm² Multi-turn dialogues refer to simulated real teaching interactions, with annotated dialogue history. They typically contain multiple rounds of "teacher question-student answer" pairings, serving as the context for generating the current round of explanation.
[0044] Using the previous example, a multi-round dialogue consists of 6 rounds.
[0045] Round 1: Framework steps: Step 1 (Identify the image and extract information).
[0046] Round 1 Dialogue: Teacher (using a model to guide): "Let's look at the question first. Observe the picture, can you tell me what shape this is? What information is labeled on the picture?" The student answered, "It's a rectangle, with a length of 8 centimeters and a width of 5 centimeters." The teacher replied, "Very good, you observed very carefully. This is a rectangle, 8 centimeters long and 5 centimeters wide. Next, to calculate its area, what do we need to know first?" Round 2: Framework steps: Step 2 (Review the area formula).
[0047] Round 2 Dialogue: Teacher (using a model to guide): "Yes, we need the area formula. Do you remember the formula for calculating the area of a rectangle?" The student answered, "Hmm... is it the length plus the width, then multiplied by 2?" The teacher replied: "'Length plus width multiplied by 2' is the formula for calculating the perimeter. Area refers to the size of a shape, and the formula is 'length multiplied by width.' Let's remember it again: Area of a rectangle = length × width." Framework Steps: Step 2 (Review the area formula - confirm understanding) Round 3: Teacher (using a model to guide): "Now, please say the correct formula for the area of a rectangle." Round 3 Dialogue: The student answered: "The area is equal to the length multiplied by the width." The teacher replied, "Correct. Now we can use this formula to calculate it. Please substitute the numbers from the problem into the formula." Round 4: Framework steps: Step 3 (substitute numerical values for calculation).
[0048] Round 4 Dialogue: Teacher (using a model): "Please substitute the length of 8 cm and the width of 5 cm into the formula 'Area = Length × Width' and calculate the result." The student answered, "8 times 5 equals 40." The teacher replied, "The calculation is accurate; the result is 40. However, is the number '40' enough for a complete answer? What's missing?" Round 5: Framework steps: Step 4 (supplementing units and final solution).
[0049] Round 5 Dialogue: Teacher (using a model to guide): "Think about it, '40' is just a number. What does it represent? In area calculations, we must add the unit." The student answered, "It's area, so the unit should be square centimeters." The teacher replied, "That's absolutely correct. Because the unit of length is centimeters, the unit of area is square centimeters. Therefore, the area of this rectangle is 40 square centimeters." Round 6 (Optional Expansion): Framework Steps: Step 5 (Quick Validation) Round 6 Dialogue: Teacher (using a model to guide): "We can quickly check the answers. The area of a rectangle is usually larger than half its perimeter. Can you calculate approximately what half the perimeter of this rectangle is?" The student answered: "The perimeter is (8+5)*2=26, and half of that is 13. 40 is greater than 13, so it seems reasonable." The teacher replied, "Yes, such quick comparisons can increase our confidence in the answers. You did very well." This dialogue example strictly follows the established explanation framework, progressing step by step. By analyzing the student's responses in each round, the model directly generates guiding questions or feedback for the next step, ultimately completing the entire explanation process from image recognition, formula review, calculation and solution to answer verification.
[0050] Explanation text label: This refers to the ideal "teacher response" labeled for the current round of "student answer" within the specific context of the "sample multi-turn dialogue." It serves as the "standard answer" for the model to learn and generate explanation text. The explanation text can end with "Do you understand?" or a question to test the user's response to the explanation.
[0051] The "teacher's reply" in the above dialogue example is a specific manifestation of the "explanation text" in a specific teaching interaction context.
[0052] In one optional implementation of this embodiment, obtaining the first training sample and the second training sample includes: Based on the sample question images and sample explanation framework, label the corresponding blackboard text. The first training sample is constructed based on the sample question image, sample explanation framework and whiteboard text label; Based on the sample question images and sample explanation framework, sample whiteboard text is generated through the first intermediate model, wherein the first intermediate model is obtained by training the initial question explanation model based on the first training samples; Based on the sample question images, sample explanation frameworks, and sample whiteboard text, a sample multi-turn dialogue is constructed, and the corresponding explanation text labels are annotated. A second training sample is constructed based on sample question images, sample explanation frameworks, sample whiteboard text, sample multi-turn dialogues, and explanation text labels. Accordingly, based on the first and second training samples, the initial problem explanation model is trained to obtain the target problem explanation model that meets the training stopping condition, including: The first intermediate model is trained based on the second training sample to obtain the target question explanation model.
[0053] The first intermediate model is an intermediate state model that is trained on the initial question explanation model based on the first training sample and has the ability to generate blackboard text based on the question image and explanation framework.
[0054] As shown above, in this scheme, firstly, the initial problem explanation model is trained based on the first training samples to obtain a first intermediate model. This first intermediate model can generate whiteboard text based on the problem image and explanation framework. Therefore, based on the sample problem image and sample explanation framework, the first intermediate model can generate whiteboard text corresponding to the sample problem image and sample explanation framework. The whiteboard text corresponding to the sample problem image and sample explanation framework generated by the first intermediate model can be directly used as the sample whiteboard text in the second training samples, or the whiteboard text generated by the model can be used as the preliminary sample whiteboard text, and then manually reviewed and corrected to determine the sample whiteboard text in the second training samples.
[0055] In this embodiment, a progressive two-stage training strategy is employed. First, the model uses a first training sample to master the core ability to generate structured whiteboard text from the question image and explanation framework. Then, based on this, a second training sample containing multi-turn dialogue scenarios is constructed, enabling the model to further learn the ability to provide coherent and accurate explanations by combining context and whiteboard content. This approach not only reduces the training difficulty of complex end-to-end models through task decomposition, improving training efficiency and stability, but also ensures consistency between the training process and the evolution of model capabilities by using the whiteboard text generated by the model itself as subsequent training data. Furthermore, it flexibly supports manual review and correction of the generated whiteboard text, balancing automation and quality control. Ultimately, the resulting question explanation model can deeply integrate visual understanding, content generation, and dialogue interaction, achieving high-quality, adaptive explanations of various questions (especially those with images), significantly improving the practicality, accuracy, and scalability of the intelligent teaching system.
[0056] In practice, the sample whiteboard text in the second training sample can also be determined manually, without any specific limitations.
[0057] In one optional implementation of this embodiment, before obtaining the first training sample, the following steps are further included: Obtain sample question images and sample question analyses; Based on the sample question images and sample question analyses, the corresponding explanation framework labels are annotated; Based on the sample question images and sample question parsing, as well as the explanation framework labels, a third training sample is constructed. Based on the third training sample, the initial question explanation model is trained to obtain a second intermediate model. Accordingly, based on the first and second training samples, the initial problem explanation model is trained to obtain the target problem explanation model that meets the training stopping condition, including: Based on the first and second training samples, the second intermediate model is trained to obtain the target question explanation model.
[0058] The second intermediate model is an intermediate state model that is trained on the initial question explanation model based on the question image, question analysis, and explanation framework labels. It has the ability to generate explanation frameworks based on question images and question analysis.
[0059] In practical applications, a large language model can be used to generate preliminary explanation framework labels based on the sample question image and the corresponding analysis. These labels can then be manually reviewed and corrected to obtain the final explanation framework labels. Alternatively, they can be determined manually without specific limitations.
[0060] The initial training model is typically a large language model, specifically a multimodal large model. A large language model is a deep learning model trained on massive amounts of text data. It can generate natural language text or understand the meaning of language text. Large language models can provide relevant knowledge on various topics by training on huge datasets. The core idea of large language models is to learn the patterns and structures of natural language through large-scale unsupervised training, simulating human language cognition and generation processes to some extent. Large language models perform well in various application scenarios, not only performing simple language tasks such as spell checking and grammar correction, but also handling complex tasks such as text summarization, machine translation, sentiment analysis, dialogue generation, and content recommendation. Through pre-training on large-scale datasets, large language models acquire powerful general modeling and generalization capabilities.
[0061] In practical applications, using a third training sample to train the initial training model allows for fine-tuning of the pre-trained large language model.
[0062] In practical applications, fine-tuning a pre-trained large language model leverages its strong generalization ability. By using only a small number of training samples, the model can be trained to possess strong processing capabilities for specific tasks, reducing the number of samples needed for training and improving training efficiency. In the embodiments of this specification, using third training samples to train the large language model enables it to generate explanatory frameworks based on the question and its corresponding parsing. Specifically, sample question images and parsings can be used as input to the large language model, which can output corresponding explanatory framework content. Then, based on the difference between the model's output explanatory framework content and the explanatory framework labels in the third training samples, the model parameters are adjusted to minimize this difference, thus training the model.
[0063] In this scheme, a pre-training stage is introduced before the original two-stage training. First, using sample question images and analysis, as well as explanation framework labels, the model is trained to automatically generate the corresponding explanation framework (second intermediate model). This design forms a progressive three-level training paradigm of "explanation framework generation → whiteboard text generation → interactive explanation generation".
[0064] In this embodiment of the specification, by enabling the model to learn from the source how to automatically plan teaching logic based on the content and analysis of the question, not only is the reliance on manually designed explanation frameworks significantly reduced and the degree of automation in question explanation is improved, but the generated explanation framework is also ensured to be highly consistent with the internal logic of the question analysis, laying a more reliable foundation for the subsequent generation of accurate blackboard writing and targeted explanation. This layered training architecture enables the model to achieve a more coherent end-to-end intelligent explanation from the understanding of the original question to the final personalized guidance when it is finally integrated. This significantly enhances the adaptability to complex and varied question types, especially novel questions that lack standard explanation templates, and improves the overall accuracy of the explanation.
[0065] In one optional implementation of this embodiment, obtaining the first training sample includes: Based on the sample question images and sample question analysis, a sample explanation framework is generated through a second intermediate model; The first training sample is constructed based on the sample question image, sample explanation framework, and corresponding whiteboard text labels.
[0066] As shown above, this scheme first acquires sample question images and sample question analyses. Then, a second intermediate model, capable of generating explanation frameworks based on the question images and corresponding analyses, generates explanation frameworks corresponding to the sample question images and sample question analyses. The sample question images and explanation frameworks generated by the second intermediate model can be directly used as sample explanation frameworks in the first training sample. Alternatively, the explanation frameworks generated by the model can be used as preliminary sample explanation frameworks, which are then manually reviewed and corrected to determine the sample explanation frameworks in the first training sample.
[0067] In this embodiment, a second intermediate model with the ability to generate explanation frameworks is introduced to automatically construct the first training samples. This significantly reduces the reliance on manually annotated explanation frameworks, greatly reduces the manpower and time costs in the data preparation stage, and improves the efficiency of the overall training process. Secondly, the explanation framework generated by the model can ensure the automatic alignment of teaching logic and question analysis, avoiding subjective biases or inconsistencies that may be introduced by manual annotation, and laying a more reliable and standardized foundation for the subsequent generation of accurate and standardized blackboard text. Finally, this design makes the entire process from the original question to the final explanation more automated.
[0068] In practical applications, the sample explanation framework can also be determined manually.
[0069] Furthermore, the sample interpretation framework in the second training sample is the same as that in the first training sample; it can be generated by the second intermediate model or determined manually.
[0070] Step 104: Based on the first training sample and the second training sample, train the initial question explanation model to obtain the target question explanation model that meets the training stopping condition.
[0071] In practical applications, the initial question explanation model is trained using the first training sample. Fine-tuning of the pre-trained large language model can leverage its strong generalization ability. By using only a small number of training samples to fine-tune the large language model, it can enable the large language model to have strong processing capabilities for specific tasks. This helps reduce the number of samples used for model training and improves model training efficiency.
[0072] In the embodiments of this specification, a first training sample is used to train the large language model, which enables the trained large language model to generate whiteboard text based on the topic and explanation framework. Specifically, the topic and explanation framework in the first training sample can be input into the large language model to obtain the whiteboard text content output by the large language model; then, based on the difference between the whiteboard text content output by the model and the whiteboard text labels, the model parameters are adjusted with the goal of minimizing the difference between the two, and the model is trained.
[0073] Training the large language model using the first training sample enables the trained model to generate explanatory text based on the topic, explanation framework, whiteboard text, and multi-turn dialogue. Specifically, the topic, explanation framework, whiteboard text, and multi-turn dialogue from the second training sample can be input into the large language model to obtain the explanatory text output by the model. Then, based on the difference between the model's output explanatory text and the explanatory text labels in the second training sample, the model parameters are adjusted to minimize the difference, and the model is trained.
[0074] In practical applications, training a large language model using the first training sample and training a large language model using the second training sample can be independent of each other. These two training methods are designed to train different capabilities of the large language model. Therefore, one can first train the large language model using the first training sample and then train it using the second training sample; or one can train the large language model using the first training sample and then train it using the second training sample; or one can train the large language model using both training samples simultaneously. There are no specific limitations on this.
[0075] In practical applications, a large language model trained with the first and second training samples can generate whiteboard text based on the question and explanation framework, and further generate corresponding explanation text based on the question, explanation framework and whiteboard text.
[0076] In practical applications, considering that users may have insufficient knowledge or weak comprehension, degraded explanation scenarios can be set up in some sample dialogues. For example, when the preset user input includes phrases like "I don't understand," "I don't comprehend," or "I can't," degraded explanations can be provided through specific examples or more detailed explanations to facilitate user understanding. By setting up degraded explanation scenarios in preset dialogues, the large language model can learn the applicable situations and strategies for degraded explanations. When encountering users expressing inability to understand in real-world applications, more detailed guidance can be provided.
[0077] In practical applications, considering that users may have a good grasp of the knowledge or strong comprehension, step-by-step explanations can be set up in some sample dialogues. For example, if the preset user input includes "I already know how to do this question", or if the user has already given the answer to the next step in their reply, the steps that the user has already mastered can be skipped and the explanation can be skipped to quickly guide the user and improve the efficiency of guidance.
[0078] In practical applications, the trained large language model (target question explanation model) can be deployed to the server of the target application, which can be an application that provides question explanation functionality to users. Users can ask questions or reply to the explanation text generated by the large language model on the client side of the target application. The trained large language model deployed on the server side of the target application can answer the questions raised by users and further output explanation text based on the user's reply.
[0079] Furthermore, when constructing the second training sample, the question analysis can also be input as important contextual information. It should be understood that the teaching guidance path and knowledge points constructed by the explanation framework and blackboard text cannot fully cover all the knowledge details and possible reasoning branches involved in solving the problem. Therefore, using question analysis containing complete standard answers and reasoning processes as supplementary input can provide the model with a more complete knowledge base. When a user's question or answer exceeds the pre-set guidance scope of the explanation framework or involves deeper principles not directly presented in the blackboard, the model can rely on authoritative information in the question analysis as a fallback, thereby ensuring that accurate and logically rigorous explanation text can still be generated in a wider range of interactive scenarios, significantly enhancing the system's robustness and the completeness of knowledge coverage.
[0080] It should be noted that the embodiments described above in this specification are for the purpose of clearly illustrating the scheme process and only use a single training sample as an example. In actual model training, a large number of training samples constructed as described will be used, and each sample will follow the same processing and training logic.
[0081] See Figure 2 , Figure 2 A flowchart illustrating a problem-solving method according to an embodiment of this specification is shown. From a hardware perspective, the method can be executed by a server. From a programming perspective, the method can be executed by a model training program hosted on the server.
[0082] like Figure 2 As shown, the method specifically includes the following steps 202-210.
[0083] Step 202: Obtain the target question image and the target explanation framework, wherein the target explanation framework is determined based on the target question image and the target question analysis.
[0084] The target explanation framework can be determined manually, or the target question image and target question analysis can be input into the target question explanation model to automatically generate the target explanation framework.
[0085] Step 204: Based on the target question image and the target explanation framework, generate the target whiteboard text using the target question explanation model. The target question explanation model is trained using the training method of the question explanation model described above.
[0086] Step 206: Obtain the target multi-turn dialogue content, which includes the historical dialogue content before the current turn and the user response content of the current turn.
[0087] Step 208: Based on the target question image, target explanation framework, target whiteboard text, and target multi-turn dialogue content, generate the target explanation text for the current round using the target question explanation model.
[0088] Step 210: Generate synchronized explanation data for the current round based on the target whiteboard text and the target explanation text.
[0089] In the embodiments of this specification, the original problem image is directly used as the input to the model, without any manual preprocessing (such as text recognition, graphic structured description, etc.). The model's visual understanding ability, acquired through training, can automatically and accurately extract key information from the image (such as geometric figures, function graphs, chart data, etc.), thus fundamentally avoiding the problems of information distortion, omissions, or descriptive biases caused by manual conversion or simplification in traditional methods. Secondly, since it does not rely on prior manual interpretation and annotation of the image, this method can handle various types of problems containing complex structures, irregular graphics, multimodal information (text-image hybrid), and even natural scene images. This breaks through the limitation of traditional methods, which can usually only handle simple, regular graphics, enabling the intelligent explanation system to cover a wider range of disciplines and problem types, from basic mathematics and geometry to physics diagrams and chemical structural formulas, significantly improving the system's universality and application value. Furthermore, in the entire explanation generation process, from image understanding and whiteboard generation to personalized explanation, everything is automatically completed by the trained target problem explanation model. This eliminates a large amount of manual work in image analysis, content annotation, and script writing for each new problem, reducing labor costs.
[0090] In one optional implementation of this embodiment, synchronous explanation data for the current round is generated based on the target whiteboard text and the target explanation text, including: Convert the text explaining the target into audio explaining the target; Convert the target whiteboard text into a whiteboard animation sequence; Synchronous explanation data is generated based on the target explanation audio and blackboard animation sequence.
[0091] Among them, the target explanation speech refers to the audio stream converted from the "target explanation text" through speech synthesis technology. It is the auditory representation of the explanation content and has a clear timeline.
[0092] A whiteboard animation sequence refers to a dynamic visual presentation sequence created by rendering the target whiteboard text using graphics and an animation engine. It is a set of instructions or renderable objects that define how the whiteboard content (such as text, formulas, and graphics) gradually appears, is highlighted, and changes. Each whiteboard animation frame represents a logical state or keyframe in the animation sequence.
[0093] In this embodiment, the target explanation text and target blackboard text generated by the model are converted into corresponding target explanation voice and blackboard animation sequences, respectively. Based on this, structured synchronous explanation data is generated, which realizes the upgrade from pure text interaction to high-quality, multimedia immersive teaching experience, significantly improving the vividness and intuitiveness of knowledge transmission and the learner's sense of immersion.
[0094] In one optional implementation of this embodiment, synchronized explanation data is generated based on the target explanation voice and blackboard animation sequence, including: Construct a time synchronization relationship between the target explanation speech and the blackboard animation sequence. The time synchronization relationship is used to indicate the correspondence between each time point of the target explanation speech and each blackboard animation frame in the blackboard animation sequence. Based on the target explanation audio, blackboard animation sequence, and time synchronization relationship, synchronized explanation data is generated.
[0095] The time synchronization mechanism establishes a mapping between the timeline of the "target explanation audio" and the visual state sequence of the "blackboard animation sequence." For example, it defines that when the audio plays to the 5.2-second mark (when explaining a formula), the corresponding blackboard animation should display and highlight that formula.
[0096] The synchronized explanation data encapsulates the "target explanation voice", "whiteboard animation sequence" and the "time synchronization relationship" between the two, which is used to achieve synchronized audio and video playback.
[0097] In this solution, independent auditory content (speech) and visual content scripts (animation sequences) are generated separately. During the content generation stage, the logical structure of the explanatory text is pre-analyzed using algorithms, and corresponding blackboard animation change nodes are matched for each explanatory segment accordingly. The correspondence between the two on the timeline is accurately calculated, forming a "time synchronization relationship." The speech, animation sequence, and synchronization relationship are then encapsulated into a structured "synchronized explanation data" package.
[0098] In the embodiments described in this specification, by pre-establishing and solidifying the audio-visual synchronization relationship during the content generation stage, frame-level precise alignment of voice narration and animation presentation can be achieved. Since the synchronization relationship has already been calculated during the data generation stage and encapsulated as metadata, the player does not need to perform complex real-time audio-visual alignment analysis or calculations at runtime. This significantly reduces the computational load on the playback end, ensuring smooth and stable playback on various terminal devices, especially suitable for performance-constrained mobile learning environments, while also reducing the risk of playback asynchrony caused by uncertainties in real-time calculations.
[0099] In one optional implementation of this embodiment, after generating the synchronized explanation data, the method further includes: Play the target explanation audio from the synchronized explanation data and monitor the current playback time. Based on the time synchronization relationship in the synchronized explanation data, determine the target whiteboard animation frame corresponding to the current playback time point; At the current playback time, the corresponding target explanation audio and whiteboard animation frames are presented synchronously.
[0100] The player parses the data packet and, based on the instructions of "time synchronization," drives the animation engine to present the corresponding whiteboard animation frames in real time at each moment of the audio playback, thereby achieving a precise synchronization effect of "writing where it is spoken and highlighting where it is written."
[0101] The current playback time point is the real-time elapsed time (usually in milliseconds) calculated from the moment the target narration begins playing. It is used to accurately track the progress of the narration playback.
[0102] A target whiteboard animation frame is a keyframe in a whiteboard animation sequence. It defines the visual state that the whiteboard content (text, formulas, graphics, etc.) should present at the current playback point, including attributes such as content, position, color, and highlight.
[0103] In this solution, the player loads the "synchronized explanation data" package, which parses out three core components: the target explanation audio (audio stream), the whiteboard animation sequence (visual instruction set), and the time synchronization relationship (mapping table). The player starts playing the target explanation audio and simultaneously starts a high-precision clock to continuously monitor the current playback time. For each (or key) current playback time, the player uses it as a "key" to query the pre-established "time synchronization relationship" mapping table. The mapping table returns the identifier or instruction of the "target whiteboard animation frame" that precisely corresponds to that time point. The player immediately sends the query result to the animation rendering engine, driving it to update the screen and present the specified whiteboard animation frame.
[0104] In the embodiments described in this specification, since the synchronization logic is pre-calculated and embedded in the data packet, the player only needs to perform efficient time querying and instruction distribution during runtime, without performing any real-time audio-visual alignment analysis or complex calculations. This greatly reduces the computational burden on the playback end, ensuring an absolutely smooth, stutter-free, and synchronization-free playback experience even on terminal devices with limited performance.
[0105] In another optional implementation of this embodiment, the audio narration and the whiteboard animation sequence are analyzed in real time during playback to dynamically establish a correspondence. Specifically, the target audio narration in the synchronized narration data is played, and the current playback time is monitored; the audio narration content corresponding to the current playback time is analyzed in real time, and the visual semantic information of each animation frame in the whiteboard animation sequence is simultaneously parsed; based on the real-time matching of the audio narration content and the visual semantic information, the target whiteboard animation frame most relevant to the current playback time is dynamically determined; and the target audio narration corresponding to the current playback time and the determined whiteboard animation frame are presented synchronously.
[0106] In one optional implementation of this embodiment, based on the target question image, target explanation framework, target whiteboard text, and multi-turn dialogue content, the target explanation text for the current round is generated through the target question explanation model, including: Based on the target question image, target explanation framework, target whiteboard text, multi-round dialogue content, and target question analysis, the target explanation text for the current round is generated through the target question explanation model.
[0107] In this approach, the question explanation is input as important contextual information. It should be understood that the teaching guidance path and key knowledge points constructed by the explanation framework and blackboard text cannot fully cover all the knowledge details and possible reasoning branches involved in solving a problem. Therefore, using the question explanation, which includes a complete standard answer and reasoning process, as supplementary input provides the model with a more comprehensive knowledge base. When a user's question or answer exceeds the pre-set guidance scope of the explanation framework or involves deeper principles not directly presented in the blackboard, the model can rely on the authoritative information in the question explanation to ensure that accurate and logically rigorous explanation text is generated even in a wider range of interactive scenarios.
[0108] The following is in conjunction with the appendix Figure 3 Taking the training method of the problem-solving model provided in this manual in a multimodal large model as an example, the training method of the problem-solving model will be further explained. Figure 3 The flowchart illustrates a training method for a problem explanation model provided in one embodiment of this specification, which specifically includes the following steps.
[0109] Step 302: Obtain the sample question image and sample question analysis; Step 304: Based on the sample question image and sample question analysis, label the corresponding explanation framework tags; Step 306: Based on the sample question images and sample question parsing, as well as the explanation framework labels, perform the first stage of training on the initial question explanation model to obtain the first model; The initial problem explanation model is a multimodal large model. The first model can generate an explanation framework based on the problem image and problem analysis.
[0110] Step 308: Based on the sample question image and sample explanation framework, label the corresponding whiteboard text. The explanatory framework label can be used directly as the sample explanatory framework, or the model can generate the sample explanatory framework based on the sample question image and sample question analysis.
[0111] Step 310: Based on the sample question image and sample explanation framework, as well as the whiteboard text labels, perform the second stage of training on the first model to obtain the second model; The second model can generate an explanation framework based on the question image and the question analysis; it can also generate whiteboard text based on the question image and the explanation framework.
[0112] Step 312: Based on the sample question image, sample explanation framework, and sample whiteboard text, construct a multi-turn dialogue and label the corresponding explanation text. The whiteboard text labels can be used directly as sample whiteboard text, or the model can generate sample whiteboard text based on sample question images and sample explanation frameworks.
[0113] Step 314: Based on the sample question images, explanation framework, whiteboard text, multi-turn dialogues, and explanation text labels, the second model is trained in the third stage to obtain the target question explanation model.
[0114] Among them, the target question explanation model can generate an explanation framework based on the question image and question analysis; it can also generate whiteboard text based on the question image and explanation framework; and it can also generate explanation text based on the sample question image, explanation framework, whiteboard text and multi-turn dialogue.
[0115] Furthermore, to improve the accuracy of the model's generated explanatory text, sample question analyses can also be used as part of the training data to train the second model in the third stage.
[0116] During the application phase, when a user has a question about a question, the system first generates a target explanation framework based on the target question image and analysis using the target question explanation model; then, it generates target whiteboard text based on the target question image and explanation framework using the target question explanation model; finally, it generates the target explanation text for the current round based on the target question image, analysis, explanation framework, whiteboard text, and multi-turn dialogue.
[0117] The target explanation text is converted into target explanation audio, and the target blackboard writing text is converted into a blackboard animation sequence. A time synchronization relationship is established between the target explanation audio and the blackboard animation sequence. The time synchronization relationship is used to indicate the correspondence between each time point of the target explanation audio and each blackboard animation frame in the blackboard animation sequence. The target explanation audio is played, and the current playback time point is monitored. Based on the time synchronization relationship, the target blackboard animation frame corresponding to the current playback time point is determined. At the current playback time point, the target explanation audio and blackboard animation frame corresponding to the current playback time point are presented synchronously.
[0118] Corresponding to the above method embodiments, this specification also provides embodiments of a training device for a problem explanation model. Figure 4 This specification illustrates a schematic diagram of a training device for a problem-solving model according to one embodiment. Figure 4 As shown, the device includes: The acquisition module 402 is configured to acquire a first training sample and a second training sample. The first training sample includes a sample question image and a sample explanation framework, as well as corresponding whiteboard text labels. The second training sample includes a sample question image, a sample explanation framework, sample whiteboard text, and a sample multi-turn dialogue, as well as corresponding explanation text labels. The sample explanation framework is determined based on the sample question image and sample question parsing. Training module 404 is configured to train the initial question explanation model based on the first training sample and the second training sample to obtain the target question explanation model that meets the training stopping condition.
[0119] Optionally, the acquisition module 402 is further configured as follows: Based on the sample question images and sample explanation framework, label the corresponding blackboard text. The first training sample is constructed based on the sample question image, sample explanation framework and whiteboard text label; Based on the sample question images and sample explanation framework, sample whiteboard text is generated through the first intermediate model, wherein the first intermediate model is obtained by training the initial question explanation model based on the first training samples; Based on the sample question images, sample explanation frameworks, and sample whiteboard text, a sample multi-turn dialogue is constructed, and the corresponding explanation text labels are annotated. A second training sample is constructed based on sample question images, sample explanation frameworks, sample whiteboard text, sample multi-turn dialogues, and explanation text labels. Accordingly, training module 404 is further configured as follows: The first intermediate model is trained based on the second training sample to obtain the target question explanation model.
[0120] Optionally, the training apparatus for the above-mentioned problem explanation model also includes a problem explanation framework training module, configured as follows: Obtain sample question images and sample question analyses; Based on the sample question images and sample question analyses, the corresponding explanation framework labels are annotated; Based on the sample question images and sample question parsing, as well as the explanation framework labels, a third training sample is constructed. Based on the third training sample, the initial question explanation model is trained to obtain a second intermediate model. Accordingly, training module 404 is further configured as follows: Based on the first and second training samples, the second intermediate model is trained to obtain the target question explanation model.
[0121] Optionally, the acquisition module 402 is further configured as follows: Based on the sample question images and sample question analysis, a sample explanation framework is generated through a second intermediate model; The first training sample is constructed based on the sample question image, sample explanation framework, and corresponding whiteboard text labels.
[0122] The above is a schematic diagram of a training device for a problem explanation model according to this embodiment. It should be noted that the technical solution of this training device for a problem explanation model and the technical solution of the training method for a problem explanation model described above belong to the same concept. For details not described in detail in the technical solution of the training device for a problem explanation model, please refer to the description of the technical solution of the training method for a problem explanation model described above.
[0123] Corresponding to the above method embodiments, this specification also provides embodiments of the topic explanation apparatus. Figure 5 A schematic diagram of a problem-solving apparatus according to one embodiment of this specification is shown. Figure 5 As shown, the device includes: The first acquisition module 502 is configured to acquire a target question image and a target explanation framework, wherein the target explanation framework is determined based on the target question image and the target question analysis. The first generation module 504 is configured to generate target whiteboard text based on the target question image and the target explanation framework, using the target question explanation model. The target question explanation model is trained using the training method of the above-mentioned question explanation model. The second acquisition module 506 is configured to acquire target multi-turn dialogue content, wherein the target multi-turn dialogue content includes historical dialogue content before the current turn and user response content in the current turn; The second generation module 508 is configured to generate the target explanation text for the current round based on the target question image, the target explanation framework, the target whiteboard text, and the target multi-turn dialogue content, through the target question explanation model. The third generation module 510 is configured to generate synchronized explanation data for the current round based on the target whiteboard text and the target explanation text.
[0124] Optionally, the third generation module 510 is further configured as follows: Convert the text explaining the target into audio explaining the target; Convert the target whiteboard text into a whiteboard animation sequence; Synchronous explanation data is generated based on the target explanation audio and blackboard animation sequence.
[0125] Convert the target explanation text into target explanation speech; Convert the target whiteboard text into a whiteboard animation sequence; The synchronized explanation data is generated based on the target explanation voice and the blackboard animation sequence.
[0126] Optionally, the third generation module 510 is further configured as follows: Construct a time synchronization relationship between the target explanation speech and the blackboard animation sequence. The time synchronization relationship is used to indicate the correspondence between each time point of the target explanation speech and each blackboard animation frame in the blackboard animation sequence. Based on the target explanation audio, blackboard animation sequence, and time synchronization relationship, synchronized explanation data is generated.
[0127] Optionally, the above-mentioned question explanation device also includes a playback module, configured as follows: Play the target explanation audio from the synchronized explanation data and monitor the current playback time. Based on the time synchronization relationship in the synchronized explanation data, determine the target whiteboard animation frame corresponding to the current playback time point; At the current playback time, the corresponding target explanation audio and whiteboard animation frames are presented synchronously.
[0128] The above is a schematic scheme of a problem-solving device according to this embodiment. It should be noted that the technical solution of this problem-solving device and the technical solution of the problem-solving method described above belong to the same concept. For details not described in detail in the technical solution of the problem-solving device, please refer to the description of the technical solution of the problem-solving method described above.
[0129] Figure 6 A structural block diagram of a computing device 600 according to one embodiment of this specification is shown. The components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and a database 650 is used to store data.
[0130] The computing device 600 also includes an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0131] In one embodiment of this specification, the above-described components of the computing device 600 and Figure 6 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 6 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0132] The computing device 600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 600 can also be a mobile or stationary server.
[0133] The processor 620 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the training method and problem explanation method of the above-mentioned problem explanation model.
[0134] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device belongs to the same concept as the training method and problem explanation method of the above-mentioned problem explanation model. For details not described in detail in the technical solution of the computing device, please refer to the description of the training method and problem explanation method of the above-mentioned problem explanation model.
[0135] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the training method and the problem explanation method of the above-described problem explanation model.
[0136] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solutions of the training method and the problem explanation method of the above-mentioned problem explanation model. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solutions of the training method and the problem explanation method of the above-mentioned problem explanation model.
[0137] An embodiment of this specification also provides a computer program, wherein when the computer program is executed in a computer, the computer is instructed to perform the steps of the training method and the problem explanation method of the above-described problem explanation model.
[0138] The above is an illustrative scheme of a computer program according to this embodiment. It should be noted that the technical solution of this computer program belongs to the same concept as the technical solutions of the training method and the problem explanation method of the above-mentioned problem explanation model. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solutions of the training method and the problem explanation method of the above-mentioned problem explanation model.
[0139] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0140] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0141] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0142] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0143] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A training method for a problem-solving explanation model, characterized in that, include: Obtain a first training sample and a second training sample, wherein the first training sample includes a sample question image and a sample explanation framework, as well as corresponding whiteboard text labels; the second training sample includes a sample question image, a sample explanation framework, sample whiteboard text and a sample multi-turn dialogue, as well as corresponding explanation text labels, wherein the sample explanation framework is determined based on the sample question image and sample question parsing. Based on the first training sample and the second training sample, the initial question explanation model is trained to obtain the target question explanation model that meets the training stopping condition.
2. The method according to claim 1, characterized in that, The acquisition of the first training sample and the second training sample includes: Based on the sample question image and the sample explanation framework, label the corresponding blackboard text. The first training sample is constructed based on the sample question image, the sample explanation framework, and the blackboard text labels; Based on the sample question image and the sample explanation framework, the sample whiteboard text is generated through a first intermediate model, wherein the first intermediate model is obtained by training the initial question explanation model based on the first training sample; Based on the sample question image, the sample explanation framework, and the sample whiteboard text, a sample multi-turn dialogue is constructed, and the corresponding explanation text tags are labeled. The second training sample is constructed based on the sample question image, the sample explanation framework, the sample whiteboard text, the sample multi-turn dialogue, and the explanation text labels. Accordingly, the step of training the initial problem explanation model based on the first training sample and the second training sample to obtain the target problem explanation model that satisfies the training stopping condition includes: The first intermediate model is trained based on the second training sample to obtain the target question explanation model.
3. The method according to claim 1, characterized in that, Before obtaining the first training sample, the process also includes: Obtain the sample question image and the sample question analysis; Based on the sample question image and the sample question analysis, the corresponding explanation framework labels are annotated; Based on the sample question image and the sample question parsing, as well as the explanation framework label, a third training sample is constructed. Based on the third training sample, the initial question explanation model is trained to obtain a second intermediate model. Accordingly, the step of training the initial problem explanation model based on the first training sample and the second training sample to obtain the target problem explanation model that satisfies the training stopping condition includes: Based on the first training sample and the second training sample, the second intermediate model is trained to obtain the target question explanation model.
4. The method according to claim 3, characterized in that, The process of obtaining the first training sample includes: Based on the sample question image and the sample question analysis, the sample explanation framework is generated through a second intermediate model; The first training sample is constructed based on the sample question image, the sample explanation framework, and the corresponding whiteboard text labels.
5. A method for explaining problems, characterized in that, include: Obtain the target question image and the target explanation framework, wherein the target explanation framework is determined based on the target question image and the target question analysis; Based on the target question image and the target explanation framework, the target whiteboard text is generated through the target question explanation model. The target question explanation model is obtained by training the question explanation model according to any one of claims 1-4. Obtain the target multi-turn dialogue content, wherein the target multi-turn dialogue content includes historical dialogue content before the current turn and user response content in the current turn; Based on the target question image, the target explanation framework, the target whiteboard text, and the target multi-turn dialogue content, the target explanation text for the current round is generated through the target question explanation model. Based on the target whiteboard text and the target explanation text, generate the synchronous explanation data for the current round.
6. The method according to claim 5, characterized in that, The step of generating synchronized explanation data for the current round based on the target whiteboard text and the target explanation text includes: Convert the target explanation text into target explanation speech; Convert the target whiteboard text into a whiteboard animation sequence; The synchronized explanation data is generated based on the target explanation voice and the blackboard animation sequence.
7. The method according to claim 6, characterized in that, The process of generating the synchronized explanation data based on the target narration voice and the blackboard animation sequence includes: Construct a time synchronization relationship between the target narration voice and the blackboard animation sequence, wherein the time synchronization relationship is used to indicate the correspondence between each time point of the target narration voice and each blackboard animation frame in the blackboard animation sequence; Based on the target narration voice, the blackboard animation sequence, and the time synchronization relationship, the synchronized narration data is generated.
8. The method according to claim 7, characterized in that, After generating the synchronized explanation data, the process also includes: Play the target explanation audio from the synchronized explanation data and monitor the current playback time. Based on the time synchronization relationship in the synchronized explanation data, determine the target whiteboard animation frame corresponding to the current playback time point; At the current playback time point, the target explanation audio and whiteboard animation frames corresponding to the current playback time point are presented synchronously.
9. A training device for a problem explanation model, characterized in that, include: The acquisition module is configured to acquire a first training sample and a second training sample. The first training sample includes a sample question image and a sample explanation framework, as well as corresponding whiteboard text labels. The second training sample includes a sample question image, a sample explanation framework, sample whiteboard text, and a sample multi-turn dialogue, as well as corresponding explanation text labels. The sample explanation framework is determined based on the sample question image and sample question parsing. The training module is configured to train the initial question explanation model based on the first training sample and the second training sample to obtain the target question explanation model that meets the training stopping condition.
10. A problem-solving explanation device, characterized in that, include: The first acquisition module is configured to acquire a target question image and a target explanation framework, wherein the target explanation framework is determined based on the target question image and the target question analysis. The first generation module is configured to generate target whiteboard text based on the target question image and the target explanation framework, using a target question explanation model. The target question explanation model is obtained by training the question explanation model according to any one of claims 1-4. The second acquisition module is configured to acquire target multi-turn dialogue content, wherein the target multi-turn dialogue content includes historical dialogue content before the current turn and user response content in the current turn; The second generation module is configured to generate the target explanation text for the current round based on the target question image, the target explanation framework, the target whiteboard text, and the target multi-turn dialogue content, using the target question explanation model. The third generation module is configured to generate synchronized explanation data for the current round based on the target whiteboard text and the target explanation text.
11. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1-8.
12. A computer-readable storage medium, characterized in that, It stores computer-executable instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1-8.
13. A computer program product, characterized in that, Includes a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1-8.