Method, apparatus, device and medium for generating a problem-solving model
By acquiring explanation information of the questions to be answered and training a problem-solving model, explanation video data is generated, which solves the problem of low efficiency in generating explanation videos in existing technologies, realizes fast and efficient generation of explanation videos, and improves user experience.
Patent Information
- Application Number
- CN202411403905.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-09
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2044-10-09
AI Technical Summary
Existing technologies for generating explanatory videos are inefficient, and a method for quickly generating explanatory videos is needed to improve students' learning efficiency.
By acquiring the questions to be answered, determining their explanation information, generating training data and training the problem-solving model, and using the trained problem-solving model to generate explanation video data, the manual writing of video code is avoided.
It enables the rapid generation of explanatory videos, improving the efficiency and quality of video generation, reducing manual operations, and enhancing the user experience.
Smart Images

Figure CN119252084B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present specification relate to the technical field of computer technology, and particularly relate to a method and apparatus for generating a problem solving model, a device, and a medium. BACKGROUND
[0002] With the development of the times, using computer technology to assist teaching has become a trend. In particular, using the form of video to explain the problem can greatly improve the learning interest and understanding ability of students, and has been pursued by people.
[0003] The generation efficiency of the explanation video can affect the learning efficiency of students, and therefore a method for quickly generating an explanation video is needed. SUMMARY
[0004] Therefore, embodiments of the present specification provide a method for generating a problem solving model. One or more embodiments of the present specification also relate to a method for solving a problem, an apparatus for generating a problem solving model, an apparatus for solving a problem, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects in the prior art.
[0005] According to a first aspect of embodiments of the present specification, a method for generating a problem solving model is provided, comprising:
[0006] obtaining a problem to be solved; the problem to be solved is a problem that can be solved using a line segment graph;
[0007] determining explanation information of the problem to be solved; the explanation information includes explanation video data for the problem to be solved;
[0008] generating training data based on the problem to be solved and the explanation information;
[0009] training a problem solving model using the training data to obtain a trained problem solving model; the trained problem solving model is used to generate explanation video data for solving a problem using a line segment graph.
[0010] According to a second aspect of embodiments of the present specification, a method for solving a problem is provided, comprising:
[0011] obtaining a target problem to be solved;
[0012] input the target to-be-solved question into the trained question-solving model to obtain target explanation information; the target explanation information includes target video data for generating an explanation video for solving the target to-be-solved question by using a line segment graph; the trained question-solving model is obtained by training a question-solving model by using training data; the training data is generated based on explanation information of a to-be-solved question that can be solved by using a line segment graph and the to-be-solved question;
[0013] render the target video data by using a video rendering tool to obtain a target explanation video for the target to-be-solved question.
[0014] According to a third aspect of an embodiment of the present specification, a device for generating a question-solving model is provided, including:
[0015] a question obtaining module configured to obtain a to-be-solved question; the to-be-solved question is a question that can be solved by using a line segment graph;
[0016] a determining module configured to determine explanation information of the to-be-solved question; the explanation information includes explanation video data for the to-be-solved question;
[0017] a data generating module configured to generate training data based on the to-be-solved question and the explanation information;
[0018] a model training module configured to train a question-solving model by using the training data to obtain a trained question-solving model; the trained question-solving model is used to generate explanation video data for solving a question by using a line segment graph.
[0019] According to a fourth aspect of an embodiment of the present specification, a device for solving a question is provided, including:
[0020] a target example question obtaining module configured to obtain a target to-be-solved question;
[0021] an input module configured to input the target to-be-solved question into a trained question-solving model to obtain target explanation information; the target explanation information includes target video data for generating an explanation video for solving the target to-be-solved question by using a line segment graph; the trained question-solving model is obtained by training a question-solving model by using training data; the training data is generated based on explanation information of a to-be-solved question that can be solved by using a line segment graph and the to-be-solved question;
[0022] a rendering module configured to render the target video data by using a video rendering tool to obtain a target explanation video for the target to-be-solved question.
[0023] According to a fifth aspect of an embodiment of the present specification, a computing device is provided, including:
[0024] a memory and a processor;
[0025] The memory is configured to store computer-executable instructions, and the processor is configured to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method.
[0026] According to a sixth aspect of an embodiment of the present specification, a computer-readable storage medium is provided, which stores computer-executable instructions, which, when executed by a processor, implement the steps of the method.
[0027] According to a seventh aspect of an embodiment of the present specification, a computer product is provided, which includes a computer program / instruction, which, when executed by a processor, implements the steps of the method.
[0028] At least one embodiment of the present specification achieves the following beneficial effects: obtaining a to-be-solved question that can be solved by using a line segment graph; determining explanation information of the to-be-solved question, the explanation information including explanation video data for the to-be-solved question; generating training data based on the to-be-solved question and the explanation information; training a question-solving model by using the training data to obtain a trained question-solving model. Thus, the trained question-solving model can be used to generate explanation video data for solving the question by using the line segment graph, and explanation video can be obtained quickly, thereby avoiding manual writing of the explanation video data, reducing manual operation, and further improving the generation efficiency of the explanation video. BRIEF DESCRIPTION OF DRAWINGS
[0029] Figure 1 is a flowchart of a method for generating a question-solving model provided by an embodiment of the present specification;
[0030] Figure 2 is a schematic diagram of an image frame of an explanation video provided by an embodiment of the present specification;
[0031] Figure 3 is a flowchart of a method for solving a question provided by an embodiment of the present specification;
[0032] Figure 4 is a structural schematic diagram of an apparatus for generating a question-solving model provided by an embodiment of the present specification;
[0033] Figure 5 is a structural schematic diagram of an apparatus for solving a question provided by an embodiment of the present specification;
[0034] Figure 6 is a structural block diagram of a computing device provided by an embodiment of the present specification. DETAILED DESCRIPTION
[0035] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present description. However, the present description can be practiced without the specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to obscure the present description.
[0036] The terminology used in this description is for the purpose of describing particular embodiments only and is not intended to be limiting of one or more embodiments of the present description. As used in this description and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0037] It will be understood that, although the terms first, second, etc. can be used herein to describe various information, these terms are not intended to denote a temporal or chronological order. Rather, these terms are used solely to distinguish one from another only. For example, without departing from the scope of one or more embodiments, first can be termed second, and similarly, second can be termed first. The term "if' as used herein can be interpreted as meaning "when" or "if," depending on the context.
[0038] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present description are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.
[0039] First, the terms involved in one or more embodiments of the present description are explained.
[0040] Large Language Model (LLM): also known as Large Language Model, Large Model, is a deep learning model that uses self-recurrent method to pre-train on large-scale corpus in the field of natural language processing, which can understand and generate natural language text. Large language model can predict the next word or sentence by learning the statistical rules and semantic information of natural language text. With the continuous expansion of input data set and parameter space, the ability of large language model will also be expanded accordingly. It can be applied to machine learning, machine translation, speech recognition, image processing and other fields.
[0041] Retrieval-Augmented Generation (RAG): A model that combines language models and information retrieval techniques. It generates answers or content by referencing information from external knowledge bases, with strong explainability and customization capabilities, suitable for question and answer systems, document generation, intelligent assistants, and other natural language processing tasks. The advantage of RAG is its strong versatility and the ability to update knowledge instantly, thereby providing more efficient and accurate information services.
[0042] Generative Pre-Trained (GPT): A deep learning model based on internet data for text generation. It is trained on a large amount of internet text and is a large language model that can generate coherent sentences or paragraphs in human style.
[0043] Prompt: A command or instruction provided by the user when interacting with a large language model. It can be a question, keyword, context information, etc., used to indicate the action the large language model needs to perform or the output it needs to generate. Users can provide clear prompts to guide the large language model to generate responses that meet expectations, improving the effectiveness and quality of interactions.
[0044] Prompt Template: A structured text used to generate prompts. Generating prompts based on prompt templates can help users interact with large language models, and large language models can better understand user intent, generating more accurate and user-demand-oriented responses.
[0045] Supervised Fine-Tuning (SFT): A machine learning strategy. SFT mainly involves using labeled datasets to further adjust the parameters of pre-trained models, which contain correct answers or labels. During the SFT process, the parameters of some layers (or all layers) of the model will be updated based on this new task, so that the model can better perform this specific task.
[0046] Manim (Mathematical Animation Engine): An open-source Python library for creating high-quality mathematical animations. Manim allows users to precisely control animation details through programming, making it ideal for education and popularization. Manim provides Dot, VGroup, Brace, and other internal classes that are helpful for solving queuing problems through drawing.
[0047] Few-Shot Learning: Few-Shot Learning is a term in the field of machine learning that refers to the ability to train a model with only a small amount of labeled data. This approach is particularly suitable for scenarios where obtaining a large amount of labeled data is costly or impractical. It is one of the key methods to address the problem of data scarcity, and has driven the application of machine learning technology in a wider range of fields, especially in situations where data acquisition is limited.
[0048] In the present specification, a method for generating a problem solving model is provided, and the present specification also relates to a method for solving a problem, an apparatus for generating a problem solving model, an apparatus for solving a problem, a computing device, a computer readable storage medium, and a computer program product, which are described in detail in the following embodiments.
[0049] Reference is made to Figure 1 , Figure 1 A flowchart of a method for generating a problem solving model according to an embodiment of the present specification is shown. From a program perspective, the execution subject of the flowchart can be a program loaded on a server or a training platform. From a hardware perspective, the execution subject of the flowchart can be a server or a training platform capable of model training. The method can specifically include the following steps.
[0050] Step 102: obtaining a problem to be solved; the problem to be solved is a problem that can be solved by using a line segment graph.
[0051] In an embodiment of the present specification, the problem to be solved can be one or more problems in mathematical applications that can be solved by using a line segment graph.
[0052] The problem that can be solved by using a line segment graph can refer to a problem involving the relationship between numerical values or the relationship between variables. For example, "the price of a banana is 1 yuan, and the price of an apple is 2 yuan, what is the total price?", for this problem, the line segments of 1 yuan and 2 yuan can be represented by line segments, and then they are added together, so that the problem solving result is obtained intuitively.
[0053] In practical applications, there can be various ways to obtain the problem to be solved. For example, the problem to be solved can be written according to expert experience; it can be generated based on a large language model; it can also be obtained from a known problem library by using a retrieval-augmented generation (RAG) model, etc., which is not limited here.
[0054] In the embodiments of the present specification, the obtained to-be-solved question can be used for training the question solving model. In order to improve the accuracy of the question solving model training, a plurality of types of questions capable of being solved by using the line segment graph can be obtained, and / or a large number of questions capable of being solved by using the line segment graph can be obtained.
[0055] Step 104: determining the explanation information of the to-be-solved question; the explanation information includes explanation video data for the to-be-solved question.
[0056] In the embodiments of the present specification, the explanation video data for the to-be-solved question included in the explanation information can be a code for generating an explanation video of the to-be-solved question, or can be other forms of digital information.
[0057] In the embodiments of the present specification, the explanation information can be obtained based on expert experience. For example, the explanation information can be obtained by manually solving the to-be-solved question. In addition, the explanation information can also be generated by using a large language model. It can be understood that in actual application, the to-be-solved question can also be solved by using a large language model to obtain a preliminary solution result, and then the preliminary solution result is adjusted by manual or preset rules to obtain the explanation information used by the training model, and the like.
[0058] Step 106: generating training data based on the to-be-solved question and the explanation information.
[0059] The training data of the embodiments of the present specification can include the to-be-solved question and the explanation information. For example, the to-be-solved question and the explanation information can be spliced to generate the training data. Specifically, the to-be-solved question can be placed in front of the explanation information, or the to-be-solved question can be placed behind the explanation information, and the like.
[0060] In the embodiments of the present specification, the training data can include a plurality of training data. For example, a plurality of to-be-solved questions can be obtained at one time, and a plurality of training data can be generated by using the obtained plurality of to-be-solved questions according to the above-mentioned method of generating training data.
[0061] For another example, the to-be-solved question can also be obtained multiple times, and each obtained to-be-solved question can be processed according to the above-mentioned method of generating training data, so that a plurality of training data can also be generated according to the to-be-solved questions obtained multiple times.
[0062] Step 108: training the question solving model by using the training data to obtain a trained question solving model; the trained question solving model is used to generate explanation video data for solving the question by using the line segment graph.
[0063] The explanation video data for solving the question by using the line segment graph generated by the trained question solving model can include a code of an explanation video for solving the question by using the line segment graph.
[0064] The code of the explanation video can be compiled to generate the explanation video. In actual application, a compilation software such as Manim can be used to compile the code of the explanation video output by the trained problem-solving model, so as to generate the explanation video.
[0065] In the embodiments of the present specification, the trained problem-solving model can generate explanation video data for solving a problem by using a line segment diagram. The trained problem-solving model can be used to solve the problem, so as to quickly obtain the code of the explanation video, avoid manual writing of the video code, improve the efficiency of obtaining explanation information, and improve the generation efficiency of the explanation video.
[0066] On the other hand, since the trained problem-solving model is trained by using a problem that can be solved by using a line segment diagram, the trained problem-solving model is specific to the problem that can be solved by using a line segment diagram, and can more accurately generate the code of the explanation video for the problem, thereby ensuring the quality of the generated explanation video.
[0067] For example, when an online teaching institution or a book publishing institution needs to generate explanation videos for a large number of problems in a database, the trained problem-solving model can be used to directly output explanation video data for a problem that can be solved by using a line segment diagram, and then the explanation video can be generated based on the explanation video data, thereby avoiding the tedious process of manually writing scripts and manually making animations, and improving the efficiency of making explanation videos.
[0068] In addition, when a teacher, a student, or a parent uses a teaching software for online learning and needs to obtain an explanation video for a problem that can be solved by using a line segment diagram, the teaching software can load the trained problem-solving model or can call the trained problem-solving model. After the user inputs the problem that can be solved by using a line segment diagram in the teaching software, the teaching software can use the trained problem-solving model to generate explanation video data for the problem. The teaching software can also use a video compilation software to compile the explanation video data, thereby obtaining and displaying the explanation video for the problem, avoiding long waiting time of the user, improving the problem-solving efficiency, and thereby improving the user experience.
[0069] Based on the method, Figure 1 The embodiments of the present specification also provide some specific embodiments of the method, which are described below.
[0070] In order to ensure the accuracy of the problem-solving model training, a large amount of training data is usually required. It is too tedious to manually generate a large amount of training data. To solve this problem, in the embodiments of the present specification, the determination of the explanation information of the problem to be solved can specifically include:
[0071] generating the explanation information of the problem to be solved by using a large language model.
[0072] The large language model is used to generate explanation information of the to-be-answered question, and then training data can be obtained based on the explanation information generated by the large language model, without generating explanation information by manual work to obtain training data, thereby improving the generation efficiency of the training data.
[0073] Specifically, the large language model can be a GPT series model, such as GPT-3.5, GPT-4, and GPT-4o. Of course, other series of models can also be used, such as the Tongyiqian model and the Antuobailing model, which are not limited here.
[0074] Generally, the common way for those skilled in the art is to collect videos that can be used as training data, and then use these videos to train the problem-solving model. It is not easy to think of using another format of information to generate training data. The large language model can process data to obtain text format information, and cannot directly generate videos. Therefore, it is even more difficult for those skilled in the art to think of using a large language model to generate training data for training a problem-solving model. However, in the present application, a large language model can be used to process a to-be-answered question to generate text format information, which can include explanation video data of the to-be-answered question. Then, the training data containing the explanation video data is used to train the problem-solving model, so that it is no longer necessary to manually generate corresponding answer videos based on the to-be-answered question, thereby reducing the cost of generating training data and improving the training efficiency of the trained problem-solving model.
[0075] In prompt engineering, a prompt can refer to a text or sentence used to guide a large language model to generate a specific response. The prompt can help the large language model better understand the user's needs, so that the large language model can generate content that better meets the user's needs.
[0076] In the embodiments of the present specification, the prompt of the large language model can be constructed to obtain explanation information corresponding to the to-be-answered question. The prompt in the embodiments of the present specification can be a text used to guide the large language model to generate explanation information of the to-be-answered question. The method for generating a problem-solving model can further include:
[0077] determining a reference example question having the same characteristics as the to-be-answered question; the characteristics include at least one of a characteristic of a number of line segments contained in a line segment diagram related to the question, a characteristic of a number of problems contained in the question, a characteristic of a number of answer results corresponding to the question, a characteristic of a number of diagrams required by the question, and a characteristic of a type of line segment diagram required by the question; and the reference example question includes an example question and video data for generating an explanation video of the example question.
[0078] The generating the explanation information of the to-be-answered question by using the large language model can specifically include:
[0079] Generating a prompt word based on the reference question and the to-be-answered question.
[0080] Inputting the prompt word into the large language model to obtain the explanation information corresponding to the to-be-answered question.
[0081] In the embodiments of the present specification, the reference question can include a question title and explanation video data used to generate the question title, wherein the explanation video data can be a code of an explanation video, or other forms of digital information.
[0082] In the embodiments of the present specification, the reference question can be selected from a reference question set. For example, a semantic matching model can be used to select a preset number of reference questions matched with the to-be-answered question from the reference question set. Alternatively, a retrieval enhancement generation model can be used to select a preset number of reference questions matched with the to-be-answered question from the reference question set. In addition, the reference questions can be selected from the reference question set in a keyword matching manner or other matching manners, and the like.
[0083] It can be understood that the reference question can also be written according to expert experience. Of course, the code of the explanation video of the existing question that can be answered by using the line segment graph can be extracted by using reverse engineering, and the obtained question and the code of the explanation video are used as the reference question, and the like.
[0084] The embodiments of the present specification can select a preset number of reference questions. For example, one reference question, three reference questions, ten reference questions, and the like are selected, which are not limited herein.
[0085] In the embodiments of the present specification, the large language model is guided to generate the explanation information by using a few-shot learning in a prompt, and a large amount of explanation information can be obtained, and then a large amount of training data can be obtained.
[0086] Further, the selected reference question can be a question having the same characteristics as the to-be-answered question. The characteristics can include at least one of a characteristic of a number of line segments contained in a line segment graph involved in the question, a characteristic of a number of problems contained in the question, a characteristic of a number of answer results corresponding to the question, a characteristic of a number of drawings required by the question, and a characteristic of a type of line segment graph required by the question.
[0087] The characteristic of the number of line segments involved in the question can represent the number of line segments contained in the line segment graph used in solving the question. For example, if the line segment graph used in solving the question contains one line segment, the characteristic of the number of line segments involved in the question can be one line segment. For example, if the line segment graph used in solving the question contains two line segments, the characteristic of the number of line segments involved in the question can be two line segments.
[0088] The characteristic of the number of questions contained in the question can represent the total number of questions contained in the question. For example, if the question contains one question, the characteristic of the number of questions contained in the question can be one question. For example, if the question contains two questions, the characteristic of the number of questions contained in the question can be two questions.
[0089] The characteristic of the number of questions contained in the question can also represent the number of question types contained in the question. For example, if the question contains one question type, such as "sum of quantities" type, the characteristic of the number of questions contained in the question can be one question type. For example, if the question contains two question types, such as "sum of quantities" type and "subtraction of quantities" type, the characteristic of the number of questions contained in the question can be two question types.
[0090] The characteristic of the number of answer results corresponding to the question can represent the total number of answer results contained in the answer results of the question. The characteristic of the number of drawings required for the question can represent the number of line segment graphs used in solving the question. For example, if one line segment graph is used in solving the question, the characteristic of the number of drawings required for the question can be one line segment graph. If multiple line segment graphs are used in solving the question, the characteristic of the number of drawings required for the question can be multiple line segment graphs.
[0091] The characteristic of the type of line segment graph required for the question can represent the type of line segment graph used in solving the question. For example, if the line segment graph used in solving the question is a line segment, the characteristic of the type of line segment graph required for the question can be a line segment. For example, if the line segment graph used in solving the question is a circle, the characteristic of the type of line segment graph required for the question can be a circle. For example, if the line segment graph used in solving the question is a rectangle, the characteristic of the type of line segment graph required for the question can be a rectangle, and so on.
[0092] Generally, in addition to the example question and the explanation video data, the reference example question can also include other content to facilitate the teaching of the large language model output explanation information, so as to assist teaching.
[0093] Optionally, the reference example question can also include at least one of the step-by-step analysis information for step-by-step solving the example question, the knowledge point information corresponding to the example question, and the answer information corresponding to the example question.
[0094] and / or,
[0095] The explanation information further includes at least one of step-by-step analysis information for step-by-step solving of the to-be-solved question, knowledge point information corresponding to the to-be-solved question, and answer information of the to-be-solved question.
[0096] The step-by-step analysis information can represent distributed explanation for the question, and the distributed explanation can include a line segment diagram corresponding to each step; the knowledge point information can represent a knowledge point examined by the question; and the answer information can represent an answer of the question.
[0097] For ease of understanding, an example of a reference example question is provided by the embodiments of the present specification, which can be as follows:
[0098]
Question
[0099] Radish bag is more than bean paste bag \(\frac{1}{4}\), exactly more than 8, how many are bean paste bags? Complete the following line segment diagram and solve.
[0100]
Knowledge Point
[0101] This question mainly examines the application of fractions and the problem solving method of line segment diagram.
[0102]
Step-by-step analysis
[0103] First, we draw two line segments, one representing the number of bean paste bags and the other representing the number of radish bags.
[0104] Second, since the radish bag is more than the bean paste bag \(\frac{1}{4}\), we can mark the part more than the bean paste bag on the line segment of the radish bag, which is exactly 8.
[0105] Third, we know that the part more than the bean paste bag in the radish bag accounts for \(\frac{1}{4}\) of the bean paste bag, so the number of bean paste bags is 4 times 8: \(8 \times 4 = 32\) (pieces).
[0106]
Answer
[0107] The number of bean paste bags is 4 times 8: \(8 \times 4 = 32\) (pieces)
[0108]
Code
[0109] from core.utils import *
[0110] class ProblemSolving(MyScene):
[0111] The question is: "There are 1 / 4 more radish buns than red bean buns, exactly 8 more. How many red bean buns are there?"
[0112] def keypoint(self):
[0113] # Showcase knowledge points
[0114] audio_text = "This question mainly tests the application of fractions and the problem-solving method of line segment diagrams."
[0115] with self.voiceover(text=audio_text):
[0116] text_keypoint = "Knowledge points: Applications of fractions and problem-solving methods using line segment diagrams"
[0117] text_keypoint = self.create_text(text_keypoint)
[0118] self.play(Write(text_keypoint))
[0119] def solve(self):
[0120] scale = self.get_scale()
[0121] # First step, we draw two line segments, one representing the number of red bean buns and the other representing the number of radish buns.
[0122] audio_text = "First, we use a line segment to represent the number of red bean buns."
[0123] with self.voiceover(text=audio_text):
[0124] # Use line segments to represent the number of red bean buns (32)
[0125] text_dousha = self.create_text("Red bean bun:")
[0126] dot_l = Dot([-1, 1, 0])
[0127] dot_r = Dot([-1 + 32 * scale, 1, 0])
[0128] line_dousha = Line(dot_l.get_center(), dot_r.get_center())
[0129] text_dousha.next_to(line_dousha, LEFT)
[0130] self.play(Create(line_dousha), Write(text_dousha), Create(dot_l),Create(dot_r))
[0131] audio_text = "Then, we use another line segment to represent the number of radish buns."
[0132] with self.voiceover(text=audio_text):
[0133] # Use a line segment to represent the number of radish buns (40)
[0134] text_luobo = self.create_text("Radish buns:")
[0135] dot_l = Dot([-1, -1, 0])
[0136] dot_r = Dot([-1 + 40 * scale, -1, 0])
[0137] line_luobo = Line(dot_l.get_center(), dot_r.get_center())
[0138] text_luobo.next_to(line_luobo, LEFT)
[0139] self.play(Create(line_luobo), Write(text_luobo), Create(dot_l),Create(dot_r))
[0140] # Step 2, since radish buns are \(\frac{1}{4}\) more than bean buns, we can mark the part that is more than bean buns on the line segment of radish buns, which is exactly 8.
[0141] audio_text = "Radish buns are \(\frac{1}{4}\) more than bean buns, which is exactly 8 more."
[0142] with self.voiceover(text=audio_text):
[0143] Divide the line segment of the red bean bun into four equal parts, each part being 8 units in length.
[0144] for distance in range(0, 32, 8):
[0145] dot = Dot([-1 + distance * scale, 1, 0])
[0146] self.play(Create(dot))
[0147] # Mark the 8 extra (32~40) on the line segment of the radish bun.
[0148] dot_l = Dot([-1 + 32 * scale, -1, 0])
[0149] dot_r = Dot([-1 + 40 * scale, -1, 0])
[0150] line_dousha_to_luobo=DashedLine(dot_l.get_center(),dot_r.get_cet
[0151] er())
[0152] self.play(Create(line_dousha_to_luobo),Create(dot_l),Create(dot_r))
[0153] brace_dousha_to_luobo = Brace(line_dousha_to_luobo, direction=UP)
[0154] text_dousha_to_luobo = self.create_text("8 more")
[0155] brace_dousha_to_luobo.put_at_tip(text_dousha_to_luobo)
[0156] self.play(Create(brace_dousha_to_luobo),Write(text_dousha_to_luobo))
[0157] brace_dousha_to_luobo=Brace(line_dousha_to_luobo,direction=DOWN)
[0158] text_dousha_to_luobo = self.create_text("More $\\frac{1}{4}$ items")
[0159] brace_dousha_to_luobo.put_at_tip(text_dousha_to_luobo)
[0160] self.play(Create(brace_dousha_to_luobo),Write(text_dousha_to_luobo))
[0161] # Third step, we know that the radish bun is larger than the red bean bun by \(\frac{1}{4}\), so the number of red bean buns is 4 times 8: \(8 \times 4 = 32\) (buns).
[0162] audio_text = "Therefore, the number of red bean buns is 4 times 8, which is 8 times 4 equals 32."
[0163] with self.voiceover(text=audio_text):
[0164] brace_dousha = Brace(line_dousha, direction=UP)
[0165] text_dousha = self.create_text("32")
[0166] brace_dousha.put_at_tip(text_dousha)
[0167] self.play(Create(brace_dousha), Write(text_dousha))
[0168] self.play_text_left(r"The number of red bean buns is 4 times 8: $8 \times 4 =
[0169] 32 (pieces)")
[0170] def get_scale(self):
[0171] # First step, collect all quantity information
[0172] dousha_count = 32
[0173] luobo_count = 32 + 8
[0174] # Second step, find the maximum value in the quantity
[0175] max_len = max(dousha_count, luobo_count)
[0176] # Calculate the scaling ratio to scale the line segment length to within 4 units
[0177] scale = 4 / max_len
[0178] return scale
[0179] Under the guidance of the prompt word, the generated explanation information of the large language model will also include information corresponding to the information included in the reference example. Thus, it can meet the requirements of step-by-step explanation, overall summary, etc. in the teaching process, and improve the teaching quality.
[0180] In the embodiments of the present specification, the large language model is used to automatically generate information including step-by-step analysis information, knowledge point information, answer information, and explanation information, which reduces the cumbersome process of manually writing the above content information, improves the efficiency of obtaining the explanation information of the to-be-solved problem, and further improves the training efficiency of the problem-solving model. Based on this, the problem-solving model training efficiency of the embodiments of the present specification is high.
[0181] In the embodiments of the present specification, the prompt word can be generated based on a prompt word template. The prompt word is generated based on the reference example and the to-be-solved problem, and specifically can include:
[0182] The prompt word is generated based on the reference example and the to-be-solved problem using the prompt word template.
[0183] The prompt word template is a structured text used to generate a prompt word. Generating a prompt word based on a prompt word template can help a large language model interact with a user, so that the large language model can better understand the user's intent and generate more accurate and user-demand-conforming replies.
[0184] In the embodiments of the present specification, the prompt word template can include a system prompt word. The system prompt word can be description information for the current task. The description information can specifically include one or more of role information of the large language model, task information indicating completion of the large language model, task requirement information, and output data format information.
[0185] The role information of the large language model can be role information that the large language model needs to play for the current task. For example, the large language model needs to play the role of a math teacher, and the role information of the large language model can be "you are a math teacher". The task information indicating completion of the large language model can be task information that the large language model needs to complete for the current task. The task requirement information can be requirement information for the current task. The task requirement information can include various information for explaining the video data, such as at least one of code type information, voice format information, and voice playback format information.
[0186] In the embodiments of the present specification, the prompt word template can also include a first to-be-filled region and a second to-be-filled region. The first to-be-filled region can be used to fill in a reference example, and the second to-be-filled region can be used to fill in a to-be-solved problem.
[0187] The server or training platform can add the system prompt word to the prompt word template, add the reference example to the first to-be-filled region, and add the to-be-solved problem to the second to-be-filled region, and the like, so that the prompt word generated by the prompt word template can include the system prompt word, the reference example, and the to-be-solved problem, and the like.
[0188] In order to improve the richness of the training data and thus improve the accuracy of the problem solving model, in the embodiments of the present specification, a plurality of categories of to-be-solved problems and explanation information corresponding to the plurality of categories of to-be-solved problems can be used to generate training data. Optionally, the to-be-solved problem includes a plurality of to-be-solved problems; and the generating training data based on the to-be-solved problem and the explanation information can specifically include:
[0189] classifying the plurality of to-be-solved problems according to a problem feature to obtain a plurality of problem sets; each of the problem sets contains a plurality of problems of the same category; and the problem feature includes at least one of a feature of a number of line segments contained in a line segment graph involved in the problem, a feature of a number of questions contained in the problem, a feature of a number of solution results corresponding to the problem, a feature of a number of drawings required by the problem, and a feature of a type of line segment graph required by the problem.
[0190] selecting at least one to-be-solved problem from each of the plurality of problem sets.
[0191] using each of the selected to-be-solved problems and the corresponding explanation information including the video data as the training data.
[0192] In the embodiments of the present specification, based on the topic characteristics, algorithms such as Naive Bayes algorithm, k-nearest neighbor algorithm, and clustering algorithm can be used to classify a plurality of to-be-solved topics.
[0193] For any topic set, a to-be-solved topic can be randomly selected from the any topic set, and then the randomly selected to-be-solved topic can be used as training data.
[0194] It can be understood that for any topic set, one to-be-solved topic can be selected, or a plurality of to-be-solved topics can be selected, which is not limited herein.
[0195] In actual application, one to-be-solved topic can have a plurality of characteristics. In this way, the same to-be-solved topic can be included in different topic sets. In order to avoid affecting the richness of the training data by selecting the same to-be-solved topic from a plurality of topic sets, in the embodiments of the present specification, if the plurality of topic sets include a first topic set and a second topic set, and at least one to-be-solved topic is selected from each topic set in the plurality of topic sets, the selection can specifically include:
[0196] A first number of first topics are selected from the first topic set.
[0197] If the first topic is included in the second topic set, a second number of second topics are selected from other topics in the second topic set; the other topics are topics in the second topic set except the first topic.
[0198] The first number of first topics can be one first topic, three first topics, or all topics in the first topic set, etc. The second number of second topics can be one second topic, three second topics, or all topics in the second topic set except the first topic, etc.
[0199] In addition, the first number and the second number can be equal or not equal.
[0200] In the embodiments of the present specification, in order to ensure the quality of the training data, the training data can be generated by selecting the explanation information meeting the problem solving requirements and the corresponding to-be-solved topic. The method for generating a problem solving model in the embodiments of the present specification can further include:
[0201] Based on the explanation video data for the to-be-solved topic included in the explanation information, an explanation video of the to-be-solved topic is generated.
[0202] It is judged whether the explanation video meets the problem solving requirements.
[0203] The generating training data based on the question to be answered and the explanation information can specifically include:
[0204] If the explanation video meets the problem solving requirement, the question to be answered and the corresponding explanation information are taken as training data.
[0205] In the embodiments of the present specification, the video compiling software can be used to compile the explanation video data of the question to be answered, to obtain the explanation video of the question to be answered.
[0206] In actual application, the explanation information can include various information, for example, step-by-step analysis information of the question to be answered, knowledge point information corresponding to the question to be answered, and answer information corresponding to the question to be answered, etc. Therefore, when the explanation video of the question to be answered is needed, the explanation video data can be filtered out from the explanation information.
[0207] Optionally, the method can further include:
[0208] The regular expression is used to determine the explanation video data contained in the explanation information.
[0209] In the embodiments of the present specification, the regular expression is a tool for describing string patterns, and can be used to find, match and replace specific strings in text. The regular expression in the embodiments of the present specification is a regular expression that can extract the explanation video data in the explanation information. For example, the regular expression can include a code identifier character, such as Python, and the information after the code identifier character can be extracted starting from the code identifier character. Commonly used regular expressions in Python include if…print, etc.
[0210] In the embodiments of the present specification, the video compiling software can be used to compile the explanation video data of the question to be answered, to obtain the explanation video of the question to be answered. Specifically, it can include:
[0211] The target video data is saved as a Python file.
[0212] The Python file is compiled using a manim command.
[0213] In the embodiments of the present specification, the video compiling software used can be Manim software. The manim (Mathematical Animation Engine) command is a Python library for making animations; the compilation of the Python file using the manim command can generate an explanation video.
[0214] It can be understood that in addition to using the Manim software to compile the explanation video, other software such as HandBrake software and FFmpeg software can also be used to compile the explanation video. This is not limited here.
[0215] Further, the embodiments of the present specification can also optimize the video compilation software, and compile the video by using the optimized video compilation software. The optimized video compilation software has an optimized logic module built-in. For the compilation of the explanation video data, a small amount of code can be used to complete the work that a large amount of code is needed before optimization. For example, if distribution analysis information needs to be displayed on the left side of the video and a line segment diagram needs to be displayed on the right side of the video, the video compilation software before optimization needs to accurately determine the absolute coordinate information of the distribution analysis information and the line segment diagram. However, based on the optimized logic module, the optimized video compilation software can execute "display the distribution analysis information on the left side and display the line segment diagram on the right side", and then automatically find the specific position on the left side and the specific position on the right side, and further display the analysis information on the left side and display the line segment diagram on the right side, without the need to accurately determine the absolute coordinate information of the distribution analysis information and the line segment diagram. Therefore, the video compilation efficiency of the optimized video compilation software is higher.
[0216] In the embodiments of the present specification, the explanation video can include images and voice synchronized with the images.
[0217] Further, the explanation video can also include video content of a step-by-step problem solving process. A part of the video display page of the explanation video is used to display blackboard content, and another part is used to display line segment diagram content. Among them, the blackboard content can be at least one of step-by-step analysis information of a to-be-solved problem, knowledge point information corresponding to the to-be-solved problem, and answer information corresponding to the to-be-solved problem.
[0218] Figure 2 is a schematic diagram of an image frame of an explanation video provided by an embodiment of the present specification. As shown in Figure 2 , the left side of the explanation video displays distribution analysis information, and the right side of the explanation video displays a line segment diagram. Figure 2
[0219] In the embodiments of the present specification, the explanation video data can be video data including problem introduction, knowledge point introduction, problem explanation, and ending. Based on the problem introduction video data, a problem introduction video can be obtained, based on the knowledge point introduction video data, a knowledge point introduction video can be obtained, based on the ending video data, an ending video can be obtained, and so on.
[0220] The topic introduction video can be a video for introducing the topic. For example, the topic introduction video can include a display of the topic, accompanied by a voice: "Let me explain the following question to you", and then the question can be announced by voice.
[0221] In addition, the explanation video data can also include video data for introducing the identity of the teaching software that loads or calls the trained problem solving model, and based on the video data, a video for introducing the identity of the teaching software that loads or calls the trained problem solving model can be obtained. For example, it can be a voice broadcast: "Hello, I am AI intelligent math teacher".
[0222] Since the explanation video is compiled from the explanation video data, whether the explanation video data meets the problem solving requirements can be determined by judging whether the explanation video meets the problem solving requirements, so as to determine whether the explanation information meets the problem solving requirements.
[0223] In the embodiments of the present specification, the explanation information capable of generating an explanation video meeting the problem solving requirements and the corresponding to-be-solved question are used as training samples, which can reduce the interference of incorrect training samples on the problem solving model, improve the accuracy of the trained problem solving model, and thus the trained problem solving model can more accurately solve the questions solved by line segment diagram.
[0224] It can be understood that after the explanation information meeting the problem solving requirements is determined according to the explanation video, the to-be-solved question corresponding to the explanation information meeting the problem solving requirements can be determined based on the explanation information meeting the problem solving requirements, so that the positive sample training data can be generated according to the determined explanation information and the to-be-solved question.
[0225] Of course, after the explanation information meeting the problem solving requirements is screened out from all the explanation information in the above manner, the explanation video data in the explanation information not meeting the problem solving requirements can also be modified to obtain correct explanation video data. The positive sample training data is generated by using the modified correct explanation information and the to-be-solved question corresponding to the modified correct explanation information.
[0226] In practical applications, the explanation video can be a question explanation using distribution analysis information, so the last frame of the explanation video often contains more content than other frames of the explanation video. In order to conveniently and accurately determine whether the explanation video meets the problem solving requirements, whether the explanation video meets the problem solving requirements can be determined based on the last frame of the explanation video. The determination of whether the explanation video meets the problem solving requirements can specifically include:
[0227] acquire a last frame image of the explanation video.
[0228] determine whether at least one of the following conditions exists in the last frame image:
[0229] the last frame image does not contain a correct answer.
[0230] the last frame image does not contain a line segment diagram.
[0231] the line segment diagram displayed in the last frame image is incomplete.
[0232] the content displayed in the last frame image is overlapped.
[0233] the line segment diagram displayed in the last frame image is inconsistent with the explanation text information contained in the explanation information.
[0234] the quantity relationship represented by the line segment diagram displayed in the last frame image is inconsistent with the quantity relationship contained in the question to be answered.
[0235] the position of the diagram parameter contained in the line segment diagram displayed in the last frame image does not conform to the preset position; the diagram parameter includes at least one of a number and a bracket.
[0236] If not, it is determined that the explanation video meets the problem solving requirements.
[0237] In the embodiments of the present specification, one or more models such as a semantic analysis model and a word vector model can be used to analyze whether the explanation content displayed in the last frame image of the question and the explanation video matches, to determine whether the explanation video contains a correct answer. Correctly answer the question.
[0238] Further, image recognition technology can be used to identify whether a line segment diagram exists in the last frame image of the explanation video, to determine whether a line segment diagram exists in the last frame image.
[0239] In practical applications, in addition to ensuring that a line segment diagram exists in the last frame image of the explanation video, it is also necessary to ensure that the line segment diagram is qualified.
[0240] Specifically, one or more of edge detection algorithms, contour detection algorithms, and hash algorithms can be used to detect the boundaries of the image to identify whether the line segment diagram is complete, to determine whether the line segment diagram is qualified. Of course, other methods can also be used to identify whether the line segment diagram is complete, which is not limited herein.
[0241] One or more of a bounding box detection method, an image segmentation method, and a deep learning model can be used to detect whether the line segment diagram in the last frame image is overlapped, to determine whether the line segment diagram is qualified.
[0242] Further, the semantic recognition model can also be used to identify whether the line segment diagram displayed in the last frame of image corresponds to the explanation text information, to determine whether the line segment diagram displayed in the last frame of image is consistent with the explanation text information contained in the explanation information, so as to determine whether the line segment diagram is qualified. The semantic recognition model can be used to detect whether the quantity relationship represented by the line segment diagram displayed in the last frame of image is consistent with the quantity relationship contained in the to-be-solved question, to determine whether the line segment diagram is qualified. Of course, other models or methods can also be used for detection and recognition, which are not limited here.
[0243] In addition, the position of the diagram parameter contained in the line segment diagram displayed in the last frame of image can also be identified by using key point matching and the like, so as to determine whether the position of the diagram parameter contained in the line segment diagram is in a preset position. If yes, it can be considered that the line segment diagram is qualified. Otherwise, it can be considered that the line segment diagram is not qualified. The diagram parameter can include at least one of a number and a bracket.
[0244] It can be understood that, in addition to using the last frame of image of the explanation video to determine whether the explanation video meets the problem-solving requirements, other frames of image of the explanation video can also be used to determine whether the explanation video meets the problem-solving requirements. The specific process of using other frames of image of the explanation video to determine whether the explanation video meets the problem-solving requirements can be referred to the related content of using the last frame of image of the explanation video to determine whether the explanation video meets the problem-solving requirements mentioned above, which is not described here.
[0245] In order to ensure the teaching quality, it is also necessary to synchronize the drawing animation displayed in the explanation video with the voice explanation, so as to avoid affecting the teaching experience. The determination of whether the explanation video meets the problem-solving requirements can specifically include:
[0246] Determining whether the drawing animation displayed in the explanation video is synchronized with the voice explanation.
[0247] If the drawing animation displayed in the explanation video is synchronized with the voice explanation, it is determined that the explanation video meets the problem-solving requirements.
[0248] Specifically, a model for determining whether the drawing animation and the voice explanation content are synchronized can be trained in advance, and then the trained model is used to identify whether the drawing animation and the voice explanation content are synchronized. Of course, related video editing software or audio analysis tools can also be used to analyze the synchronization of video and audio.
[0249] In the embodiments of the present specification, the training data for training the problem-solving model can be data in a predetermined format, so as to facilitate the problem-solving model to understand the training data, improve the training efficiency of the problem-solving model, and improve the model accuracy.
[0250] Optionally, the training data is generated based on the question to be answered and the explanation information, and specifically can include:
[0251] The question to be answered and the explanation information are generated into training data conforming to the training data format based on the training data format.
[0252] In the embodiments of the present specification, the question to be answered and the explanation information of the question to be answered can be spliced in a preset order to generate training data.
[0253] Specifically, the question to be answered can be placed in front of the explanation information, and a first preset character is used to separate the question to be answered and the explanation information. In addition, a second preset character can also be set before the question to be answered, so as to facilitate the problem solving model to understand the training data and improve the training efficiency and model accuracy of the problem solving model. The first preset character and the second preset character can be set according to actual needs, which are not limited here.
[0254] For the convenience of understanding, the embodiments of the present specification are specifically described for the training data format.
[0255] Optionally, the training data format can include a role definition part, a question part, and an analysis part.
[0256] The role definition part includes information for defining the role played by the problem solving model.
[0257] The question part includes information of the question to be answered.
[0258] The analysis part includes knowledge point information, step-by-step analysis information, answer information, and explanation video data corresponding to the question to be answered.
[0259] In the embodiments of the present specification, the training data format can include multiple identification words with different meanings, and the problem solving model can learn the training data from different angles according to the identification words.
[0260] Specifically, the identification word can include a first identification word such as [SYSTEM], which can be used to define the role of the model and explain the task of the model this time. The first identification word can be located at the starting position of the role definition part. The identification word can also include a second identification word such as [USER], which can be the question to be answered. The identification word can also include a third identification word such as [ASSISTANT], which can be the analysis part of the question to be answered. Of course, the above identification words can also be other fields of words or texts, which are not limited here.
[0261] Further, the problem solving model can understand the task of the content after the first identification word this time, and therefore, the first identification word can be used to distinguish different business scenarios. In addition, according to the first identification word, the problem solving model can also determine the identity of the model performing the task this time, for example, the problem solving model is a math teacher. The second identification word and the third identification word can be understood as a group of dialogues. The second identification word can be a question asked by the user as user, and the content after the second identification word can be the content of the question asked by the user. The third identification word can be the answer content made by the problem solving model as assistant to the question asked by user, and the content after the third identification word can be the answer content of the assistant. The above three identification words are spliced in order, which can be used as training data to train the problem solving model.
[0262] It can be understood that [ASSISTANT] can be the first preset character mentioned in the foregoing. [USER] can be the second preset character mentioned in the foregoing.
[0263] For example, the format and content of the training data can be as follows:
[0264] [SYSTEM] You are a math teacher, please write a code using Manim to make a teaching video for the following question.
[0265] [USER] Xiaoli has 8 apples, and Xiaoming has 2 more apples than Xiaoli. How many apples do Xiaoli and Xiaoming have in total?
[0266] [ASSISTANT]
Knowledge point
Step-by-step analysis
Answer information
Code information
[0267] Generally, we want the trained problem solving model to be able to output accurate explanation information for the question after obtaining the question. Specifically, when training the problem solving model using the training data, the problem solving model can output prediction data according to the role definition part and the question part proposed by the user, and the problem solving model can also use the analysis part of the question as reference data, and then obtain the trained problem solving model.
[0268] In the embodiments of the present specification, the problem solving model can be trained using supervised fine-tuning. In the training process, the problem solving model can adjust its hyperparameters according to the training data to obtain the trained problem solving model.
[0269] The training target of the problem solving model is to minimize the difference between the reference data and the predicted data, that is, to minimize the loss value loss between the reference data and the predicted data. Therefore, the parameters of the problem solving model can be continuously adjusted during the training process to reduce the loss value loss. When the loss value loss no longer decreases or decreases to within a preset range, it can be indicated that the problem solving model training is completed.
[0270] To improve the training efficiency, the problem solving model in the embodiments of the present specification can be a pre-trained model, such as a YunLLM70B model, a Multimodal-CoT model, or a Mixtral8x7B model, etc. Of course, it can also be other types of pre-trained models, which are not specifically limited here.
[0271] In the embodiments of the present specification, the obtained training data can be used to train the problem solving model.
[0272] Suppose multiple training data are obtained, each training data can include a to-be-solved question, knowledge point information, step-by-step analysis information, answer information, and the code of the corresponding explanation video, etc. In addition, the training data should include various types of to-be-solved questions and other information corresponding to the various types of to-be-solved questions, so as to improve the generalization ability of the problem solving model.
[0273] After obtaining the training data, data cleaning and formatting can also be performed on the training data to ensure the consistency and standardization of the data, thereby facilitating the training of the problem solving model. Specifically, tokenization, noise removal processing, and formatting processing of the explanation video data can be performed on the training data, etc. If the knowledge point information, step-by-step analysis information, answer information, and corresponding explanation video data in the training data are generated by using a large language model, since the data generated by the large language model is usually standardized and uniform in format, data cleaning and formatting processing can not be performed on the training data at this time.
[0274] When the problem solving model is initially trained, the problem solving model can be initialized, and some training parameters can be set. For example, for the YunLLM70B model, the YunLLM70B model can be loaded, and training parameters such as learning rate, batch size, and training round number can be configured.
[0275] In addition, the training data can be used to fine-tune the problem solving model. Specifically, the to-be-solved question can be used as input, and the generated accurate knowledge point information, step-by-step analysis information, answer information, and explanation video data can be used as the target. During the training process, the problem solving model continuously adjusts the parameters through back propagation to minimize the difference between the generated explanation video data and the labeled explanation video data.
[0276] During the problem-solving model training process, the performance of the model can also be evaluated periodically or irregularly using the validation set to adjust the problem-solving model parameters and training strategies, thereby ensuring the stability and accuracy of the problem-solving model.
[0277] After the problem-solving model training is completed, the model parameters and structure of the trained problem-solving model can be saved for subsequent use and deployment.
[0278] In practical applications, model evaluation is an important step to verify the quality and reliability of the problem-solving model. As an implementation, the performance of the trained problem-solving model can be evaluated by the following steps.
[0279] A test set is obtained. Specifically, 100 questions can be randomly selected from the test question library, and it is determined that the questions in the test set do not overlap with the training data. In addition, the test set usually needs to ensure that the questions cover various types and complexities of questions that can be solved using line segment graphs, so as to comprehensively evaluate the performance of the problem-solving model.
[0280] The trained problem-solving model is used to generate corresponding explanation video data for each question in the test set. The generation process is similar to the training phase. Specifically, the question can be input into the trained problem-solving model, and the trained problem-solving model can output knowledge point information, step-by-step analysis information, and explanation video data.
[0281] The generated explanation video data is input into a video compilation software to compile an explanation video.
[0282] Further, the performance of the trained problem-solving model can be evaluated according to the compiled explanation video. Specifically, the explanation video can be evaluated using pre-determined evaluation criteria. The evaluation criteria can include the accuracy of the video, such as whether the problem-solving steps and results are correct, the intuitiveness of the video, such as whether the video is clear and easy to understand, the liveliness of the video, such as whether the explanation process is lively and interesting, and the technicality of the video, such as whether the code generation is efficient and the video playback is smooth.
[0283] In addition, the generated video can also be evaluated using the method mentioned earlier to determine whether the explanation video meets the problem-solving requirements. Professional teachers and technical experts can also be organized to manually evaluate the generated video and record the evaluation results. Of course, the above-mentioned methods can be combined, which is not limited here.
[0284] Further, the evaluation results can be statistically analyzed to evaluate the performance of the trained model. According to the analysis results, the advantages and disadvantages of the trained problem-solving model are determined, thereby providing a basis for further optimizing the problem-solving model.
[0285] Through the above evaluation steps, it is ensured that the problem solving model after training has the ability to generate high-quality explanation video data, and reliable technical support is provided for actual application.
[0286] Corresponding to the method embodiments described above, the specification also provides an embodiment of a method for solving a problem, Figure 3 is a flowchart of a method for solving a problem provided by an embodiment of the specification. As shown in Figure 3 , the method can include:
[0287] Step 302: Obtain a target problem to be solved.
[0288] In an embodiment of the specification, the target problem to be solved can be a problem input by a user, specifically a problem that can be solved using a line segment graph.
[0289] Step 304: Input the target problem to be solved into the trained problem solving model to obtain target explanation information; the target explanation information includes target video data for generating an explanation video for solving the target problem to be solved using a line segment graph; the trained problem solving model is obtained by training a problem solving model using training data; the training data is generated based on explanation information of a problem to be solved that can be solved using a line segment graph and the problem to be solved.
[0290] The process of obtaining the trained problem solving model can refer to the previously mentioned method of generating a problem solving model, which is not repeated here.
[0291] In an embodiment of the specification, based on the target problem to be solved, the trained problem solving model can directly obtain the target explanation information. Compared with obtaining the target explanation information using a large language model, since the trained problem solving model is obtained by training a problem that can be solved using a line segment graph, it is more accurate to generate explanation information for a problem that can be solved using a line segment graph, and thus the quality of the generated explanation video can be guaranteed.
[0292] In addition, when using the trained problem solving model to solve a problem, there is no need to preset templates to obtain a large number of prompt words similar to those used by the large language model, which can also improve the problem solving efficiency.
[0293] Step 306: Render the target video data using a video rendering tool to obtain a target explanation video for the target problem to be solved.
[0294] In an embodiment of the specification, the target video data can be a code of the target explanation video for the target problem to be solved, or other forms of digital information.
[0295] In an embodiment of the present specification, the target explanation video includes images and voice synchronized with the images; the target explanation video includes video content of step-by-step explanation of the problem solving process, a part of the video display page of the target explanation video is used to display the board content, and another part is used to display the line segment graph content. The board content can be at least one of step-by-step analysis information of the target problem to be solved, knowledge point information corresponding to the target problem to be solved, and answer information corresponding to the target problem to be solved.
[0296] The process of screening target video data from target explanation information can refer to the specific process of screening explanation video from explanation information in the embodiment of the method of generating a problem solving model, which is not repeated here.
[0297] The process of rendering target video data by using a video rendering tool can refer to the specific process of compiling explanation video data by using a video compiling software in the embodiment of the method of generating a problem solving model, which is not repeated here.
[0298] In a specific embodiment, the target video data can also be the code of the animation in the explanation video. In actual application, the explanation text information in the target explanation information can be converted into voice by using a text-to-speech (TTS) technology, the code of the animation in the explanation video can be rendered into an animation by using a video compiling software, and then the animation and the voice can be synthesized by using an audio synthesis technology, so as to obtain the target explanation video of the target problem to be solved.
[0299] Corresponding to the above-mentioned embodiment of the method of generating a problem solving model, the present specification also provides an embodiment of an apparatus for generating a problem solving model, Figure 4 A structural schematic diagram of an apparatus for generating a problem solving model provided by an embodiment of the present specification is shown. As shown in the figure, Figure 4 The apparatus includes:
[0300] The problem obtaining module 402 is configured to obtain a problem to be solved; the problem to be solved is a problem that can be solved by using a line segment graph.
[0301] The determining module 404 is configured to determine explanation information of the problem to be solved; the explanation information includes explanation video data for the problem to be solved.
[0302] The data generating module 406 is configured to generate training data based on the problem to be solved and the explanation information.
[0303] The model training module 408 is configured to train a problem solving model by using the training data to obtain a trained problem solving model; the trained problem solving model is used to generate explanation video data for solving a problem by using a line segment graph.
[0304] Optionally, the determining module 404 can be specifically configured to:
[0305] generating the explanation information of the to-be-solved question by using a large language model.
[0306] Optionally, the apparatus can further include:
[0307] a question example determining module configured to determine a reference question example having the same characteristics as the to-be-solved question; the characteristics include at least one of a characteristic of a number of line segments contained in a line segment diagram involved in the question, a characteristic of a number of problems contained in the question, a characteristic of a number of solution results corresponding to the question, a characteristic of a number of diagrams required by the question, and a characteristic of a type of line segment diagram required by the question; the reference question example includes a question example and video data used to generate an explanation video of the question example.
[0308] The generating of the explanation information of the to-be-solved question by using a large language model can specifically include:
[0309] generating prompt words based on the reference question example and the to-be-solved question.
[0310] inputting the prompt words into the large language model to obtain the explanation information corresponding to the to-be-solved question.
[0311] Optionally, the determining of the reference question example having the same characteristics as the to-be-solved question can specifically include:
[0312] selecting a preset number of reference question examples matched with the to-be-solved question from a reference question example set by using a semantic matching model.
[0313] Optionally, the to-be-solved question includes a plurality of to-be-solved questions; and the data generating module 406 can be specifically configured to:
[0314] classifying the plurality of to-be-solved questions according to question characteristics to obtain a plurality of question sets; each of the question sets contains a plurality of questions of the same category; and the question characteristics include at least one of a characteristic of a number of line segments contained in a line segment diagram involved in the question, a characteristic of a number of problems contained in the question, a characteristic of a number of solution results corresponding to the question, a characteristic of a number of diagrams required by the question, and a characteristic of a type of line segment diagram required by the question.
[0315] selecting at least one to-be-solved question from each of the plurality of question sets.
[0316] taking each of the selected to-be-solved questions and the explanation information corresponding thereto as the training data.
[0317] Optionally, the apparatus can further include:
[0318] The explanation video generation module is configured to generate an explanation video of the to-be-solved question based on explanation video data of the to-be-solved question contained in the explanation information.
[0319] The judgment module is configured to judge whether the explanation video meets a question-solving requirement.
[0320] The data generation module 406 can be specifically configured to:
[0321] If the explanation video meets the question-solving requirement, the to-be-solved question and the corresponding explanation information are taken as training data.
[0322] Optionally, the judgment module can be specifically configured to:
[0323] Obtain a last frame image of the explanation video.
[0324] Judge whether at least one of the following situations exists in the last frame image:
[0325] There is no correct answer in the last frame image.
[0326] There is no line segment graph in the last frame image.
[0327] The line segment graph displayed in the last frame image is incomplete.
[0328] There is overlap in the content displayed in the last frame image.
[0329] The line segment graph displayed in the last frame image is inconsistent with explanation text information contained in the explanation information.
[0330] The quantity relationship represented by the line segment graph displayed in the last frame image is inconsistent with a quantity relationship contained in the to-be-solved question.
[0331] The position of a diagram parameter contained in the line segment graph displayed in the last frame image does not conform to a preset position; the diagram parameter includes at least one of a number and a bracket.
[0332] If not, it is determined that the explanation video meets the question-solving requirement.
[0333] Optionally, the explanation information further includes at least one of step-by-step analysis information for step-by-step solving of the to-be-solved question, knowledge point information corresponding to the to-be-solved question, and answer information of the to-be-solved question.
[0334] Optionally, the explanation video data is video data including question introduction, knowledge point introduction, question explanation, and ending.
[0335] The above is a schematic scheme of the problem solving model generating device of the embodiment. It should be noted that the technical scheme of the problem solving model generating device is the same as the technical scheme of the problem solving model generating method described above, and the details of the technical scheme of the problem solving model generating device not described in detail can be seen from the description of the technical scheme of the problem solving model generating method.
[0336] Corresponding to the above-mentioned method embodiment for answering questions, the present specification also provides a device embodiment for answering questions, Figure 5 A structural schematic diagram of a device for answering questions provided by an embodiment of the present specification is shown. As shown in the figure, Figure 5 The device comprises:
[0337] A target question obtaining module 502 is configured to obtain a target question to be answered.
[0338] An input module 504 is configured to input the target question to be answered into a trained problem solving model to obtain target explanation information; the target explanation information contains target video data for generating an explanation video of the target question to be answered using a line segment graph; the trained problem solving model is obtained by training a problem solving model using training data; and the training data is generated based on explanation information of a question to be answered that can be answered using a line segment graph and the question to be answered.
[0339] A rendering module 506 is configured to render the target video data using a video rendering tool to obtain a target explanation video of the target question to be answered.
[0340] The above is a schematic scheme of the device for answering questions of the embodiment. It should be noted that the technical scheme of the device for answering questions is the same as the technical scheme of the method for answering questions described above, and the details of the technical scheme of the device for answering questions not described in detail can be seen from the description of the technical scheme of the device for answering questions.
[0341] Figure 6 A structural block diagram of a computing device 600 according to an embodiment of the present specification is shown. The components of the computing device 600 include but are not limited to a memory 610 and a processor 620. The processor 620 is connected with the memory 610 through a bus 630, and a database 650 is used to save data.
[0342] The computing device 600 also includes an access device 640 that enables the computing device 600 to communicate via one or more networks 660. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or combinations of such networks, such as the Internet. The access device 640 can include one or more of any type of network interface (for example, a network interface card (NIC)) such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC).
[0343] In one embodiment of the present specification, the above-mentioned components of the computing device 600 and other components not shown in the above-mentioned components can be connected to each other, for example, through a bus. It should be understood that Figure 6 Figure 6 The computing device structure diagram shown is merely for the purpose of example, and is not a limitation on the scope of the present specification. Other components can be added or replaced by those skilled in the art as needed.
[0344] The computing device 600 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (for example, a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, and the like), a mobile phone (for example, a smartphone), a wearable computing device (for example, a smart watch, smart glasses, and the like), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 600 can also be a mobile or stationary server.
[0345] The processor 620 is configured to execute computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned method for generating a problem solving model, and / or the steps of the above-mentioned method for solving a problem.
[0346] The above describes a schematic scheme of the computing device of the embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the problem solving model method described above, and / or the technical scheme of the problem solving method described above belong to the same concept. Details of the technical scheme of the computing device which are not described in detail can be referred to the description of the technical scheme of the problem solving model method described above, and / or the technical scheme of the problem solving method described above.
[0347] An embodiment of the present specification also provides a computer readable storage medium storing computer executable instructions. When the computer executable instructions are executed by a processor, the steps of the above method for generating a problem solving model, and / or the steps of the above problem solving method are implemented.
[0348] The above describes a schematic scheme of the computer readable storage medium of the embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the problem solving model method described above, and / or the technical scheme of the problem solving method described above belong to the same concept. Details of the technical scheme of the storage medium which are not described in detail can be referred to the description of the technical scheme of the problem solving model method described above, and / or the technical scheme of the problem solving method described above.
[0349] An embodiment of the present specification also provides a computer program. When the computer program is executed in a computer, the computer executes the steps of the above method for generating a problem solving model, and / or the steps of the above problem solving method.
[0350] The above describes a schematic scheme of the computer program of the embodiment. It should be noted that the technical scheme of the computer program and the technical scheme of the problem solving model method described above, and / or the technical scheme of the problem solving method described above belong to the same concept. Details of the technical scheme of the computer program which are not described in detail can be referred to the description of the technical scheme of the problem solving model method described above, and / or the technical scheme of the problem solving method described above.
[0351] The above describes specific embodiments of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order and still accomplish desirable results. Additionally, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order to achieve desirable results. In certain implementations, multitasking and parallel processing can be advantageous.
[0352] The computer readable medium can include any entity or apparatus capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, software distribution medium, etc. It should be noted that the computer readable medium can include appropriate additions or deletions according to the requirements of patent practice, for example, according to the patent practice in some regions, the computer readable medium does not include electrical carrier signals and telecommunication signals.
[0353] It should be noted that for the foregoing method embodiments, in order to facilitate description, they are all expressed as a combination of a series of actions, but those skilled in the art should know that the embodiments of the present specification are not limited by the order of the described actions, because according to the embodiments of the present specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of the present specification.
[0354] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0355] The preferred embodiments of the present specification disclosed above are only used to help explain the present specification. The alternative embodiments do not describe all the details and limit the invention to the specific embodiments described. Obviously, according to the content of the embodiments of the present specification, many modifications and changes can be made. The present specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present specification, so that those skilled in the art can well understand and use the present specification. The present specification is limited only by the claims and their entire scope and equivalents.
Claims
1. A method of generating a problem-solving model, characterized by, The method comprises the following steps: obtaining a to-be-solved question; the to-be-solved question is a question that can be solved by using a line segment graph; determining the explanation information of the to-be-solved question; the explanation information contains explanation video data for the to-be-solved question, and the explanation video data for the to-be-solved question is used to generate the code of the explanation video of the to-be-solved question; based on the to-be-solved question and the explanation information, generate training data; training the problem solving model using the training data to obtain a trained problem solving model; the trained problem solving model is used to generate explanation video data for solving a question by using a line segment graph, and the explanation video data for solving a question by using a line segment graph includes the code of the explanation video for solving a question by using a line segment graph; the method further comprises: based on the explanation video data for the to-be-solved question contained in the explanation information, generate the explanation video of the to-be-solved question; obtain the last frame image of the explanation video, and determine whether the following at least one condition exists in the last frame image: there is no correct answer in the last frame image; there is no line segment graph in the last frame image; the line segment graph displayed in the last frame image is incomplete; there is overlap in the content displayed in the last frame image; the line segment graph displayed in the last frame image is inconsistent with the explanation text information contained in the explanation information; the quantity relationship represented by the line segment graph displayed in the last frame image is inconsistent with the quantity relationship contained in the to-be-solved question; the position of the diagram parameter contained in the line segment graph displayed in the last frame image does not conform to the preset position; the diagram parameter includes at least one of a number and a parenthesis; if the above at least one condition does not exist, it is determined that the explanation video meets the problem solving requirements; based on the to-be-solved question and the explanation information, generate training data, specifically including: if the explanation video meets the problem solving requirements, the to-be-solved question and the corresponding explanation information are used as training data.
2. The method of claim 1, wherein, determining the explanation information of the to-be-solved question, specifically including: generating the explanation information of the to-be-solved question by using a large language model.
3. The method of claim 2, wherein, the method further comprises: determining a reference example question having the same characteristics as the to-be-solved question; the characteristics include at least one of the characteristics of the number of line segments contained in the line segment graph related to the question, the characteristics of the number of problems contained in the question, the characteristics of the number of answer results corresponding to the question, the characteristics of the number of diagrams required by the question, and the characteristics of the type of line segment graph required by the question; the reference example question includes an example question and video data for generating an explanation video of the example question; generating the explanation information of the to-be-solved question by using a large language model, specifically including: based on the reference example question and the to-be-solved question, generate prompt words; input the prompt words into the large language model to obtain the explanation information corresponding to the to-be-solved question.
4. The method of claim 3, wherein, determining a reference example question having the same characteristics as the to-be-solved question, specifically including: using a semantic matching model to select a preset number of reference example questions matched with the to-be-solved question from a reference example question set.
5. The method of claim 1, wherein, The to-be-solved question includes a plurality of to-be-solved questions; and the training data is generated based on the to-be-solved question and the explanation information, and specifically includes: classifying the plurality of to-be-solved questions according to question characteristics to obtain a plurality of question sets; each question set includes a plurality of questions of the same category; the question characteristics include at least one of a characteristic of a number of line segments included in a line segment graph involved by the question, a characteristic of a number of problems included in the question, a characteristic of a number of solution results corresponding to the question, a characteristic of a number of drawings required by the question, and a characteristic of a type of line segment graph required by the question; selecting at least one to-be-solved question from each of the plurality of question sets; using the selected to-be-solved questions and the corresponding explanation information including the explanation video data as the training data.
6. The method of claim 1, wherein, The explanation information further includes at least one of step-by-step analysis information for solving the to-be-solved question step by step, knowledge point information corresponding to the to-be-solved question, and answer information of the to-be-solved question.
7. The method of claim 1, wherein, The explanation video data is video data including question introduction, knowledge point introduction, question explanation, and ending.
8. A method of solving a problem, characterized by, includes: obtaining a target to-be-solved question; inputting the target to-be-solved question into a trained question solving model to obtain target explanation information; The target explanation information includes target video data for generating an explanation video for solving the target to-be-solved question using a line segment graph, and the target video data includes code for the explanation video for solving the target to-be-solved question using a line segment graph; The trained question solving model is obtained by training a question solving model using training data; the training data is generated based on explanation information of a to-be-solved question that can be solved using a line segment graph and the to-be-solved question; and the trained question solving model is obtained by training the method of any one of claims 1 to 7; rendering the target video data using a video rendering tool to obtain a target explanation video for the target to-be-solved question.
9. The method of claim 8, wherein, The target explanation video includes images and voice synchronized with the images; the target explanation video includes video content for step-by-step explanation of the solving process; a part of a video display page of the target explanation video is used to display board content, and another part is used to display line segment graph content.
10. An apparatus for generating a problem-solving model, the apparatus comprising: includes: a question obtaining module configured to obtain a to-be-solved question; The to-be-solved question is a question that can be solved using a line segment graph; a determining module configured to determine explanation information of the to-be-solved question; The explanation information includes explanation video data for the to-be-solved question, and the explanation video data for the to-be-solved question is used to generate code for an explanation video of the to-be-solved question; a data generating module configured to generate training data based on the to-be-solved question and the explanation information; a model training module configured to train a question solving model using the training data to obtain a trained question solving model; The post-training problem solving model is configured to generate an explanation video data for solving the problem by using the line segment diagram, the explanation video data for solving the problem by using the line segment diagram comprising a code of the explanation video for solving the problem by using the line segment diagram; The device further comprises: An explanation video generation module configured to generate an explanation video for the problem to be solved based on the explanation video data for the problem to be solved contained in the explanation information; A judgment module configured to acquire a last frame image of the explanation video and determine whether at least one of the following conditions exists in the last frame image: There is no correct answer in the last frame image; There is no line segment diagram in the last frame image; The line segment diagram displayed in the last frame image is incomplete; There is overlap in the content displayed in the last frame image; The line segment diagram displayed in the last frame image is inconsistent with the explanation text information contained in the explanation information; The quantity relationship represented by the line segment diagram displayed in the last frame image is inconsistent with the quantity relationship contained in the problem to be solved; The position of the diagram parameter contained in the line segment diagram displayed in the last frame image does not conform to a preset position; the diagram parameter comprises at least one of a number and a parenthesis; If none of the at least one condition exists, it is determined that the explanation video meets the problem solving requirement; The data generation module is specifically configured to: If the explanation video meets the problem solving requirement, the problem to be solved and the corresponding explanation information are taken as training data.
11. An apparatus for solving a problem, characterized by Comprise: A target example problem acquisition module configured to acquire a target problem to be solved; An input module configured to input the target problem to be solved to a post-training problem solving model to obtain target explanation information; The target explanation information contains target video data for generating an explanation video for solving the target problem to be solved by using a line segment diagram, the target video data comprising a code of the explanation video for solving the target problem to be solved by using the line segment diagram; The post-training problem solving model is obtained by training a problem solving model using training data; the training data is generated based on explanation information and a problem to be solved that can be solved by using a line segment diagram; the post-training problem solving model is obtained by training the method of any one of claims 1 to 7; A rendering module configured to render the target video data using a video rendering tool to obtain a target explanation video for the target problem to be solved.
12. A computing device, comprising: Comprise: A memory and a processor; The memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions, which realize the steps of the method of any one of claims 1 to 9 when executed by the processor.
13. A computer-readable storage medium, characterized in that, The memory has stored computer executable instructions, which realize the steps of the method of any one of claims 1 to 9 when executed by the processor.
14. A computer program product, characterised in that, Comprise computer programs / instructions, which realize the steps of the method of any one of claims 1 to 9 when executed by the processor.
Citation Information
Patent Citations
Question explanation method and device, electronic equipment and storage medium
CN118506620A