Method and apparatus for generating a solution model for solving a planar geometry problem
By training a problem-solving model using a large language model and training data, videos explaining plane geometry problems are generated and rendered, solving the problem of time-consuming manual recording and achieving efficient automated video generation.
Patent Information
- Application Number
- CN202411403903.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-09
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-10-09
AI Technical Summary
Current technologies require manual recording of videos explaining plane geometry problems, which is time-consuming and labor-intensive, and lack efficient automated generation methods.
The solution information for the questions to be solved is obtained by using a large language model. The solution model is trained by training data, the explanatory video data is generated, and the target explanatory video is rendered by video rendering tools to avoid manual recording.
It reduces the cost of generating explanatory videos, improves generation efficiency, and achieves automated explanatory video generation.
Smart Images

Figure CN119252086B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of automation technology, and in particular to a method, apparatus, device, medium, and product for generating a solution model for a plane geometry problem. Background Technology
[0002] With the rapid development of artificial intelligence and natural language processing technologies, large language models and natural language processing can understand and generate natural language text, which can be used to assist teaching. They can generate detailed problem-solving steps and processes based on questions encountered during the teaching process. However, if explanations are to be given via video, manual recording of the videos is still required, which is time-consuming and labor-intensive.
[0003] Therefore, there is a need for a faster method to generate explanatory videos for these issues. Summary of the Invention
[0004] In view of the above, embodiments of this specification provide a method for generating a solution model for a plane geometry problem. One or more embodiments of this specification also relate to an apparatus for generating a solution model for a plane geometry problem, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a method for generating a solution model for a plane geometry problem is provided, comprising:
[0006] Obtain a number of unsolved problems; the unsolved problems are problems related to plane geometry.
[0007] Using a large language model, solution information for the several unanswered questions is obtained; the solution information includes video data used to generate explanation videos for the unanswered questions;
[0008] Training data is obtained based on the aforementioned questions to be answered and the corresponding answer information containing the video data.
[0009] The training data is used to train a problem-solving model to obtain a post-trained problem-solving model; the post-trained problem-solving model is used to generate explanatory video data, and the explanatory video data is used to generate explanatory videos explaining problems about plane geometry.
[0010] According to a second aspect of the embodiments of this specification, a method for solving a problem is provided, comprising:
[0011] Obtain the target unsolved problem related to plane geometry;
[0012] The target unsolved question is input into the trained problem-solving model to obtain target problem-solving information; the target problem-solving information includes target video data used to generate an explanation video for the target unsolved question; the trained problem-solving model is obtained by training the problem-solving model with training data; the training data is generated using the solution information of several unsolved questions obtained from the large model and the several unsolved questions.
[0013] The target video data is rendered using a video rendering tool to obtain a target explanation video for the target question to be answered.
[0014] According to a third aspect of the embodiments of this specification, an apparatus for generating a solution model for a plane geometry problem is provided, comprising:
[0015] The acquisition module is configured to acquire a number of unsolved problems; the unsolved problems are problems related to plane geometry.
[0016] The solution information acquisition module is configured to use a large language model to obtain solution information for the plurality of unanswered questions; the solution information includes video data for generating explanatory videos for the unanswered questions;
[0017] The training data acquisition module is configured to acquire training data based on the plurality of unanswered questions and the corresponding answer information containing the video data.
[0018] The model training module is configured to train a problem-solving model using the training data to obtain a trained problem-solving model; the trained problem-solving model is used to generate explanatory video data, and the explanatory video data is used to generate explanatory videos explaining problems about plane geometry.
[0019] According to a fourth aspect of the embodiments of this specification, an apparatus for answering questions is provided, comprising:
[0020] The target unsolved problem acquisition module is configured to acquire target unsolved problems related to plane geometry.
[0021] The solution module is configured to input the target unsolved question into the trained problem-solving model to obtain target solution information; the target solution information includes target video data for generating an explanation video of the target unsolved question; the trained problem-solving model is obtained by training the problem-solving model with training data; the training data is generated using the solution information of several unsolved questions obtained from the large model and the several unsolved questions.
[0022] The rendering module is configured to use a video rendering tool to render the target video data to obtain a target explanation video for the target question to be answered.
[0023] According to a fifth aspect of the embodiments of this specification, a computing device is provided, comprising:
[0024] Memory and processor;
[0025] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, they implement the steps of the method for generating a solution model for a plane geometry problem or the steps of the method for solving the problem.
[0026] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer-executable instructions, which, when executed by a processor, implement the steps of the method for generating a solution model for a plane geometry problem or the steps of the method for solving the problem.
[0027] According to a seventh aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the method for generating a solution model for a plane geometry problem or the steps of the method for solving the problem described above.
[0028] At least one embodiment provided in this specification can achieve the following beneficial effects: A large language model is used to process several unsolved problems concerning plane geometry to obtain solution information containing video data used to generate explanatory videos for the unsolved problems; training data is obtained based on the several unsolved problems and the solution information; and a problem-solving model is trained using the training data to obtain a trained problem-solving model. Thus, videos can be generated based on the video code obtained from the large language model, facilitating the training of the problem-solving model using training data containing videos. This avoids the need for manual recording of corresponding explanatory videos based on the unsolved problems to obtain training data, reducing the cost of generating training data and improving the training efficiency of the trained problem-solving model.
[0029] At least one other embodiment provided in this specification can achieve the following beneficial effects: By training a problem-solving model to solve a target problem about plane geometry, problem-solving information containing target video data is obtained. This problem-solving information containing target video data can then be rendered using a video rendering tool to obtain a target explanation video corresponding to the target problem. Therefore, explanation videos for geometry problems can be generated based on the trained problem-solving model, avoiding the need for manual recording of corresponding explanation videos, reducing the cost of generating explanation videos, and improving the efficiency of generating explanation videos for geometry problems. Attached Figure Description
[0030] Figure 1 This is a flowchart illustrating a method for generating a solution model for a plane geometry problem, provided in one embodiment of this specification.
[0031] Figure 2 This is a flowchart illustrating a method for solving a problem according to one embodiment of this specification;
[0032] Figure 3 This is a schematic diagram of the structure of a device for generating a solution model for a plane geometry problem, provided in one embodiment of this specification.
[0033] Figure 4 This is a schematic diagram of the structure of a device for solving problems provided in one embodiment of this specification;
[0034] Figure 5 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0035] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0036] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0037] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0038] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0039] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0040] Large Language Models (LLMs), also known as large-scale language models or big models, are a type of natural language processing technique based on deep learning. They can better understand natural language and generate high-quality text based on given context. Examples include models such as GPT-4, PaLM Galactica, and LaMA. Large language models typically contain hundreds of billions (or more) of training parameters, which are trained on a vast amount of text data. Through training, large language models can learn the grammatical structure of a language and the relationships between words, thus enabling them to perform a wide range of tasks.
[0041] GPT-4: GPT-4 is one of a series of generative pre-trained transform models (GPT) developed by OpenAI. It has been trained on a large amount of Internet text and is a large language model that can generate coherent sentences or paragraphs in a human style.
[0042] Semantic matching models are a natural language processing technique used to determine the semantic similarity between two or more text segments. This type of model is crucial for understanding and processing the complexity and diversity of human language.
[0043] Supervised Fine-Tuning (SFT): SFT is a machine learning strategy. SFT primarily involves further parameter tuning of a pre-trained model using labeled datasets containing correct answers or labels. During SFT, the parameters of some (or all) layers of the model are updated based on this new task, allowing the model to perform better for that specific task.
[0044] Plane geometry: A branch of mathematics that studies plane figures and their properties, such as the properties of geometric figures on a plane, such as points, lines, angles, polygons, and circles, and the relationships between them, as well as the problem-solving process for figures such as triangles, circles, and polygons.
[0045] Manim, the Mathematical Animation Engine, is an open-source Python library for creating high-quality mathematical animations. Manim allows users to precisely control the details of animations programmatically, making it ideal for education and science popularization. Manim provides numerous inner classes such as Dot, VGroup, and Brace, which are helpful for solving plane geometry problems through drawing.
[0046] Few-Shot Learning: Few-shot learning is a term in machine learning that refers to the ability to train a model with only a small amount of labeled data. This method is particularly suitable for scenarios where obtaining large amounts of labeled data is costly or infeasible. It is one of the key methods for solving the problem of data scarcity, driving the application of machine learning techniques in a wider range of fields, especially when data access is limited.
[0047] Retrieval-Augmented Generation (RAG) technology combines information retrieval and natural language processing to improve the quality of a model's responses by integrating a large-scale external knowledge base with the generative model. RAG retrieves relevant information from an external knowledge source before generating a response, then uses this retrieved content as context to generate a more accurate and up-to-date answer. This provides a more reliable and targeted service, particularly effective in answering plane geometry questions.
[0048] This specification provides a method for generating a solution model for a plane geometry problem. It also relates to an apparatus for generating a solution model for a plane geometry problem, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0049] See Figure 1 , Figure 1 This diagram illustrates a flowchart of a method for generating a solution model for a plane geometry problem according to an embodiment of this specification. From a programming perspective, the executor of the process can be a program hosted on a server or training platform. From a hardware perspective, the executor of the process can be a server or training platform capable of training the model. Figure 1 As shown, the specific steps include:
[0050] Step 102: Obtain several unsolved problems; the unsolved problems are problems related to plane geometry.
[0051] The questions to be answered in the embodiments of this specification can be questions on plane geometry problems generated based on a large language model; they can also be questions on plane geometry problems obtained from a known question bank using a retrieval-augmented generation model (RAG); they can also be questions on plane geometry problems written based on expert experience; or they can be questions on plane geometry problems obtained through other models or processing methods. The method of obtaining the questions to be answered is not specifically limited here.
[0052] The plane geometry problems described in this specification can be problems concerning geometric figures such as lines, angles, polygons, and circles. Specifically, they can be problems concerning the position, size, shape, properties, and relationships between geometric figures. Some typical plane geometry problems include calculating lengths, areas, or angles; proving two lines are parallel or perpendicular; exploring the properties of triangles, quadrilaterals, or other polygons; circle-related problems such as chords, tangents, arc lengths, and sector areas; proportional problems of similar figures; applications of geometric transformations (translation, rotation, reflection); solving problems related to symmetry; solving problems using the Pythagorean theorem; and solving geometric problems using coordinate geometry methods.
[0053] In practical applications, in order to improve the accuracy of problem-solving model training, a large number of unsolved problems targeting plane geometry problems can be used to train the problem-solving model, resulting in a trained problem-solving model that can accurately solve plane geometry problems. For example, the number of unsolved problems can be greater than or equal to 1000, or it can be 2000, 500, etc., without specific limitations.
[0054] Step 104: Using a large language model, obtain the solution information for the several questions to be answered; the solution information includes video data used to generate explanation videos for the questions to be answered.
[0055] The large language model in this embodiment can learn the grammatical structure and word relationships of a language, thereby generating solution information to explain a given question. Video data can represent video code used to generate the explanation video, digitized information capable of generating the explanation video, or the explanation video itself. In addition to the video data used to generate the explanation video for the question, the solution information may also include at least one of the following: step-by-step analysis information of the question, knowledge point information corresponding to the question, and answer information corresponding to the question. It is understood that the solution information contains information related to the question and used to answer it.
[0056] In practical applications, the solution information for the unanswered questions can also be obtained based on expert experience. For example, a human can answer several unanswered questions to obtain the solution information corresponding to the several unanswered questions; or a combination of large language models and human intervention can be used to answer several unanswered questions to obtain the solution information corresponding to the several unanswered questions. Here, no specific limitation is made on the method of obtaining the solution information for several unanswered questions.
[0057] In practical applications, one question can be input into the large language model at a time. After the large language model outputs the answer information, another question can be input, and so on, until all questions are input into the large language model and the answer information is obtained. Alternatively, multiple questions can be input at once. Specifically, several questions can be input into the large language model at once, or multiple questions can be input in batches until all questions are input into the large language model and the answer information is obtained.
[0058] Step 106: Based on the several unanswered questions and the corresponding answer information containing the video data, training data is obtained.
[0059] In the embodiments of this specification, the training data may include the question to be solved and the corresponding solution information. Specifically, the question to be solved and the corresponding solution information may be concatenated to generate training data. For example, the solution information containing video data obtained in the above manner may be concatenated before the question to be solved, or the solution information containing video data obtained in the above manner may be concatenated after the question to be solved, resulting in several training data sets.
[0060] Step 108: Train the problem-solving model using the training data to obtain the trained problem-solving model; the trained problem-solving model is used to generate explanatory video data, and the explanatory video data is used to generate explanatory videos explaining problems about plane geometry.
[0061] In the embodiments of this specification, during the training process, the problem-solving model can continuously adjust its parameters through backpropagation to minimize the difference between the generated code and the labeled code, thereby improving the problem-solving ability of the model. When the difference between the generated code and the labeled code meets the requirements, the training of the problem-solving model can be considered complete.
[0062] In the embodiments of this specification, the problem-solving model can be a pre-trained model, such as the YunLLM 70B model, BERT model, RoBERTa model, XLNet model, Electra model, etc. These models have been pre-trained with a large amount of data. Supervised fine-tuning can be used to fine-tune the problem-solving model in the embodiments of this specification. Specifically, as described above, the problem-solving model can be trained using plane geometry problem-solving data to obtain a trained problem-solving model that is better adapted to the application scenarios of plane geometry problems.
[0063] In practical applications, training parameters such as learning rate, batch size, and number of training epochs can be configured for supervised fine-tuning. If the supervised fine-tuning training consists of multiple epochs, the training data can be grouped, such as by random grouping or equal grouping. The training data in each epoch after grouping can contain the same training data or different amounts of training data; there are no specific limitations here. For example, if there are 2000 unanswered questions that need to be trained for 10 epochs, they can be evenly divided into 200 unanswered questions per epoch; or 300 or 400 unanswered questions can be randomly assigned to each epoch.
[0064] In practical applications, the problem-solving model can be a model with a structural framework that has not been pre-trained, or it can be trained using a large amount of training data corresponding to plane geometry problems to obtain a trained problem-solving model. This trained model can then process plane geometry problems and obtain accurate solution information, making it well-suited for plane geometry problem-solving scenarios.
[0065] In practical applications, a validation set evaluation model can be used to evaluate the problem-solving model after completing a preset number of training rounds. The evaluation data can be randomly selected from the training data; alternatively, a portion of the training data obtained above can be used as test data to test the problem-solving model, while the remaining portion can be used as training data to train the model; alternatively, a preset number of data points can be obtained from a preset question bank as test data to test the model; or new questions can be obtained, and corresponding training data can be obtained using the above methods as test data to test the model. If the performance of the problem-solving model meets the requirements after testing, the trained model can be obtained; if the performance does not meet the requirements after testing, the model can continue to be trained.
[0066] In this embodiment, a large language model is used to process several unsolved problems related to plane geometry, resulting in solution information including video data used to generate explanatory videos for the unsolved problems. Based on the unsolved problems and the solution information, training data is obtained. This training data is then used to train a problem-solving model, resulting in a trained problem-solving model. The training of the problem-solving model in this application requires video training. Typically, those skilled in the art would collect videos that can be used as training data and then use these videos to train the problem-solving model; it is less likely to consider using other information formats to generate training data. While a large language model can process data to obtain text-formatted information, it cannot directly generate videos. Therefore, those skilled in the art would find it even less likely to consider using a large language model to generate training data for training the problem-solving model. However, in this application, a large language model can be used to process the unsolved problems to generate text-formatted information. This text-formatted information can include explanatory video data for the unsolved problems. The training data containing the explanatory video data is then used to train the problem-solving model, eliminating the need for manual recording of corresponding solution videos based on the unsolved problems, reducing the cost of generating training data, and improving the training efficiency of the trained problem-solving model.
[0067] based on Figure 1 In addition to the method described herein, this specification also provides some specific implementation schemes of the method, which will be described below.
[0068] To ensure that the large language model can accurately output solution information for the questions to be solved and to improve the accuracy of the training data, the embodiments of this specification may also provide relevant reference examples to the large language model so that it can output more accurate solution information. Optionally, the method described in the embodiments of this specification may further include:
[0069] For any one of the plurality of unanswered questions, a reference example question with the same characteristics as the unanswered question is determined; the characteristics include at least one of the following: the shape characteristics of the plane geometric image involved in the question, the number of questions contained in the question, the number of corresponding solution results of the question, and the number of drawings required for the question; the reference example question includes the example question and video data used to generate the explanation video of the example question.
[0070] The process of using a large language model to obtain solution information for the several unsolved questions may specifically include:
[0071] The reference example questions are provided to the large language model so that the large language model can obtain the solution information corresponding to any of the unsolved questions.
[0072] In this embodiment of the specification, for any unsolved question, a preset number of reference examples with the same characteristics as the unsolved question can be obtained. This allows the large language model to refer to the questions and corresponding solution information of the preset number of reference examples to solve any unsolved question, thereby improving the accuracy of the solution information and ultimately improving the problem-solving accuracy of the trained problem-solving model. The preset number can be set according to actual needs, such as 3, 4, 5, etc., and is not specifically limited here.
[0073] The features in the embodiments of this specification can represent the question type of an unsolved problem concerning plane geometry. Specific plane geometry types can be classified from different perspectives, such as from the perspective of shapes, including circles, rectangles, polygons, triangles, etc. From the perspective of question type, they can include single questions, multiple questions, single cases, multiple cases, etc. From the perspective of the number of diagrams, they can include one diagram, multiple diagrams, etc.
[0074] The features in the embodiments of this specification may also include the question type corresponding to the problem, such as side length problems, area problems, angle problems, etc., and may also include other features, such as the knowledge points tested in the problem, the number of solution methods corresponding to the problem, etc., which will not be listed here. The shape feature of the plane geometric image involved in the problem can indicate that the problem contains at least one shape such as a circle, quadrilateral, triangle, polygon, sector, etc. The feature of the number of questions contained in the problem can indicate that there are one, two, three, etc., questions that need to be solved. The feature of the number of solution results corresponding to the problem can indicate that there are one, two, three, etc., answers to the problem. The feature of the number of drawings required for the problem can indicate that there are one, two, three, etc., figures that need to be drawn in the process of solving the problem.
[0075] Reference examples may also include plane geometry problems, knowledge points, solution steps, corresponding diagrams for the solution steps, answers, and video codes.
[0076] The following is an example of a reference problem:
[0077]
topic
[0078] A 15-meter-long steel bar is used to form two equilateral triangles with a side length ratio of 3:2. The side lengths of the larger and smaller equilateral triangles are ( ) meters and ( ) meters respectively. (Fill in the larger one first, then the smaller one.)
[0079] [Key Knowledge Points]
[0080] This question mainly tests the properties of equilateral triangles and the application of proportions.
[0081] [Step-by-step analysis]
[0082] The first step is to draw two triangles.
[0083] The second step is to divide the total length of 15 meters into $3 + 2 = 5$ (parts) according to the problem. Each part is $15 \div 5 = 3$ (meters).
[0084] Third, the large equilateral triangle occupies 3 parts, so the perimeter of the large triangle is $3 \times 3 = 9$ (meters).
[0085] Fourth step, the small equilateral triangle occupies 2 parts, so the perimeter of the small triangle is $2 \times 3 = 6$ (meters).
[0086] Step 5: The side length of the large triangle is 9 ÷ 3 = 3 meters, and the side length of the small triangle is 6 ÷ 3 = 2 meters.
[0087]
Answer
[0088] How many portions are there in total: $3 + 2 = $5 (portions)
[0089] Length of each portion: $15 \div 5 = 3$ (meters)
[0090] The perimeter of the large triangle is: 3 × 3 = 9 (meters).
[0091] The perimeter of the small triangle is: 2 × 3 = 6 (meters).
[0092] The side length of the large triangle is: $9 \div 3 = 3$ (meters)
[0093] The side length of the small triangle: $6 \div 3 = 2$ (meters)
[0094]
Code
[0095] from core.utils import *
[0096] class ProblemSolving(MyScene):
[0097] The question is: "A 15-meter-long steel bar is used to form two equilateral triangles with a side length ratio of 3:2. The side lengths of the larger and smaller equilateral triangles are ( ) meters and ( ) meters respectively. (Fill in the larger one first, then the smaller one.)"
[0098] def keypoint(self):
[0099] # Showcase knowledge points
[0100] audio_text = "This question mainly tests the properties of equilateral triangles and the application of proportions."
[0101] with self.voiceover(text=audio_text):
[0102] text_keypoint = "Knowledge Point: Properties of Equilateral Triangles and Applications of Proportions"
[0103] text_keypoint = self.create_text(text_keypoint)
[0104] self.play(Write(text_keypoint))
[0105] def solve(self):
[0106] scale = self.get_scale()
[0107] # First step, draw two triangles.
[0108] audio_text = "Let's draw a diagram to represent the two equilateral triangles formed by a 15-meter-long steel bar."
[0109] with self.voiceover(text=audio_text):
[0110] # Since we need to draw two shapes simultaneously, we need to reduce the scaling ratio by half again.
[0111] scale / = 2
[0112] # Draw a large equilateral triangle with a side length of 3
[0113] triangle1 = Triangle(radius=3 * math.sqrt(3) / 3 * scale, color=RED)
[0114] # The two shapes are arranged vertically, so the first shape is shifted upwards by 3.
[0115] triangle1.shift(UP * 3 * scale)
[0116] self.play(Create(triangle1))
[0117] # Draw a small equilateral triangle with a side length of 2
[0118] triangle2 = Triangle(radius=2 * math.sqrt(3) / 3 * scale, color=BLUE)
[0119] # The two shapes are arranged vertically, so the second shape is shifted downwards by 2.
[0120] triangle2.shift(DOWN * 2 * scale)
[0121] self.play(Create(triangle2))
[0122] # Second step, according to the problem, we know that the ratio of the side lengths of the two equilateral triangles is 3:2. So we can divide the total length of 15 meters into $3 + 2 = 5$ (parts), and each part is $15 \div 5 = 3$ (meters).
[0123] audio_text =“According to the question, we can divide the total length of 15 meters into 3 + 2 = 5 parts, and each part is 15 ÷ 5 = 3 meters.”
[0124] with self.voiceover(text=audio_text):
[0125] # Mark the large equilateral triangle as 3 parts
[0126] triangle1_text = self.create_text("3 copies").next_to(triangle1, direction=DOWN)
[0127] self.play(Write(triangle1_text))
[0128] # Mark the small equilateral triangle as 2 parts
[0129] triangle2_text = self.create_text("2 copies").next_to(triangle2, direction=DOWN)
[0130] self.play(Write(triangle2_text))
[0131] self.play_text_left(r"There are a total of $3 + 2 = 5$ (pieces)")
[0132] self.play_text_left(r"Length of each piece: $15 \div 5 = 3$ (meters)")
[0133] # Step 3: The large equilateral triangle occupies 3 parts, so the perimeter of the large triangle is $3 \times 3 = 9$ (meters).
[0134] The audio text reads: "The side length of the large equilateral triangle is 3 parts, so the perimeter of the large equilateral triangle is 3×3=9 meters."
[0135] with self.voiceover(text=audio_text):
[0136] self.play_text_left(r"The perimeter of the large triangle is: $3 times 3 = 9 (meters)")
[0137] # Step 4: The small equilateral triangle occupies 2 parts, so the perimeter of the small triangle is $2 \times 3 = 6$ (meters).
[0138] The audio text reads: "The side length of the small equilateral triangle is 2 parts, so the perimeter of the small equilateral triangle is 2 × 3 = 6 meters."
[0139] with self.voiceover(text=audio_text):
[0140] self.play_text_left(r"The perimeter of the small triangle is 2 times 3 = 6 meters")
[0141] # Step 5: The side length of the large triangle is $9 \div 3 = 3$ (meters), and the side length of the small triangle is $6 \div 3 = 2$ (meters).
[0142] The audio text reads: "The perimeter of the large equilateral triangle is 9 meters, so the side length of the large equilateral triangle is 9 ÷ 3 = 3 meters."
[0143] with self.voiceover(text=audio_text):
[0144] self.play_text_left(r"The side length of the large triangle is: $9 \div 3 = 3$ (meters)")
[0145] The audio text reads: "The perimeter of the small equilateral triangle is 6 meters, so the side length of the small equilateral triangle is 6 ÷ 3 = 2 meters."
[0146] with self.voiceover(text=audio_text):
[0147] self.play_text_left(r"Side length of the small triangle: $6 \div 3 = 2$ (meters)") .
[0148] As mentioned above, the reference examples can contain information such as the question, knowledge points, step-by-step analysis, answer, and code, which helps the large language model understand the specific problem-solving process and thus solve the problem to be solved according to the solution method of the reference examples.
[0149] In practical applications, few-shot learning within prompt engineering can be used to guide large models in generating problem-solving information, avoiding the inability to obtain a large number of reference examples. Specifically, prompt words can be used to generate solution information for the unsolved problem based on reference examples. Prompt words can be used to guide large models in generating specific response text or sentences; specifically, they can be text used to guide large models in generating solution information for the unsolved problem. Prompt words can guide large language models to better understand user needs, thus facilitating the generation of content that better meets user requirements. The prompt words and their format can be set according to actual needs and are not specifically limited here.
[0150] In practical applications, prompts can include system prompts, reference examples, and information about the questions to be solved. System prompts can contain descriptive information specific to the task at hand, which may include one or more of the following: role information of the large model, task instructions for the large model to complete, task requirements, and output data format information.
[0151] The following is an example of a system suggestion word:
[0152] If you were an AI-powered teacher, please use manim to write a piece of code to create a math video for the following question. Requirements are as follows:
[0153] Create a `ProblemSolving` class that inherits from `MyScene`, and write a `self.solve` method.
[0154] First, use `audio_text` to generate a conversational narration. Then, use `self.voiceover` to play the synthesized audio. Note that LaTeX cannot be used in `audio_text`; the formulas need to be converted to Chinese.
[0155] For plane geometry problems, graphical aids can be used for calculation.
[0156] * Call the built-in `Dot`, `VGroup`, `Brace`, etc. of `manim` to draw the graph, and add brackets using `Brace` on each `VGroup`.
[0157] *Call `self.create_text` to create a `Text` object; do not use the built-in `Text` object in `manim`.
[0158] *The solution text for the problem should be presented in conjunction with a diagram.
[0159] * Only output the [Question], [Knowledge Points], [Step-by-Step Explanation], [Answer], and [Code]. Do not output any other content.
[0160] As illustrated above, "*Create a `ProblemSolving` class, inheriting from `MyScene`, and write a `self.solve` method." indicates the code type information of the solution video code; "*First, use `audio_text` to generate a conversational explanation. Then use `self.voiceover` to play the synthesized audio." indicates the audio format information and playback format information in the solution video; "*For plane geometry problems, graphical aids in calculation can be used." indicates the content contained in the solution information used to solve the problem in the solution video; "*Call the built-in `Dot`, `VGroup`, `Brace`, etc. of `manim` to draw graphics, and add brackets using `Brace` on each `VGroup`." indicates the information of the software tools used to draw the graphical objects in the problem.
[0161] In practical applications, additional information such as reference examples and unanswered questions can be added to the system's prompts to generate new prompts. The large language model can then output the answer information for the unanswered questions based on the content of the prompts. The answer information can include the question, knowledge points, step-by-step explanations, the answer, and video code.
[0162] The sample problems can include a predetermined number of selected plane geometry problems, along with the corresponding knowledge points, step-by-step explanations, answers, and codes. The problem information for the unsolved problems can include the problem itself and the user's requirements, such as: "A 15-meter-long steel bar is used to form two equilateral triangles with a side length ratio of 3:2. The side lengths of the larger and smaller equilateral triangles are ( ) meters and ( ) meters respectively. (Fill in the larger one first, then the smaller one)." Please generate the [Problem], [Knowledge Points], [Step-by-Step Explanation], [Answer], and [Code] according to the above requirements.
[0163] To facilitate understanding of this solution, the embodiments of this specification also provide specific details regarding the determination of reference examples that share the same characteristics as any unsolved problem.
[0164] Optionally, in the embodiments of this specification, determining a reference example with the same characteristics as any one of the plurality of unsolved questions may specifically include:
[0165] Using a semantic matching model, a predetermined number of target reference questions are selected from the set of reference questions that match any of the questions to be solved.
[0166] In the embodiments of this specification, the semantic matching model can match each reference example in the reference example set with any question to be solved, obtain the corresponding similarity, randomly select a preset number of reference examples from those with similarity greater than a preset similarity threshold as target reference examples, or sort the reference examples from largest to smallest based on similarity, and select the preset number of reference examples at the top of the sort as target reference examples. The semantic matching model can be a model based on RAG technology for semantic recognition, such as the Word2Vec model, GloVe model, Sentence-BERT model, Siamese network model, dual-tower model, etc.
[0167] In practical applications, each reference example in the reference example set can be matched with any question to be solved using keyword matching to determine a preset number of target reference examples with high matching degree; alternatively, predefined rules can be used to select reference examples with high matching degree as target reference examples; other methods can also be used to select target reference examples, such as n-gram models, deep learning methods, etc. Here, no specific limitation is made on the method of obtaining target reference examples.
[0168] The set of reference examples can be a preset number of plane geometry problems and their corresponding answer information selected through a preset method. Specifically, it can be typical elementary school plane geometry problems, typical middle school plane geometry problems, etc. The preset method can be manual selection, expert experience, model selection, etc., and the preset number can be 10, 20, 25, 30, etc.
[0169] Optionally, in the embodiments of this specification, training data is obtained based on the plurality of unsolved questions and the corresponding solution information containing the video data, which may specifically include:
[0170] The several unanswered questions are classified according to their characteristics to obtain multiple question sets; each question set contains several questions of the same category; the question characteristics include at least one of the following: the shape characteristics of the plane geometric image involved in the question, the number of questions contained in the question, the number of corresponding solution results for the question, and the number of drawings required for the question;
[0171] Select at least one question to be answered from each of the multiple question sets;
[0172] The selected questions to be answered and the corresponding answer information containing the video data are used as the training data.
[0173] In the embodiments described in this specification, at least one problem to be solved and its corresponding solution information can be selected from the problem sets of each category as training data. This allows the training data to include problems of various types, enabling the problem-solving model to complete training based on training data of different types. This improves the accuracy of the trained problem-solving model in solving various types of plane geometry problems and avoids the problem-solving model being limited to solving only a single type of plane geometry problem. The training data may or may not include reference examples.
[0174] In the embodiments of this specification, classification models, clustering models, and other models can be used to classify several unanswered questions, thereby dividing unanswered questions with the same characteristics into one category and obtaining multiple question sets.
[0175] In practical applications, the questions to be answered may have one or more different characteristics; the questions contained in one set of questions and those contained in another set of questions may overlap or may be completely different, without specific limitations here.
[0176] In this embodiment, at least one unanswered question is selected from each of the multiple question sets. Specifically, each major category of question features may have multiple subcategories. For example, shape features may include multiple subcategories such as circle, square, polygon, and triangle. At least one unanswered question can be selected from each subcategory. If the selected unanswered question is present in other major categories, it can be removed, and at least one unanswered question can be selected from the subcategories of the next major category. Alternatively, at least one unanswered question can be selected from a major category. If the selected unanswered question is present in other major categories, it can be removed, and at least one unanswered question can be selected from the next major category. This method allows for the selection of unanswered questions and corresponding answer information as training data, avoiding duplicate unanswered questions in the training data. This ensures that the training data contains a more comprehensive range of planar geometric types, thereby improving the accuracy of model training.
[0177] Of course, the repetition of training data can also be disregarded, and a certain number of questions can be randomly selected from various question sets as training data. No specific limitations are made here. In the embodiments of this specification, the training data can be data in a preset format obtained by processing answer information such as [question], [knowledge point], [step-by-step analysis], [answer], and [code]. As one implementation method, the preset format training data may include a role definition section, a question section, and an analysis section.
[0178] The role definition section includes information for defining the role played by the problem-solving model.
[0179] The question section includes information about the question to be answered.
[0180] The analysis section includes the knowledge point information, step-by-step analysis information, answer information, and code information of the solution video corresponding to the question to be answered.
[0181] In the embodiments of this specification, the training data in the preset format may include multiple identifiers with different meanings. The problem-solving model can learn the training data from different perspectives based on the identifiers. As one implementation, the identifiers may include a first identifier, such as [SYSTEM], which can be used to define the model's role and describe the model's current task. The first identifier can be located at the beginning of the role definition section. The identifiers may also include a second identifier, such as [USER], which can be a question posed by the user. The identifiers may also include a third identifier, such as [ASSISTANT], which can be the parsing part of the question. Of course, the above identifiers can also be words or text from other fields, and are not specifically limited here. The problem-solving model can understand its current task based on the content following the first identifier; therefore, the first identifier can be used to distinguish different business scenarios. Based on the first identifier, the problem-solving model can also clarify the model's identity in performing this task; for example, the problem-solving model is a math teacher. The second and third identifiers can be understood as a dialogue; the second identifier can be understood as the user asking a question, and the content following the second identifier is the content of the user's question. The third identifier can be understood as the answer content of the problem-solving model as the assistant to the user's question. In other words, the content after the third identifier is the assistant's answer content. By concatenating these three identifiers in order, they can be used as training data to train the problem-solving model.
[0182] To facilitate understanding of the preset format of training data, this specification provides an example of training data for a plane geometry problem, which can be described as follows:
[0183] [SYSTEM] You are a math teacher. Please use manim to write a piece of code to create an instructional video for the following questions.
[0184] [USER] A 15-meter-long steel bar is used to form two equilateral triangles with a side length ratio of 3:2. The side lengths of the larger and smaller equilateral triangles are ( ) meters and ( ) meters respectively. (Fill in the larger one first, then the smaller one).
[0185] [ASSISTANT] [Knowledge Point] xxx, [Step-by-Step Analysis] xxx, [Answer Information] xxx, and [Code Information] xxx.
[0186] To make the training data more accurate, the method described in the embodiments of this specification may optionally include:
[0187] Based on the video data corresponding to any one of the several unanswered questions, generate an explanation video for that one unanswered question;
[0188] Determine whether the instructional video meets the problem-solving requirements;
[0189] The training data obtained based on the plurality of unanswered questions and the corresponding answer information containing the video data specifically includes:
[0190] If the explanation video meets the problem-solving requirements, then any unsolved problem and its corresponding solution information will be used as training data.
[0191] In the embodiments of this specification, if the explanation video does not meet the problem-solving requirements, it indicates that the problem-solving information generated by the large language model is not accurate enough. In order to enable the trained model to solve problems more accurately, training data in which the explanation video does not meet the problem-solving requirements may not be used. Specifically, any unsolved question and its corresponding solution information can be removed from the training data; alternatively, the selection of the problem-solving information corresponding to the explanation video as training data can be determined based on whether the explanation video meets the problem-solving requirements. For example, if it meets the requirements, the question and solution information corresponding to the explanation video are selected as training data; if not, they are not selected.
[0192] The requirements for solving the problem can be related to the questions, knowledge points, step-by-step explanations, and answers in the video, such as the correctness of the answers. In practical applications, the video can be annotated manually or using a model to demonstrate whether it meets the requirements for solving the problem.
[0193] In practical applications, video code can be rendered using an audio / video rendering engine to generate explanatory videos. This engine can generate explanatory videos with audio narration based on the video code. Alternatively, an audio synthesis engine and a video rendering engine can be used separately to process the problem-solving code and generate explanatory videos. Specifically, a video rendering engine can be used to convert the video code into a visual initial video without audio, including the drawing of geometric figures and step-by-step demonstrations. An audio synthesis engine can then be used to generate synchronized audio narration based on the visual initial video without audio. This audio narration is then combined with the initial video to generate the explanatory video. This allows the explanatory video to not only contain visual video but also relevant audio narration, enhancing the vividness and intuitiveness of the problem-solving process.
[0194] To facilitate understanding of this solution, specific details regarding the above-mentioned problem-solving requirements are also provided in the embodiments of this specification.
[0195] Optionally, the determination of whether the explanation video meets the problem-solving requirements in the embodiments of this specification may specifically include:
[0196] Obtain the last frame of the explanatory video;
[0197] Determine whether at least one of the following conditions exists in the last frame image:
[0198] The geometric shape displayed in the last frame of the image is not the same type as the geometric shape contained in the question to be answered;
[0199] There are no geometric shapes in the last frame of the image;
[0200] The content displayed in the last frame of the image overlaps.
[0201] If one or more of the above situations exist in the explanation video in the embodiments of this specification, it can be determined that the explanation video does not meet the problem-solving requirements.
[0202] In practical applications, one can first determine if the last frame of the image contains geometric shapes. If it does, then determine if the type of geometric shape displayed in the last frame matches the type of geometric shape in the problem to be solved. If they match, then determine if there is any overlap between the content displayed in the last frame. Other judgment orders can also be followed, such as first determining if the last frame contains geometric shapes; if so, determining if there is any overlap between the content displayed in the last frame; if not, then determining if the type of geometric shape displayed in the last frame matches the type of geometric shape in the problem to be solved. Alternatively, these problem-solving requirements can be judged simultaneously; if any one of them is not met, the judgment of other problem-solving requirements can be terminated. A predetermined number of problem-solving requirements can also be judged.
[0203] In the embodiments of this specification, image recognition models, image segmentation models, image semantic segmentation models, and other image processing models can be used to determine whether the explanatory video meets the problem-solving requirements. Alternatively, expert experience can be used for judgment, and the settings can be configured according to actual needs.
[0204] In practical applications, it is also possible to determine whether the answer in the displayed answer image frame is correct; to determine whether the step-by-step analysis is logical based on each image frame; to determine whether the knowledge points in the image frame displaying knowledge points are the knowledge points to be tested in any of the questions to be answered; to determine whether the content of the image frame matches the audio content; to determine whether the diagram displayed in the answer video is consistent with the solution text information contained in the answer video; to determine whether the graphics displayed in the video are consistent with the image objects in the question; and to make judgments on other problem-solving requirements, which will not be listed here.
[0205] In the embodiments of this specification, one or more models, such as semantic analysis models and word vector models, can be used to determine whether the explanatory video meets the problem-solving requirements.
[0206] Optionally, the explanatory video described in the embodiments of this specification is a video that includes an introduction to the topic, an introduction to the knowledge points, an explanation of the topic, and a conclusion.
[0207] The solution videos in this embodiment may contain corresponding audio and visual information. For example, in the problem introduction section, the video displays the problem and is accompanied by the voice: "Hi, I'm Xiaoyuan AI Smart Teacher, let me explain this problem to you." In the knowledge point introduction section, the video screen displays: "Knowledge Point: XXX," accompanied by the voice: "This problem mainly tests XXX." The problem explanation can be divided into two parts: the left side displays the whiteboard, such as calculation formulas, and the right side draws diagrams, such as the graphic objects in the problem, accompanied by voice explaining the formulas and diagrams for each step. A concluding remark can be displayed on the screen and read aloud, such as "This problem has been explained; have you learned it?" This embodiment may optimize the audio and video rendering engine, making it easier to generate videos for elementary school plane geometry problems, thereby improving the quality of the generated videos.
[0208] Following the same approach, this manual also provides methods for solving problems using the trained model described above. The following, in conjunction with the appendix... Figure 2 This document uses the solution methods provided in this manual as an example to illustrate the application of these methods in plane geometry problems. Figure 2 This diagram illustrates a flowchart of a method for solving a problem according to an embodiment of this specification. From a programming perspective, the executing entity of the process can be a program within a server, a problem-solving terminal, or a problem-solving platform. The server or problem-solving terminal may contain or be able to call a problem-solving model for solving plane geometry problems. From a hardware perspective, the executing entity can be a server, a problem-solving terminal, or a problem-solving platform. This executing entity can be the same as or different from the executing entity in the above embodiments. Figure 2 As shown, the specific steps include:
[0209] Step 202: Obtain the target unsolved problem regarding plane geometry.
[0210] In the embodiments of this specification, the target unsolved question can be a target unsolved question input by the user through a user terminal. Specifically, it can be a text description of the target unsolved question; or it can be an uploaded image containing the target unsolved question, which can be uploaded locally or taken and uploaded directly by the user. Large language models or image recognition models can be used to identify the target unsolved question input by the user to determine the problem to be solved. The target unsolved question can be a question related to plane geometry.
[0211] Step 204: Input the target problem to be solved into the trained problem-solving model to obtain the target problem-solving information.
[0212] The target problem-solving information includes target video data used to generate explanation videos for the target unsolved problems; the trained problem-solving model is obtained by training the problem-solving model with training data; the training data is generated using the solution information of several unsolved problems obtained from the large model and the several unsolved problems.
[0213] The post-trained problem-solving model used in the embodiments of this specification can be generated in the manner described in the above embodiments, and will not be repeated here.
[0214] In practical applications, training data can be obtained by cleaning and formatting several unanswered questions and their corresponding unanswered information, which can ensure the consistency and standardization of the data. Specifically, this can include word segmentation of text data, noise removal, and formatting of code data.
[0215] Step 206: Render the target video data using a video rendering tool to obtain the target explanation video for the target question to be answered.
[0216] In the embodiments of this specification, the video rendering tool can generate a target explanation video with voice narration based on the target video code. Specifically, it can convert the target video code into a visual initial video, including the drawing of geometric shapes and step demonstrations; generate a voice narration synchronized with the video based on the initial video; and synthesize the voice narration with the initial video to obtain the target explanation video. The target explanation video can be displayed on the user's terminal for easy viewing.
[0217] The trained problem-solving model in the embodiments of this specification can generate video code for solving plane geometry problems. By using the trained problem-solving model to solve plane geometry problems, the video code for the solution can be obtained quickly, which improves the efficiency of generating video solutions for plane geometry problems. Furthermore, by providing corresponding explanatory videos to solve the problems, the vividness and intuitiveness of the problem-solving process can be enhanced, which helps users understand the problem.
[0218] based on Figure 2 In addition to the method described herein, this specification also provides some specific implementation schemes of the method, which will be described below.
[0219] Optionally, the target explanation video described in the embodiments of this specification may include images and audio synchronized with the images; the target explanation video includes video content explaining the problem-solving process step by step, and a portion of the video display page of the target explanation video is used to display blackboard content, while another portion is used to display drawing content.
[0220] In practical applications, the video can be divided into two sections: the left side displays the written content on the whiteboard, and the right side displays the diagram; or vice versa. Alternatively, the video can be divided into two sections: the top section displays the written content on the whiteboard, and the bottom section displays the diagram; or vice versa. The written content can include formulas for calculations and annotations for explanation; the diagrams can include graphs related to the problem, auxiliary lines added during the solution process, and numbers labeled on the graphs. Other information can also be included in the written content and diagrams, depending on the specific problem and explanation requirements; no specific limitations are imposed here.
[0221] Corresponding to the above-described method embodiment for generating solution models for plane geometry problems, this specification also provides an embodiment of an apparatus for generating solution models for plane geometry problems. Figure 3 This specification illustrates a schematic diagram of an apparatus for generating a solution model for a plane geometry problem, according to one embodiment of the present specification. Figure 3 As shown, the device includes:
[0222] The acquisition module 302 is configured to acquire a number of unsolved problems; the unsolved problems are problems related to plane geometry.
[0223] The solution information obtaining module 304 is configured to use a large language model to obtain solution information for the plurality of unanswered questions; the solution information includes video data for generating explanatory videos for the unanswered questions;
[0224] The training data acquisition module 306 is configured to obtain training data based on the plurality of unanswered questions and the corresponding answer information containing the video data.
[0225] The model training module 308 is configured to train a problem-solving model using the training data to obtain a trained problem-solving model; the trained problem-solving model is used to generate explanatory video data, and the explanatory video data is used to generate explanatory videos explaining problems about plane geometry.
[0226] based on Figure 3 The embodiments of this specification also provide some specific implementation schemes of the method, which are described below.
[0227] Optionally, the device may further include a reference example determination module, which can be configured to:
[0228] For any one of the plurality of unanswered questions, a reference example question with the same characteristics as the unanswered question is determined; the characteristics include at least one of the following: the shape characteristics of the plane geometric image involved in the question, the number of questions contained in the question, the number of corresponding solution results of the question, and the number of drawings required for the question; the reference example question includes the example question and video data used to generate the explanation video of the example question.
[0229] The module for obtaining the solution information can be configured as follows:
[0230] The reference example questions are provided to the large language model so that the large language model can obtain the solution information corresponding to any of the unsolved questions.
[0231] Optionally, the device may further include a reference example determination module, which can be configured to:
[0232] Using a semantic matching model, a predetermined number of target reference questions are selected from the set of reference questions that match any of the questions to be solved.
[0233] Optionally, the training data acquisition module can be configured as follows:
[0234] The several unanswered questions are classified according to their characteristics to obtain multiple question sets; each question set contains several questions of the same category; the question characteristics include at least one of the following: the shape characteristics of the plane geometric image involved in the question, the number of questions contained in the question, the number of corresponding solution results for the question, and the number of drawings required for the question;
[0235] Select at least one question to be answered from each of the multiple question sets;
[0236] The selected questions to be answered and the corresponding answer information containing the video data are used as the training data.
[0237] Optionally, the device may further include a problem-solving requirement judgment module, which can be configured as follows:
[0238] Based on the video data corresponding to any one of the several unanswered questions, generate an explanation video for that one unanswered question;
[0239] Determine whether the instructional video meets the problem-solving requirements;
[0240] The training data acquisition module can be configured as follows:
[0241] If the explanation video meets the problem-solving requirements, then any unsolved problem and its corresponding solution information will be used as training data.
[0242] Optionally, the device can be configured to:
[0243] Obtain the last frame of the explanatory video;
[0244] Determine whether at least one of the following conditions exists in the last frame image:
[0245] The geometric shape displayed in the last frame of the image is not the same type as the geometric shape contained in the question to be answered;
[0246] There are no geometric shapes in the last frame of the image;
[0247] The content displayed in the last frame of the image overlaps.
[0248] Optionally, the solution information may also include at least one of the following: step-by-step analysis information for solving the question in steps, knowledge point information corresponding to the question in steps, and answer information for the question in steps.
[0249] Optionally, the instructional video may include an introduction to the question, an introduction to the knowledge points, an explanation of the question, and a conclusion.
[0250] The above is a schematic scheme of an apparatus for generating a solution model for a plane geometry problem according to this embodiment. It should be noted that the technical solution of this apparatus for generating a solution model for a plane geometry problem belongs to the same concept as the technical solution of the method for generating a solution model for a plane geometry problem described above. Details not described in detail in the technical solution of the apparatus for generating a solution model for a plane geometry problem can be found in the description of the technical solution of the method for generating a solution model for a plane geometry problem described above.
[0251] Corresponding to the above-described method embodiments for solving problems, this specification also provides an embodiment of an apparatus for solving problems. Figure 4 A schematic diagram of a device for solving problems according to one embodiment of this specification is shown. Figure 4 As shown, the device includes:
[0252] The target unsolved problem acquisition module 402 is configured to acquire target unsolved problems related to plane geometry problems.
[0253] The solution module 404 is configured to input the target unsolved question into the trained problem-solving model to obtain target solution information; the target solution information includes target video data used to generate an explanation video for the target unsolved question; the trained problem-solving model is obtained by training the problem-solving model with training data; the training data is generated using the solution information of several unsolved questions obtained from the large model and the several unsolved questions.
[0254] The rendering module 406 is configured to render the target video data using a video rendering tool to obtain a target explanation video for the target question to be answered.
[0255] based on Figure 4 The embodiments of this specification also provide some specific implementation schemes of the method, which are described below.
[0256] Optionally, the target explanation video includes images and audio synchronized with the images; the target explanation video includes video content explaining the problem-solving process step by step, and a portion of the video display page of the target explanation video is used to display blackboard content, while another portion is used to display drawing content.
[0257] The above is an illustrative scheme of an apparatus for answering questions according to this embodiment. It should be noted that the technical solution of this apparatus for answering questions belongs to the same concept as the technical solution of the method for answering questions described above. For details not described in detail in the technical solution of the apparatus for answering questions, please refer to the description of the technical solution of the method for answering questions described above.
[0258] Figure 5 A structural block diagram of a computing device according to one embodiment of this specification is shown. The components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and a database 550 is used to store data.
[0259] The computing device 500 also includes an access device 540, which enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 540 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0260] In one embodiment of this specification, the above-described components of the computing device 500 and Figure 5 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 5 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0261] The computing device 500 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 500 can also be a mobile or stationary server.
[0262] The processor 520 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the method for generating a solution model for a plane geometry problem or the steps of the method for solving the problem.
[0263] The above is a schematic scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device belongs to the same concept as the above-described method for generating a solution model for a plane geometry problem or a method for solving the problem. For details not described in detail in the technical solution of the computing device, please refer to the description of the above-described method for generating a solution model for a plane geometry problem or a method for solving the problem.
[0264] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the method for generating a solution model for a plane geometry problem or the steps of the method for solving the problem.
[0265] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solution of the method for generating a solution model for a plane geometry problem or the method for solving the problem described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the method for generating a solution model for a plane geometry problem or the method for solving the problem described above.
[0266] An embodiment of this specification also provides a computer program product, wherein when the computer program is executed in a computer, the computer is instructed to perform the steps of the method for generating a solution model for a plane geometry problem or the steps of the method for solving the problem.
[0267] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product belongs to the same concept as the above-described method for generating a solution model for a plane geometry problem or a method for solving the problem. For details not described in detail in the technical solution of the computer program, please refer to the description of the above-described method for generating a solution model for a plane geometry problem or a method for solving the problem.
[0268] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0269] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0270] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0271] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0272] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A method for generating a solution model for a plane geometry problem, characterized in that, include: Obtain a number of questions to be answered; The problem to be solved is a problem related to plane geometry. Using a large language model, solution information for the several unanswered questions is obtained; the solution information includes video data used to generate explanation videos for the unanswered questions; Training data is obtained based on the aforementioned questions to be answered and the corresponding answer information containing the video data. The problem-solving model is trained using the training data to obtain the trained problem-solving model; The trained problem-solving model is used to generate explanatory video data, and the explanatory video data is used to generate explanatory videos explaining problems about plane geometry. The method further includes: the training data is obtained when the explanation video corresponding to the question to be solved meets the problem-solving requirements; determining whether the explanation video meets the problem-solving requirements specifically includes: determining whether the last frame of the explanation video has at least one of the following conditions: the geometric shape displayed in the last frame is inconsistent with the type of geometric shape contained in the question to be solved; there is no geometric shape in the last frame; or the content displayed in the last frame overlaps.
2. The method according to claim 1, characterized in that, The method further includes: For any one of the plurality of unanswered questions, a reference example question with the same characteristics as the unanswered question is determined; the characteristics include at least one of the following: the shape characteristics of the plane geometric image involved in the question, the number of questions contained in the question, the number of corresponding solution results of the question, and the number of drawings required for the question; the reference example question includes the example question and video data used to generate the explanation video of the example question. The process of using a large language model to obtain solution information for the several unsolved questions specifically includes: The reference example questions are provided to the large language model so that the large language model can obtain the solution information corresponding to any of the unsolved questions.
3. The method according to claim 2, characterized in that, For any one of the plurality of unsolved questions, determining a reference example question that has the same characteristics as that unsolved question specifically includes: Using a semantic matching model, a predetermined number of target reference questions are selected from the set of reference questions that match any of the questions to be solved.
4. The method according to claim 1, characterized in that, Based on the aforementioned questions to be answered and the corresponding answer information containing the video data, training data is obtained, specifically including: The several unanswered questions are classified according to their characteristics to obtain multiple question sets; each question set contains several questions of the same category; the question characteristics include at least one of the following: the shape characteristics of the plane geometric image involved in the question, the number of questions contained in the question, the number of corresponding solution results for the question, and the number of drawings required for the question; Select at least one question to be answered from each of the multiple question sets; The selected questions to be answered and the corresponding answer information containing the video data are used as the training data.
5. The method according to claim 1, characterized in that, The method further includes: Based on the video data corresponding to any one of the several unanswered questions, generate an explanation video for that one unanswered question; Determine whether the instructional video meets the problem-solving requirements; The training data obtained based on the plurality of unanswered questions and the corresponding answer information containing the video data specifically includes: If the explanation video meets the problem-solving requirements, then any unsolved problem and its corresponding solution information will be used as training data.
6. The method according to claim 1, characterized in that, The solution information also includes at least one of the following: step-by-step analysis information for solving the question in steps, knowledge point information corresponding to the question in steps, and answer information for the question in steps.
7. The method according to claim 1, characterized in that, The instructional video includes an introduction to the question, an introduction to the knowledge points, an explanation of the question, and a conclusion.
8. A method for solving problems, characterized in that, include: Obtain the target unsolved problem related to plane geometry; The target problem to be solved is input into the trained problem-solving model to obtain the target problem-solving information; The target problem-solving information includes target video data used to generate an explanation video for the target problem to be solved; The post-trained problem-solving model is obtained by training the problem-solving model with training data; the training data is generated using the solution information of several unsolved problems obtained from the large model and the several unsolved problems themselves; the training data is obtained when the explanation video corresponding to the unsolved problem meets the problem-solving requirements; determining whether the explanation video meets the problem-solving requirements specifically includes: determining whether the last frame of the explanation video has at least one of the following conditions: the geometric shape displayed in the last frame is inconsistent with the type of geometric shape contained in the unsolved problem; there is no geometric shape in the last frame; the content displayed in the last frame overlaps; The target video data is rendered using a video rendering tool to obtain a target explanation video for the target question to be answered.
9. The method according to claim 8, characterized in that, The target explanation video includes images and audio synchronized with the images; the target explanation video includes video content explaining the problem-solving process step by step, and a part of the video display page of the target explanation video is used to display the blackboard content, and another part is used to display the drawing content.
10. An apparatus for generating a solution model for a plane geometry problem, characterized in that, include: The retrieval module is configured to retrieve a number of questions to be answered. The problem to be solved is a problem related to plane geometry. The solution information acquisition module is configured to use a large language model to obtain solution information for the plurality of unanswered questions; the solution information includes video data for generating explanatory videos for the unanswered questions; The training data acquisition module is configured to acquire training data based on the plurality of unanswered questions and the corresponding answer information containing the video data. The model training module is configured to train a problem-solving model using the training data to obtain a trained problem-solving model. The trained problem-solving model is used to generate explanatory video data, and the explanatory video data is used to generate explanatory videos explaining problems about plane geometry. The device further includes: the training data is obtained when the explanation video corresponding to the question to be solved meets the problem-solving requirements; determining whether the explanation video meets the problem-solving requirements specifically includes: determining whether the last frame of the explanation video has at least one of the following conditions: the geometric shape displayed in the last frame is inconsistent with the type of geometric shape contained in the question to be solved; there is no geometric shape in the last frame; or the content displayed in the last frame overlaps.
11. A device for solving problems, characterized in that, include: The target unsolved problem acquisition module is configured to acquire target unsolved problems related to plane geometry. The solution module is configured to input the target problem to be solved into the trained problem-solving model to obtain the target solution information; The target problem-solving information includes target video data used to generate an explanation video for the target problem to be solved; The post-trained problem-solving model is obtained by training the problem-solving model with training data; the training data is generated using the solution information of several unsolved problems obtained from the large model and the several unsolved problems themselves; the training data is obtained when the explanation video corresponding to the unsolved problem meets the problem-solving requirements; determining whether the explanation video meets the problem-solving requirements specifically includes: determining whether the last frame of the explanation video has at least one of the following conditions: the geometric shape displayed in the last frame is inconsistent with the type of geometric shape contained in the unsolved problem; there is no geometric shape in the last frame; the content displayed in the last frame overlaps; The rendering module is configured to use a video rendering tool to render the target video data to obtain a target explanation video for the target question to be answered.
12. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, they implement the steps of the method for generating a solution model for a plane geometry problem as described in any one of claims 1 to 7, or the steps of the method for solving a problem as described in any one of claims 8 to 9.
13. A computer-readable storage medium, characterized in that, It stores computer-executable instructions, which, when executed by a processor, implement the steps of the method for generating a solution model for a plane geometry problem as described in any one of claims 1 to 7, or the steps of the method for solving a problem as described in any one of claims 8 to 9.
14. A computer program product, characterized in that, It includes a computer program / instruction that, when executed by a processor, implements the steps of the method for generating a solution model for a plane geometry problem as described in any one of claims 1 to 7, or the steps of the method for solving a problem as described in any one of claims 8 to 9.
Citation Information
Patent Citations
Corpus data construction method and device, equipment and storage medium
CN117973513A
Question explanation method and device, electronic equipment and storage medium
CN118506620A