Method and apparatus for generating a problem-solving model for solving coordinate system problems

By utilizing large language models to acquire and train problem-solving models to generate problem-solving videos, the problem of low efficiency in the manual production of existing teaching videos is solved, enabling the rapid generation of high-quality problem-solving videos and improving the efficiency of online teaching.

CN119293183BActive Publication Date: 2025-12-26BEIJING YUANLI WEILAI SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411403821.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-09
Publication Date
2025-12-26
Estimated Expiration
2044-10-09

AI Technical Summary

Technical Problem

Most existing instructional videos are manually produced, which is inefficient and cannot quickly generate problem-solving videos.

Method used

The solution information is obtained by using a large language model, and the problem-solving model is trained to generate video solutions for explaining coordinate system problems. The solution information of the problem to be solved is obtained by using a large language model, and the problem-solving model is trained using training data to generate video data for the solution video.

Benefits of technology

This improves the efficiency of generating problem-solving videos, reduces manual operations, and increases the production efficiency of problem-solving videos to meet the needs of online teaching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119293183B_ABST
    Figure CN119293183B_ABST
Patent Text Reader

Abstract

Embodiments of the present specification provide a method and device for generating a problem solving model for solving coordinate system problems, wherein the method can include: obtaining a plurality of pieces of answer information of to-be-answered questions obtained by using a large language model; the to-be-answered questions are questions related to coordinate systems; the answer information includes video data for generating a problem solving video of the to-be-answered questions; based on the plurality of to-be-answered questions and the corresponding answer information including the video data, training data is obtained; the problem solving model is trained by using the training data to obtain a trained problem solving model; the trained problem solving model can generate video data of a problem solving video for explaining the coordinate system problems. Embodiments of the present specification can quickly generate video data of a problem solving video of a question related to a coordinate system by using the trained problem solving model, and then the problem solving video can be quickly obtained, avoiding the case of manually writing video data, and improving the generation efficiency of the problem solving video.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present specification relate to the technical field of computer technology, and particularly relate to a method and device for generating a problem-solving model for solving coordinate system problems, equipment and medium. BACKGROUND

[0002] With the development of artificial intelligence and big data technology, using computer technology to assist teaching has become a trend. In particular, using the form of animation and video to explain the problems can greatly improve the learning interest and understanding ability of students. However, most of the existing teaching videos are artificially produced, and the work efficiency is low.

[0003] Therefore, it is necessary to provide a method for quickly generating a problem-solving video. SUMMARY

[0004] Therefore, the embodiments of the present specification provide a method for generating a problem-solving model for solving coordinate system problems. One or more embodiments of the present specification also relate to a device for generating a problem-solving model for solving coordinate system problems, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects in the prior art.

[0005] According to a first aspect of the embodiments of the present specification, a method for generating a problem-solving model for solving coordinate system problems is provided, comprising:

[0006] Obtaining solving information of a plurality of to-be-solved problems obtained by using a large language model; the to-be-solved problems are problems related to coordinate systems; the solving information includes video data for generating a problem-solving video of the to-be-solved problems;

[0007] Based on the plurality of to-be-solved problems and the corresponding solving information containing the video data, training data is obtained;

[0008] Training the problem-solving model using the training data to obtain a trained problem-solving model; the trained problem-solving model can generate video data for explaining a problem-solving video about coordinate system problems.

[0009] According to a second aspect of the embodiments of the present specification, a device for generating a problem-solving model for solving coordinate system problems is provided, comprising:

[0010] The solving information acquisition module is configured to obtain solving information of a plurality of to-be-solved problems obtained by using a large language model; the to-be-solved problems are problems related to coordinate systems; the solving information includes video data for generating a problem-solving video of the to-be-solved problems;

[0011] The training data obtaining module is configured to obtain training data based on the several to-be-answered questions and corresponding answer information containing the video data.

[0012] The model training module is configured to train a problem-solving model using the training data to obtain a trained problem-solving model, and the trained problem-solving model can generate video data of a problem-solving video for explaining a problem related to a coordinate system.

[0013] According to a third aspect of an embodiment of the present specification, a computing device is provided, comprising:

[0014] a memory and a processor;

[0015] The memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions, and the computer executable instructions, when executed by the processor, implement the steps of the above method.

[0016] According to a fourth aspect of an embodiment of the present specification, a computer readable storage medium is provided, which stores computer executable instructions, and the instructions, when executed by a processor, implement the steps of the above method.

[0017] According to a fifth aspect of an embodiment of the present specification, a computer product is provided, comprising a computer program / instruction, and the computer program / instruction, when executed by a processor, implements the steps of the above method.

[0018] An embodiment of the present specification realizes the following beneficial effects: obtaining answer information of several to-be-answered questions obtained by using a large language model, the to-be-answered questions being questions related to a coordinate system, and the answer information including video data of a problem-solving video for generating the to-be-answered questions; obtaining training data based on the several to-be-answered questions and corresponding answer information containing the video data; training a problem-solving model using the training data to obtain a trained problem-solving model. Thus, the trained problem-solving model can be used to generate video data of a problem-solving video for explaining a problem related to a coordinate system, and the problem-solving video can be obtained quickly, avoiding the case of manually writing video data, reducing manual operation, and thus improving the generation efficiency of the problem-solving video. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 is a flowchart of a method for generating a problem-solving model for answering a coordinate system problem provided by an embodiment of the present specification;

[0020] Figure 2 is a schematic diagram of an image frame of a problem-solving video provided by an embodiment of the present specification;

[0021] Figure 3 is another schematic diagram of an image frame of a problem-solving video provided by an embodiment of the present specification;

[0022] Figure 4 is a process diagram of a method for generating a problem-solving model for solving a coordinate system problem according to an embodiment of the present specification;

[0023] Figure 5 is a structural diagram of an apparatus for generating a problem-solving model for solving a coordinate system problem according to an embodiment of the present specification;

[0024] Figure 6 is a structural block diagram of a computing device according to an embodiment of the present specification. DETAILED DESCRIPTION

[0025] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present specification. However, the present specification can be practiced without the specific details, other than in the examples described herein. Those of ordinary skill in the art, and others, can readily ascertain combinations and sub-combinations of the various elements disclosed in the present specification, without departing from the scope of the present specification. Accordingly, the present specification is not limited to the examples described herein.

[0026] The terminology used in one or more embodiments of the present specification is for the purpose of describing particular embodiments only and is not intended to be limiting of one or more embodiments of the present specification. As used in one or more embodiments of the present specification and the accompanying claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in one or more embodiments of the present specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0027] It will be understood that, although the terms first, second, etc. can be used herein to describe various information, these terms are not intended to denote a temporal or chronological order. Rather, these terms are used only to distinguish one from another. For example, without departing from the scope of one or more embodiments of the present specification, first can be termed second, and similarly, second can be termed first. The term "if' as used herein can be interpreted as meaning "when" or "in response to determining" depending on the context.

[0028] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present specification are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0029] First, the terms used in this specification are explained.

[0030] Large Language Model (LLM): Also known as Large Language Model, Large Model, is a deep learning model that uses self-recurrent method to pre-train on large corpus in natural language processing field, which can understand and generate natural language text. Large language model can predict the next word or sentence by learning the statistical rules and semantic information of natural language text. With the expansion of input data set and parameter space, the ability of large language model will also expand accordingly. It can be applied to machine learning, machine translation, speech recognition, image processing and other fields.

[0031] Coordinate system related questions: Coordinate system related questions are a classic type of questions in primary school mathematics. This question refers to the use of coordinate system to solve mathematical problems, especially those related to position and graphics. For example, "Xiao Min sees Xiao Fang in the direction of north 60^o west, what direction does Xiao Fang see Xiao Min?"

[0032] Retrieval-Augmented Generation (RAG): A model that combines language models and information retrieval techniques. It generates answers or content by referencing information from external knowledge bases, with strong interpretability and customization capabilities, suitable for question and answer systems, document generation, intelligent assistants and other natural language processing tasks. The advantage of RAG is that it is versatile and can update knowledge instantly, thus providing more efficient and accurate information services.

[0033] GPT (Generative Pre-Trained ): A deep learning model based on Internet data for text generation. It is trained on a large amount of Internet text and is a large language model that can generate coherent sentences or paragraphs in human style.

[0034] Prompt: A command or instruction provided by the user when interacting with a large language model. It can be a question, keyword, context information, etc., used to indicate the action the large language model needs to perform or the output it needs to generate. Users can provide clear prompts to guide the large language model to generate responses that meet expectations, thereby improving the effectiveness and quality of interaction.

[0035] Prompt Template: A structured text used to generate prompts. Generating prompts based on prompt templates can help users interact with large language models, and large language models can better understand user intent and generate more accurate and user-demand-oriented responses.

[0036] Supervised Fine-Tuning (SFT): SFT is a machine learning strategy. SFT mainly involves further parameter adjustment of a pre-trained model using a labeled dataset containing correct answers or labels. During the SFT process, the parameters of some layers (or all layers) of the model will be updated based on this new task so that the model can better perform this specific task.

[0037] Manim (Mathematical Animation Engine): It is an open-source Python library for creating high-quality mathematical animations. Manim allows users to precisely control the details of animations through programming, making it ideal for education and popular science. Manim provides Dot, VGroup, Brace, and other internal classes that are helpful for solving queuing problems through drawing.

[0038] Few-Shot Learning: Few-shot learning is a term in the field of machine learning that refers to the ability to train a model with only a small amount of labeled data. This method is particularly suitable for scenarios where obtaining a large amount of labeled data is costly or impossible. It is one of the key methods to solve the problem of data scarcity, and it has promoted the application of machine learning technology in a wider range of fields, especially in situations where data acquisition is limited.

[0039] In this specification, a method for generating a problem-solving model for solving coordinate system problems, an apparatus for generating a problem-solving model for solving coordinate system problems, a computing device, a computer-readable storage medium, and a computer program product are provided, which are described in detail one by one in the following embodiments.

[0040] Reference Figure 1 , Figure 1 A flowchart of a method for generating a problem-solving model for solving coordinate system problems according to an embodiment of the present specification is shown. From a program perspective, the execution subject of the flow can be a program running on a server or a training platform. From a hardware perspective, the execution subject of the flow can be a server or a training platform capable of model training. The method can specifically include the following steps.

[0041] Step 102: Obtain the solution information of a plurality of to-be-solved problems obtained by using a large language model; the to-be-solved problems are problems related to coordinate systems; the solution information includes video data for generating a solution video of the to-be-solved problems.

[0042] In actual application, a plurality of to-be-solved questions can be acquired first, and then the solution information of the plurality of to-be-solved questions is acquired. The to-be-solved questions can be acquired in various ways, such as being written according to expert experience, being generated based on a large language model, being acquired from a known question library by using a retrieval-augmented generation (RAG) model, and the like, which are not limited here.

[0043] In the embodiments of the present specification, the to-be-solved question can be one or more questions related to a coordinate system in mathematical application.

[0044] The question related to the coordinate system can refer to a mathematical question solved by using the coordinate system. For example, “the auditorium is 30° east-south of the stadium, then what direction is the stadium from the auditorium?”, for this question, a first coordinate system can be established with the stadium as the origin. Taking the north as the upper, the south as the lower, the west as the left, and the east as the right of the coordinate system as an example, a line segment can be drawn to indicate the position of the auditorium 30° east-south of the stadium. Then a second coordinate system can be established with the auditorium as the origin, and based on the north as the upper, the south as the lower, the west as the left, and the east as the right of the coordinate system, it can be known that the stadium is in the west-north 30° direction from the auditorium.

[0045] In the embodiments of the present specification, the video data of the problem-solving video can be the code of the problem-solving video of the to-be-solved question, or other forms of digital information. The problem-solving video can be a video for solving the to-be-solved question.

[0046] In the embodiments of the present specification, the solution information of the plurality of to-be-solved questions is generated by using the large language model, and then the training data can be obtained based on the solution information generated by the large language model, without manually generating the training data, thereby improving the generation efficiency of the training data.

[0047] Step 104: obtaining training data based on the plurality of to-be-solved questions and the corresponding solution information containing the video data.

[0048] The training data of the embodiments of the present specification can include the to-be-solved question and the solution information. For example, the to-be-solved question and the solution information can be spliced to generate the training data. Specifically, the to-be-solved question can be placed in front of the solution information, or the to-be-solved question can be placed behind the solution information, and the like.

[0049] In the embodiments of the present specification, the training data can include a plurality of training data. For example, a plurality of to-be-solved questions can be acquired at one time, and a plurality of training data can be generated by using the acquired plurality of to-be-solved questions according to the above-mentioned method of generating the training data.

[0050] For example, the to-be-answered question can be acquired multiple times, and multiple training data can be generated according to the to-be-answered question acquired multiple times according to the above method of generating training data.

[0051] Step 106: training the problem-solving model by using the training data to obtain a trained problem-solving model; the trained problem-solving model can generate video data of a problem-solving video for explaining the problem related to the coordinate system.

[0052] The video data of the problem-solving video for explaining the problem related to the coordinate system generated by the trained problem-solving model can include a code of the problem-solving video for explaining the problem related to the coordinate system.

[0053] The code of the problem-solving video can be compiled to generate the problem-solving video. In actual application, a compiling software such as Manim can be used to compile the code of the problem-solving video output by the trained problem-solving model to generate the problem-solving video.

[0054] Generally, the common method for those skilled in the art is to collect videos that can be used as training data, and then use these videos to train the problem-solving model; it is not easy to think of using information in another format to generate training data. A large language model can process data to obtain information in text format, and cannot directly generate videos, so it is even more difficult for those skilled in the art to think of using a large language model to generate training data for training the problem-solving model. However, in the present application, a large language model can be used to process the to-be-answered question to generate information in text format, which can include video data of a problem-solving video for the to-be-answered question, and then use the training data containing the video data of the problem-solving video to train the problem-solving model, so that it is no longer necessary to manually create the corresponding problem-solving video based on the to-be-answered question, thereby reducing the cost of generating training data and improving the training efficiency of the trained problem-solving model.

[0055] On the other hand, since the trained problem-solving model is obtained by training the question related to the coordinate system, it is more accurate to generate the code of the problem-solving video for the question related to the coordinate system, thereby ensuring the quality of the problem-solving video for the question related to the coordinate system.

[0056] For example, when an online teaching institution or a book publishing institution needs to generate problem-solving videos for a large number of questions in the database, the trained problem-solving model can be directly used to output the video data of the problem-solving video for the question related to the coordinate system, and then the problem-solving video can be generated based on the video data, thereby avoiding the tedious process of manually writing scripts and manually making animations, and improving the efficiency of making problem-solving videos.

[0057] Also, when a user such as a teacher, a student, or a parent uses teaching software for online learning and needs to obtain a problem-solving video of a question related to a coordinate system, the teaching software can load or call the trained problem-solving model. After the user inputs the question related to the coordinate system in the teaching software, the teaching software can generate video data of the problem-solving video of the question by using the trained problem-solving model. The teaching software can also compile the video data by using a video compiling software, thereby obtaining and displaying the problem-solving video of the question, avoiding long waiting of the user, improving problem-solving efficiency, and thereby improving user experience.

[0058] Based on the method, Figure 1 The embodiments of the present specification also provide some specific embodiments of the method, which are described below.

[0059] In the embodiments of the present specification, the large language model used can be a model of the GPT series, such as GPT-3.5, GPT-4, and GPT-4o, etc. Of course, other series of models can also be used, such as the Tongyi Qianwen model, the Ant Hundredling large model, etc., which are not limited here.

[0060] In prompt engineering (Prompt Engineering), a prompt can refer to a text or a sentence used to guide a large language model to generate a specific response. The prompt can help the large language model better understand the user's needs, so that the large language model can generate content that better meets the user's needs.

[0061] In the embodiments of the present specification, the answer information of the question to be answered can be obtained by constructing a prompt of a large language model. The prompt of the embodiments of the present specification can be a text used to guide the large language model to generate the answer information of the question to be answered. The method can also include:

[0062] Obtain a plurality of reference example questions of different types related to the coordinate system; the different types include at least one of the direction type and the number pair type; the reference example question includes an example question and video data for generating a problem-solving video of the example question.

[0063] Based on the reference example question, generate a prompt word template containing the reference example question.

[0064] For any one of the plurality of questions to be answered, add the any one question to be answered to the prompt word template to obtain a prompt word for the any one question to be answered.

[0065] Input the prompt word into the large language model to obtain the answer information corresponding to the any one question to be answered.

[0066] In the embodiments of the present specification, the reference example question can be selected from a reference example question set. For example, a semantic matching model can be used to select a reference example question related to the coordinate system from the reference example question set. Alternatively, a retrieval enhancement generation model can be used to select a reference example question related to the coordinate system from the reference example question set. In addition, the reference example question can be selected from the reference example question set in a keyword matching manner or other matching manners, and the like.

[0067] It can be understood that the reference example question can also be written according to expert experience. Of course, the existing problem solving video related to the coordinate system can also be used to extract the code of the problem solving video by using reverse engineering, so as to obtain the question and the code of the problem solving video as the reference example question, and the like.

[0068] The embodiments of the present specification can select a preset number of reference example questions. For example, one reference example question, three reference example questions, ten reference example questions, and the like are selected, which are not limited herein.

[0069] In the embodiments of the present specification, a few-shot learning in a prompt can be used to guide a large language model to generate answer information, a large amount of answer information can be obtained, and a large amount of training data can be obtained.

[0070] In the embodiments of the present specification, the reference example question can include a question title and video data of a problem solving video used to generate the question title, wherein the video data of the problem solving video can be code of the problem solving video.

[0071] Further, in addition to the question title and the video data of the problem solving video, the reference example question can also include other content, so that the answer information output by the large language model also includes the above-mentioned other content, thereby assisting teaching.

[0072] Optionally, the reference example question further includes at least one of step-by-step analysis information used to step-by-step answer the question title, knowledge point information corresponding to the question title, and answer information corresponding to the question title.

[0073] And / or, the answer information obtained by using the large language model further includes at least one of step-by-step analysis information used to step-by-step answer the to-be-answered question, knowledge point information corresponding to the to-be-answered question, and answer information corresponding to the to-be-answered question.

[0074] The step-by-step analysis information can represent a distributed explanation for the question, and the distributed explanation can include a line segment graph corresponding to each step. The knowledge point information can represent a knowledge point examined by the question. The answer information can represent an answer to the question.

[0075] For ease of understanding, the embodiments of the present specification provide an example of a reference example, which can be as follows:

[0076]

Question

[0077] The school auditorium is 30° east-south of the stadium. Then the stadium is in the ( ) direction of the auditorium.

[0078] A. 60° east-south B. 30° west-north C. 60° west-north

[0079]

Knowledge point

[0080] This question mainly examines the understanding and application of direction angle.

[0081]

Step-by-step analysis

[0082] First, we first establish a coordinate system with the stadium as the observation point, and the stadium is at the center of the coordinate system.

[0083] Second, the auditorium is 30° east-south of the stadium. According to the above north-south, left-right, east, we draw a line segment to represent the position of the auditorium, which is 30° east-south of the stadium.

[0084] Third, finally, we establish a coordinate system with the auditorium as the center. According to the image and the nature of the relative position, we can know that the stadium is in the 30° west-north direction of the auditorium.

[0085]

Answer

[0086] B. 30° west-north

[0087]

Code

[0088] ...python

[0089] {{example_code}} ...

[0090] The embodiments of the present specification also provide another example of a reference example, which can be as follows:

[0091]

Question

[0092] On a square grid, there is a triangle with three vertices at A(2,2), B(5,2), and C(2,4). Please answer: (1) What kind of triangle is it? (2) If the side length of a square represents 10 centimeters, what is the area of this triangle in square centimeters?

[0093]

Knowledge point

[0094] This question mainly examines the identification and area calculation of triangles in the coordinate system.

[0095] [Step-by-step analysis]

[0096] The first step, according to the problem, is to establish a coordinate system and draw these three points in the coordinate system, with the positions of the three vertices being A(2,2), B(5,2) and C(2,4).

[0097] The second step is to observe the positions of these three points. We can see that the y-coordinates of A and B are the same, and the x-coordinates of A and C are the same. This means that AB is horizontal and AC is vertical. Therefore, triangle ABC is a right triangle with the right angle at point A.

[0098] The third step is to connect AB, BC, and AC on the coordinate system, and then calculate the lengths of AB and AC. The length of AB is 5-2=3 units, and the length of AC is 4-2=2 units.

[0099] Fourth step, according to the formula for calculating the area of ​​a triangle, the area is equal to half multiplied by the base multiplied by the height. We know that the area of ​​triangle ABC is 1 / 2 * AB * AC = 1 / 2 * 3 * 2 = 3 units.

[0100] Step 5: If the side length of one square represents 10 centimeters, then the area of ​​the triangle is 3 * 10. 2 = 300 square centimeters.

[0101]

Answer

[0102] (1) This is a right triangle.

[0103] (2) If the side length of a square represents 10 centimeters, the area of ​​the triangle is 300 square centimeters.

[0104] ...python

[0105] {{example_code}} ...

[0106] In this embodiment of the specification, the reference example problem may include one or more of the following: step-by-step analysis information, knowledge point information, and answer information. Similarly, the prompt words for the problem to be solved may include one or more of the aforementioned step-by-step analysis information, knowledge point information, and answer information. Guided by the prompt words, the solution information generated by the large language model will also include information corresponding to that included in the reference example problem. This satisfies the requirements for step-by-step explanation and overall summarization in the teaching process, thereby improving teaching quality.

[0107] In the embodiments of the present specification, the large language model is used to automatically generate information including step-by-step analysis information, knowledge point information, answer information, and code of problem solving video, thereby reducing the cumbersome process of manually writing the above content information, improving the efficiency of obtaining the answer information of the to-be-solved problem, improving the efficiency of obtaining the training data, and thereby improving the training efficiency of the problem solving model.

[0108] In the embodiments of the present specification, the prompt word can be generated based on the prompt word template.

[0109] The prompt word template is a structured text used to generate the prompt word. Generating the prompt word based on the prompt word template can help the large language model interact with the user, so that the large language model can better understand the user's intention and generate more accurate and user-demand-conforming replies.

[0110] In the embodiments of the present specification, the prompt word template can include a system prompt word. The system prompt word can be description information for the current task. The description information can specifically include one or more of role information of the large language model, task information indicating the large language model to complete, task requirement information, and output data format information.

[0111] The role information of the large language model can be the role information that the large language model needs to play for the current task. For example, the large language model needs to play the role of a math teacher, and the role information of the large language model can be "you are a math teacher". The task information indicating the large language model to complete can be the task information that the large language model needs to complete for the current task. The task requirement information can be the requirement information for the current task. The task requirement information can include various information of the video data of the problem solving video, such as at least one of code type information, voice format information, and voice playing format information.

[0112] In the embodiments of the present specification, the prompt word template can further include a first to-be-filled area and a second to-be-filled area. The first to-be-filled area can be used to fill the reference example, and the second to-be-filled area can be used to fill the to-be-solved problem.

[0113] The server or training platform can add the system prompt word to the prompt word template, add the reference example to the first to-be-filled area, and add the to-be-solved problem to the second to-be-filled area, and the prompt word generated by the prompt word template can include the system prompt word, the reference example, and the to-be-solved problem.

[0114] To ensure the accuracy of the obtained answer information, the embodiments of the present specification can also obtain the prompt word of the large prediction model based on different types of reference examples. In this way, for any to-be-answered question, the large language model can generate answer information more accurately based on the reference examples of the type corresponding to the any to-be-answered question. Optionally, the generating a prompt word template containing the reference example based on the reference example can specifically include:

[0115] Selecting a first number of first target reference examples belonging to the orientation type from the plurality of reference examples.

[0116] Generating a first prompt word template containing the first number of first target reference examples based on the first number of first target reference examples.

[0117] The adding the any to-be-answered question to the prompt word template to obtain a prompt word for the any to-be-answered question can specifically include:

[0118] If the any to-be-answered question is a question of the orientation type, adding the any to-be-answered question to the first prompt word template to obtain a prompt word for the any to-be-answered question.

[0119] In actual application, the reference example of the orientation type can be a reference example related to orientation. For example, the reference example mentioned in the foregoing is "The auditorium is 30° east-south of the stadium. What direction is the stadium from the auditorium?"

[0120] To facilitate understanding of the present solution, the embodiments of the present specification also provide specific content about the above-mentioned first prompt word template. The following is an example of a first prompt word template:

[0121] You are a math teacher,

[0122] Please write a code using manim to make a teaching video for the following question.

[0123] Requirements are as follows:

[0124] * Create a `ProblemSolving` class that inherits `MyScene` and writes a `self.solve` method.

[0125] * First, use `audio_text` to generate a spoken explanation. Then use `self.voiceover` to play the synthesized audio. Note that latex cannot be used in audio_text, and formulas need to be converted to Chinese.

[0126] * Call the `Axes` class built-in `manim` to create a coordinate system, call the `MyAngle`, `Line`, `Arrow`, etc. class to mark the angle relationship.

[0127] * Call `self.create_text` to create a `Text` object, do not use the built-in `Text` of `manim`.

[0128] * Use `self.play_text_left` to show the formula, and need to combine the text explanation, the text explanation should be as concise as possible.

[0129] * When the geographical direction appears in the question, call `calculate_direction_vector` to calculate the direction vector.

[0130] * Only output

question

knowledge point

step-by-step analysis

answer

code

[0131] Reference Example 1: ...

[0132] Reference Example 2: ...

[0133] Enter a question, and you need to generate

question

knowledge point

step-by-step analysis

answer

code

[0134] As in the above example, "You are a math teacher" can represent the role information of the large language model; "Please use manim to write a piece of code to make a teaching video for the following question" can represent the task information that the large language model needs to complete; "Use `self.play_text_left` to show the formula, and need to combine the text explanation, the text explanation should be as concise as possible" can represent the requirement information of the task. "* Create a `ProblemSolving` class that inherits `MyScene`, and write the `self.solve` method." can represent the code type information of the data; "First, use `audio_text` to generate a spoken explanation. Then use `self.voiceover` to play the synthesized audio." can represent the voice format information in the data, and the voice playback format information.

[0135] In the embodiment of the present specification, the first prompt word template can further include: instruction information indicating that the large language model establishes a coordinate system and marks the angle relationship. As in the above example, "* Call the `Axes` class built-in `manim` to create a coordinate system, call the `MyAngle`, `Line`, `Arrow`, etc. class to mark the angle relationship."

[0136] Among them, reference example 1 and reference example 2 can be reference examples of the orientation type.

[0137] Optionally, the generating a prompt word template containing the reference example based on the reference example can specifically include:

[0138] Selecting a second number of second target reference examples belonging to the number pair type from the plurality of reference examples.

[0139] Generating a second prompt word template containing the second number of second target reference examples based on the second number of second target reference examples.

[0140] The adding the any to-be-solved question to the prompt word template to obtain a prompt word for the any to-be-solved question can specifically include:

[0141] If the any to-be-solved question is a question of the number pair type, the any to-be-solved question is added to the second prompt word template to obtain a prompt word for the any to-be-solved question.

[0142] In actual application, the reference example of the number pair type can be a reference example of the relevant position relationship between numbers. For example, "On a grid map, there is a triangle, and the positions of the three vertices are A(2, 2), B(5, 2), and C(2, 4). If the side length of a grid represents 10 centimeters, what is the area of this triangle in square centimeters?"

[0143] In order to facilitate the understanding of the present scheme, the present specification also provides specific content about the second prompt word template. The following is an example of a second prompt word template:

[0144] You are a math teacher,

[0145] Please write a code using manim to make a teaching video for the following question.

[0146] Requirements are as follows:

[0147] * Create a `ProblemSolving` class that inherits `MyScene` and writes the `self.solve` method.

[0148] * First, use `audio_text` to generate a spoken explanation. Then use `self.voiceover` to play the synthesized audio. Note that latex cannot be used in audio_text, and formulas need to be converted to Chinese.

[0149] * Call the `manim` built-in graphical object `Axes` to create a coordinate system and add arrows, tick marks, and numbers to the coordinate system.

[0150] * Call `Dot` to create a dot to represent the position of the object, and use different colors to distinguish different objects in the figure.

[0151] * Call `self.create_text` to create a `Text` object, and do not use the `manim` built-in `Text`.

[0152] * Use `self.play_text_left` to show the formula, and combine it with the text explanation. The text explanation should be as concise as possible.

[0153] * Only output

Question

Knowledge Point

Step-by-step Analysis

Answer

Code

[0154] Reference Example 3: ...

[0155] Reference Example 4: ...

[0156] Enter a question, and generate

Question

Knowledge Point

Step-by-step Analysis

Answer

Code

[0157] Among them, Reference Example 3 and Reference Example 4 can be reference examples of number pairs.

[0158] The second prompt word template provided by the embodiments of the present specification and the content part of the first prompt word template provided are the same, the difference includes the use of different reference examples, and the second prompt word template can include instruction information indicating that the large language model establishes a coordinate system and marks coordinate system parameters, wherein the coordinate system parameters include at least one of the information of arrows, tick marks, numbers, shapes used to represent objects in the question, and colors. The instruction information indicating that the large language model establishes a coordinate system and marks the coordinate system parameters can be "* Call the `manim` built-in graphical object `Axes` to create a coordinate system and add arrows, tick marks, and numbers to the coordinate system." and "* Call `Dot` to create a dot to represent the position of the object, and use different colors to distinguish different objects in the figure."

[0159] In practical applications, if the to-be-solved question is a direction type question, the direction type question can be added to the first prompt word template to obtain the prompt word of the large language model. If the to-be-solved question is a number pair type question, the number pair type question can be added to the second prompt word template to obtain the prompt word of the large language model.

[0160] In the embodiments of the present specification, in order to ensure the quality of the training data, the answer information of the video data meeting the problem solving requirements and the corresponding to-be-solved problems can be selected to generate the training data. The method for generating the problem solving model for solving the coordinate system problem in the embodiments of the present specification can further include:

[0161] For any to-be-solved problem in the plurality of to-be-solved problems, the video data contained in the answer information corresponding to the any to-be-solved problem is compiled by using the video compiling software to obtain a problem solving video.

[0162] It is determined whether the problem solving video meets the problem solving requirements.

[0163] The training data is obtained based on the plurality of to-be-solved problems and the corresponding answer information containing the video data, and specifically can include:

[0164] If the problem solving video meets the problem solving requirements, the any to-be-solved problem and the answer information corresponding to the any to-be-solved problem are taken as the training data.

[0165] In actual application, the answer information can include a plurality of different information, for example, step-by-step analysis information of the to-be-solved problem, knowledge point information corresponding to the to-be-solved problem, and answer information corresponding to the to-be-solved problem, etc. Therefore, when the problem solving video of the to-be-solved problem needs to be generated, the video data in the answer information can be screened out from the answer information.

[0166] Optionally, the video data contained in the answer information corresponding to the any to-be-solved problem is compiled by using the video compiling software to obtain a problem solving video, and specifically can include:

[0167] The video code contained in the answer information is determined by using a regular expression.

[0168] The video code is saved as a Python file.

[0169] The Python file is compiled by using a manim command to obtain a problem solving video.

[0170] In the embodiments of the present specification, the regular expression is a tool for describing string patterns, and the regular expression can be used to find, match and replace specific strings in text. The regular expression in the embodiments of the present specification is a regular expression capable of extracting the video data in the answer information. For example, the regular expression can contain a code identifier character, such as Python, and the information after the code identifier character can be extracted by taking the code identifier character as the start. The commonly used regular expressions in Python include if…print, etc.

[0171] In the embodiments of this specification, the video compilation software used can be Manim. The manim (MathematicalAnimation Engine) command is a Python library for creating animations; using manim to compile Python files can generate the target explanatory video.

[0172] It is understandable that, besides using Manim software to compile and explain the videos, other software such as HandBrake and FFmpeg can also be used. No specific restrictions are imposed here.

[0173] Furthermore, the embodiments of this specification can also optimize the video compilation software, using the optimized software to compile videos. The optimized video compilation software incorporates optimized logic modules, enabling the compilation of video data to accomplish tasks that previously required a large amount of code with only a small amount of code. For example, if distribution analysis information needs to be displayed on the left side of the video and a line segment diagram on the right side, the previous video compilation software required precise determination of the absolute coordinates of the distribution analysis information and the line segment diagram. However, the optimized video compilation software, based on the optimized logic modules, can execute "display distribution analysis information on the left and display a line segment diagram on the right," automatically finding the specific positions on the left and right sides, and thus displaying the analysis information on the left and the line segment diagram on the right, without needing to precisely determine the absolute coordinates of the distribution analysis information and the line segment diagram. Therefore, the optimized video compilation software has higher video compilation efficiency.

[0174] In the embodiments described in this specification, the problem-solving video may include images and audio synchronized with the images.

[0175] Furthermore, problem-solving videos can also include video content explaining the solution process step by step. A portion of the video display page is used to show the whiteboard content, and another portion is used to show the line segment diagram. The whiteboard content can include step-by-step explanations, key knowledge points, and answer information.

[0176] Figure 2 This is a schematic diagram of image frames from one embodiment of a problem-solving video provided in this specification. Figure 2 As shown, Figure 2 The left side of the problem-solving video displays distribution analysis information, while the right side displays a line segment diagram.

[0177] Figure 3 This is a schematic diagram of image frames from another problem-solving video provided in one embodiment of this specification. For example... Figure 3 As shown, Figure 3 The answer information is displayed on the left side of the problem-solving video, and a line segment diagram is displayed on the right side of the video.

[0178] In the embodiments of the present specification, the video data of the problem solving video can be video data including a problem introduction, a knowledge point introduction, a problem explanation, and a closing speech. Based on the problem introduction video data, a problem introduction video can be obtained, based on the knowledge point introduction video data, a knowledge point introduction video can be obtained, based on the closing speech video data, a closing speech video can be obtained, and the like.

[0179] The problem introduction video can be a video that introduces the problem. For example, the problem introduction video can include a display of the problem, accompanied by a voice: “Let me explain the following problem to you”, and then the problem can be broadcasted by voice. The knowledge point introduction video can be a video that introduces the knowledge point information related to the problem. The knowledge point information related to the problem can be the distribution analysis information related to the problem. The closing speech video can be a video for ending the current video explanation, such as a voice playing: “This problem explanation is complete, have you learned it, student?”.

[0180] In addition, the video data of the problem solving video can also include video data introducing the identity of the teaching software loading or calling the trained problem solving model, based on which a video introducing the identity of the teaching software loading or calling the trained problem solving model can be obtained. For example, it can be a voice broadcast: “Hello, I am AI intelligent math teacher”.

[0181] Since the problem solving video is compiled from the video data of the problem solving video, by judging whether the problem solving video meets the problem solving requirement, it can be determined whether the video data of the problem solving video meets the problem solving requirement, and thus whether the answer information meets the problem solving requirement.

[0182] In the embodiments of the present specification, the answer information capable of generating a problem solving video meeting the problem solving requirement and the corresponding to-be-solved problem are used as training samples, which can reduce the interference of incorrect training samples on the problem solving model, improve the accuracy of the trained problem solving model, and thus the trained problem solving model can more accurately solve the problems related to the coordinate system.

[0183] It can be understood that after the answer information meeting the problem solving requirement is determined according to the problem solving video, the to-be-solved problem corresponding to the answer information meeting the problem solving requirement can be determined based on the answer information meeting the problem solving requirement, so that the positive sample training data can be generated according to the determined answer information and the to-be-solved problem.

[0184] Of course, after the answer information meeting the problem solving requirement is screened out from all the answer information in the above manner, the video data in the answer information not meeting the problem solving requirement can also be modified to obtain correct video data. The positive sample training data is generated by using the modified correct answer information and the to-be-solved problem corresponding to the modified correct answer information.

[0185] For ease of understanding, the embodiments of the present specification also provide specific contents for judging whether the problem-solving video meets the problem-solving requirements.

[0186] The judging whether the problem-solving video meets the problem-solving requirements can specifically include:

[0187] Judging whether the problem-solving video has at least one of the following situations:

[0188] The content of the problem-solving video fails to correctly answer the question.

[0189] The problem-solving video does not draw a coordinate graph.

[0190] There is image overlap in the coordinate graph drawn in the problem-solving video.

[0191] The coordinate graph drawn in the problem-solving video is incomplete.

[0192] The coordinate graph drawn in the problem-solving video is inconsistent with the explanation content.

[0193] The position of the point representing the number pair in the coordinate graph drawn in the problem-solving video is inconsistent with that in the question.

[0194] The position of the point in the coordinate graph drawn in the problem-solving video is inconsistent with the coordinate value of the point.

[0195] The coordinate graph drawn in the problem-solving video does not mark the scale or arrow.

[0196] The angle or direction of the marked position in the coordinate graph drawn in the problem-solving video is inconsistent with that in the question.

[0197] The drawing animation in the problem-solving video is not synchronized with the voice explanation.

[0198] If not, it is determined that the problem-solving video meets the problem-solving requirements.

[0199] In the embodiments of the present specification, one or more models such as a semantic analysis model and a word vector model can be used to analyze whether the question and the explanation content of the problem-solving video match, to judge whether the explanation process of the problem-solving video is correct and whether the explanation result is correct, and the like, so as to judge whether the problem-solving video correctly answers the question. Specifically, each frame of image in the problem-solving video can be identified, or one or more frames of image can be randomly selected from the problem-solving video for identification, which is not specifically limited herein.

[0200] Further, image recognition technology can be used to identify whether there is a coordinate graph in the video frame of the problem-solving video, to judge whether a line segment graph is drawn in the problem-solving video. If there is a coordinate graph in one or more image frames in the problem-solving video, it can be considered that a coordinate graph is drawn in the problem-solving video, otherwise, it can be considered that no coordinate graph is drawn in the problem-solving video.

[0201] In practical applications, in addition to ensuring that the coordinate graph is drawn in the problem-solving video, it is also necessary to ensure that the coordinate graph is qualified.

[0202] Specifically, one or more of the bounding box detection method, image segmentation method, and deep learning model can be used to detect whether the coordinate graph in the problem-solving video has image overlap, so as to determine whether the coordinate graph is qualified.

[0203] One or more of the edge detection algorithm, contour detection algorithm, and hash algorithm can be used to detect the boundary of the image to identify whether the coordinate graph is complete, so as to determine whether the coordinate graph is qualified.

[0204] Further, a semantic recognition model can also be used to identify whether the coordinate graph drawn in the explanation process corresponds to the explanation content, so as to determine whether the coordinate graph drawn in the problem-solving video is consistent with the explanation content, thereby determining whether the coordinate graph is qualified.

[0205] Key point matching and other methods can be used to identify whether the positions of points representing pairs of numbers in the coordinate graph drawn in the problem-solving video are consistent with the positions expressed in the question, whether the positions of points in the coordinate graph drawn in the problem-solving video are consistent with the coordinate values of the points, whether the coordinate graph drawn in the problem-solving video is marked with scales or arrows, and whether the angle or direction of the marked position in the coordinate graph drawn in the problem-solving video is inconsistent with the question, and so on, to determine whether the coordinate graph is qualified.

[0206] In order to ensure the quality of teaching, it is also necessary to synchronize the drawing animation displayed in the problem-solving video with the voice explanation to avoid affecting the teaching experience. Specifically, a model for determining whether the drawing animation and the voice explanation content are synchronized can be trained in advance, and then the trained model can be used to identify whether the drawing animation and the voice explanation content are synchronized. Of course, related video editing software or audio analysis tools can also be used to analyze the synchronization of the video and the audio.

[0207] In practical applications, the problem-solving video can be a question explanation using distributed analysis information, so the last frame of the problem-solving video often contains more content than other frames of the problem-solving video. In order to conveniently and accurately determine whether the problem-solving video meets the problem-solving requirements, the last frame of the problem-solving video can be used to determine whether the problem-solving video meets the problem-solving requirements.

[0208] For the specific process of determining whether the last frame of the problem-solving video has the above-mentioned situation, please refer to the previously mentioned content related to determining whether the problem-solving video meets the problem-solving requirements, which will not be repeated here.

[0209] In the embodiments of the present specification, the training data for training the problem solving model can be data in a predetermined format, so as to facilitate the problem solving model to understand the training data, improve the training efficiency of the problem solving model and the model accuracy.

[0210] Optionally, the training data is obtained based on the plurality of problems to be solved and the corresponding solution information containing the video data, and specifically can include:

[0211] Based on the training data format, the plurality of problems to be solved and the corresponding solution information containing the video data are generated into training data conforming to the training data format.

[0212] In the embodiments of the present specification, the problem to be solved and the solution information of the problem to be solved can be spliced in a preset order to generate training data.

[0213] Specifically, the problem to be solved can be placed in front of the solution information, and a first preset character is used to separate the problem to be solved and the solution information. In addition, a second preset character can also be set before the problem to be solved, so as to facilitate the problem solving model to understand the training data, improve the training efficiency of the problem solving model and the model accuracy. The first preset character and the second preset character can be set according to actual needs, which are not limited here.

[0214] For the convenience of understanding, the embodiments of the present specification are specifically described for the training data format.

[0215] Optionally, the training data format can include a role definition part, a problem part and an analysis part.

[0216] The role definition part includes information for defining the role played by the problem solving model.

[0217] The problem part includes information of the problem to be solved.

[0218] The analysis part includes knowledge point information, step-by-step analysis information, answer information and video data of the problem solving video corresponding to the problem to be solved.

[0219] In the embodiments of the present specification, the training data format can include a plurality of identification words with different meanings, and the problem solving model can learn the training data from different angles according to the identification words.

[0220] Specifically, the identification words can include a first identification word, such as [SYSTEM], which can be used to define the role of the model and explain the task of the model this time. The first identification word can be located at the starting position of the role definition part. The identification words can also include a second identification word, such as [USER], which can be the question to be answered. The identification words can also include a third identification word, such as [ASSISTANT], which can be the analysis part of the question to be answered. Of course, the above identification words can also be other field words or texts, which are not limited here.

[0221] Further, the problem-solving model can understand the task of the model this time according to the content after the first identification word, so the first identification word can be used to distinguish different business scenarios. In addition, according to the first identification word, the problem-solving model can also determine the identity of the model performing the task this time, for example, the problem-solving model is a math teacher. The second identification word and the third identification word can be understood as a group of dialogues. The second identification word can be a question asked by the user as user, and the content after the second identification word can be the content of the question asked by the user. The third identification word can be the answer content of the problem-solving model as assistant to the question asked by the user, and the content after the third identification word can be the answer content of the assistant. The above three identification words can be spliced in order as training data for training the problem-solving model.

[0222] It can be understood that [ASSISTANT] can be the first preset character mentioned in the foregoing. [USER] can be the second preset character mentioned in the foregoing.

[0223] For example, the format and content of the training data can be as follows:

[0224] [SYSTEM] You are a math teacher, please use manim to write a code to make a teaching video for the following question.

[0225] [USER] The school auditorium is 30° east-south of the stadium, then the stadium is in which direction of the auditorium.

[0226] [ASSISTANT] [Knowledge points] xxx, [Step-by-step analysis] xxx, [Answer information] xxx, and [Code information] xxx.

[0227] Generally, we hope that the trained problem-solving model can output accurate answer information for the question after obtaining the question. Specifically, when training the problem-solving model using the training data, the problem-solving model can output prediction data according to the role definition part and the question part proposed by the user, and the problem-solving model can also use the analysis part of the question as reference data to adjust the model parameters, thereby obtaining the trained problem-solving model.

[0228] In the embodiments of the present specification, the problem-solving model can be trained by using supervised fine-tuning. In the training process, the problem-solving model can adjust its hyperparameters according to the training data to obtain the trained problem-solving model.

[0229] The goal of training the problem-solving model is to minimize the difference between the reference data and the predicted data, that is, to minimize the loss value loss between the reference data and the predicted data. Therefore, the parameters of the problem-solving model can be continuously adjusted to reduce the loss value loss during the training process. When the loss value loss no longer decreases or decreases to within a preset range, it can be indicated that the training of the problem-solving model is complete.

[0230] To improve the training efficiency, the problem-solving model in the embodiments of the present specification can be a pre-trained model, such as a YunLLM70B model, a Multimodal-CoT model, or a Mixtral8x7B model, etc. It can also be other types of pre-trained models, which are not specifically limited here.

[0231] In the embodiments of the present specification, the obtained training data can be used to train the problem-solving model.

[0232] Suppose multiple training data are obtained, each training data can include a to-be-answered question, knowledge point information, step-by-step analysis information, answer information, and the code of the corresponding problem-solving video, etc. In addition, the training data should include various types of to-be-answered questions and other information corresponding to the various types of to-be-answered questions to improve the generalization ability of the problem-solving model.

[0233] After obtaining the training data, data cleaning and formatting can also be performed on the training data to ensure the consistency and standardization of the data, thereby facilitating the training of the problem-solving model. Specifically, the training data can be segmented, de-noised, and formatted for problem-solving video data, etc. If the knowledge point information, step-by-step analysis information, answer information, and corresponding problem-solving video data in the training data are generated by using a large language model, since the data generated by the large language model is usually standardized and uniform in format, data cleaning and formatting processing of the training data can not be performed at this time.

[0234] When the problem-solving model is initially trained, the problem-solving model can be initialized, and some training parameters can be set. For example, for the YunLLM70B model, the YunLLM70B model can be loaded, and training parameters such as learning rate, batch size, and training round number can be configured.

[0235] In addition, the problem-solving model can be fine-tuned using the training data. Specifically, the problem to be solved can be taken as input, and the generated accurate knowledge point information, step-by-step analysis information, answer information, and video data of the problem-solving video can be taken as the target. During the training process, the problem-solving model continuously adjusts the parameters through backpropagation to minimize the difference between the generated video data of the problem-solving video and the labeled video data of the problem-solving video.

[0236] During the training process of the problem-solving model, the model performance can also be evaluated periodically or irregularly using a validation set to adjust the parameters and training strategies of the problem-solving model, thereby ensuring the stability and accuracy of the problem-solving model.

[0237] After the training of the problem-solving model is completed, the model parameters and structure of the trained problem-solving model can be saved for subsequent use and deployment.

[0238] In practical applications, model evaluation is an important step to verify the quality and reliability of the video data of the problem-solving video generated by the trained problem-solving model. As an implementation, the performance of the trained problem-solving model can be evaluated by the following steps.

[0239] A test set is obtained. Specifically, 100 questions can be randomly selected from the test question library, and it is determined that the questions in the test set do not overlap with the training data. In addition, it is also necessary to ensure that the questions in the test set cover various types and complexities of questions that can be solved using line segment graphs, so as to comprehensively evaluate the performance of the problem-solving model.

[0240] The trained problem-solving model is used to generate the corresponding video data of the problem-solving video for each question in the test set. The generation process is similar to the training phase. Specifically, the question can be input into the trained problem-solving model, and the trained problem-solving model can output knowledge point information, step-by-step analysis information, and video data of the problem-solving video, etc.

[0241] The generated video data of the problem-solving video is input into a video compiling software to compile the problem-solving video.

[0242] Further, the performance of the trained problem-solving model can be evaluated according to the compiled problem-solving video. Specifically, the problem-solving video can be evaluated using pre-determined evaluation criteria. The evaluation criteria can include the accuracy of the video, such as whether the problem-solving steps and results are correct, the intuitiveness of the video, such as whether the video is clear and easy to understand, the liveliness of the video, such as whether the explanation process is lively and interesting, and the technicality of the video, such as whether the code generation is efficient and the video playback is smooth, etc.

[0243] In addition, the generated video can be evaluated by using the method of judging whether the problem solving video meets the problem solving requirements mentioned above. Professional teachers and technical experts can also be organized to manually evaluate the generated video and record the evaluation results. Of course, the above-mentioned several methods can be combined, which is not limited here.

[0244] Further, the evaluation results can be statistically analyzed to evaluate the performance of the trained model. According to the analysis results, the advantages and disadvantages of the trained problem solving model are determined, thereby providing a basis for further optimizing the problem solving model.

[0245] Through the above evaluation steps, it is ensured that the trained problem solving model has the ability to generate high-quality video data of problem solving videos, and provides reliable technical support for actual application.

[0246] Figure 4 is a process diagram of a method for generating a problem solving model for solving coordinate system problems provided by an embodiment of the present specification. As shown in Figure 4 The method for generating a problem solving model for solving coordinate system problems can include:

[0247] Step 402: Obtain a reference example.

[0248] In an embodiment of the present specification, the reference example can be a plurality of reference examples of different types related to the coordinate system.

[0249] The reference example can be selected from the reference example set. For example, a semantic matching model can be used, or a retrieval enhancement generation model can be used to select a reference example related to the coordinate system from the reference example set.

[0250] It can be understood that the reference example can also be written according to expert experience. Of course, the code of the problem solving video can also be extracted by using reverse engineering for the existing problem solving video related to the coordinate system, so as to obtain the obtained problem and the code of the problem solving video as a reference example, etc.

[0251] Step 404: Obtain solution information.

[0252] In an embodiment of the present specification, the solution information can be solution information of a plurality of to-be-solved problems obtained by using a large language model.

[0253] The solution information can include video data for generating a problem solving video of a to-be-solved problem. The video data of the problem solving video can be the code of the problem solving video.

[0254] The large language model is used to generate answer information of a plurality of to-be-solved questions, and then training data can be obtained based on the answer information generated by the large language model, without manually generating the training data, so that the generation efficiency of the training data can be improved.

[0255] Step 406: compiling the problem-solving video.

[0256] In actual application, the compiled software such as Manim can be used to compile the code of the problem-solving video in the answer information obtained in step 404, so as to generate the problem-solving video.

[0257] Step 408: labeling the problem-solving video.

[0258] In the embodiments of the present specification, labeling the problem-solving video can mean judging whether the problem-solving video meets the problem-solving requirements. Then, the problem-solving video that does not meet the problem-solving requirements can be removed, and the question and the problem-solving information corresponding to the problem-solving video meeting the problem-solving requirements are used as training samples to train the problem-solving model.

[0259] Step 410: model fine-tuning.

[0260] In the embodiments of the present specification, the problem-solving model can be trained by using supervised fine-tuning. In the training process, the problem-solving model can adjust its hyperparameters according to the training data to obtain the trained problem-solving model.

[0261] The video data of the problem-solving video for explaining the problem of the coordinate system generated by the trained problem-solving model can include the code of the problem-solving video for explaining the problem of the coordinate system.

[0262] In the embodiments of the present specification, the video data of the problem-solving video for explaining the problem of the coordinate system can be directly obtained by using the trained problem-solving model. Compared with obtaining the video data by using the large language model, since the trained problem-solving model is obtained by training the reference example questions related to the coordinate system, the video data of the question related to the coordinate system can be more accurately generated, so that the quality of the generated problem-solving video can be ensured.

[0263] In addition, when the trained problem-solving model is used to solve the question, a large number of prompt words similar to those used by the large language model described above do not need to be preset, and the problem-solving efficiency can be improved.

[0264] Corresponding to the method embodiment for generating the problem-solving model for solving the problem of the coordinate system, the present specification also provides a device embodiment for generating the problem-solving model for solving the problem of the coordinate system, Figure 5 A structural schematic diagram of a device for generating a problem-solving model for solving the problem of the coordinate system is shown in an embodiment of the present specification. As shown in Figure 5 The device comprises:

[0265] The answer information obtaining module 502 is configured to obtain answer information of a plurality of to-be-solved problems obtained by using a large language model; the to-be-solved problems are problems related to a coordinate system; and the answer information includes video data for generating a solving video of the to-be-solved problems.

[0266] The training data obtaining module 504 is configured to obtain training data based on the plurality of to-be-solved problems and corresponding answer information including the video data.

[0267] The model training module 506 is configured to train a solving model by using the training data to obtain a trained solving model. The trained solving model can generate video data of a solving video for explaining a problem related to a coordinate system.

[0268] Optionally, the apparatus can further include:

[0269] The reference example obtaining module is configured to obtain a plurality of reference examples of different types related to a coordinate system; the different types include at least one of a direction type and a number pair type; and the reference examples include example questions and video data for generating a solving video of the example questions.

[0270] The template generating module is configured to generate a prompt word template including the reference examples based on the reference examples.

[0271] The prompt word obtaining module is configured to add any to-be-solved problem in the plurality of to-be-solved problems to the prompt word template to obtain a prompt word for the any to-be-solved problem.

[0272] The prompt word inputting module is configured to input the prompt word to the large language model to obtain answer information corresponding to the any to-be-solved problem.

[0273] Optionally, the template generating module can be specifically configured to:

[0274] select a first number of first target reference examples belonging to the direction type from the plurality of reference examples.

[0275] generate a first prompt word template including the first number of first target reference examples based on the first number of first target reference examples.

[0276] The prompt word obtaining module can be specifically configured to:

[0277] if the any to-be-solved problem is a problem of the direction type, add the any to-be-solved problem to the first prompt word template to obtain a prompt word for the any to-be-solved problem.

[0278] Optionally, the template generation module can be specifically configured to:

[0279] select, from the several reference examples, a second number of second target reference examples belonging to the pair type.

[0280] generate, based on the second number of second target reference examples, a second prompt word template containing the second number of second target reference examples.

[0281] The prompt word acquisition module can be specifically configured to:

[0282] If the any to-be-solved question is a question of the pair type, add the any to-be-solved question to the second prompt word template to obtain a prompt word for the any to-be-solved question.

[0283] Optionally, the first prompt word template further includes instruction information indicating that a large language model establishes a coordinate system and marks an angle relationship.

[0284] Optionally, the second prompt word template further includes instruction information indicating that a large language model establishes a coordinate system and marks coordinate system parameters. The coordinate system parameters include at least one of an arrow, a scale line, a number, and information for representing a shape and a color of an object in a question.

[0285] Optionally, the reference examples further include at least one of step-by-step analysis information for step-by-step solving the example questions, knowledge point information corresponding to the example questions, and answer information corresponding to the example questions.

[0286] And / or, the solving information obtained by the large language model further includes at least one of step-by-step analysis information for step-by-step solving the to-be-solved questions, knowledge point information corresponding to the to-be-solved questions, and answer information corresponding to the to-be-solved questions.

[0287] Optionally, the device can further include:

[0288] The compiling module is configured to, for any to-be-solved question in the several to-be-solved questions, compile video data contained in the solving information corresponding to the any to-be-solved question by using video compiling software to obtain a question-solving video.

[0289] The judging module is configured to judge whether the question-solving video meets a question-solving requirement.

[0290] The training data acquisition module 504 can be specifically configured to:

[0291] If the problem-solving video meets the problem-solving requirements, the any to-be-answered question and the corresponding answer information of the any to-be-answered question are taken as training data.

[0292] Optionally, the judging module can be specifically configured to:

[0293] judge whether the problem-solving video has at least one of the following situations:

[0294] the content of the problem-solving video fails to correctly answer the question.

[0295] the problem-solving video does not draw a coordinate graph.

[0296] there is image overlap in the coordinate graph drawn in the problem-solving video.

[0297] the coordinate graph drawn in the problem-solving video is incomplete.

[0298] the coordinate graph drawn in the problem-solving video is inconsistent with the explanation content.

[0299] the position of a point representing a number pair in the coordinate graph drawn in the problem-solving video is inconsistent with that in the question.

[0300] the position of a point in the coordinate graph drawn in the problem-solving video is inconsistent with the coordinate value of the point.

[0301] the coordinate graph drawn in the problem-solving video does not mark scales or arrows.

[0302] the angle or direction of the marked position in the coordinate graph drawn in the problem-solving video is inconsistent with that in the question.

[0303] the drawing animation in the problem-solving video is not synchronized with the voice explanation.

[0304] If not, it is determined that the problem-solving video meets the problem-solving requirements.

[0305] Optionally, the compiling module can be specifically configured to:

[0306] determine the video code contained in the answer information by using a regular expression.

[0307] save the video code as a Python file.

[0308] compile the Python file by using a manim command to obtain a problem-solving video.

[0309] Optionally, the training data acquisition module 504 can be specifically configured to:

[0310] Based on a training data format, the several to-be-answered questions and corresponding answer information containing the video data are generated into training data conforming to the training data format.

[0311] Optionally, the training data format comprises a role definition part, a question part and an analysis part.

[0312] The role definition part comprises information for defining a role played by the question-solving model.

[0313] The question part comprises information of the to-be-answered questions.

[0314] The analysis part comprises knowledge point information, step-by-step analysis information, answer information and video data of a question-solving video corresponding to the to-be-answered questions.

[0315] Figure 6 A structural block diagram of a computing device 600 according to one embodiment of the present specification is shown. The components of the computing device 600 include but are not limited to a memory 610 and a processor 620. The processor 620 is connected with the memory 610 through a bus 630, and a database 650 is used to save data.

[0316] The computing device 600 further comprises an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 640 can include one or more of any type of network interface (for example, a network interface card (NIC)) such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC).

[0317] In one embodiment of the present specification, the above-mentioned components of the computing device 600 andFigure 6 Other components not shown in FIG. 6 can also be connected to each other through the bus. It should be understood that, Figure 6 The illustrated computing device structural block diagram is merely for the purpose of example, and is not a limitation on the scope of the present specification. Other components can be added or replaced by those skilled in the art as needed.

[0318] The computing device 600 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other type of mobile device, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 600 can also be a mobile or stationary server.

[0319] The processor 620 is configured to execute computer-executable instructions, which, when executed by the processor, implement the steps of the method for generating a problem-solving model for solving a coordinate system problem.

[0320] The above is a schematic scheme of the computing device of the embodiment. It should be noted that the technical scheme of the computing device belongs to the same concept as the technical scheme of the method for generating a problem-solving model for solving a coordinate system problem and / or the technical scheme of the method for solving a problem, and details of the technical scheme of the computing device that are not described in detail can be referred to the description of the technical scheme of the method for generating a problem-solving model for solving a coordinate system problem and / or the method for solving a problem.

[0321] An embodiment of the present specification also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method for generating a problem-solving model for solving a coordinate system problem.

[0322] The above is a schematic scheme of the computer-readable storage medium of the embodiment. It should be noted that the technical scheme of the storage medium belongs to the same concept as the technical scheme of the method for generating a problem-solving model for solving a coordinate system problem, and details of the technical scheme of the storage medium that are not described in detail can be referred to the description of the technical scheme of the method for generating a problem-solving model for solving a coordinate system problem.

[0323] An embodiment of the present specification also provides a computer program, which, when executed in a computer, causes the computer to perform the steps of the method for generating a problem-solving model for solving a coordinate system problem.

[0324] The above is a schematic solution of the computer program of the embodiment. It should be noted that the technical solution of the computer program and the technical solution of the problem solving model method described above, and / or the technical solution of the problem solving method described above belong to the same concept. The technical solution of the computer program is not described in detail. The details can be seen from the description of the technical solution of the problem solving model method and / or the technical solution of the problem solving method.

[0325] The above describes specific embodiments of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different than the order in the embodiments and still achieve the desired result. In addition, the processes depicted in the accompanying drawings do not necessarily require the particular order shown or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing are also possible or advantageous.

[0326] The computer instructions include computer program code, which can be in the form of source code, object code, executable code, or some intermediate form. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the computer readable medium can include appropriate additions or subtractions according to the requirements of patent practice. For example, in some regions, according to the patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.

[0327] It should be noted that for the foregoing method embodiments, in order to facilitate description, they are all expressed as a combination of a series of actions, but those skilled in the art should know that the embodiments of the present specification are not limited by the order of the described actions, because according to the embodiments of the present specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the embodiments of the present specification.

[0328] In the above embodiments, the description of each embodiment is focused on, and the parts not described in detail in a certain embodiment can be seen from the related description of other embodiments.

[0329] The preferred embodiments of the present specification disclosed above are only used to help explain the present specification. Alternative embodiments do not describe all the details and limit the present application to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the embodiments of the present specification. The present specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present specification, so that those skilled in the art can well understand and utilize the present specification. The present specification is limited only by the claims and their full scope and equivalents.

Claims

1. A method of generating a problem-solving model for solving coordinate system problems, characterized by, The method comprises the following steps: obtaining answer information of a plurality of to-be-answered questions obtained by using a large language model; the to-be-answered question is a question related to a coordinate system; the answer information includes video data for generating a problem-solving video of the to-be-answered question, and the video data of the problem-solving video of the to-be-answered question includes a code of the problem-solving video of the to-be-answered question; based on the plurality of to-be-answered questions and the corresponding answer information containing the video data, training data is obtained; training a problem-solving model using the training data to obtain a trained problem-solving model; the trained problem-solving model can generate video data of a problem-solving video for explaining a problem related to a coordinate system; the video data of the problem-solving video for explaining the problem related to the coordinate system includes a code of the problem-solving video for explaining the problem related to the coordinate system; the method further comprises: obtaining a plurality of reference examples related to a coordinate system of different types; the different types include at least one of an orientation type and a number pair type; the reference example includes an example question and video data for generating a problem-solving video of the example question; based on the reference example, a prompt word template containing the reference example is generated; for any to-be-answered question in the plurality of to-be-answered questions, the any to-be-answered question is added to the prompt word template to obtain a prompt word for the any to-be-answered question; inputting the prompt word into the large language model to obtain answer information corresponding to the any to-be-answered question.

2. The method of claim 1, wherein, The generation of the prompt word template containing the reference example based on the reference example specifically comprises: selecting a first number of first target reference examples belonging to the orientation type from the plurality of reference examples; based on the first number of first target reference examples, a first prompt word template containing the first number of first target reference examples is generated; the any to-be-answered question is added to the first prompt word template to obtain a prompt word for the any to-be-answered question. The generation of the prompt word template containing the reference example based on the reference example specifically comprises:

3. The method of claim 1, wherein, selecting a second number of second target reference examples belonging to the number pair type from the plurality of reference examples; based on the second number of second target reference examples, a second prompt word template containing the second number of second target reference examples is generated; the any to-be-answered question is added to the second prompt word template to obtain a prompt word for the any to-be-answered question. The first prompt word template further includes instruction information indicating that the large language model establishes a coordinate system and marks an angle relationship. ​ 4. The method of claim 2, wherein, ​ 5. The method of claim 3, wherein, The second prompt word template further includes instruction information indicating that the large language model establishes a coordinate system and labels coordinate system parameters; the coordinate system parameters include at least one of an arrow, a scale line, a number, a shape for representing an object in the question, and color information.

6. The method of claim 1, wherein, The reference example question further includes at least one of step-by-step analysis information for step-by-step solving of the example question, knowledge point information corresponding to the example question, and answer information corresponding to the example question; And / or, the solving information obtained by the large language model further includes at least one of step-by-step analysis information for step-by-step solving of the to-be-solved question, knowledge point information corresponding to the to-be-solved question, and answer information corresponding to the to-be-solved question.

7. The method of claim 1, wherein, The method further includes: For any to-be-solved question in the plurality of to-be-solved questions, compiling video data contained in the solving information corresponding to the any to-be-solved question by using video compiling software to obtain a problem-solving video; determining whether the problem-solving video meets the problem-solving requirements; The training data is obtained based on the plurality of to-be-solved questions and the corresponding solving information containing the video data, specifically including: If the problem-solving video meets the problem-solving requirements, the any to-be-solved question and the solving information corresponding to the any to-be-solved question are used as training data.

8. The method of claim 7, wherein, The determination of whether the problem-solving video meets the problem-solving requirements specifically includes: determining whether the problem-solving video has at least one of the following situations: The content of the problem-solving video cannot correctly solve the question; The coordinate graph is not drawn in the problem-solving video; There is image overlap in the coordinate graph drawn in the problem-solving video; The coordinate graph drawn in the problem-solving video is incomplete; The coordinate graph drawn in the problem-solving video is inconsistent with the explanation content; The position of the point representing the number pair in the coordinate graph drawn in the problem-solving video is inconsistent with that in the question; The position of the point in the coordinate graph drawn in the problem-solving video is inconsistent with the coordinate value of the point; The coordinate graph drawn in the problem-solving video is not labeled with scales or arrows; The angle or direction of the marked position in the coordinate graph drawn in the problem-solving video is inconsistent with that in the question; The drawing animation in the problem-solving video is not synchronized with the voice explanation; If not, it is determined that the problem-solving video meets the problem-solving requirements.

9. The method of claim 7, wherein, The video data contained in the solving information corresponding to the any to-be-solved question is compiled by using video compiling software to obtain a problem-solving video, specifically including: determining the video code contained in the solving information by using a regular expression; save the video code as a Python file; The Python file is compiled using the manim command to obtain a problem-solving video.

10. The method of claim 1, wherein, The training data is obtained based on the plurality of to-be-solved questions and the corresponding solving information containing the video data, specifically including: Based on the training data format, the plurality of to-be-solved questions and the corresponding solving information containing the video data are generated into training data conforming to the training data format.

11. The method of claim 10, wherein, The training data format includes a role definition part, a question part, and an analysis part; The role definition part includes information for defining a role played by the problem solving model; The question part includes information of the question to be answered; The analysis part includes knowledge point information corresponding to the question to be answered, step-by-step analysis information, answer information, and video data of a problem solving video.

12. An apparatus for generating a problem-solving model for solving coordinate system problems, characterized by: Comprise: The answer information acquisition module is configured to acquire answer information of a plurality of questions to be answered obtained by using a large language model; The question to be answered is a question related to a coordinate system; the answer information includes video data for generating a problem solving video of the question to be answered, and the video data of the problem solving video of the question to be answered includes code of the problem solving video of the question to be answered; The training data acquisition module is configured to obtain training data based on the plurality of questions to be answered and corresponding answer information containing the video data; The model training module is configured to train a problem solving model using the training data to obtain a trained problem solving model; The trained problem solving model can generate video data of a problem solving video for explaining a problem related to a coordinate system; The video data of the problem solving video for explaining the problem related to the coordinate system includes code of the problem solving video for explaining the problem related to the coordinate system; The device further comprises: The reference example question acquisition module is configured to acquire a plurality of reference example questions of different types related to a coordinate system; the different types include at least one of a direction type and a number pair type; the reference example question includes an example question and video data for generating a problem solving video of the example question; The template generation module is configured to generate a prompt word template containing the reference example question based on the reference example question; The prompt word acquisition module is configured to add any question to be answered in the plurality of questions to be answered to the prompt word template to obtain a prompt word for the any question to be answered; The prompt word input module is configured to input the prompt word into the large language model to obtain answer information corresponding to the any question to be answered.

13. A computing device, comprising: Comprise: Memory and processor; The memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions, which realize the steps of the method in any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, It stores computer executable instructions, which realize the steps of the method in any one of claims 1 to 11 when executed by the processor.

15. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instructions realize the steps of the method in any one of claims 1 to 11 when executed by the processor.

Citation Information

Patent Citations

  • Calculation question answering and model training method and device, equipment and medium

    CN118133009A

  • Question explanation method and device, electronic equipment and storage medium

    CN118506620A