Teaching video generation method, computer device, and program product
Patent Information
- Application Number
- CN202610652137.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-12
- Publication Date
- 2026-09-08
AI Technical Summary
[0004]然而,目前技术无法实现静态几何题目向动态教学视频的低成本高质量转化
[0021]The teaching video generation method, computer equipment, and program products provided in this application automatically generate corresponding dynamic graphics by acquiring a structured sequence of problem-solving steps, extracting the dynamic parameters and range of motion of geometric elements, and simultaneously generating synthesized audio explanations. This process can achieve automated production from static geometry problems to dynamic teaching videos without human intervention, thereby reducing video production costs while ensuring strict consistency between graphic changes and problem-solving logic, as well as accurate audio-visual synchronization. This enables low-cost, high-quality conversion of static geometry problems into dynamic teaching videos.
Smart Images

Figure CN122714616A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of teaching video generation technology, and in particular to a teaching video generation method, computer equipment, and program product. Background Technology
[0002] In the field of mathematics education technology, especially in the area of automated generation of geometry teaching content, how to efficiently and cost-effectively create intuitive and dynamic explanations for a massive number of geometry problems is an important issue currently facing the challenge of improving teaching effectiveness and students' depth of understanding.
[0003] Currently, the production of explanations for geometry problems mainly relies on the following technical solutions: First, a solution that automatically generates static vector graphics based on natural language descriptions; second, a solution that uses teachers to record screen handwriting or professional animation software (such as GeoGebra and Manim) for manual production; and third, a solution that uses general video generators to directly generate explanation videos from large models.
[0004] However, current technology cannot achieve low-cost, high-quality conversion of static geometry problems into dynamic instructional videos. Summary of the Invention
[0005] Therefore, it is necessary to provide a teaching video generation method, computer equipment, and program product that can achieve low-cost, high-quality transformation of static geometry problems into dynamic teaching videos, addressing the aforementioned technical problems.
[0006] Firstly, this application provides a method for generating instructional videos, the method comprising:
[0007] Obtain a sequence of solution steps for geometry problems with a preset structure format; Extract the dynamic parameters of the geometric elements and the range of motion of the dynamic parameters from the sequence of problem-solving steps; Based on dynamic parameters and the range of motion of the dynamic parameters, dynamic graphics corresponding to the sequence of problem-solving steps are generated. Generate audio explanations corresponding to the sequence of problem-solving steps, synthesize dynamic graphics and audio explanations, and output dynamic geometry teaching videos.
[0008] In one embodiment, the dynamic parameters of the geometric elements and the range of motion of the dynamic parameters are extracted from the sequence of problem-solving steps, including: By using predefined operation instruction templates, operation instructions in the problem-solving step sequence can be identified. Based on the recognized operation commands, the dynamic parameters of the geometric elements and the range of motion of the dynamic parameters are extracted.
[0009] In one embodiment, extracting the dynamic parameters of the geometric elements and the range of motion of the dynamic parameters from the sequence of problem-solving steps further includes: When there are unrecognized problem-solving step sequences in the operation instruction template, the unrecognized problem-solving step sequences are semantically annotated using a pre-trained recognition language model to identify the dynamic parameters and movement range of the problem-solving step sequences.
[0010] In one embodiment, semantic annotation of unrecognized problem-solving step sequences is performed using a pre-trained recognition language model, including: By recognizing the language model, semantic role labeling is performed on the unrecognized problem-solving step sequences to identify the agent role, predicate description, and patient role contained in the problem-solving step sequences; Based on the results of semantic role labeling, determine the dynamic type and dynamic parameters of the operation corresponding to the sequence of problem-solving steps; Among them, the operation dynamic type is used to characterize the transformation category of geometric elements in the problem-solving step sequence and to determine the parameter extraction method corresponding to the dynamic parameters.
[0011] In one embodiment, the method further includes: Determine the confidence level of the identified dynamic parameters and evaluate the confidence level of the dynamic parameters. The confidence level evaluation combines semantic confidence level, geometric rationality confidence level, and system feasibility confidence level. If the confidence level of the dynamic parameters is lower than the preset threshold, a static graphic is generated based on the dynamic parameters and the static graphic is marked. Among them, static graphics are used to represent the static display of the steps corresponding to the problem-solving steps sequence, and markers are used to indicate the review of the step-by-step teaching of the problem-solving steps sequence.
[0012] In one embodiment, the method further includes, prior to synthesizing motion graphics and audio narration: Get the duration of the audio explanation; Using the duration of the speech as a constraint, the interpolation step size is determined based on the range of motion of the dynamic parameters, and duration-matched animation clips are generated so that the animation duration of the dynamic graphics matches the duration of the speech narration.
[0013] In one embodiment, generating a voice explanation corresponding to the sequence of problem-solving steps includes: Each step of the problem-solving sequence is divided into an explanation script; The explanation script is converted from text to speech and synthesized to obtain the audio explanation corresponding to the sequence of problem-solving steps.
[0014] In one embodiment, the synthesis of dynamic graphics and voice narration further includes: Extract the display text for each step in the sequence of problem-solving steps; Generate a display page corresponding to each step. The display page includes a graphic area for displaying dynamic graphics and a whiteboard area for displaying the text. Use the display page as a background and overlay dynamic graphics onto the graphic area.
[0015] In one embodiment, the display page adopts a column layout, with a graphic area on one side and a whiteboard area on the other side; The whiteboard area displays the presentation text for each step in an append-only manner, and highlights the presentation text for the current step.
[0016] In one embodiment, obtaining a sequence of solution steps for a geometry problem with a preset structural format includes: Obtain the text description of the geometry problem and generate the initial geometric figure based on the text description; Using a pre-trained parsing language model, a sequence of problem-solving steps with a preset structure format is generated based on text descriptions and initial geometric figures.
[0017] In one embodiment, the parsing language model is obtained by fine-tuning a mathematical vertical domain model.
[0018] In one embodiment, generating an initial geometry based on a text description includes: Analyze the static geometry in the text description to determine the type, positional relationships, and constraints of the static geometry; Based on the type, positional relationship, and constraints of the static geometric body, a geometric coordinate system is established, and the coordinate data of the static geometric body is determined to obtain the initial geometric figure.
[0019] Secondly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.
[0020] Thirdly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0021] The teaching video generation method, computer equipment, and program products provided in this application automatically generate corresponding dynamic graphics by acquiring a structured sequence of problem-solving steps, extracting the dynamic parameters and range of motion of geometric elements, and simultaneously generating synthesized audio explanations. This process can achieve automated production from static geometry problems to dynamic teaching videos without human intervention, thereby reducing video production costs while ensuring strict consistency between graphic changes and problem-solving logic, as well as accurate audio-visual synchronization. This enables low-cost, high-quality conversion of static geometry problems into dynamic teaching videos. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 A flowchart illustrating a teaching video method provided in an embodiment of this application; Figure 2 A flowchart illustrating the steps for extracting dynamic parameters of geometric elements and the range of motion of these dynamic parameters, provided in an embodiment of this application; Figure 3 A flowchart illustrating another step for extracting dynamic parameters of geometric elements and the range of motion of the dynamic parameters, provided in an embodiment of this application; Figure 4 A flowchart illustrating the confidence assessment steps included in a dynamic geometry teaching video provided in this application embodiment; Figure 5 A flowchart illustrating the steps of synthesizing dynamic graphics and voice narration, provided for an embodiment of this application; Figure 6 This is a schematic diagram of the structure of an instructional video generation device provided in an embodiment of this application; Figure 7 This is a schematic diagram of the internal structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0025] In one exemplary embodiment, Figure 1This is a flowchart illustrating a teaching video method provided in an embodiment of this application, such as... Figure 1 As shown, a teaching video method is provided. This example illustrates the method's application to a terminal. It is understood that this method can also be applied to a server, or to a system including both a terminal and a server, and is implemented through interaction between the terminal and the server. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, and tablets. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. In this embodiment, the method includes the following steps S101 to S104. Wherein: S101. Obtain a sequence of solution steps for geometry problems with a preset structure format.
[0026] In this context, "geometry problem" refers to textual descriptions or structured data containing mathematical elements such as geometric figures, relationships between points, lines, and planes, and geometric transformations, such as a plane geometry proof or drawing problem. "Preset structure format" refers to organizing solution steps in a machine-readable manner with clearly defined fields or tags, such as structured storage according to step number, operation type, and operation object. "Solution step sequence" refers to a collection of multiple steps arranged in a logical order, where each step describes a specific geometric operation or reasoning action.
[0027] For example, a terminal can receive a sequence of solution steps for a geometry problem from local storage or a web server. This sequence of solution steps may be pre-generated by other modules (such as a large language model) and stored and transmitted in a preset structured format (such as JSON or XML). The terminal can parse this structured data to obtain an ordered list of steps, each step containing an operation description and possible related parameters.
[0028] By acquiring a sequence of problem-solving steps with a preset structure, the terminal can obtain machine-understandable problem-solving logic, providing a clear instruction basis for the subsequent automatic generation of dynamic graphics and voice explanations. This avoids the ambiguity of directly analyzing steps from natural language problems, thereby enabling low-cost and high-quality conversion of static geometry problems into dynamic teaching videos.
[0029] As an example, in a specific application scenario, the sequence of problem-solving steps obtained by the terminal could be a list of steps obtained after analyzing a problem such as "rotate triangle ABC 90 degrees clockwise around point C". For example, step 1: "Connect OA, and draw OB perpendicular to AC", step 2: "Let point P start from point A and move along side AB towards point B", step 3: "Rotate triangle ABC 90 degrees clockwise around point C", etc. These steps can be stored in a structured form for easy subsequent processing.
[0030] Optionally, the sequence of solution steps can be obtained by reading from a local file, downloading from a cloud server, or by receiving and structuring the user's input of the solution approach in real time. The preset structure format can also be a custom text format, such as one step per line, with specific delimiters indicating the operation type and parameters.
[0031] S102. Extract the dynamic parameters of the geometric elements and the range of motion of the dynamic parameters from the sequence of problem-solving steps.
[0032] Geometric elements refer to basic objects that constitute geometric figures, such as points, lines, surfaces, angles, and circles, as well as composite figures formed by their combinations. Dynamic parameters refer to variable objects that describe changes in geometric elements over time or during operation, such as the position coordinates of a moving point, the angle of rotation, and the distance of translation. The range of motion refers to the range or discrete set of values that dynamic parameters are allowed to take during their changes, such as an angle from 0 degrees to 90 degrees, or a time parameter t from 0 to 1.
[0033] For example, the terminal can analyze the acquired sequence of problem-solving steps with a preset structured format, identify which geometric elements change in each step, and extract the specific numerical values or variables describing these changes. For instance, for a step like "point P moves from A to B," the terminal can extract the path parameters of the moving point P and its start and end points, thereby obtaining dynamic parameters (such as position coordinates) and its range of motion (linear interpolation from the coordinates of point A to the coordinates of point B). The terminal can parse the step text using predefined rules or machine learning models to extract the required dynamic parameters and range of motion.
[0034] By accurately extracting the dynamic parameters and range of motion of geometric elements from the sequence of problem-solving steps, the terminal can transform abstract problem-solving steps into quantifiable graphic transformation instructions, providing precise control data for the subsequent generation of dynamic graphics consistent with the problem-solving logic. This enables the low-cost, high-quality transformation of static geometry problems into dynamic teaching videos.
[0035] As an example, taking the aforementioned rotating triangle as an example, the terminal can extract the geometric element as triangle ABC, the dynamic parameters as the rotation angle, the movement range as 0 to 90 degrees (clockwise), and the rotation center as point C from the step "rotate triangle ABC 90 degrees clockwise around point C". These dynamic parameters can be used to drive the subsequent generation of graphic animations.
[0036] Alternatively, dynamic parameters can be extracted using keyword matching based on natural language processing, such as identifying actions like rotation and translation through regular expressions and capturing the subsequent angle or distance values. More complex semantic parsing models can also be employed to gain a deeper understanding of the step text, addressing complex steps with varying expressions.
[0037] S103. Based on the dynamic parameters and the range of motion of the dynamic parameters, generate dynamic graphics corresponding to the sequence of problem-solving steps.
[0038] Dynamic graphics refer to vector animations or image sequences that can display the movement of geometric elements over time, such as a point moving along a line segment or a triangle rotating around a point. Dynamic graphics corresponding to a problem-solving step sequence refer to animations that present the geometric changes described in each step of the problem-solving process, allowing viewers to intuitively see the evolution of the graphics.
[0039] For example, the terminal can invoke a graphics rendering engine (such as a rendering module based on libraries like Manim or GeoGebra) to generate a series of consecutive frames based on the dynamic parameters and their range of motion extracted from each problem-solving step, simulating the smooth change of geometric elements according to specified parameters and ranges. For a sequence containing multiple problem-solving steps, the terminal can sequentially generate animation clips corresponding to each step and record the duration and content of each clip.
[0040] By generating dynamic graphics based on dynamic parameters and range of motion, the terminal can automatically create animation content that strictly matches the problem-solving logic, ensuring that every geometric operation is accurately visualized. This avoids the high cost and error-proneness of manually creating animations, thus enabling the low-cost, high-quality transformation of static geometry problems into dynamic teaching videos.
[0041] As an example, continuing with the rotating triangle, the terminal can use a mathematical plotting library to calculate the coordinates of the triangle vertices in each frame based on the extracted dynamic parameters (rotation angle from 0 to 90 degrees) and the range of motion, generating a smooth rotating animation lasting 2 seconds, in which the triangle rotates continuously from the initial position to the final position, visually displaying the rotation trajectory.
[0042] Optionally, when generating dynamic graphics, the terminal can automatically adjust the smoothness and frame rate of the animation based on the complexity of the problem-solving steps. For example, for fast movements, the frame rate can be reduced to decrease computation, while for fine operations, interpolation points can be added to make the animation smoother. It can also support user-defined animation styles, such as color and line type.
[0043] S104. Generate audio explanations corresponding to the sequence of problem-solving steps, synthesize dynamic graphics and audio explanations, and output dynamic geometry teaching videos.
[0044] Among them, "audio explanation" refers to audio narration generated using text-to-speech technology, corresponding to the problem-solving steps, used to explain the geometric operations and thought process at each step. "Dynamic geometry teaching video" refers to a playable video file obtained by synchronously synthesizing dynamic graphics and audio explanations, which may include visuals, sound, and possibly subtitles.
[0045] For example, the terminal can perform speech synthesis (such as TTS) on the text in the sequence of problem-solving steps to generate a speech segment corresponding to each step. Simultaneously, previously generated dynamic graphic segments can be aligned and synthesized with their corresponding speech segments according to the step order. During synthesis, the terminal can ensure that the playback duration of the speech matches the duration of the graphic animation, and adjust the animation timeline if necessary. Finally, all step segments are spliced together into a complete video file, and a dynamic geometry teaching video is output.
[0046] By automatically generating voice explanations and synthesizing them with dynamic graphics, the terminal can automate the entire process from problem-solving steps to complete teaching videos. This allows the video content to have both intuitive graphic changes and synchronized voice explanations, which can improve teaching effectiveness and content production efficiency. Thus, it can achieve low-cost, high-quality transformation of static geometry problems into dynamic teaching videos.
[0047] As an example, the terminal can convert the text corresponding to the solution step "rotate triangle ABC 90 degrees clockwise around point C" into speech "We rotate triangle ABC 90 degrees clockwise around point C" using a preset TTS engine, and obtain the speech duration as 2.5 seconds. Subsequently, the terminal can extend the previously generated 2-second rotation animation to 2.5 seconds through timeline resampling, achieving audio-visual synchronization. Finally, the audio-visual segments of all steps are spliced together to output a complete MP4 video file.
[0048] Optionally, when generating audio explanations, the terminal can refine and expand the original step text to make it more conversational. For example, "rotate 90 degrees" can be expanded to "Students, please look, now we will rotate this triangle 90 degrees clockwise around point C." When synthesizing videos, subtitles can also be overlaid, and the whiteboard content can be displayed in columns on one side of the screen.
[0049] In this embodiment, through steps S101 to S104, the terminal can acquire a structured sequence of problem-solving steps, extract the dynamic parameters and movement range of geometric elements, and then automatically generate corresponding dynamic graphics, while simultaneously generating synthesized audio explanations. This process can achieve automated production from static geometry problems to dynamic teaching videos without manual intervention, thereby reducing video production costs while ensuring strict consistency between graphic changes and problem-solving logic, as well as accurate audio-visual synchronization. This enables low-cost, high-quality conversion of static geometry problems into dynamic teaching videos.
[0050] In one exemplary embodiment, Figure 2 This application provides a flowchart illustrating the steps for extracting dynamic parameters of geometric elements and the range of motion of these dynamic parameters, as shown in the embodiments of this application. Figure 2 As shown, in step S102, the dynamic parameters of the geometric elements and the range of motion of the dynamic parameters are extracted from the sequence of problem-solving steps, including: S201. Using a predefined operation instruction template, identify the operation instructions in the problem-solving step sequence. S202. Based on the recognized operation instructions, extract the dynamic parameters of the geometric elements and the range of motion of the dynamic parameters.
[0051] Among them, the predefined operation instruction templates can refer to a pre-constructed set of rules used to match common geometric operation expressions. They can exist in the form of regular expressions, keyword libraries, or template strings, such as templates like "connect point A and point B" or "draw the perpendicular line to line L". Operation instructions can refer to descriptions of specific operations performed on geometric elements in the problem-solving steps, such as actions like connecting, rotating, translating, and drawing perpendicular lines. They can be the core semantic units that drive changes in the figure.
[0052] For example, after acquiring a structured sequence of problem-solving steps, the terminal can parse the text content of each step. The terminal can match the step text with a predefined library of operation instruction templates, for example, by using regular expressions to search the text for keywords such as rotation and translation, and whether it conforms to a specific sentence structure. Once a match is successful, the terminal can identify the operation instruction type corresponding to that step. Based on the parameter extraction rules corresponding to that operation instruction type, the terminal can extract specific geometric element names (e.g., "triangle ABC"), action parameters (e.g., 90 degrees), direction descriptions (e.g., clockwise), and implicit range of motion (e.g., the continuous change interval from the initial state to the target state) from the step text.
[0053] By using predefined operation instruction templates for rapid identification and parameter extraction, the terminal can efficiently and accurately process a large number of common and well-defined geometric operation steps, avoiding complex semantic parsing overhead and improving processing speed and stability. This enables low-cost, high-quality conversion of static geometry problems into dynamic teaching videos.
[0054] As an example, in one specific implementation, the terminal's built-in instruction template library includes rotation templates. When the terminal processes the step "rotate triangle ABC 90 degrees clockwise around point C," the template matches successfully, recognizing the operation instruction as rotation, and extracting the geometric elements as triangle ABC, with point C as the center of rotation, clockwise direction, and an angle of 90 degrees. Combining this with geometric common sense, the terminal can determine the range of motion as a continuous angular variation from 0 degrees to 90 degrees for subsequent rendering.
[0055] Optionally, predefined operation instruction templates can be dynamically expanded and updated according to actual application scenarios. For example, when the system encounters new operation expressions, new operation instruction templates can be generated and added to the library through manual annotation or machine learning, enabling the system to continuously learn and adapt. In addition, operation instruction template matching can adopt a priority strategy, prioritizing the matching of more specific and longer templates to avoid mismatches caused by short templates.
[0056] In this embodiment, by using predefined operation instruction templates for identification and extraction, the terminal can quickly respond to high-frequency standard instructions in the problem-solving step sequence, thereby improving the efficiency and accuracy of dynamic parameter extraction, ensuring the real-time and smoothness of the video generation process, and thus enabling low-cost and high-quality conversion of static geometry problems into dynamic teaching videos.
[0057] In one exemplary embodiment, Figure 3 A flowchart illustrating another step for extracting dynamic parameters of geometric elements and the range of motion of these dynamic parameters, as provided in this embodiment of the application, is shown below. Figure 3 As shown, combined with Figure 2 In step S102, such as Figure 3 As shown, step S102, which involves extracting the dynamic parameters of the geometric elements and their range of motion from the solution sequence, also includes: S203. When there is an unrecognized problem-solving step sequence in the operation instruction template, the unrecognized problem-solving step sequence is semantically annotated using a pre-trained recognition language model to identify the dynamic parameters of the problem-solving step sequence and the range of motion of the dynamic parameters.
[0058] Among these, "unrecognized problem-solving step sequences" refers to text steps that cannot be successfully matched using predefined regular expressions, keyword libraries, or template strings. This can be due to reasons such as complex expression, novel sentence structure, the presence of pronouns, or omissions. "Pre-trained recognition language model" refers to a deep learning model, such as BERT or GPT, that has been pre-trained on a large amount of mathematical text or general corpus and fine-tuned for geometric operation recognition tasks. This model is capable of understanding the deep semantics of the text. "Semantic annotation" refers to the process of assigning a semantic role label to each word or phrase in the text, such as agent, action, patient, time, place, and manner, thereby structurally understanding the meaning of the sentence.
[0059] For example, when a terminal fails to match a problem-solving step using a predefined operation instruction template, it can send that step to a pre-trained language recognition model for processing. This model can perform deep semantic understanding of the step, outputting structured semantic annotations, such as annotating the core action, the agent of the action, the receiver of the action, and related modifiers. Based on these semantic annotations and predefined mapping rules, the terminal can infer the geometric operation type (e.g., rotation, translation, drawing a perpendicular line) corresponding to the step, and extract the involved geometric elements, specific dynamic parameters (e.g., angle, distance), and implicit range of motion (e.g., continuous change from the current position to the target position). In this way, the terminal can handle complex or novel expressions that templates cannot cover.
[0060] As an example, in one specific implementation, the terminal receives the text of a problem-solving step: "Draw a straight line through point A such that the line is parallel to line l and tangent to circle O." This problem-solving step is complex and difficult to cover with a simple regular expression template. In this case, the terminal can feed the text into a semantic role labeling model based on BERT fine-tuning. After analysis, the model outputs the labeling results: the action is "make," the object is "a straight line," and the additional conditions include "through point A" (agent or condition), "parallel to line l" (attribute 1), and "tangent to circle O" (attribute 2). Based on this structured information and a geometric knowledge base, the terminal can infer that this is a complex operation of "drawing a parallel and tangent straight line" and extract the involved geometric elements: point A, line l, circle O, and dynamic parameters—the generation process of the straight line (from nothing to something, ultimately satisfying two constraints). Its range of motion can be understood as an iterative process of gradually adjusting from an initial guessed position to a position that satisfies all constraints.
[0061] Optionally, the language recognition model is not limited to semantic role labeling; it can also directly employ sequence-to-sequence generation models. For example, the text of problem-solving steps can be directly mapped into structured operation instructions and parameter lists. For instance, training a T5 or BART model with the input "Draw a straight line through point A such that the line is parallel to line l and tangent to circle O" will output "Operation type: Draw a straight line; Constraints: [through point A, parallel to line l, tangent to circle O]". Furthermore, after the model outputs, the terminal can also combine a geometric rule engine to verify the reasonableness of the extracted parameters, such as checking whether point A lies on line l, to ensure the accuracy of the parsing results.
[0062] In this embodiment, by using a pre-trained language recognition model to process complex steps not covered by the template, the terminal can achieve a deep understanding and accurate parsing of various flexible expressions in natural language, ensuring that dynamic information in the problem-solving steps is completely and accurately extracted, thereby enabling low-cost and high-quality conversion of static geometry problems into dynamic teaching videos.
[0063] In an exemplary embodiment, step S203, which involves semantically annotating the unrecognized problem-solving step sequence using a pre-trained recognition language model, includes: By recognizing the language model, semantic role labeling is performed on the unrecognized problem-solving step sequences to identify the agent role, predicate description, and patient role contained in the problem-solving step sequences; Based on the results of semantic role labeling, determine the dynamic type and dynamic parameters of the operation corresponding to the sequence of problem-solving steps; Among them, the operation dynamic type is used to characterize the transformation category of geometric elements in the problem-solving step sequence and to determine the parameter extraction method corresponding to the dynamic parameters.
[0064] Semantic role labeling refers to the process of assigning semantic role labels related to the predicate verb to each word or phrase in the text. This includes labeling who performs the action (agent), what the action itself is (predicate), who receives the action (patient), and additional components such as time, place, and manner. The agent role refers to the executor or initiator of the action; in geometric operations, this usually refers to the geometric object performing the operation or the implicit subject of the operation. The predicate description refers to the core action or state change vocabulary in the sentence, such as rotation, connection, and drawing a perpendicular line. The patient role refers to the recipient of the action or the object affected, such as a rotated triangle or a connected line segment. Operation dynamic type refers to the category of geometric transformation determined based on the predicate description and context, such as translation, rotation, scaling, and reflection. The operation dynamic type characterizes how to subsequently analyze and quantify dynamic changes. Parameter extraction methods refer to the different strategies or rules used for different operation dynamic types to extract specific numerical parameters (such as angle, distance, and direction) from the text and determine the range of variation for these parameters.
[0065] For example, when a terminal inputs an unrecognized sequence of problem-solving steps into a language recognition model, the model can perform deep semantic analysis on the step text, outputting semantic role labels for each component. Based on these labels, the terminal identifies the core predicate description of the sentence, as well as the agent and patient roles associated with that predicate. Based on the identified predicate description, the terminal can determine the operation dynamic type corresponding to the step using mapping rules; for example, if the predicate is rotation, the operation dynamic type is rotation transformation. The terminal can then extract specific dynamic parameter values from the step text based on the preset parameter extraction method corresponding to this operation dynamic type, combined with the patient role and other labeled additional components (such as angle and direction), and infer the range of motion of the parameters.
[0066] As an example, in one specific implementation, the terminal receives an unrecognized step: "Rotate triangle ABC 90 degrees clockwise around point C." The recognition language model performs semantic role labeling on this step, outputting: the predicate description is rotation, the patient is triangle ABC, and the additional components include around point C (labeled as the reference point), clockwise (labeled as the direction), and 90 degrees (labeled as the angle). Based on the predicate (rotation), the terminal determines the operation dynamic type to be a rotation transformation. According to the parameter extraction method for rotation transformations, it extracts geometric elements from the patient (triangle ABC), and from the additional components, it extracts the rotation center as point C, the rotation direction as clockwise, and the rotation angle as 90 degrees, defining the range of motion as a continuous variation from 0 degrees to 90 degrees.
[0067] Optionally, the language recognition model can be a semantic role labeling model, which can be based on pre-trained models such as BERT or RoBERTa and fine-tuned on mathematical domain corpora labeled with semantic roles. For cases where the referent is ambiguous or the subject is omitted, the language recognition model can infer from the context, for example, treating implicit "we" or "please" as the default agent role. Furthermore, the mapping relationship between operation dynamic types and parameter extraction methods can be stored in a configuration table for easy dynamic updates and expansion.
[0068] In this embodiment, fine-grained semantic parsing is achieved through semantic role labeling. The terminal can penetrate complex sentence structures and accurately identify the core actions and participants in the steps, providing accurate structured information for subsequent parameter extraction and graph generation. This enables low-cost, high-quality conversion of static geometry problems into dynamic teaching videos.
[0069] In one exemplary embodiment, Figure 4 A flowchart illustrating the confidence assessment steps included in a dynamic geometry teaching video provided in this application embodiment is shown below. Figure 4 As shown, the method also includes: S401. Determine the confidence level of the identified dynamic parameters and evaluate the confidence level of the dynamic parameters. The confidence level evaluation combines semantic confidence level, geometric rationality confidence level, and system feasibility confidence level. S402. If the confidence level of the dynamic parameters is lower than the preset threshold, then a static graphic is generated based on the dynamic parameters and the static graphic is marked. Among them, static graphics are used to represent the static display of the steps corresponding to the problem-solving steps sequence, and markers are used to indicate the review of the step-by-step teaching of the problem-solving steps sequence.
[0070] Among these, the confidence level of dynamic parameters refers to a quantitative measure of the accuracy and reliability of dynamic parameters extracted from the problem-solving steps, reflecting the degree to which the parameter can correctly drive geometric changes. Confidence assessment refers to the process of performing multi-dimensional verification of dynamic parameters and calculating their overall confidence level. Semantic confidence level refers to the probability value output by the recognition language model or rule template when parsing the step text, reflecting the reliability of the semantic understanding of the text. Geometric rationality confidence level refers to the verification result of whether the dynamic parameters conform to Euclidean geometry axioms and the initial conditions of the problem, such as checking whether the sum of the interior angles of a triangle is 180 degrees, or whether a point lies on a specified line segment. System feasibility confidence level refers to the pre-check result of the graphics rendering engine on the executability of the current dynamic parameter instructions, such as whether the parameters are within the numerical range supported by the engine, or whether the graphic elements exist. A preset threshold refers to a configurable numerical limit used to determine whether dynamic parameters are reliable; if the value falls below this threshold, a degradation process can be triggered. Static graphics refer to geometric images that do not contain dynamic changes and only show the final state of the current step. The marker can refer to a special identifier attached to a static graphic or step metadata, used to prompt human reviewers that the step needs to be reviewed.
[0071] For example, after extracting dynamic parameters from the problem-solving steps, the terminal may not directly use them for graphics generation. Instead, it may first perform a multi-dimensional confidence assessment. The terminal obtains the semantic confidence score generated during the semantic parsing stage, and simultaneously calls the geometric rule engine to check whether the parameters conform to geometric constraints and calculate the geometric rationality confidence score. It then pre-checks the executability of the instructions with the rendering engine and obtains the system feasibility confidence score. The terminal can combine these three confidence scores according to preset weights or rules to obtain a total confidence score. If this score is lower than a preset threshold, the terminal determines that the dynamic parameters of the current step are unreliable and unsuitable for generating dynamic graphics. In this case, the terminal can generate a static graphic corresponding to that step, that is, only display the final state of the geometric elements after the step is completed, and add a mark indicating that it needs to be reviewed to the graphic or the corresponding step data for subsequent manual review and processing.
[0072] As an example, in one specific implementation, the terminal extracts the rotation angle of 90 degrees from the problem-solving step "rotate triangle ABC 90 degrees around point C," with a semantic confidence level of 0.95. However, geometric plausibility verification reveals that, according to the initial conditions of the problem, triangle ABC is a right triangle, and point C is the right-angle vertex. After rotation, the new triangle overlaps with the original figure, resulting in a geometric plausibility confidence level of only 0.3. System feasibility pre-check shows that the rendering engine can execute this instruction, with a confidence level of 1.0. The terminal calculates the overall confidence level using a weighted average (e.g., semantic 0.3, geometric 0.5, system 0.2): 0.95 × 0.3 + 0.3 × 0.5 + 1.0 × 0.2 = 0.285 + 0.15 + 0.2 = 0.635. If the preset threshold is 0.7, the confidence level is determined to be below the threshold. Instead of generating a rotation animation, the terminal generates a static graphic displaying the rotated triangle, adds a red "requires review" marker in the corner of the graphic, and generates a review work order record.
[0073] Optionally, the dimensions of confidence assessment can be further expanded, for example, by adding context consistency confidence to check the logical coherence between the parameters of this step and the preceding and following steps. The overall confidence score can be calculated using weighted summation, product, or more complex machine learning models. The preset threshold can be dynamically adjusted according to business needs; for example, the threshold can be increased for scenarios with high security requirements. In addition to generating static graphics and labels, the degradation process can also pause the automatic generation of subsequent steps, waiting for manual review of the results before continuing.
[0074] In this embodiment, through confidence assessment and degradation processing mechanisms, the terminal can automatically identify and isolate possible parsing errors or logical contradictions, ensuring that the final output teaching video content is accurate and reliable. At the same time, through marking to guide human intervention for optimization, a closed-loop improvement process of human-machine collaboration is formed, thereby enabling low-cost and high-quality transformation of static geometry problems into dynamic teaching videos.
[0075] In some exemplary embodiments, the confidence assessment of dynamic parameters can be implemented using a comprehensive scoring function. Specifically, the terminal calculates the comprehensive confidence of the dynamic parameters. The overall confidence score is composed of semantic confidence score. Geometric rationality confidence level and system feasibility confidence The weighted summation yields the following formula:
[0076] in, , , These are the weight coefficients corresponding to semantic confidence, geometric rationality confidence, and system feasibility confidence, respectively, and they satisfy the following conditions: The specific values of each weight coefficient can be preset or dynamically adjusted according to the importance attached to the three dimensions in the actual application scenario. Semantic confidence. This can be derived from the probability values output by the language model during the text parsing process, reflecting the model's grasp of the text; geometric plausibility confidence score. The dynamic parameters can be obtained by verifying whether they conform to Euclidean geometry axioms and the initial conditions of the problem through a geometric rule engine. For example, it can be used to verify whether the sum of the interior angles of a triangle is 180 degrees, or whether a point lies on a specified line segment. The system feasibility confidence level can also be assessed. The rendering engine can perform an executability pre-check on the dynamic parameter instructions to generate the image, such as checking whether the parameters are within the range of values supported by the engine and whether geometric elements exist. The terminal compares the calculated overall confidence score S with a preset threshold T. If S < T, the confidence score of the dynamic parameters is deemed insufficient, triggering a downgrade process, which generates a static graphic and adds a verification mark.
[0077] In one exemplary embodiment, the method further includes, prior to synthesizing motion graphics and audio narration: Get the duration of the audio explanation; Using the duration of the speech as a constraint, the interpolation step size is determined based on the range of motion of the dynamic parameters, and duration-matched animation clips are generated so that the animation duration of the dynamic graphics matches the duration of the speech narration.
[0078] The audio duration of the narration refers to the total duration of the audio file generated through text-to-speech technology, measured in seconds or milliseconds. Using audio duration as a constraint means using the audio duration as a time reference or limitation for animation generation, requiring the animation playback duration to strictly match the audio duration. Adjusting the animation duration of motion graphics refers to modifying the timeline of already generated or currently being generated animation segments, such as by changing the playback speed, inserting or deleting intermediate frames, to make its total duration equal to the audio duration. Synchronization between motion graphics and audio narration means that in the final video playback, the changes in geometric shapes on the screen are precisely aligned in time with the audio narration, so that viewers do not perceive any delay or misalignment where the screen is faster or slower than the sound.
[0079] For example, after generating the audio narration for each step, the terminal obtains the precise duration data of the audio file. When generating the corresponding motion graphics, the terminal passes the audio duration as an input parameter to the graphics rendering engine. Based on this duration and the motion range of the motion parameters, the rendering engine calculates the time interval of each frame of the animation, thereby generating an animation clip with a total duration exactly matching the audio duration. If the animation has already been pre-generated, the terminal can also post-process and adjust the existing animation according to the audio duration, for example, by using timeline resampling technology to stretch or compress the animation's timeline to match its duration with the audio duration. In this way, the terminal ensures that the visual changes and audio narration are presented synchronously in the final synthesized video.
[0080] Specifically, timeline resampling can be performed based on the audio duration and the original duration of the motion graphics to determine the difference step size. The original duration of the motion graphics refers to the total playback time occupied when the motion graphics are generated at the default motion speed and frame rate without considering audio duration constraints. Timeline resampling refers to resampling an existing animation frame sequence on the timeline. By proportionally stretching or compressing the timeline and increasing or decreasing the number of intermediate frames, the total playback duration of the animation can be changed while maintaining the relative order of the motion trajectories of geometric elements and key change points in the animation.
[0081] After generating a motion graphic, or during the generation process, the terminal obtains the current original duration of the motion graphic. Simultaneously, the terminal can obtain the duration of the corresponding audio narration. The terminal compares these two durations; if they are inconsistent, the motion graphic undergoes timeline resampling. Specifically, the terminal can calculate the ratio of the audio duration to the original duration as a resampling factor. If the audio duration is longer than the original duration, the terminal generates new intermediate frames by interpolating between existing frames, increasing the total number of frames and slowing down the animation; if the audio duration is shorter than the original duration, the terminal reduces the total number of frames by uniformly extracting some frames, speeding up the animation. After resampling, the new duration of the motion graphic equals the audio duration, thus achieving audio-visual duration matching.
[0082] As an example, in one specific implementation, the terminal generates an animation of a rotating triangle, originally 2 seconds long and consisting of 50 frames. The corresponding audio narration is 2.5 seconds long. The terminal calculates a resampling factor of 2.5 / 2 = 1.25, meaning the animation needs to be stretched to 1.25 times its original length. The terminal uses a linear interpolation algorithm, inserting a certain number of new frames between every two frames, increasing the total number of frames from 50 to 50 × 1.25 = 62.5 frames, rounded down to 63 frames. The newly generated animation has a playback duration of 63 frames / 25 frames / second = 2.52 seconds, approximately matching the audio duration. In the adjusted animation, the starting and ending positions of the triangle rotation remain unchanged, and the movement is smoother and slower, consistent with the rhythm of the audio.
[0083] In a specific example, the terminal can use a TTS engine to convert the step text "We will rotate triangle ABC 90 degrees clockwise around point C" into speech, obtaining a speech duration of 2.5 seconds. When generating the rotation animation corresponding to this step, the terminal passes the target duration of 2.5 seconds to the rendering engine. Based on the rotation angle range from 0 degrees to 90 degrees and the set frame rate (e.g., 25 frames / second), the rendering engine calculates that a total of 2.5 seconds × 25 frames / second = 62.5 frames need to be generated, rounded to 63 frames. Simultaneously, it calculates the angle increment between each frame as 90 degrees / 62 ≈ 1.45 degrees / frame, thus generating a smooth rotation animation that plays for exactly 2.5 seconds, achieving audio-visual synchronization.
[0084] Optionally, adjusting the animation duration can include dynamic frame interpolation or frame extraction. For example, if the audio duration is longer than the original animation duration, the terminal can insert additional intermediate frames between the keyframes of the animation using an interpolation algorithm, making the animation slower and longer; if the audio duration is shorter than the original animation duration, the terminal can evenly extract some intermediate frames, making the animation faster and shorter. Furthermore, keyframe-based timeline mapping technology can be used to ensure that key change points in the animation (such as rotation to position, a point reaching its endpoint) are time-aligned with corresponding keywords in the audio (such as rotating 90 degrees, moving to point B), further improving the precision of synchronization. Optionally, timeline resampling can employ various interpolation algorithms, such as linear interpolation and spline interpolation, to ensure the smoothness of the animation. For animations containing key events (such as a point reaching a specific position), resampling can prioritize ensuring the time alignment of keyframes to avoid misalignment between key events and audio keywords. Additionally, resampling can be combined with the dynamic rendering process, directly calculating the parameter increment for each frame based on the audio duration during rendering, thereby avoiding the additional computational overhead of post-processing resampling.
[0085] In this embodiment, by adjusting the duration of dynamic graphics through timeline resampling technology, the terminal can flexibly adapt to different durations of voice explanations without changing the movement logic of geometric elements, ensuring the accuracy of audio-visual synchronization. This not only improves the efficiency of video production but also guarantees the professionalism of the final output content and user experience, thereby enabling the low-cost, high-quality conversion of static geometry problems into dynamic teaching videos.
[0086] In one exemplary embodiment, generating a voice explanation corresponding to the sequence of problem-solving steps includes: Each step of the problem-solving sequence is divided into an explanation script; The explanation script is converted from text to speech and synthesized to obtain the audio explanation corresponding to the sequence of problem-solving steps.
[0087] Here, "each step text" refers to the original text description corresponding to each step in the problem-solving sequence, such as "connect AB," "draw a perpendicular line," or "rotate the triangle 90 degrees around point C," usually presented in natural language. "Explanation script" refers to text content derived from the step text that is suitable for audio reading. It may be obtained by refining, expanding, or reorganizing the original step text to make it more suitable for auditory comprehension. "Text-to-speech synthesis" refers to the process of inputting the text-based explanation script into a speech synthesis engine to automatically generate a corresponding audio file containing human-like voice content. "Audio explanation" refers to the final generated audio explanation corresponding to the problem-solving steps, used in the video to explain the operation and thought process of the current step to the learner.
[0088] For example, after acquiring the structured sequence of problem-solving steps, the terminal processes the text of each step. Based on preset rules or models, the terminal transforms the original step text into an explanation script more suitable for voice playback, such as adding guiding phrases, expanding abbreviations, and converting symbols into natural language expressions. The terminal can then input the generated explanation script into a text-to-speech synthesis engine, which generates corresponding audio files and records the duration of each step's audio. By traversing all steps, the terminal ultimately obtains a set of step-by-step voice explanations corresponding to the entire problem-solving sequence.
[0089] As an example, in one specific implementation, the sequence of problem-solving steps processed by the terminal includes a step text: "Rotate triangle ABC 90 degrees clockwise around point C." The terminal converts this text into an explanation script: "We rotate triangle ABC 90 degrees clockwise around point C," adding the subject "we" to make the language more conversational. Subsequently, the terminal inputs this explanation script into a text-to-speech synthesis engine to generate a 2.5-second audio file as the voice explanation for this step.
[0090] Optionally, the generated explanation script can include various polishing methods, such as converting the mathematical symbol "⊥" to "perpendicular to", "△ABC" to "triangle ABC", or adding conjunctions such as "next" or "then" before each step to make the entire explanation sound more coherent. The text-to-speech synthesis engine can employ various technologies, such as parameter-based synthesis or waveform concatenation-based engines, and can also select different timbres, speech rates, and intonations to adapt to different teaching scenarios and user preferences. In addition, the generated speech explanation can be stored step by step, facilitating duration matching and alignment during subsequent step-by-step synthesis.
[0091] In this embodiment, by converting the step text into an explanation script and performing speech synthesis, the terminal can automatically generate a natural and fluent voice explanation that is synchronized with the problem-solving logic, avoiding the high cost and tedious post-production of manual recording. This enables the low-cost and high-quality conversion of static geometry problems into dynamic teaching videos.
[0092] In one exemplary embodiment, Figure 5 This application provides a flowchart illustrating the steps of synthesizing dynamic graphics and voice narration in an embodiment of the present application. Figure 5 As shown, step S104, which involves synthesizing dynamic graphics and audio narration, also includes: S501. Divide the text of each step in the problem-solving sequence into display text; S502. Generate a display page corresponding to each step. The display page includes a graphic area for displaying dynamic graphics and a whiteboard area for displaying display text. S503. Use the display page as the background and overlay dynamic graphics on the graphic area.
[0093] The presentation text refers to text extracted from the problem-solving steps and suitable for static presentation on the video screen. It can be a concise summary of the steps, such as connecting A and B, drawing a perpendicular line, or rotating 90 degrees. The presentation page refers to a static image or interface layout generated for each problem-solving step, specifying the overall structure and element placement of the video screen. The graphic area refers to a specific area reserved on the presentation page for displaying dynamic geometric figures, which can occupy a major portion of the screen. The blackboard area refers to a specific area reserved on the presentation page for displaying the presentation text, usually located on one side of the screen, simulating a classroom blackboard or whiteboard. Using the presentation page as a background and overlaying dynamic graphics onto the graphic area means that during video compositing, the presentation page is used as the underlying image, and then the dynamic graphics are overlaid on the graphic area of the presentation page in the form of a video stream or animation sequence to form a complete image.
[0094] For example, when processing each problem-solving step, the terminal can extract display text for whiteboard presentation from the step text. This can be concise text retaining the core operation description. The terminal can generate a display page image containing a graphic area and a whiteboard area based on a preset layout template. The whiteboard area contains the display text for that problem-solving step. During the video compositing stage, the terminal can use this display page as a background layer and the previously generated dynamic graphic of the corresponding step as a foreground layer, precisely placing the dynamic graphic within the graphic area of the background page for overlay. In this way, the final image contains both static whiteboard text and dynamic geometric changes, presented collaboratively without interference.
[0095] As an example, in one specific implementation, the sequence of problem-solving steps processed by the terminal may include a step text: "Rotate triangle ABC 90 degrees clockwise around point C." The terminal extracts the display text "Rotation center: C; Rotation angle: 90° (clockwise)" from this text. Subsequently, the terminal generates a display page with a graphic area on the left and a whiteboard area on the right, displaying the aforementioned text. When synthesizing the video, the terminal overlays the previously generated triangle rotation animation as a dynamic graphic onto the left graphic area of the display page, forming the final image. Viewers can see the textual explanation of the steps on the right and the dynamic process of the triangle rotation on the left.
[0096] Optionally, the layout ratio and position of the graphic area and the blackboard area can be adjusted according to actual needs. For example, the graphic area can occupy two-thirds of the screen, and the blackboard area can occupy one-third, or it can be dynamically adjusted according to the characteristics of the question. The display page can be generated separately for each step, or after the basic layout is generated in the first step, subsequent steps can only update the content of the blackboard area to save computing resources. In addition, the blackboard area can be updated incrementally, that is, the display text of the previous steps is retained, and the text of the current step is highlighted, simulating the experience of writing out the solution process step by step in the classroom.
[0097] In this embodiment, by generating a display page that includes a graphic area and a whiteboard area, and overlaying dynamic graphics on the graphic area, the terminal can achieve separate layout and collaborative display of graphic demonstrations and text explanations in teaching videos. This allows learners to focus on the dynamic changes and key knowledge points simultaneously, improving the teaching effect and viewing experience of the video. As a result, it is possible to achieve low-cost and high-quality transformation of static geometry problems into dynamic teaching videos.
[0098] In one exemplary embodiment, the display page adopts a column layout, with a graphic area on one side and a whiteboard area on the other side; The whiteboard area displays the presentation text for each step in an append-only manner, and highlights the presentation text for the current step.
[0099] The layout includes several key elements: **Columnar Layout:** This refers to dividing the visible area of a page into two or more parallel columns, each containing different types of content, creating a partitioned information display. **Graphic Area:** This refers to a column within the columnar layout specifically designed to present dynamic geometric shapes. It can occupy a large portion of the page to ensure clear visibility of graphic changes. **Blackboard Area:** This refers to a column within the columnar layout specifically designed to display textual explanations of the problem-solving steps. It can be located on the side of the page, simulating the writing area of a blackboard in a real classroom. **Addition Method:** This involves gradually accumulating the text for each step in the blackboard area as the problem-solving process progresses. New steps are added after existing text, retaining previous content to form a complete record of the problem-solving process. **Highlighting:** This involves using special visual styles to highlight the text for the current step, such as changing the text color, bolding the font, adding a background color, or adding a border, making it easily identifiable and focusable among numerous historical steps.
[0100] For example, when generating the display page for each step, the terminal can use a preset column layout template to divide the page into a graphic area and a whiteboard area. In the whiteboard area, the terminal can maintain a progressively accumulating list of text. For the first step, the terminal writes the display text for that step at the beginning of the whiteboard area. For subsequent steps, the terminal identifies the current step's number and appends the display text for that step below the existing text in the whiteboard area, while keeping the text content of all previous steps unchanged. Furthermore, the terminal can apply a preset highlighting style to the newly added text for the current step, such as setting it to bold yellow font or adding a light background, to visually distinguish it from the ordinary text of previous steps. In this way, as the steps progress, the whiteboard area gradually accumulates the complete solution process, and the current operation step remains prominently displayed.
[0101] As an example, in a specific implementation, the terminal processes a triangle rotation problem in three steps. First, the displayed text is "1. Connect OA." The terminal generates a display page, showing a triangle in the left graphic area and "1. Connect OA" in the right whiteboard area, both highlighted. Second, the displayed text is "2. Construct OB perpendicular to AC." The terminal updates the display page, overlaying a new animation in the left graphic area, and appending "2. Construct OB perpendicular to AC" below "1. Connect OA" in the right whiteboard area, highlighting "2. Construct OB perpendicular to AC" in bold yellow, while "1. Connect OA" reverts to normal black font. Third, the displayed text is "3. Rotate the triangle 90 degrees around point C." The terminal updates again, appending "3. Rotate the triangle 90 degrees around point C" to the whiteboard area and highlighting it. The text for the first two steps is retained in its normal style. In the final video, the right whiteboard fully records the three-step solution process, with the currently explained third step always prominently displayed.
[0102] Optionally, the left-right proportion of the column layout can be dynamically adjusted according to the content. For example, when the graphics are more complex, the area occupied by the graphics can be increased; when the text is longer, the whiteboard area can be appropriately increased. Adding text can be done through simple text additions, or by using a mind map or tree structure to display the logical relationships between steps. Highlighting methods are not limited to color and bolding; they can also include dynamic flashing, underlining, and adding an arrow to the left. Furthermore, for text lists exceeding the length of the whiteboard area, scrollbars or pagination can be automatically enabled to ensure that all historical text remains accessible.
[0103] In this embodiment, a column layout is used to achieve the coordinated display of graphics and blackboard writing in different sections. The accumulation of problem-solving process and the current focus are presented through appending and highlighting. The terminal can provide learners with a clear information hierarchy and a coherent learning path, simulating the teaching experience of a real classroom. This enables the low-cost and high-quality transformation of static geometry problems into dynamic teaching videos.
[0104] In an exemplary embodiment, obtaining a sequence of solution steps for a geometry problem with a preset structural format includes: Obtain the text description of the geometry problem and generate the initial geometric figure based on the text description; Using a pre-trained parsing language model, a sequence of problem-solving steps with a preset structure format is generated based on text descriptions and initial geometric figures.
[0105] The text description of the geometry problem refers to a mathematical geometry problem presented in natural language, which may include known conditions, geometric elements and their relationships, and the objective to be solved or proven, such as "In triangle ABC, AB=AC, point D is the midpoint of BC, connect AD, prove that AD is perpendicular to BC." The initial geometric figure refers to a basic vector graphic generated from the static geometric solids and constraints in the text description to represent the initial state of the problem, containing geometric objects such as points, lines, and surfaces and their precise coordinates. The pre-trained analytical language model refers to a large language model fine-tuned from a large corpus of mathematical geometry, capable of understanding geometric problems and generating solution strategies. The pre-set structural format refers to the organization of solution steps in a machine-readable manner with clearly defined fields or tags, such as structured storage according to step number, operation type, operation object, dynamic parameters, etc. The solution step sequence refers to a set of multiple steps arranged in a logical order, where each step describes a specific geometric operation or reasoning action.
[0106] For example, the terminal can acquire a text description of a geometric problem input by the user or read from a problem bank. The terminal parses the text description, identifying the static geometric objects it contains, such as points, lines, and triangles, as well as their positional relationships and constraints, such as equal length, perpendicularity, and parallelism. Based on this information, the terminal establishes a geometric coordinate system, calculates the coordinates of each geometric object, and generates an initial geometric figure that is completely consistent with the initial conditions of the problem. The terminal can input the problem text description and the generated initial geometric figure data (such as coordinate information) together into a pre-trained parsing language model. This parsing language model can be specifically trained in the field of mathematical geometry, enabling it to understand the problem requirements and deduce the solution steps. The model outputs a structured sequence of solution steps, each step containing information such as operation type, operation object, and parameters, organized in a preset format for use by subsequent modules.
[0107] As an example, in one specific implementation, the terminal receives the problem text: "In right triangle ABC, ∠C = 90°. Rotate triangle ABC 90 degrees clockwise around point C to obtain triangle A'B'C." The terminal can first parse the text, identify the static geometric figures: points A, B, and C, and the constraint: ∠C = 90°, generate an initial triangle shape, and determine the coordinates of each point, for example, setting C as the origin, A on the x-axis, and B on the y-axis. Subsequently, the terminal can input the problem text and the initial shape coordinate data into an analytical language model (such as a mathematical version model based on ChatGLM or GPT) that has been fine-tuned with a mathematical problem-solving corpus. The analytical language model, after inference, outputs a structured sequence of solution steps, for example, represented in JSON format. This structured sequence of solution steps clearly expresses the operation type and parameters of each step.
[0108] Optionally, when generating the initial geometry, a rule engine combined with computational geometry algorithms can be used to ensure the accuracy and efficiency of the drawing. The parsing language model can be a general-purpose LLM fine-tuned with mathematical instructions, or it can be a general-purpose large model that generates structured output through few-shot learning. The preset structure format of the generated solution step sequence can be common data exchange formats such as JSON, XML, and YAML, facilitating parsing and processing by subsequent modules. In addition, when generating the step sequence, the model can also refer to the coordinate information in the initial figure to avoid generating operations that contradict the figure, such as not generating a conclusion that is not isosceles in an isosceles triangle.
[0109] In this embodiment, by generating initial geometric figures based on text descriptions, and then combining the graphic data with an analytical language model to generate structured problem-solving steps, the terminal can ensure that the generated problem-solving steps are strictly consistent with the geometric constraints of the problem, avoiding geometric ambiguities that may arise from pure text parsing. This provides an accurate and reliable logical foundation for subsequent dynamic video generation, thereby enabling low-cost and high-quality conversion of static geometric problems into dynamic teaching videos.
[0110] In one exemplary embodiment, the parsed language model is obtained by fine-tuning a mathematical vertical domain model.
[0111] A mathematics vertical domain model refers to a specialized base model obtained by further pre-training on a large-scale mathematical corpus (such as mathematics textbooks, exercise solutions, and geometry proofs) based on a general-purpose large-scale language model. Its parameters have been optimized for mathematical semantics and logic. Fine-tuning refers to further training the model on the basis of the mathematics vertical domain model using labeled data for specific tasks (such as generating geometric problem-solving steps) to make the model output more consistent with the preset structural format and problem-solving logic.
[0112] For example, the parsing language model used by the terminal can be a specially customized LLM model. A mathematics vertical domain model is constructed, which can be obtained by further pre-training on a general LLM model using massive amounts of text such as mathematics textbooks, mathematical papers, and problem-solving solutions, giving it a deeper understanding of mathematical terminology, symbols, and logical relationships. Subsequently, based on this vertical domain model, it can be fine-tuned using a large amount of manually annotated geometry problems and paired data of structured solution steps. Through fine-tuning, the mathematics vertical domain model learns how to combine the input geometry problem text with the geometric conditions in the problem to output solution steps conforming to a preset structural format (such as a JSON-formatted sequence of steps). The final deployed parsing language model is this finely tuned model.
[0113] In this embodiment, by using an analytical language model fine-tuned from a mathematical vertical domain model, the terminal can acquire semantic understanding and logical reasoning capabilities specifically optimized for mathematical geometry tasks. This can improve the accuracy, professionalism, and structure of the generated problem-solving step sequences, thereby enabling low-cost, high-quality conversion of static geometry problems into dynamic teaching videos.
[0114] In one exemplary embodiment, generating an initial geometry based on a text description includes: Analyze the static geometry in the text description to determine the type, positional relationships, and constraints of the static geometry; Based on the type, positional relationship, and constraints of the static geometric body, a geometric coordinate system is established, and the coordinate data of the static geometric body is determined to obtain the initial geometric figure.
[0115] In this context, static geometric objects refer to geometric objects that do not change over time or exist as initial conditions in the problem description, such as fixed points, line segments, triangles, and circles. These can be the basic elements for constructing the initial figure. The type of static geometric object refers to its category attribute, such as point, line, line segment, ray, angle, triangle, quadrilateral, and circle. Different types of entities have different geometric representations. Positional relationships refer to the relative orientation of static geometric objects in space, such as a point on a straight line, intersecting line segments, two parallel or perpendicular lines, or a point inside or outside a circle. Constraints refer to mathematical restrictions imposed on static geometric objects, such as equal line segment lengths, angles of specific values, or a point being the midpoint of a line segment. These conditions determine the accuracy of the figure construction. A geometric coordinate system refers to the reference frame established to determine the position of geometric objects, usually a two-dimensional Cartesian coordinate system, but a three-dimensional coordinate system can also be established as needed for the problem. Coordinate data can refer to the specific positional values of a geometric object in the selected coordinate system, such as the coordinates of a point, the equation parameters of a line, the center coordinates and radius of a circle, etc.
[0116] For example, after acquiring the text description of a geometry problem, the terminal can perform natural language parsing to identify all the static geometric objects contained within and determine the type of each entity. Simultaneously, the terminal can extract the positional relationships and constraints between these entities from the text. For instance, it can identify that "triangle ABC" means that points A, B, and C form a triangle; that "AB=AC" means that line segments AB and AC are equal in length; and that "point D is the midpoint of BC" means that point D is located at the midpoint of line segment BC. Based on this information and combined with geometric modeling rules, the terminal selects a suitable geometric coordinate system, setting the origin at a convenient location, such as the midpoint of the base of an isosceles triangle or a right-angle vertex. The terminal can then use computational geometry algorithms to solve for the coordinates of each entity that satisfies all constraints, such as calculating the coordinates of triangle vertices based on side lengths and angles, and calculating the coordinates of the midpoint using the midpoint formula. Finally, the terminal generates an initial geometric figure containing all static geometric objects and their precise coordinate data, serving as the basis for subsequent dynamic graphic generation.
[0117] Optionally, various strategies can be employed when establishing a geometric coordinate system. For example, the simplest coordinate system can be chosen based on the characteristics of the problem, or the figure can be centered on the canvas. For complex figures with numerous constraints, the terminal can use a geometric constraint solver to numerically solve for coordinates that satisfy all conditions. When parsing static geometry, a rule-based entity recognition method or a well-trained named entity recognition model can be used to handle diverse textual descriptions. For geometric figures that are not explicitly given but are implicit (such as the altitude and median of a triangle), the terminal can temporarily omit them and dynamically add them in subsequent problem-solving steps based on operation instructions.
[0118] In this embodiment, by parsing the type, positional relationship, and constraints of static geometric objects from the text description, and establishing a coordinate system and calculating coordinate data based on this, the terminal can automatically generate precise geometric figures that are strictly consistent with the initial conditions of the problem, providing an accurate spatial reference for subsequent dynamic changes. This enables the low-cost, high-quality conversion of static geometric problems into dynamic teaching videos.
[0119] In some exemplary embodiments, the methods in the above embodiments can be implemented in conjunction with more technical details.
[0120] For example, in the basic graphics construction module, when the terminal parses static geometry from a text description, it can extract entities based on a rule engine or a large model and calculate coordinates using computational geometry algorithms. Specifically, the terminal can set key parameters such as the coordinate system origin position and scaling ratio to ensure that the generated initial geometry can adapt to display interfaces of different sizes. The input is a natural language description of the problem, and the output is a set of initial geometric objects and their coordinate data.
[0121] In some exemplary implementations, a finely tuned large mathematical vertical domain model can be used for inference and formatted output. Key parameters set in the terminal include model temperature and maximum generation length to control the diversity and length of generation steps. The input consists of problem text and basic graphical data, and the output is a structured sequence of solution steps, including operation type labels such as "add auxiliary lines" and "define moving points".
[0122] In some exemplary implementations, dynamic parameter extraction and constraint analysis can employ a hybrid semantic parsing and multi-dimensional verification architecture, specifically comprising three layers of processing logic: The first layer employs a dual-track parallel recognition mechanism. The terminal utilizes a rule engine and a large model, forming a dual-track channel. The fast channel leverages regular expressions and a keyword template library to identify high-frequency standard commands, such as "connect AB" and "draw a perpendicular line," within milliseconds, thus covering approximately 70% of basic operations. For complex long sentences that fail to match the template, the deep channel invokes a pre-trained language model for semantic role labeling, accurately extracting the agent, predicate, patient, and adverbial, resolving the problems of unclear referents or omitted subjects.
[0123] The second layer is an adaptive mapping rule base. The terminal has a built-in natural language-mathematical model mapping knowledge base, stored in key-value pair format, such as {Action: "Rotation", Model: "RotationMatrix", Params: ["Center", "Angle", "Direction"]}. This rule base supports both expert preset and online learning modes. The terminal can automatically discover new expressions from manually corrected logs and update the rule table.
[0124] The third layer is a three-dimensional confidence assessment model. The terminal constructs a scoring function that integrates semantic confidence, geometric rationality confidence, and system feasibility confidence. Semantic confidence comes from the probability distribution values output by the model; geometric rationality checks whether the parameters conform to Euclidean geometric axioms, such as checking the sum of the interior angles of a triangle and whether the coordinates overflow the canvas; system feasibility is pre-checked by the rendering engine for the executability of the instructions. If the overall confidence is lower than a preset threshold, the terminal automatically downgrades to a static display and generates a manual review work order. The input is the structured problem-solving steps, and the output is a set of dynamic parametric equations and the range of parameter changes, for example, t∈[0,1]. Key parameters include action type, parameter range, and confidence threshold.
[0125] In some exemplary implementations, the dynamic rendering engine can drive changes in geometric elements based on parametric equations to generate a sequence of vector animation frames. The terminal can employ mathematical plotting libraries such as GeoGebra, Matplotlib, or Manim, combined with speech duration parameters to calculate the interpolation step size. The inputs are the base graphics, the dynamic parametric model, and the target audio duration; the output is a sequence of vector animation clips. Key parameters include frame rate and interpolation step size.
[0126] In some exemplary implementations, the multimodal step-by-step synthesis steps can generate speech, whiteboard text, and animation separately, and perform duration alignment and layer synthesis according to the steps. The terminal generates text-to-speech and motion graphics in parallel, adopts an intelligent column layout, and uses tools such as FFmpeg for multi-layer overlay. The input consists of problem-solving steps text, animation clips, and speech streams, and the output consists of single-step audiovisual segments and complete teaching videos. Key parameters include speech duration and layout ratio.
[0127] In some exemplary implementations, in order to achieve audio-visual synchronization, the terminal adopts a duration-driven dynamic rendering strategy. The terminal can pass the voice duration as a time constraint parameter to the rendering engine. The rendering engine calculates the interpolation step size of each frame based on the motion range of geometric parameters and the target duration, and directly generates animation clips with precise duration matching, avoiding post-processing stretching or freezing of frames, and ensuring the natural and smooth animation speed.
[0128] In some exemplary implementations, the display page can be generated using an intelligent column layout algorithm. The terminal can divide the screen into two columns: the left column is a graphic demonstration area, occupying two-thirds of the screen width, used to display dynamic geometric figures to ensure visual focus; the right column is a whiteboard explanation area, occupying one-third of the screen width, used to display the explanation text. The whiteboard area uses an append-only update method, retaining the previous solution steps and highlighting the text of the current step, such as by making it bold in yellow, simulating the experience of a teacher writing the solution process on the right side of the blackboard in a real classroom.
[0129] In some exemplary implementations, when generating geometry teaching videos, the terminal can use the display page as the underlying background, overlay dynamic animation clips onto the geometric areas of the background, and add subtitles on top, with the audio track playing synchronized audio narration. Since the duration consistency between animation and audio has been ensured in duration-driven rendering, this step mainly completes multi-layer rendering and compositing to output standard audiovisual segments.
[0130] In some exemplary implementations, when synthesizing motion graphics and audio narration, the terminal can splice all aligned segments in timeline order and overlay the generated subtitle track to output the final MP4 video file and accompanying lesson plan document.
[0131] In some exemplary implementations, the following example is taken: "In △ABC, ∠C=90°, rotate △ABC 90° clockwise around point C to obtain △A'B'C". The terminal first performs basic construction, generating a right triangle ABC, where C is the origin, A(3,0), and B(0,4). Step analysis identifies the action as rotation, with point C as the center, an angle of 90°, and a clockwise direction. Parameter extraction constructs a rotation matrix, extracting parameter θ within the range [0, 90°]. In the resource generation stage, the terminal performs text-to-speech synthesis on the script "We rotate triangle ABC 90° clockwise around point C", obtaining the speech duration. Using this duration as the target, the rotation angle increment per frame is calculated, generating a smooth rotation animation. The whiteboard displays "1. Rotation center: C; 2. Rotation angle: 90°", highlighting the current step. In the synthesis stage, a video clip of the triangle smoothly rotating to its new position is generated, the voice explanation ends synchronously, and the text on the right side of the screen is highlighted synchronously. In the final output video, students can intuitively see the trajectory of the triangle, perfectly matching the rhythm of the audio explanation.
[0132] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0133] The teaching video generation apparatus provided in the embodiments of this application is described below. The teaching video generation apparatus has the same inventive concept as the teaching video generation method described above. The solution to the problem provided by the apparatus is similar to the solution described in the method described above. Therefore, the specific limitations of one or more teaching video generation apparatus embodiments provided below can be referred to the limitations of the teaching video generation method above. The teaching video generation apparatus described below and the teaching video generation method described above can be referred to each other, and will not be repeated here.
[0134] In one exemplary embodiment, Figure 6 This is a schematic diagram of the structure of a teaching video generation device provided in an embodiment of this application, such as... Figure 6 As shown, the teaching video generation device 60 includes: a problem-solving step acquisition module 610, a dynamic parameter extraction module 620, a dynamic graphics generation module 630, and a teaching video output module 640, wherein: The problem-solving step acquisition module 610 is used to acquire a sequence of problem-solving steps with a preset structure format for geometry problems.
[0135] The dynamic parameter extraction module 620 is used to extract the dynamic parameters of geometric elements and the range of motion of the dynamic parameters from the sequence of problem-solving steps.
[0136] The dynamic graphics generation module 630 is used to generate dynamic graphics corresponding to the sequence of problem-solving steps based on dynamic parameters and the range of motion of the dynamic parameters.
[0137] The teaching video output module 640 is used to generate audio explanations corresponding to the sequence of problem-solving steps, and synthesize dynamic graphics and audio explanations to output dynamic geometry teaching videos.
[0138] In an exemplary embodiment, the dynamic parameter extraction module 620 is used to identify the operation instructions in the problem-solving step sequence from the problem-solving step sequence using a predefined operation instruction template; and to extract the dynamic parameters of the geometric elements and the range of motion of the dynamic parameters based on the identified operation instructions.
[0139] In an exemplary embodiment, the dynamic parameter extraction module 620 is further configured to, when there is an unrecognized problem-solving step sequence in the operation instruction template, perform semantic annotation on the unrecognized problem-solving step sequence using a pre-trained recognition language model, and identify the dynamic parameters of the problem-solving step sequence and the range of motion of the dynamic parameters.
[0140] In an exemplary embodiment, the dynamic parameter extraction module 620 is used to perform semantic role labeling on the unrecognized problem-solving step sequence by recognizing a language model, so as to identify the agent role, predicate description, and patient role contained in the problem-solving step sequence; based on the result of semantic role labeling, the operation dynamic type and dynamic parameters corresponding to the problem-solving step sequence are determined; wherein, the operation dynamic type is used to characterize the transformation category of geometric elements in the problem-solving step sequence and to determine the parameter extraction method corresponding to the dynamic parameters.
[0141] In an exemplary embodiment, the device further includes a confidence assessment module. The confidence assessment module determines the confidence level of the identified dynamic parameters and performs a confidence assessment on the dynamic parameters, which integrates semantic confidence, geometric rationality confidence, and system feasibility confidence. If the confidence level of the dynamic parameters is lower than a preset threshold, a static graph is generated based on the dynamic parameters, and the static graph is marked. The static graph represents the static display of the steps corresponding to the problem-solving sequence, and the marking indicates that the step-by-step instruction of the problem-solving sequence needs to be reviewed.
[0142] In an exemplary embodiment, the teaching video output module 640 is further configured to obtain the audio duration of the audio explanation; using the audio duration as a constraint, determine the interpolation step size based on the motion range of the dynamic parameters, and generate an animation clip with matching duration so that the animation duration of the dynamic graphics matches the audio duration of the audio explanation.
[0143] In an exemplary embodiment, the teaching video output module 640 is used to divide the text of each step contained in the problem-solving step sequence into an explanation script; and to perform text-to-speech synthesis on the explanation script to obtain the audio explanation corresponding to the problem-solving step sequence.
[0144] In an exemplary embodiment, the teaching video output module 640 is used to divide the text of each step in the problem-solving step sequence into display text; generate a display page corresponding to each step, the display page including a graphic area for displaying dynamic graphics and a whiteboard area for displaying display text; and use the display page as a background to overlay dynamic graphics on the graphic area.
[0145] In one exemplary embodiment, the display page adopts a column layout, with a graphic area on one side and a whiteboard area on the other side; the whiteboard area displays the display text for each step in an append-only manner, and highlights the display text for the current step.
[0146] In an exemplary embodiment, the problem-solving step acquisition module 610 is used to acquire the text description of the geometry problem and generate an initial geometric figure based on the text description; using a pre-trained parsing language model, a sequence of problem-solving steps with a preset structural format is generated based on the text description and the initial geometric figure.
[0147] In one exemplary embodiment, the parsed language model is obtained by fine-tuning a mathematical vertical domain model.
[0148] In an exemplary embodiment, the problem-solving step acquisition module 610 is used to parse the static geometry in the text description, determine the type, positional relationship and constraints of the static geometry; based on the type, positional relationship and constraints of the static geometry, establish a geometric coordinate system, determine the coordinate data of the static geometry, and obtain the initial geometric figure.
[0149] Each module in the aforementioned instructional video generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0150] In one exemplary embodiment, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the instructional video generation methods described above.
[0151] In one exemplary embodiment, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the teaching video generation methods described above.
[0152] In one exemplary embodiment, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the instructional video generation methods described above.
[0153] Indicatively, such as Figure 7 As shown, Figure 7 This is a schematic diagram of the internal structure of a computer device 700 provided in an embodiment of this application. The computer device 700 can be provided as a server. (Refer to...) Figure 7 The computer device 700 includes a processor 702, which further includes one or more processors, and a memory resource represented by a memory 701 for storing instructions executable by the processor 702, such as a computer program. The computer program stored in the memory 701 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processor 702 is configured to execute instructions to perform the instructional video generation method of any of the above embodiments. The computer device 700 can operate on an operating system stored in the memory 701, such as Windows Server™, Mac OS X™, Unix™, Linux™, Free BSD™, or similar.
[0154] The computer device 700 may also include a power supply component 703 configured to perform power management of the computer device 700, a wired or wireless network interface 704 configured to connect the computer device 700 to a network, and an input / output (I / O) interface 705. Wireless operation may be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a method for generating instructional videos. The display unit 707 of the computer device is used to form a visually visible image and may be a display screen, a projection device, or a virtual reality imaging device. The display screen may be an LCD screen or an e-ink display screen. The input device 706 of the computer device may be a touch layer covering the display screen, buttons, a trackball, or a touchpad located on the computer device casing, or an external keyboard, touchpad, or mouse, etc.
[0155] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0156] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0157] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0158] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0159] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for generating instructional videos, characterized in that, The method includes: Obtain a sequence of solution steps for geometry problems with a preset structure format; From the sequence of problem-solving steps, extract the dynamic parameters of the geometric elements and the range of motion of the dynamic parameters; Based on the dynamic parameters and the range of motion of the dynamic parameters, a dynamic graphic corresponding to the sequence of problem-solving steps is generated. Generate audio explanations corresponding to the sequence of problem-solving steps, and synthesize the dynamic graphics and audio explanations to output a dynamic geometry teaching video.
2. The method according to claim 1, characterized in that, The step of extracting the dynamic parameters of the geometric elements and the range of motion of the dynamic parameters from the sequence of problem-solving steps includes: Using a predefined operation instruction template, the operation instructions in the problem-solving step sequence are identified. Based on the recognized operation instructions, the dynamic parameters of the geometric elements and the range of motion of the dynamic parameters are extracted.
3. The method according to claim 2, characterized in that, The step of extracting the dynamic parameters of the geometric elements and the range of motion of the dynamic parameters from the sequence of problem-solving steps further includes: When there is a problem-solving step sequence that is not recognized by the operation instruction template, the unrecognized problem-solving step sequence is semantically annotated by a pre-trained recognition language model to identify the dynamic parameters of the problem-solving step sequence and the range of motion of the dynamic parameters.
4. The method according to claim 3, characterized in that, The step of semantically annotating unrecognized problem-solving step sequences using a pre-trained recognition language model includes: Using the language recognition model, semantic role labeling is performed on the unrecognized problem-solving step sequence to identify the agent role, predicate description, and patient role contained in the problem-solving step sequence; Based on the results of semantic role labeling, the operation dynamic type and dynamic parameters corresponding to the problem-solving step sequence are determined; The operation dynamic type is used to characterize the transformation category of geometric elements in the problem-solving step sequence and to determine the parameter extraction method corresponding to the dynamic parameter.
5. The method according to claim 2, characterized in that, The method further includes: The confidence level of the identified dynamic parameters is determined, and the confidence level of the dynamic parameters is evaluated. The confidence level evaluation integrates semantic confidence level, geometric rationality confidence level, and system feasibility confidence level. If the confidence level of the dynamic parameter is lower than a preset threshold, a static graphic is generated based on the dynamic parameter, and the static graphic is marked. The static graphics are used to represent the static display of the steps corresponding to the problem-solving step sequence, and the markers are used to indicate the review of the step-by-step teaching of the problem-solving step sequence.
6. The method according to claim 1, characterized in that, Before synthesizing the dynamic graphics and the audio narration, the method further includes: Obtain the duration of the audio narration; Using the speech duration as a constraint, and based on the motion range of the dynamic parameters, an interpolation step size is determined, and a duration-matched animation clip is generated so that the animation duration of the dynamic graphic matches the speech duration of the speech narration.
7. The method according to claim 1, characterized in that, The process of synthesizing the dynamic graphics and the voice narration also includes: Extract display text from the text of each step in the sequence of problem-solving steps; Generate a display page corresponding to each of the above steps. The display page includes a graphic area for displaying dynamic graphics and a whiteboard area for displaying the display text. The display page is used as the background, and the dynamic graphics are overlaid on the graphic area.
8. The method according to claim 7, characterized in that, The display page adopts a column layout, with a graphic area on one side and a whiteboard area on the other side; The whiteboard area displays the presentation text for each step in an append-only manner, and highlights the presentation text for the current step.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.