Test question explanation video generation method and system and electronic equipment

By analyzing test questions to generate explanation video scripts and identifying truncation symbols, and combining large language models and deep learning technology, the streaming generation and real-time push of test question explanation videos have been realized. This solves the problems of insufficient generation efficiency and real-time performance in existing technologies and improves students' ability to obtain accurate explanations instantly.

CN121999653APending Publication Date: 2026-05-08IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2025-12-22
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing test question explanation video generation technologies are insufficient in terms of real-time performance and efficiency to meet students' needs for timely and accurate explanations, especially when students encounter difficulties while doing homework online, as existing technologies cannot quickly provide targeted explanation videos.

Method used

The system generates video scripts by analyzing test questions and identifies truncation symbols during the streaming generation process. Based on the script before the truncation symbols, it generates video frames for test question explanation. It uses a large language model for real-time analysis and generation, and combines natural language processing and deep learning technologies to achieve streaming generation and output of explanation videos.

Benefits of technology

It improves the efficiency and real-time performance of test question explanation video generation, enabling the output of explanation videos while simultaneously capturing and rendering the generated scripts during the generation process, thus achieving instant push of test question explanation videos and enhancing students' learning experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121999653A_ABST
    Figure CN121999653A_ABST
Patent Text Reader

Abstract

The invention provides a test question explanation video generation method and system and electronic equipment, and the method comprises the steps: carrying out the analysis of a test question, generating an explanation video script in a streaming manner, and enabling the explanation video script to comprise a script identifier, and enabling the script identifier to be used for identifying a test question explanation stage and / or explanation video elements in the explanation video script; in the process of streaming generation of the explanation video script, identifying a truncation symbol from the generated explanation video script; the truncation symbols comprise preset script identifiers and / or punctuation marks used for representing the ending of the explanation steps; and under the condition that any truncation symbol is identified from the generated explanation video script, generating a test question explanation video frame according to the explanation video script before the identified truncation symbol. According to the method, the real-time performance and efficiency of generating the test question explanation video can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, system, and electronic device for generating test question explanation videos. Background Technology

[0002] With the rapid rise of online education and the widespread application of intelligent tutoring systems, the demand for test question explanation videos in the education sector has exploded.

[0003] While the technology for automatically generating test question explanation videos has made some progress, its real-time generation capability needs improvement, failing to meet students' needs for immediate and accurate explanations. For example, when students encounter difficult problems while doing homework online, they expect to quickly access targeted explanation videos, a feat that current technology cannot adequately achieve. Summary of the Invention

[0004] To address the aforementioned technical issues, this application provides a method, system, and electronic device for generating test question explanation videos, which can improve the real-time performance and efficiency of generating test question explanation videos.

[0005] The first aspect of this application provides a method for generating test question explanation videos, including: The test questions are analyzed and a streaming explanation video script is generated. The explanation video script contains a script identifier, which is used to identify the test question explanation stage and / or explanation video elements in the explanation video script. During the streaming generation of the explanation video script, truncation symbols are identified from the generated explanation video script; the truncation symbols include pre-defined script identifiers and / or punctuation marks used to indicate the end of the explanation steps; If any truncation symbol is identified from the generated explanation video script, a question explanation video frame is generated based on the explanation video script before the identified truncation symbol.

[0006] In some implementations, if any truncation symbol is identified from the generated instructional video script, the method further includes: Based on the identified truncation symbol, a video rendering data frame is generated from the explanation video script. The video rendering data frame includes multiple fields and the corresponding field content for each field. The multiple fields include the test question explanation stage, the oral content, the screen display content, the interactive actions, and the interactive questions. The step of generating test question explanation video frames based on the explanation video script preceding the identified truncation symbol includes: The video frames for explaining the test questions are generated based on the video rendering data frames.

[0007] In some implementations, generating video rendering data frames based on the explanatory video script preceding the identified truncation symbol includes: From the generated explanatory video script, extract the explanatory video script that is after the previously extracted explanatory video script and before the currently identified truncation symbol, and generate video rendering data frames based on the extracted explanatory video script.

[0008] In some implementations, generating the test question explanation video frame based on the video rendering data frame includes: The system generates video explanation audio based on the spoken content and / or interactive questions, generates screen display images based on the screen display content, and renders a virtual human image based on the interactive actions.

[0009] In some implementations, the process of parsing the test questions and generating a streaming explanation video script includes: The test questions are input into a large language model, which then parses the test questions and streams an explanation video script.

[0010] The second aspect of this application provides a test question explanation video generation system, including: The test question analysis unit is used to analyze the test questions and generate a streaming explanation video script. The explanation video script contains a script identifier, which is used to identify the test question explanation stage and / or explanation video elements in the explanation video script. The script processing unit is used to identify truncation symbols from the generated explanation video script during the process of the test question analysis unit streaming the explanation video script; the truncation symbols include pre-set script identifiers and / or punctuation marks used to indicate the end of the explanation steps; The video generation unit is used to generate test question explanation video frames based on the explanation video script before the identified truncation symbol when any truncation symbol is identified from the generated explanation video script.

[0011] In some implementations, the question parsing unit parses the questions and generates a streaming explanation video script, including: The test questions are input into a large language model, which then parses the test questions and streams an explanation video script.

[0012] In some implementations, the script processing unit is further used for: If any truncation symbol is identified from the generated explanation video script, a video rendering data frame is generated based on the explanation video script before the identified truncation symbol. The video rendering data frame includes multiple fields and the field content corresponding to each field. The multiple fields include the test question explanation stage, oral content, screen display content, interactive actions, and interactive questions. The video generation unit generates test question explanation video frames based on the explanation video script preceding the identified truncation symbol, including: The video generation unit generates video frames for explaining the test questions based on the video rendering data frames.

[0013] In some implementations, the script processing unit generates video rendering data frames based on the explanatory video script preceding the identified truncation symbol, including: From the generated explanatory video script, extract the explanatory video script that is after the previously extracted explanatory video script and before the currently identified truncation symbol, and generate video rendering data frames based on the extracted explanatory video script.

[0014] A third aspect of this application provides an electronic device, including a memory and a processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the above-mentioned method for generating test question explanation videos by running the program in the memory.

[0015] The test question explanation video generation method provided in this application generates an explanation video script by parsing the test questions, and can also generate the explanation video script in a streaming manner during the process of parsing the test questions. In the process of generating the explanation video script in a streaming manner, truncation symbols are identified from the generated explanation video script, and test question explanation video frames are generated based on the explanation video script before the identified truncation symbols.

[0016] The above solution can generate and output the explanation video script while simultaneously extracting the already generated explanation video script and using the extracted explanation video script to generate test question explanation video frames. In other words, it realizes streaming test question analysis and test question explanation video generation processing, thereby improving the efficiency and real-time performance of test question explanation video generation. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating the method for generating test question explanation videos provided in this application embodiment.

[0019] Figure 2 This is a schematic diagram illustrating the processing steps of the test question explanation video generation method provided in this application embodiment.

[0020] Figure 3 This is a schematic diagram of the structure of a test question explanation video generation system provided in an embodiment of this application.

[0021] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0022] With the rapid rise of online education and the widespread application of intelligent tutoring systems, the demand for test explanation videos in the education sector has exploded. Online education breaks the time and space limitations of traditional education, allowing students to access learning resources anytime, anywhere. This makes a wealth of high-quality teaching videos a key element in the development of online education. Intelligent tutoring systems are dedicated to providing personalized learning support for students. One of the core functions of intelligent tutoring systems is to provide timely, detailed, and accurate test explanation videos to address the problems students encounter during problem-solving.

[0023] Traditional test question explanation videos are primarily produced manually. First, teachers need to thoroughly research and prepare the test questions, outlining problem-solving strategies and preparing the content, including planning the order of explanation of knowledge points and highlighting key and difficult points. After thorough preparation, teachers choose a suitable location for recording, possibly using cameras, screen recording software, etc. During recording, teachers need to explain the test questions completely, ensuring accuracy and fluency. After recording, the video enters the post-production editing stage, where editors trim the video, removing unnecessary segments, adjusting the pacing, adding subtitles, etc., to improve the video's quality and viewing experience.

[0024] For example, when recording a video explaining a high school math function problem, the teacher first writes down the solution steps and key points for each step in detail on a lesson plan. Then, the teacher explains and records the problem in front of a blackboard or electronic whiteboard in the recording studio. After recording, the editor cuts out parts of the video where the teacher pauses too long or makes mistakes, and adds clear subtitles of mathematical formulas, ultimately creating an explanation video that students can watch and learn from.

[0025] Teachers then upload the video explanations of the test questions generated in this way to the internet for students to watch or download. This traditional production method has been widely used in the education field in the past, especially in offline educational institutions for creating teaching materials and in schools for teachers to record teaching videos for students to review after class. However, the drawback of generating test question explanation videos in this way is that which question's explanation video a student can see, or when they can see the required explanation video, depends entirely on whether and when the teacher records the explanation video, making it difficult to meet students' needs for accessing the test question explanation videos anytime, anywhere. Moreover, this method requires a lot of manpower and time, making it extremely inefficient.

[0026] To address the aforementioned shortcomings, a technology for automatically generating test question explanation videos has been proposed. This technology aims to receive test questions uploaded by students in real time and generate explanation videos instantly. Currently, this technology has made some progress, but several limitations remain. Existing technologies often require pre-setting detailed explanation templates for different types of questions, and generating test questions and explanations involves filling in these templates. This approach lacks adaptability to diverse questions and complex knowledge points.

[0027] Existing test question explanation video generation technologies also need improvement in terms of real-time video generation, failing to meet students' needs for immediate and accurate explanations. In practical applications, such as when students encounter difficulties while doing online homework, they expect to quickly obtain targeted explanation videos, which current technology cannot adequately achieve. For example, when a student wants a specific test question to be explained in a video, they upload the question to the system, then wait for a period of time. During this waiting period, the system analyzes and solves the question in the background, then processes the information from the solution and generates the corresponding explanation video according to a pre-defined strategy. Only after the video is generated is it pushed to the student for viewing. This entire process requires a considerable amount of time for students to wait, which seriously affects their learning efficiency and experience, failing to meet their learning needs during a busy study schedule.

[0028] To address the aforementioned technical issues, this application provides a novel method for generating test question explanation videos, which can improve the efficiency and real-time performance of generating such videos.

[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0030] This application first provides a method for generating test question explanation videos. This method can be executed on devices or equipment with data processing capabilities, such as computers, servers, smart terminals, handheld devices, and wearable devices. Specifically, it can be executed on a processor similar to the aforementioned devices or equipment.

[0031] For example, the test question explanation video generation method provided in this application embodiment can be written as a computer program capable of implementing the processing procedure of the test question explanation video generation method. By running the computer program through the processor in the above-mentioned device or equipment, the test question explanation video generation method provided in this application embodiment can be executed in the above-mentioned device or equipment.

[0032] In some embodiments, the above-described method for generating test question explanation videos can be run as a standalone computer program in the above-described apparatus or device. When it is necessary to generate a test question explanation video for a specific test question, the computer program is run, and the information of the test question to be explained is input into the computer program. The computer program can then generate a test question explanation video corresponding to the test question to be explained by executing the processing steps of the above-described method for generating test question explanation videos.

[0033] Alternatively, in other embodiments, the provided test question explanation video generation method is configured to run on the aforementioned device or equipment as a test question explanation video generation service program, and to provide a service call interface in other computer programs. When a user encounters a test question that requires generating a corresponding test question explanation video while using other computer programs running on the aforementioned device or equipment, they can call the test question explanation video generation service program on the aforementioned device or equipment through the service call interface. At this time, the test question explanation video generation service program executes the processing procedure of the aforementioned test question explanation video generation method, generates a test question explanation video corresponding to the test question to be explained, and pushes the generated explanation video to the aforementioned other computer programs so that the user can watch the explanation video in other computer programs.

[0034] For example, the test question explanation video generation method provided in this application can be applied to a server, which can be a cloud server. A front-end interface runs on a device or equipment that communicates with the server. Users upload test questions through the front-end interface running locally on the device or equipment. At this time, the test question explanation video generation program on the server can be called through the front-end interface. By executing the processing procedure of the test question explanation video generation method provided in this application, the test question explanation video is generated and pushed to the front-end interface for display.

[0035] The embodiments of this application do not limit the execution environment, application scenarios, or implementation forms of the test question explanation video generation method provided in this application. The embodiments of this application mainly focus on the specific processing steps of the test question explanation video generation method provided in this application. The test question explanation video generation method provided in this application can be implemented in any scenario, any execution environment, or in any form by executing the specific processing steps described in the embodiments of this application.

[0036] See Figure 1 As shown in the embodiments of this application, the method for generating test question explanation videos includes: S101. Analyze the test questions and generate a streaming explanation video script.

[0037] The video script includes a script identifier, which is used to identify the question explanation stage and / or explanation video elements in the video script.

[0038] The aforementioned test questions refer to those to be explained, that is, those requiring analysis and the generation of corresponding videos explaining the solution process. These test questions can be from any subject and of any type; this embodiment does not impose any limitations.

[0039] Analyzing the test questions involves identifying the relevant knowledge points, the thought process for solving the problem, the methods used, and a summary of the solution. This analysis forms the basis for creating video explanations of the test questions.

[0040] In this embodiment, natural language processing technology is used to understand the content of the test questions, and knowledge graph and deep learning technologies are combined to analyze the test questions, and an explanation video script is generated based on the analysis results.

[0041] The aforementioned video explanation script refers to a script that generates a video explanation of the test question based on the test question analysis results. In this embodiment, the video explanation script is output in string form; that is, the generated video explanation script is a string composed of text and symbols. By performing image rendering and video generation according to the video explanation script, the video explanation of the test question can be obtained.

[0042] In this embodiment, test question explanation videos are generated according to four stages: reading the question, analysis, standard answer, and summary. Correspondingly, the scripts for the explanation videos generated from analyzing the test questions also include scripts for each of these four stages.

[0043] To facilitate the differentiation of different question explanation stages from the generated explanation video script, this embodiment sets script identifiers in the generated explanation video script. These script identifiers include identifiers for identifying question explanation stages; that is, they include script identifiers used to identify the script content of different question explanation stages in the explanation video script. For example, through...<phase_start>

XX

XX

[0044] Furthermore, in this embodiment, the script identifier set in the explanation video script also includes an identifier for identifying explanation video elements, that is, it also includes a script identifier for identifying the script content of the corresponding explanation video elements in the explanation video script.

[0045] The aforementioned explanatory video elements refer to the video elements presented in the test question explanation video, including spoken content, on-screen content, interactive actions, and interactive questions.

[0046] For example, through<display_start> and<display_end> These respectively mark the beginning and end of the displayed content.<display_start> and<display_end> The script content between these lines is the content displayed on the screen; through<action_emphasize> and< / action_emphasize> These clearly indicate the start and end of the "emphasis" interaction; through...<action_please> Interactive actions that indicate "Please answer"; through<action_display> Identify the content to be broadcast; through<question_start> and<question_end> Mark the start and end of the interactive question respectively.

[0047] Taking the question "Fill in the blanks with symbols greater than, less than, or equal to: 89 yuan ( ) 98 yuan" as an example, the following is the explanation video script generated by this embodiment to analyze the question: <phase_start> [Read the question]<action_display> Let's look at the following question together: Fill in the blanks with signs indicating greater than, less than, or equal to. Which is greater, 89 yuan or 98 yuan?<phase_end> <phase_start>[Analysis] Think outside the box! Where is the "solution code" hidden in this problem? When we compare the size of two numbers,<action_emphasize> When the units are the same< / action_emphasize> So, isn't the key just to look at the numbers themselves?<display_start> To compare the size of two numbers, if the units are the same, simply compare their numerical values.<display_end><action_display> Look, the unit here is "yuan", so we can just compare the two numbers 89 and 98.<question_start> {\"Explanation Content\": \"<action_please> Which number is larger, 89 or 98? "Question": "Which number is larger, 89 or 98?" "Options": ["They are the same", "98", "89", "Don't know"], "Correct Answer": "98"}<question_end><action_please> Which number is larger, 89 or 98?<display_start> 89 yuan < 98 yuan<display_end><action_display> Let's take a look. Both 89 and 98 are two-digit numbers. First, compare the tens digit. 8 is smaller than 9, so 89 is smaller than 98. Therefore, 89 yuan is less than 98 yuan.<phase_end> <phase_start> [Standard Answer]<display_start> ["<"]<display_end><action_display> Therefore, the answer is less than sign.<phase_end> <phase_start> [Summary] Let's review the thought process behind this problem. When comparing the amounts in RMB,<display_start> Directly comparing the size of numbers<display_end><action_display> When the units are the same, simply compare the numerical values; the larger the value, the more money is involved, and the smaller the value, the less money is involved.<phase_end> As can be seen from the above examples, the explanation video script generated in step S101 of this application embodiment contains script identifiers. Through these script identifiers, the question explanation stage and / or explanation video elements can be distinguished from the generated explanation video script.

[0048] Furthermore, in this embodiment, when generating the explanation video script by parsing the test questions, a streaming generation method is adopted, that is, the explanation video script is generated while the already generated explanation video script is output, so that the generated explanation video script is output in the form of a string stream.

[0049] In some embodiments, it is disclosed that by training a large language model, the large language model can perform the above-described step S101. That is, the test question is input into the large language model, which parses the test question and streams and outputs an explanation video script, and the explanation video script carries the aforementioned script identifier. The large language model can be any type of large language model, and the training of the large language model can also employ conventional large model training methods; this embodiment does not impose any limitations.

[0050] S102. During the process of streaming the explanation video script, identify truncation symbols from the generated explanation video script.

[0051] The truncation symbols include pre-defined script identifiers and / or punctuation marks used to indicate the end of the explanation steps.

[0052] Specifically, in this embodiment, during the process of streaming the explanation video script in step S101, the already generated explanation video script is simultaneously streamed.

[0053] The streaming process includes identifying truncation symbols from the generated instructional video script. These truncation symbols are used to indicate the end of an instructional step. When a truncation symbol is identified from the string stream of the stream-generated instructional video script, it indicates the end of an instructional step. At this point, the string stream of the generated instructional video script is truncated at the position of the truncation symbol to facilitate the processing of the instructional video script segment corresponding to the instructional step before the truncation symbol.

[0054] In this embodiment, certain specific script identifiers and / or punctuation marks are pre-set as truncation marks. When these script identifiers and / or punctuation marks are identified in the generated explanatory video script, the truncation operation on the generated explanatory video script is performed.

[0055] In this embodiment, the aforementioned truncation symbol specifically includes<phase_end> ,<action_emphasize> ,< / action_emphasize> ,<display_start> ,<question_start> ,".","?"wait.

[0056] Following the processing method of step S102 above, during the process of generating the narration video script in step S101, truncation symbols are identified from the generated narration video script.

[0057] S103. If any truncation symbol is identified from the generated explanation video script, generate a test question explanation video frame based on the explanation video script before the identified truncation symbol.

[0058] Specifically, when any truncation symbol is identified from the generated explanation video script through step S102, the generated explanation video script is truncated at the identified truncation symbol, and the explanation video script before the truncation symbol is rendered to obtain the test question explanation video frame.

[0059] The video frame for explaining the test questions, obtained by rendering the video script before the truncation symbol, can be a single video frame or multiple consecutive video frames.

[0060] After generating the test question explanation video frame in step S103, the generated test question explanation video frame can be pushed to the user terminal for playback, so that the user can watch the generated explanation video.

[0061] During the execution of step S103, which renders the video script before the truncation symbol to obtain the test question explanation video frames, step S101 continues to generate the explanation video script, and step S102 continues to identify the truncation symbol from the already generated explanation video script. This forms a streaming execution of S101, S102, and S103. This achieves the generation of explanation videos based on the explanation video script during the streaming generation of the explanation video script, allowing users to watch the already generated test question explanation video content while continuously generating subsequent test question explanation video content.

[0062] As shown in Figure 2, the above-described process of generating the explanation video script, identifying truncation symbols from the generated explanation video script, and generating test question explanation video frames based on the explanation video script before the identified truncation symbols is illustrated in Figure 2.

[0063] As can be seen from the above description, the test question explanation video generation method provided in this application generates an explanation video script by parsing the test questions, and can generate the explanation video script in a streaming manner during the parsing process. In the process of generating the explanation video script in a streaming manner, truncation symbols are identified from the generated explanation video script, and test question explanation video frames are generated based on the explanation video script before the identified truncation symbols.

[0064] The above solution can generate and output the explanation video script while simultaneously extracting the already generated explanation video script and using the extracted explanation video script to generate test question explanation video frames. In other words, it realizes streaming test question analysis and test question explanation video generation processing, thereby improving the efficiency and real-time performance of test question explanation video generation.

[0065] In another embodiment, it is disclosed that during the process of streaming the explanatory video script and identifying truncation symbols from the generated explanatory video script, if any truncation symbol is identified from the generated explanatory video script, the explanatory video script before the truncation symbol is formally transformed, specifically by generating video rendering data frames based on the explanatory video script before the identified truncation symbol.

[0066] The aforementioned video rendering data frames refer to data frames that include the question explanation stage and explanation video elements required for rendering and generating the question explanation video.

[0067] In this embodiment, the video rendering data frame described above consists of multiple fields and the corresponding field content of each field.

[0068] The multiple fields are the question explanation stage, the oral content, the screen display content, the interactive actions, and the interactive questions. By determining the corresponding field content for each field, the video rendering data frame mentioned above can be formed.

[0069] For example, the video rendering data frame mentioned above can be in the following form: {"step": "", "read": "", "showV0": "", "question": {}, "action": ""} In this context, “step” represents the test question explanation stage field; “read” represents the oral content field; “showV0” represents the screen display content field; “question” represents the interactive question field; and “action” represents the interactive action field. The quotation marks after each field are used to specify the specific content corresponding to the field.

[0070] For example, when any truncation symbol is detected in the stream-generated explanatory video script, the script content corresponding to each of the above fields is identified from the explanatory video script before the truncation symbol. Based on the identified script content, the video rendering data frame described above can be constructed. If no script content corresponding to a certain field is identified in the explanatory video script, the field content corresponding to that field in the generated video rendering data frame will be empty.

[0071] In another embodiment, when a truncation symbol is identified from an already generated explanatory video script, an explanatory video script that is located after the previously truncated explanatory video script and before the currently identified truncation symbol is extracted from the already generated explanatory video script, and a video rendering data frame is generated based on the extracted explanatory video script.

[0072] In other words, each time a truncation symbol is detected from the generated instructional video script, the instructional video script that precedes the currently detected truncation symbol and follows the previously detected truncation symbol is extracted, and then a video rendering data frame is generated based on the extracted instructional video script. Specifically, in the case where a truncation symbol is detected for the first time from the generated instructional video script, a video rendering data frame is generated based on all instructional video scripts preceding the detected truncation symbol.

[0073] The above process ensures that each video rendering data frame is generated based on the newly generated explanation video script, thus achieving the effect of streaming the generation of explanation video frames. This allows users to simultaneously watch the generated video content while the test question explanation video is being generated in a streaming manner.

[0074] Taking the example of the question "Fill in the blanks with symbols greater than, less than, or equal to: 89 yuan ( ) 98 yuan", by executing the question explanation video generation method provided in this embodiment, during the process of parsing the question and streaming the explanation video script, when generating "<phase_start> [Read the question]<action_display> Let's look at the following question together: Fill in the blanks with signs indicating greater than, less than, or equal to. Which is greater, 89 yuan or 98 yuan?<phase_end> "When creating these instructional video scripts, it's possible to identify them from the already generated instructional video scripts."<phase_end> This truncation symbol indicates that the video script will be cut off at this point, and the truncation symbol will be displayed.<phase_end> The previous explanation video script was converted into video rendering data frames: {"step": "readQuestion", "read": "Let's look at the question together: Fill in the blanks with signs indicating greater than, less than, or equal to. Which is larger, 89 yuan or 98 yuan?", "showV0": "", "question": {}, "action": "<action_display> Then, based on the video rendering data frame, image rendering is performed to generate test question explanation video frames, and the test question explanation video frames are pushed to the user terminal for playback.

[0075] During the above process, the script for the instructional video continues to be generated. When the script is generated...<phase_start> [Analysis] "Get your brains working! Where is the 'solution code' hidden in this question?" When creating these explanation video scripts, the truncation symbol "?" can be identified. The explanation video script is then truncated at this point, and the script before the truncation symbol "?" is converted into a video rendering data frame: "{"step": "analyse", "read": "Get your brains working! Where is the 'solution code' hidden in this question?", "showV0": "", "question": {}, "action": ""}". Then, based on this video rendering data frame, image rendering is performed to generate the question explanation video frame, which is then pushed to the user's device for playback.

[0076] Throughout the above process, the script for the explanatory video continues to be generated. When generating the script for "When we compare the size of two numbers,"...<action_emphasize> "When creating these instructional video scripts, it's possible to identify..."<action_emphasize> This truncation symbol indicates that the video script will be cut off at this point, and the truncation symbol will be displayed.<action_emphasize> The previous explanation video script is converted into a video rendering data frame: "{"step": "analyse", "read": "When we compare the size of two numbers,", "showV0": "", "question": {}, "action": ""}"; then, based on this video rendering data frame, image rendering is performed to generate the question explanation video frame, and the question explanation video frame is pushed to the user's terminal for playback.

[0077] During the above process, the script for the explanatory video continues to be generated. When generating "if the units are the same..."< / action_emphasize> "When creating these instructional video scripts, it is possible to identify..."< / action_emphasize> This truncation symbol indicates that the video script will be cut off at this point, and the truncation symbol will be displayed.< / action_emphasize> The previous explanation video script was converted into video rendering data frames: {"step": "analyse", "read": "Under the condition of the same units,", "showV0": "", "question": {}, "action": "<action_emphasize> Then, based on the video rendering data frame, image rendering is performed to generate test question explanation video frames, and the test question explanation video frames are pushed to the user terminal for playback.

[0078] During the above process, the script for the instructional video continues to be generated. When the script is generated...<display_start> To compare the size of two numbers, if the units are the same, simply compare their numerical values.<display_end><action_display> "Look, here the unit is 'yuan', so we can directly compare the numbers 89 and 98." When creating these explanation video scripts, the truncation symbol "." can be identified in the generated explanation video script. The explanation video script is then truncated at this point, and the explanation video script before the truncation symbol "." is converted into video rendering data frames: {"step": "analyse", "read": "Look, here the unit is 'yuan', so we can directly compare the numbers 89 and 98.", "showV0": "Compare the size of two numbers; if the units are the same, compare the numerical values ​​directly", "question": {}, "action": "<action_display> Then, based on the video rendering data frame, image rendering is performed to generate test question explanation video frames, and the test question explanation video frames are pushed to the user terminal for playback.

[0079] During the above process, the script for the instructional video continues to be generated. When the script is generated...<question_start> {\"Explanation Content\": \"<action_please> Which number is larger, 89 or 98? "Question": "Which number is larger, 89 or 98?" "Options": ["They are the same", "98", "89", "Don't know"], "Correct Answer": "98"}<question_end><action_please> Which number is larger, 89 or 98? When creating these video scripts, the "?" punctuation mark can be identified in the generated script. The script is then truncated at this point, and the part before the "?" is converted into video rendering data frames: {"step": "analyse", "read": "Which number is larger, 89 or 98?", "showV0": "Compare the size of two numbers; if the units are the same, compare the numerical values ​​directly", "question": {"Explanation content": "Which number is larger, 89 or 98?", "Question": "Which number is larger, 89 or 98?", "Options": ["Same size", "98", "89", "Don't know"], "Correct answer": "98"}, "action":<action_please> Then, based on the video rendering data frame, image rendering is performed to generate test question explanation video frames, and the test question explanation video frames are pushed to the user terminal for playback.

[0080] Similarly, by following the above method, it is possible to generate test question explanation video frames in a streaming manner and push the real-time generated test question explanation video frames to the user's terminal for playback, so that the user can watch the generated test question explanation video stream while the test question explanation video stream is being continuously generated.

[0081] In another embodiment of the application, it is disclosed that when generating a video frame for explaining test questions based on the generated video rendering data frame, video explanation audio is generated based on the spoken content and / or interactive questions recorded in the video rendering data frame, screen display images are generated based on the screen display content recorded in the video rendering data frame, and virtual human figures are rendered based on the interactive actions recorded in the video rendering data frame.

[0082] Specifically, this embodiment uses a variety of rendering techniques to generate video frames for explaining test questions.

[0083] In terms of audio rendering, speech synthesis technology is used to generate natural and fluent voice narration based on the text content recorded in the video rendering data frames. For example, by selecting a suitable speech synthesis engine, the spoken content and / or interactive questions recorded in the video rendering data frames can be converted into audio with different timbres and intonations to meet the needs of different users.

[0084] The whiteboard rendering utilizes computer graphics technology to accurately depict text, formulas, and graphs related to problem-solving. For math problems, it can precisely render various mathematical symbols and formulas. For example, when rendering a quadratic function formula like "y = ax^2 + bx + c", it ensures the accurate position and size of the symbols and displays the solution steps progressively according to the explanation's pace. Similarly, for video rendering, the computer graphics technology described above can be used to render corresponding screen images or graphics from the recorded data frames.

[0085] Interactive actions recorded in the video rendering data frames are displayed by rendering virtual human figures. Virtual human figure rendering utilizes 3D modeling and animation techniques to create realistic virtual human figures that can perform corresponding actions and expressions based on the content being explained. Through a skeletal animation system, the virtual human's head rotation and limb movements are realized. For example, when explaining key content, the virtual human will attract the user's attention by emphasizing its tone and increasing the size of its gestures.

[0086] The combined application of multiple rendering technologies can ensure that all the various explanatory video elements in the video rendering data frames are reflected in the generated test question explanation video frames, thereby improving the quality of the generated test question explanation video.

[0087] Corresponding to the above-mentioned method for generating test question explanation videos, this application also provides a test question explanation video generation system, see [link to system]. Figure 3 As shown, the system includes: The test question analysis unit 100 is used to analyze the test questions and generate a streaming explanation video script. The explanation video script includes a script identifier, which is used to identify the test question explanation stage and / or explanation video elements in the explanation video script. The script processing unit 110 is used to identify truncation symbols from the generated explanation video script during the process of the test question analysis unit streaming the explanation video script; the truncation symbols include pre-set script identifiers and / or punctuation marks used to indicate the end of the explanation steps. The video generation unit 120 is used to generate a test question explanation video frame based on the explanation video script before the identified truncated symbol when any truncation symbol is identified from the generated explanation video script.

[0088] In some embodiments, the test question analysis unit 100 analyzes the test questions and generates a streaming explanation video script, including: The test questions are input into a large language model, which then parses the test questions and streams an explanation video script.

[0089] In some embodiments, the script processing unit 110 is further configured to: If any truncation symbol is identified from the generated explanation video script, a video rendering data frame is generated based on the explanation video script before the identified truncation symbol. The video rendering data frame includes multiple fields and the field content corresponding to each field. The multiple fields include the test question explanation stage, oral content, screen display content, interactive actions, and interactive questions. The video generation unit 120 generates test question explanation video frames based on the explanation video script preceding the identified truncation symbol, including: The video generation unit 120 generates video frames for explaining the test questions based on the video rendering data frames.

[0090] In some embodiments, the script processing unit 120 generates video rendering data frames based on the explanatory video script preceding the identified truncation symbol, including: From the generated explanatory video script, extract the explanatory video script that is after the previously extracted explanatory video script and before the currently identified truncation symbol, and generate video rendering data frames based on the extracted explanatory video script.

[0091] In some embodiments, the video generation unit 120 generates test question explanation video frames based on the video rendering data frames, including: The system generates video explanation audio based on the spoken content and / or interactive questions, generates screen display images based on the screen display content, and renders a virtual human image based on the interactive actions.

[0092] The test question explanation video generation system provided in this embodiment belongs to the same application concept as the test question explanation video generation method provided in the above embodiments of this application. It can execute the test question explanation video generation method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects of executing the method. Technical details not described in detail in this embodiment can be found in the specific processing content of the test question explanation video generation method provided in the above embodiments of this application, and will not be repeated here.

[0093] The functions implemented by each of the above units can be implemented by the same or different processors, and this application embodiment does not limit this.

[0094] It should be understood that the units in the above system can be implemented by a processor calling software. For example, the system includes a processor connected to memory, which stores instructions. The processor calls the instructions stored in memory to implement any of the above methods or to implement the functions of each unit in the system. The processor can be a general-purpose processor, such as a CPU or microprocessor, and the memory can be internal or external to the device. Alternatively, the units in the system can be implemented as hardware circuits. By designing the hardware circuits, some or all of the unit functions can be implemented. The hardware circuit can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, which implements some or all of the unit functions by designing the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a PLD, such as an FPGA, which can include a large number of logic gates. The connection relationships between the logic gates are configured through configuration files to implement some or all of the unit functions. All units in the above system can be implemented entirely by a processor calling software, entirely by hardware circuits, or partially by a processor calling software with the remaining parts implemented by hardware circuits.

[0095] In this application embodiment, a processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a CPU, microprocessor, GPU, or DSP. In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. These logical relationships are fixed or reconfigurable. For example, the processor may be a hardware circuit implemented as an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the processor loading instructions to implement the functions of some or all of the above units. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, or DPU.

[0096] As can be seen, each unit in the above system can be one or more processors (or processing circuits) configured to implement the above methods, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor types.

[0097] Furthermore, the units in the above systems can be integrated in whole or in part, or they can be implemented independently. In one implementation, these units are integrated together and implemented in the form of a System-on-Chip (SoC). The SoC may include at least one processor for implementing any of the above methods or implementing the functions of the units of the device. The at least one processor may be of different types, such as CPU and FPGA, CPU and artificial intelligence processor, CPU and GPU, etc.

[0098] This application also proposes a control device, which includes a processor and an interface circuit. The processor in the control device is connected to an input / output component through the interface circuit of the control device.

[0099] The input / output component specifically refers to a functional component that enables users to input test questions to be explained and to output test question explanation videos to users. For example, the input / output component may include a handwriting tablet, a touch module integrated with a display screen, a keyboard, a mouse, and a display screen.

[0100] The aforementioned interface circuit can be any interface circuit capable of implementing data communication functions, such as a USB interface circuit, a Type-C interface circuit, a serial port circuit, a PCIe circuit, etc.

[0101] The processor in this control device is also a circuit with signal processing capabilities. It executes the arbitrary question explanation video generation method described in the above embodiments to parse the questions selected or input by the user through the aforementioned input / output components, generate corresponding question explanation videos, and play and output the generated question explanation videos through the aforementioned input / output components. The specific implementation of this processor can be found in the processor implementation methods described above; this application embodiment does not impose strict limitations.

[0102] When the control device is applied to a handheld terminal device, the input / output components connected to the control device can be input / output components on the handheld terminal device, such as touch display components. Meanwhile, the processor of the control device can be the CPU or GPU built into the handheld terminal device, and the interface circuit of the control device can be the interface circuit between the input / output components of the handheld terminal device and the CPU or GPU processor.

[0103] This application also provides a test question explanation device, which includes an interactive function component and a processor connected to the interactive function component.

[0104] The interactive component is used to receive user-inputted questions and output the generated question explanation video. The processor is configured to parse the test questions received through the interactive function component and generate test question explanation videos by executing any of the test question explanation video generation methods described in any of the above embodiments.

[0105] The aforementioned interactive functional components can be any functional components that support users in inputting test questions and playing and displaying test question explanation videos, such as keyboards, touch screens, monitors, etc.

[0106] For details on the specific processing steps of the processor described above, please refer to the description of the above method embodiments. For details on the specific implementation of the processor, please refer to the description of the above embodiments.

[0107] The equipment explained in this test question can specifically be smart terminal devices, such as handheld terminals, wearable terminals, and computers.

[0108] Another embodiment of this application also provides an electronic device, see [link to relevant documentation] Figure 4 As shown, the device includes: Memory 200 and processor 210; The memory 200 is connected to the processor 210 and is used to store programs; The processor 210 is used to implement the test question explanation video generation method disclosed in any of the above embodiments by running the program stored in the memory 200.

[0109] Specifically, the aforementioned electronic device may also include: a bus, a communication interface 220, an input device 230, and an output device 240.

[0110] The processor 210, memory 200, communication interface 220, input device 230, and output device 240 are interconnected via a bus. Among them: A bus can include a pathway for transmitting information between various components of a computer system.

[0111] Processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0112] Processor 210 may include a main processor, as well as a baseband chip, modem, etc.

[0113] The memory 200 stores a program that executes the technical solution of this invention, and may also store an operating system and other key business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 200 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.

[0114] Input device 230 may include a device for receiving user input data and information, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor.

[0115] Output device 240 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.

[0116] The communication interface 220 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.

[0117] The processor 210 executes the program stored in the memory 200 and calls other devices, which can be used to implement each step of any of the test question explanation video generation methods provided in the above embodiments of this application.

[0118] Specifically, the electronic device can be a smart terminal device, such as a handheld terminal, a wearable terminal, or a computer.

[0119] This application also proposes a chip, which includes a processor and a data interface. The processor reads and runs a program stored in a memory through the data interface to execute the test question explanation video generation method described in any of the above embodiments. For the specific processing procedure and its beneficial effects, please refer to the above embodiments of the test question explanation video generation method.

[0120] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the test question explanation video generation method described in any of the above embodiments of this specification.

[0121] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0122] Furthermore, embodiments of this application may also be storage media storing a computer program, which, when run by a processor, causes the processor to execute the steps in the test question explanation video generation method described in any of the above embodiments of this specification.

[0123] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0124] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0125] The steps in the methods of the various embodiments of this application can be adjusted, merged, or deleted in order according to actual needs, and the technical features described in each embodiment can be replaced or combined.

[0126] The modules and sub-modules in the various embodiments of the present application's devices and terminals can be merged, divided, and deleted according to actual needs.

[0127] It should be understood that the disclosed terminals, devices, and methods can be implemented in other ways, given the several embodiments provided in this application. For example, the terminal embodiments described above are merely illustrative. For instance, the division of modules or sub-modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple sub-modules or modules may be combined or integrated into another module, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0128] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.

[0129] Furthermore, the functional modules or sub-modules in the various embodiments of this application can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or in the form of software functional modules or sub-modules.

[0130] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0131] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software unit executed by a processor, or a combination of both. The software unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0132] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0133] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for generating test question explanation videos, characterized in that, include: The test questions are analyzed and a streaming explanation video script is generated. The explanation video script contains a script identifier, which is used to identify the test question explanation stage and / or explanation video elements in the explanation video script. During the streaming generation of the explanation video script, truncation symbols are identified from the generated explanation video script; the truncation symbols include pre-defined script identifiers and / or punctuation marks used to indicate the end of the explanation steps; If any truncation symbol is identified from the generated explanation video script, a question explanation video frame is generated based on the explanation video script before the identified truncation symbol.

2. The method according to claim 1, characterized in that, In the event that any truncation symbol is identified from the generated instructional video script, the method further includes: Based on the identified truncation symbol, a video rendering data frame is generated from the explanation video script. The video rendering data frame includes multiple fields and the corresponding field content for each field. The multiple fields include the test question explanation stage, the oral content, the screen display content, the interactive actions, and the interactive questions. The step of generating test question explanation video frames based on the explanation video script preceding the identified truncation symbol includes: The video frames for explaining the test questions are generated based on the video rendering data frames.

3. The method according to claim 2, characterized in that, The process of generating video rendering data frames based on the explanatory video script preceding the identified truncation symbol includes: From the generated explanatory video script, extract the explanatory video script that is after the previously extracted explanatory video script and before the currently identified truncation symbol, and generate video rendering data frames based on the extracted explanatory video script.

4. The method according to claim 2, characterized in that, The step of generating test question explanation video frames based on the video rendering data frames includes: The system generates video explanation audio based on the spoken content and / or interactive questions, generates screen display images based on the screen display content, and renders a virtual human image based on the interactive actions.

5. The method according to claim 1, characterized in that, The process of analyzing the test questions and generating streaming explanation video scripts includes: The test questions are input into a large language model, which then parses the test questions and streams an explanation video script.

6. A test question explanation video generation system, characterized in that, include: The test question analysis unit is used to analyze the test questions and generate a streaming explanation video script. The explanation video script contains a script identifier, which is used to identify the test question explanation stage and / or explanation video elements in the explanation video script. The script processing unit is used to identify truncation symbols from the generated explanation video script during the process of the test question analysis unit streaming the explanation video script; the truncation symbols include pre-set script identifiers and / or punctuation marks used to indicate the end of the explanation steps; The video generation unit is used to generate test question explanation video frames based on the explanation video script before the identified truncation symbol when any truncation symbol is identified from the generated explanation video script.

7. The system according to claim 6, characterized in that, The test question analysis unit analyzes the test questions and generates a streaming explanation video script, including: The test questions are input into a large language model, which then parses the test questions and streams an explanation video script.

8. The system according to claim 6, characterized in that, The script processing unit is also used for: If any truncation symbol is identified from the generated explanation video script, a video rendering data frame is generated based on the explanation video script before the identified truncation symbol. The video rendering data frame includes multiple fields and the field content corresponding to each field. The multiple fields include the test question explanation stage, oral content, screen display content, interactive actions, and interactive questions. The video generation unit generates test question explanation video frames based on the explanation video script preceding the identified truncation symbol, including: The video generation unit generates video frames for explaining the test questions based on the video rendering data frames.

9. The system according to claim 8, characterized in that, The script processing unit generates video rendering data frames based on the explanatory video script preceding the identified truncation symbol, including: From the generated explanatory video script, extract the explanatory video script that is after the previously extracted explanatory video script and before the currently identified truncation symbol, and generate video rendering data frames based on the extracted explanatory video script.

10. An electronic device, characterized in that, Including memory and processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the test question explanation video generation method as described in any one of claims 1 to 5 by running the program in the memory.