Question explanation video generation method and apparatus, electronic device and storage medium
Through the generative model, the problem of low video generation efficiency in the existing technology is solved, efficient generation is achieved and explanation effect is improved, and users can interact with the video, making the understanding process more convenient.
Patent Information
- Application Number
- PCT/CN2024/107378
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-05
- Filing Date
- 2024-07-24
- Publication Date
- 2025-07-10
AI Technical Summary
In the prior art, the generation efficiency of question explanation videos is low, and explanation videos need to be recorded separately for each question, resulting in inefficiency.
The generative model automatically generates the title explanation video, including text data and animation data, and uses the generative model to generate interactive explanation videos, which contain interactive copy components and animation elements.
The efficiency of the video of the title explanation is improved, the cost of manual recording is reduced, and the explanation effect is improved through multiple forms (text, animation, dubbing). Users can interact with the video to facilitate understanding of the explanation process.
Smart Images

Figure CN2024107378_10072025_PF_FP_ABST
Abstract
Description
Title: Video Generation Method, Device, Electronic Device, and Storage Medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed on January 5, 2024, with application number 202410023815.5 and invention name “Method, device, electronic device and storage medium for generating topic explanation videos”. The entire contents of the application are incorporated by reference into this application. Technical Field
[0003] The present disclosure relates to the field of computer technology, and in particular to a method, device, electronic device, and storage medium for generating a topic explanation video. Background Art
[0004] Currently, students can watch explanation videos provided in supplementary teaching materials to learn the explanation process for questions they don't know how to answer. However, in existing technology, most explanation videos are recorded in advance by teachers, and a separate explanation video needs to be recorded for each question, which has the disadvantage of low video generation efficiency.
[0005] Summary of the Invention
[0006] The embodiments of the present disclosure provide a method, device, electronic device, and storage medium for generating a question explanation video, which can automatically generate a question explanation video based on a generative model, and has high video generation efficiency.
[0007] In a first aspect, an embodiment of the present disclosure provides a method for generating a topic explanation video, comprising:
[0008] Obtaining stem data of a first question, obtaining solution data for the first question based on the stem data of the first question, and determining a question type of the first question;
[0009] generating explanation data for the first question using a generative model based on the problem-solving data and the question type; the explanation data including text data for explaining the first question and animation data for explaining the first question;
[0010] An interactive explanation video of the first topic is generated based on the explanation data; the interactive explanation video displays text corresponding to the text data and an animation view corresponding to the animation data; the text includes an interactive text component, and / or the animation view includes an interactive animation element.
[0011] In a second aspect, an embodiment of the present disclosure provides a device for generating a topic explanation video, comprising:
[0012] a data acquisition unit, configured to acquire stem data of a first question, acquire solution data for the first question based on the stem data of the first question, and determine a question type of the first question;
[0013] a first generating unit configured to generate explanation data for the first question using a generative model based on the problem-solving data and the question type; the explanation data comprising text data for explaining the first question and animation data for explaining the first question;
[0014] The second generating unit is used to generate an interactive explanation video of the first topic based on the explanation data; the interactive explanation video displays the copy corresponding to the text data and the animation view corresponding to the animation data; the copy includes an interactive copy component, and / or the animation view includes an interactive animation element.
[0015] In a third aspect, an embodiment of the present disclosure provides an electronic device, comprising: a processor; and a memory configured to store computer-executable instructions, wherein the computer-executable instructions, when executed, enable the processor to implement the steps of the method described in the first aspect above.
[0016] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, which is used to store computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the method described in the first aspect are implemented.
[0017] In one or more embodiments of the present disclosure, first, the stem data of the first question is obtained, and based on the stem data of the first question, the solution data of the first question is obtained and the question type of the first question is determined. Then, based on the solution data and the question type of the first question, the explanation data of the first question is generated through a generative model, and the explanation data includes text data for explaining the first question and animation data for explaining the first question. Finally, based on the explanation data of the first question, an interactive explanation video of the first question is generated, and the interactive explanation video displays the text corresponding to the above text data and the animation view corresponding to the above animation data. The text includes an interactive text component, and / or the animation view includes an interactive animation element. It can be seen that through this embodiment, the explanation video of the question can be automatically generated based on the generative model, and the video generation efficiency is high, which effectively alleviates the defect of the prior art that a corresponding explanation video needs to be recorded separately for each question and the video generation efficiency is low. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate one or more embodiments of the present disclosure or technical solutions in the prior art, the following briefly introduces the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments described in the present disclosure. Those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0019] FIG1 is a flow chart of a method for generating a topic explanation video according to an embodiment of the present disclosure;
[0020] FIG2 is a schematic diagram showing the principle of generating a topic explanation video according to an embodiment of the present disclosure;
[0021] FIG3a is a schematic diagram of an interactive explanation video provided by an embodiment of the present disclosure;
[0022] FIG3 b is a schematic diagram of an interactive explanation video provided by another embodiment of the present disclosure;
[0023] FIG3c is a schematic diagram of an interactive explanation video provided by another embodiment of the present disclosure;
[0024] FIG4 is a schematic diagram of the structure of a device for generating a topic explanation video according to an embodiment of the present disclosure;
[0025] FIG5 is a schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0026] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of the present disclosure, the technical solutions in one or more embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in one or more embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on one or more embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present disclosure.
[0027] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0028] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0029] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0030] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0031] The disclosed embodiment provides a method for generating a question explanation video, which can automatically generate a question explanation video based on a generative model, with high video generation efficiency. The method for generating a question explanation video in this embodiment can be applied to a backend server and executed by the server.
[0032] FIG1 is a flow chart of a method for generating a topic explanation video according to an embodiment of the present disclosure. As shown in FIG1 , the process includes:
[0033] Step S102, obtaining the stem data of the first question, obtaining the solution data of the first question based on the stem data of the first question, and determining the question type of the first question;
[0034] Step S104: generating explanation data for the first question using a generative model based on the problem-solving data and the problem type; the explanation data includes text data for explaining the first question and animation data for explaining the first question;
[0035] Step S106: Generate an interactive explanation video for the first question based on the explanation data; the interactive explanation video displays text corresponding to the text data and an animation view corresponding to the animation data; the text includes an interactive text component, and / or the animation view includes an interactive animation element.
[0036] In this embodiment, first, the stem data of the first question is obtained, and based on the stem data of the first question, the solution data of the first question is obtained and the question type of the first question is determined. Then, based on the solution data and the question type of the first question, the explanation data of the first question is generated through a generative model. The explanation data includes text data for explaining the first question and animation data for explaining the first question. Finally, based on the explanation data of the first question, an interactive explanation video of the first question is generated. The interactive explanation video displays the text corresponding to the above text data and the animation view corresponding to the above animation data. The text includes an interactive text component, and / or the animation view includes an interactive animation element. It can be seen that through this embodiment, the explanation video of the question can be automatically generated based on the generative model, and the video generation efficiency is high, which effectively alleviates the defect of the prior art that a corresponding explanation video needs to be recorded separately for each question, and the video generation efficiency is low.
[0037] The following is a detailed introduction to the method flow in Figure 1.
[0038] In the above step S102, the stem data of the first question is obtained. In an example, the stem data of the first question can be: "500 grams of salt water with a concentration of 20% becomes 400 grams after evaporation. What is the salt concentration in the salt water at this time?"
[0039] In one embodiment, obtaining the stem data of the first question includes: obtaining a question image of the first question, recognizing the question image of the first question, and obtaining the stem data of the first question.
[0040] In this embodiment, a specific application can be installed and run in the user's terminal device, and the application can communicate with the server that executes the method in this embodiment. The user can use the application to take a picture of the first question, for example, by taking a picture of the teaching materials containing the first question, to obtain a question image of the first question, and upload the question image of the first question to the server through the application. The server recognizes the question image, for example, by performing OCR (Optical Character Recognition), to obtain the question stem data of the first question. The question stem data of the first question may include the text portion of the question stem of the first question, or may include the text portion and the image portion of the question stem of the first question.
[0041] In step S102, the solution data for the first question is obtained based on the stem data of the first question, and the question type of the first question is determined. The solution data for the first question includes at least the answer and analysis of the first question. For example, in the above example, the answer to the first question is "concentration is 25%." The analysis is: "According to the equation salt mass = brine mass * concentration, the brine mass is 500 * 20% = 100 grams. Since the water content decreases before and after evaporation, but the salt mass remains unchanged, the salt concentration in the brine after evaporation is 100 / 400 = 25%."
[0042] In this embodiment, multiple question types can be pre-set, such as application questions, reasoning questions, and short-answer questions. Within each question type, multiple sub-types can be set. For example, application questions include concentration questions, travel distance questions, chicken-and-rabbit-in-a-cage questions, and drainage and water release questions. Reasoning questions include image reasoning and numerical reasoning. Short-answer questions include drawing questions, text questions, and numerical questions. Next, among the various pre-set question types, the question type of the first question is determined. For example, the question type of the first question is determined to be application question - concentration question.
[0043] In the above step S104, based on the solution data of the first question and the question type of the first question, explanation data for the first question is generated by a generative model. The explanation data may be data generated according to a preset data format, and the explanation data includes text data for explaining the first question and animation data for explaining the first question. A generative model refers to a type of model that can generate corresponding results based on input data. In this embodiment, the generative model includes but is not limited to an LLM (Large Language Model) and may also be a cross-modal model. In this embodiment, explanation data for the first question can be generated by a generative model.
[0044] In step S106, an interactive explanation video for the first question is generated based on the explanation data for the first question. The interactive explanation video displays text corresponding to the text data and an animation view corresponding to the animation data. The text includes interactive text components, and / or the animation view includes interactive animation elements. Examples of interactive text components include "previous page" and "next page" buttons, and examples of interactive animation elements include cars and balls in the animation view.
[0045] In this step, an interactive explanation video for the first question is generated based on the explanation data for the first question. The interactive explanation video includes multiple explanation screens, which constitute the interactive explanation video. In each explanation screen, at least the text corresponding to the above-mentioned text data is displayed, and an animation view corresponding to the above-mentioned animation data may also be displayed. In one embodiment, the text includes an interactive text component, in another embodiment, the animation view includes an interactive animation element, and in yet another embodiment, the text includes an interactive text component, and the animation view includes an interactive animation element.
[0046] The interactive text component can realize at least one of the functions of selecting and turning the explanation screen, and displaying knowledge point information in the explanation screen. The interactive animation element can realize at least one of the functions of triggering the animation element to adjust the animation effect and displaying knowledge point information.
[0047] It can be seen that through this embodiment, on the one hand, it is possible to automatically and efficiently generate interactive explanation videos of questions based on the generative model; on the other hand, through the interactive explanation videos of the questions, it is also possible to explain the questions to users in the dual forms of text and animation, thereby improving the effect of the question explanation; on still another hand, users can interact with the explanation video during the process of explaining the question through the interactive text components and / or interactive animation elements in the interactive explanation video, and realize at least one of the functions of selecting the explanation screen, turning pages, adjusting animation effects, displaying knowledge point information, etc., so as to facilitate users to better understand the explanation process of the question.
[0048] In one embodiment, obtaining the solution data of the first question according to the stem data of the first question includes:
[0049] According to the question stem data of the first question, searching the question bank for the solution data of the first question;
[0050] or,
[0051] The problem-solving model is used to generate problem-solving data for the first question based on the stem data of the first question.
[0052] In this embodiment, after obtaining the stem data of the first question, in one case, the solution data of the first question can be retrieved from a pre-created question bank based on the stem data of the first question. For example, questions whose stem data are the same as the stem data of the first question are retrieved from the pre-created question bank, and the solution data of the same questions are used as the solution data of the first question. The solution data of the first question at least includes the answer and analysis of the first question.
[0053] In another case, a pre-created problem-solving model is used to generate problem-solving data for the first problem based on the stem data of the first problem. The problem-solving model can be a generative model that has the ability to generate corresponding results based on input data. In this embodiment, the generative model can be used to generate problem-solving data for the first problem based on the stem data of the first problem. The problem-solving data for the first problem at least includes the answer and analysis of the first problem.
[0054] In one embodiment, a question whose stem data is the same as the stem data of the first question can be first retrieved from a pre-created question bank. If found, the solution data of the same question will be used as the solution data of the first question. If not found, the solution data of the first question will be generated based on the stem data of the first question through a solution model.
[0055] It can be seen that through this embodiment, the solution data of the first question can be obtained efficiently and quickly based on the question stem data of the first question by retrieval or through a generative model.
[0056] In one embodiment, determining the topic type of the first topic includes:
[0057] The question type discrimination model is used to determine the question type of the first question according to the question stem data of the first question and / or the solution data of the first question.
[0058] In this embodiment, a question type discrimination model is created in advance. The model can also be a generative model or a classification model. At least one of the stem data of the first question and the solution data of the first question is input into the model. The model processes the input data and outputs the question type of the first question.
[0059] In one embodiment, the question type discrimination model is a generative model, and a variety of question types are pre-configured for the generative model, such as application questions, reasoning questions, short-answer questions, etc. Under each question type, a variety of sub-types can also be set. For example, application questions include concentration questions, travel distance questions, chicken and rabbit in the same cage questions, drainage and water release questions, etc., reasoning questions include image reasoning, numerical reasoning, etc., and short-answer questions include drawing short answers, text short answers, numerical simple answers, etc. At least one of the stem data of the first question and the solution data of the first question is input into the generative model. The generative model performs semantic understanding on the input data, and outputs the question type of the first question based on the understood semantic information and the pre-configured semantic information of various question types, such as outputting "application question-concentration question". Among them, the semantic information of various question types can be identified by the generative model, or pre-configured in the generative model.
[0060] In another embodiment, the question type discrimination model is a classification model, and a plurality of question types are pre-configured for the classification model, such as application questions, reasoning questions, short-answer questions, etc. Under each question type, a plurality of sub-types can also be set, for example, application questions include concentration questions, travel distance questions, chicken and rabbit in the same cage questions, drainage and water release questions, etc., reasoning questions include image reasoning, numerical reasoning, etc., and short-answer questions include drawing short answers, text short answers, numerical simple answers, etc. At least one of the stem data of the first question and the solution data of the first question is input into the classification model, and the classification model processes the input data, extracts keywords from the input data, and matches the identified keywords with the keywords of each question type, and outputs the question type of the first question based on the matching result, for example, outputting "application question-concentration question". Among them, the keywords of each question type can be identified by the classification model, or pre-configured in the classification model.
[0061] In another embodiment, if the first question is located in a pre-created question bank, a question type can be set for the first question when the first question is entered into the question bank, thereby obtaining the question type of the first question from the question bank.
[0062] It can be seen that through this embodiment, the question type discrimination model can be used to accurately determine the question type of the first question based on the question stem data of the first question and / or the solution data of the first question, thereby improving the accuracy and efficiency of determining the question type.
[0063] In one embodiment, generating explanation data for the first question using a generative model based on the problem-solving data and the question type includes:
[0064] Determine the explanation method requirements and explanation data requirements corresponding to the first question based on the question type;
[0065] Through the generative model, the explanation data for the first question is generated according to the problem-solving data, explanation method requirements and explanation data requirements.
[0066] In this embodiment, the explanation method requirements and explanation data requirements corresponding to the first question are first determined based on the question type. In this embodiment, explanation method requirements and explanation data requirements corresponding to different question types are pre-established, with each question type corresponding to one explanation method requirement and one explanation data requirement. Based on this correspondence, the explanation method requirements and explanation data requirements corresponding to the first question are searched based on the question type of the first question.
[0067] The explanation method requirements for the first question are used to indicate the requirements for the explanation method of the first question, such as the tone, wording, and method of explanation. The explanation method refers to the method used to explain the first question, such as explanation through drawing, reasoning, or equations. The explanation method includes two dimensions. The first dimension refers to the explanation method for the question type of the first question, and the second dimension refers to the explanation method for the first question under the question type of the first question. For example, if the first dimension indicates that the explanation method is the equation method, the second dimension indicates that the explanation method is the quadratic equation method.
[0068] The explanation data requirements corresponding to the first question are used to indicate the requirements for the explanation data for the first question, such as the requirements for the data volume of each explanation screen for the first question, the requirements for the data format of the explanation data used for the first question, etc.
[0069] Next, the generative model generates explanation data for the first problem based on the solution data, explanation method requirements, and explanation data requirements for the first problem. For example, the solution data, explanation method requirements, and explanation data requirements for the first problem are input into the generative model, which processes the input data and outputs explanation data for the first problem.
[0070] It can be seen that through this embodiment, the explanation method requirements and explanation data requirements corresponding to the first question can be determined according to the question type of the first question, and then the explanation data of the first question can be generated according to the problem-solving data, explanation method requirements and explanation data requirements of the first question through the generative model, which has the advantage of efficiently and quickly generating the problem-solving data of the first question.
[0071] In one embodiment, the generative model generates explanation data for the first question based on the problem-solving data, the explanation method requirements, and the explanation data requirements, including:
[0072] Through the generative model, the explanation content of the first question is determined according to the problem-solving data and the explanation method requirements;
[0073] Through the generative model, explanation data for the first question is generated according to the explanation data requirements and the explanation content of the first question; the explanation data is used to represent the explanation content of the first question.
[0074] In this embodiment, a generative model is first used to determine the explanation content of the first question based on the solution data of the first question and the explanation method requirements of the first question. For example, the explanation content of the first question is determined based on the solution data of the first question and the requirements for the explanation tone, explanation words, and explanation method of the first question. The explanation content of the first question is the content that will eventually be presented to the user through a video, and the user can learn the solution process of the first question through the explanation content.
[0075] In one example, the first question includes multiple explanation stages, such as the question reading stage, analysis stage, problem-solving stage, and summary stage. In the explanation method requirements, explanation method sub-requirements are recorded for each explanation stage. The explanation method sub-requirements for each explanation stage can include requirements for explanation tone, explanation vocabulary, and explanation methods. Based on this, according to the problem-solving data of the first question and the explanation method sub-requirements for each explanation stage, the explanation sub-content of each explanation stage of the first question is determined. The explanation sub-content of each explanation stage together constitutes the explanation content of the first question. The explanation content of the first question is used to represent the content of the problem-solving data of the first question.
[0076] In one example, the solution data for the first question includes at least the answer and explanation for the first question. The answer to the first question is: "The concentration is 25%." The explanation is: "According to the equation: salt mass = brine mass * concentration, the brine mass is 500 * 20% = 100 grams. Since the amount of water decreases before and after evaporation but the mass of salt remains unchanged, the concentration of salt in the brine after evaporation is 100 / 400 = 25%." In this example, the explanation for the first question may be: "Using a gentle tone and the solution method for concentration questions, while expressing the fact that the mass of salt remains unchanged, explain that according to the equation: salt mass = brine mass * concentration, the brine mass is 500 * 20% = 100 grams. Since the amount of water decreases before and after evaporation but the mass of salt remains unchanged, the concentration of salt in the brine after evaporation is 100 / 400 = 25%. The explanation may be illustrated in the form of a list or animation."
[0077] Then, using the generative model, explanation data for the first question is generated based on the explanation data requirements and the explanation content of the first question. This can generate explanation data for each stage of the explanation of the first question. The explanation data includes text data used to explain the first question and animation data used to explain the first question. The explanation data is used to represent the explanation content of the first question. In this embodiment, since an interactive explanation video for the first question is also generated based on the explanation data, the user can obtain the explanation content of the first question through the interactive explanation video, thereby clearly understanding the solution process of the first question.
[0078] It can be seen that through this embodiment, the generative model can be used to determine the explanation content of the first question based on the problem-solving data and the explanation method requirements. Then, the explanation data for the first question is generated based on the explanation data requirements and the explanation content of the first question, thereby improving the efficiency of generating explanation data through the generative model.
[0079] In one embodiment, the generative model generates explanation data for the first question based on the explanation data requirements and the explanation content of the first question, including:
[0080] Determine, by means of a generative model, the explanation content and the expression method of each explanation screen for the first question based on the explanation content quantity requirement for the explanation screen in the explanation data requirement and the explanation content of the first question;
[0081] Through the generative model, explanation data for the first question is generated according to the data format requirements for the explanation data, the explanation content of the explanation screen of the first question, and the expression method of the explanation content in the explanation data requirements; the text data in the explanation data includes the display text data corresponding to the explanation screen of the first question and the broadcast text data corresponding to the explanation screen of the first question; the animation data in the explanation data includes the sub-animation data corresponding to the explanation screen of the first question.
[0082] In this embodiment, after determining the explanation content of the first question based on the solution data and explanation method requirements of the first question, the generative model is first used to determine the explanation content and the expression method of the explanation content for each explanation screen of the first question based on the explanation content volume requirements for the explanation screen in the explanation data requirements and the explanation content of the first question.
[0083] In this embodiment, the interactive explanation video is composed of multiple explanation screens, and the explanation data requirements include a requirement on the amount of explanation content for each explanation screen. For example, the explanation content of each explanation screen is required to be no more than three sentences, so as to facilitate user understanding. Based on this, the explanation content of the first question is first divided according to the requirement on the amount of explanation content for the explanation screen in the explanation data requirements and the explanation content of the first question, and the explanation content of each explanation screen of the first question is determined.
[0084] After determining the explanation content for each explanation screen for the first question, the presentation method for the explanation content in each explanation screen can also be determined. In this embodiment, the presentation method for the explanation content in each explanation screen can be determined based on the various presentation methods provided in the explanation data requirements. For example, the explanation data requirements may provide multiple presentation methods such as text, animation, table, and dubbing. For each explanation screen, the appropriate presentation method is determined based on the explanation content in that explanation screen. For example, for one explanation screen, the explanation content can be expressed in the form of text and animation, while for another explanation screen, the explanation content can be expressed in the form of text and dubbing.
[0085] The explanation data requirements also include data format requirements for the explanation data, such as format requirements for text, animation, tables, and audio for dubbing. Based on these requirements, a generative model is used to generate explanation data for the first question based on the data format requirements for the explanation data, the explanation content of the explanation screen for the first question, and the expression of the explanation content. It should be understood that the expression of the explanation content is related to the format requirements for the explanation data. When the expression is textual, the format requirements include the text format requirements, and so on.
[0086] When generating the explanation data for the first question based on the data format requirements for the explanation data, the explanation content of the explanation screen of the first question and the expression method of the explanation content, for each explanation screen, the explanation data of the explanation screen is generated based on the explanation content of the explanation screen, the expression method of the explanation content of the explanation screen and the data format requirements for the expression method, and the explanation data complies with the expression method and the data format requirements for the expression method.
[0087] For example, for a certain explanation screen, after the explanation content is determined, the expression of the explanation content includes text and animation, and the explanation data of the explanation screen is generated according to the data format requirements of the relevant text and animation.
[0088] As previously described, the explanation data includes text data for explaining the first question and animation data for explaining the first question. In this embodiment, the text data in the explanation data includes display text data corresponding to the explanation screen for the first question and broadcast text data corresponding to the explanation screen for the first question. The display text data and broadcast text data of each explanation screen together constitute the aforementioned text data. The animation data in the explanation data includes sub-animation data corresponding to the explanation screen for the first question. The sub-animation data of each explanation screen together constitute the aforementioned animation data. In this embodiment, a table is also a specific form of animation.
[0089] In this embodiment, for the explanation screen, its corresponding explanation data includes text data and animation data. The text data includes display text data for display and broadcast text data for broadcasting in the form of dubbing. The animation data includes sub-animation data corresponding to the explanation screen. Of course, the expression method of each explanation screen can include one or more of text, animation, and dubbing. For the explanation screen, the corresponding explanation content can be displayed in at least one form of displaying text, displaying animation, and playing dubbing.
[0090] It can be seen that through this embodiment, the explanation content and the expression method of the explanation content of each explanation screen of the first question can be determined according to the explanation content volume requirement for the explanation screen in the explanation data requirement and the explanation content of the first question, so as to achieve the effect of content splitting and expression method determination for each explanation screen. Then, according to the data format requirement for the explanation data, the explanation content of the explanation screen and the expression method of the explanation content in the explanation data requirement, the explanation data of the explanation screen is generated, so as to achieve the effect of generating explanation data that matches its explanation content, expression method and data format requirements for each explanation screen. In addition, the text data in the explanation data includes display text data corresponding to the explanation screen and broadcast text data corresponding to the explanation screen, and the animation data in the explanation data includes sub-animation data corresponding to the explanation screen. Therefore, for the explanation screen, the corresponding explanation content can be displayed in at least one form of displaying text, displaying animation and playing dubbing, thereby improving the effect of users watching interactive explanation videos to understand the problem-solving process.
[0091] As previously mentioned, the explanation data requirements have corresponding data format requirements for text, animation, and dubbing. In one embodiment, the sub-animation data corresponding to the explanation screen is written in DSL (Domain Specific Language). Accordingly, the explanation data for the first question is generated by a generative model based on the data format requirements for the explanation data, the explanation content of the explanation screen for the first question, and the expression method of the explanation content, including:
[0092] Using a generative model, the content of the explanation for the first question and the way it is expressed are used to determine the content to be expressed through animation.
[0093] Through the generative model, sub-animation data corresponding to the explanation screen of the first question is generated according to the format requirements for DSL in the data format requirements and the explanation content expressed through animation.
[0094] In this embodiment, a generative model is first used to determine the explanation content to be expressed through animation based on the explanation content and the expression method of the explanation content in the explanation screen for the first question. For example, if the explanation screen contains three explanation contents, each labeled with a corresponding expression method, the explanation content to be expressed through animation is determined based on the expression method of each explanation content. Then, the explanation content to be expressed through animation is compiled according to the format requirements for the DSL language in the data format requirements, generating the sub-animation data corresponding to the explanation screen.
[0095] In this embodiment, the DSL language used to generate the animation data is a pre-edited and defined language. Pre-set fields in the language represent various information such as animation elements, positions, movement speeds, and directions of the animation elements. Subsequently, the DSL language can be parsed according to the parsing rules corresponding to the DSL language to determine various information such as the positions, movement speeds, and directions of the animation elements, thereby facilitating the generation of animation views in interactive explanation videos.
[0096] It can be seen that through this embodiment, the explanation content expressed through animation can be determined based on the explanation content and the expression method of the explanation content of the explanation screen of the first question, and the sub-animation data corresponding to the explanation screen of the first question is generated through the format requirements of the customized DSL language and the explanation content expressed through animation, so as to achieve the effect of generating sub-animation data through the customized language and then combining them to obtain animation data. Subsequently, the customized language can be parsed according to the parsing rules corresponding to the customized language to determine various information such as the position, movement speed, and movement direction of the animation elements, so as to generate animation views in interactive explanation videos.
[0097] Similar to the process of generating animation data, in this embodiment, a generative model is also used to determine the explanation content to be expressed through display text based on the explanation content and the expression method of the explanation content on the explanation screen for the first question. The generative model is also used to generate display text data corresponding to the explanation screen for the first question based on the format requirements for display text data and the explanation content to be expressed through display text in the data format requirements. Furthermore, a generative model is also used to determine the explanation content to be expressed through broadcast text based on the explanation content and the expression method of the explanation content on the explanation screen for the first question. The generative model is also used to generate broadcast text data corresponding to the explanation screen for the first question based on the format requirements for broadcast text data and the explanation content to be expressed through broadcast text in the data format requirements.
[0098] After generating the display text data, broadcast text data, and sub-animation data corresponding to each explanation screen, in one embodiment, generating an interactive explanation video for the first topic based on the explanation data includes:
[0099] Generate the animation view in the explanation screen according to the sub-animation data;
[0100] Generate the text for the explanation screen based on the displayed text data;
[0101] Generate audio data corresponding to the explanation screen according to the broadcast text data;
[0102] Assemble the text, animation view and explanation audio to obtain an interactive explanation video for the first question.
[0103] In this embodiment, first, based on the sub-animation data of the explanation screen, an animation view in the explanation screen is generated. The animation view includes animation elements, such as a car, a track, a balloon, and other elements. These animation elements have their own position, movement direction, movement speed, and other information. Then, based on the display text data of the explanation screen, the text in the explanation screen is generated. The text is used to express the text content of the display text data. Next, based on the broadcast text data of the explanation screen, the explanation audio data corresponding to the explanation screen is generated. The explanation audio data is used to broadcast the broadcast text data. Finally, the text, animation view, and explanation audio are assembled to obtain an interactive explanation video for the first question. For example, the text, animation view, and explanation audio in the same explanation screen are assembled to obtain the explanation screen. Each explanation screen constitutes an interactive explanation video for the first question.
[0104] It can be seen that through this embodiment, it is possible to generate an animation view in the explanation screen based on the sub-animation data of the explanation screen, generate a text in the explanation screen based on the display text data of the explanation screen, generate explanation audio data corresponding to the explanation screen based on the broadcast text data of the explanation screen, assemble the text, animation view and explanation audio to obtain an interactive explanation video for the first question, and based on the output results of the generative model, generate an interactive explanation video for the first question efficiently and quickly.
[0105] In one embodiment, explanation audio data corresponding to the explanation screen is generated based on the broadcast text data, including: converting the broadcast text data into corresponding audio data through Text To Speech (TTS) technology, and using the converted audio data as the explanation audio data corresponding to the explanation screen.
[0106] In one embodiment, the animation view includes interactive animation elements. Through the interactive animation elements, at least one of the functions of adjusting the animation effect and displaying knowledge point information can be achieved. For example, the interactive animation element is a car. By triggering the car, the car's position can be adjusted and the knowledge point information associated with the car can be displayed. Based on this, when the above-mentioned sub-animation data is written in DSL language, the animation view in the explanation screen is generated based on the sub-animation data, including:
[0107] Parsing the sub-animation data according to the parsing rules corresponding to the DSL to determine the first animation element for user interaction represented by the sub-animation data, the display mode of the first animation element, and the interaction mode of the first animation element;
[0108] An interactive animation element in the explanation screen is generated according to the first animation element, the display mode of the first animation element, and the interaction mode of the first animation element.
[0109] In this embodiment, the sub-animation data corresponding to the explanation screen is parsed according to the parsing rules corresponding to the DSL language to determine the first animation element for user interaction represented by the sub-animation data, the display mode of the first animation element, and the interaction mode of the first animation element. Of course, the second animation element not used for user interaction represented by the sub-animation data and the display mode of the second animation element can also be determined. The display mode of the animation element includes at least one of the display position, movement direction, and movement speed. The interaction mode of the animation element includes click interaction, voice interaction, long press interaction, drag interaction, etc.
[0110] Then, for the first animation element, an interactive animation element in the explanation screen is generated according to the first animation element, the display mode of the first animation element and the interaction mode of the first animation element. That is, according to the display mode and interaction mode of the first animation element, the display and interaction of the first animation element are controlled in the explanation screen, and the first animation element after display and interaction is the interactive animation element.
[0111] Of course, for the second animation element, a non-interactive animation element can be generated in the explanation screen according to the second animation element and the display method of the second animation element. That is, according to the display method of the second animation element, the display of the second animation element is controlled in the explanation screen, and the second animation element after display is a non-interactive animation element.
[0112] It can be seen that through this embodiment, the sub-animation data can be parsed according to the parsing rules corresponding to the DSL, and the first animation element for user interaction, the display method of the first animation element, and the interaction method of the first animation element represented by the sub-animation data can be determined. According to the first animation element, the display method of the first animation element, and the interaction method of the first animation element, the interactive animation element in the explanation screen is generated, so that the user can interact with the interactive explanation video through the animation element, thereby improving the user experience of watching the interactive explanation video.
[0113] In one embodiment, the copy includes an interactive copy component. Through the interactive copy component, at least one of the functions of selecting the explanation screen, turning the page, and displaying knowledge point information in the explanation screen can be realized. For example, if the interactive copy component is "Next Page", then by triggering the interactive copy component, a video page turning effect can be achieved. Based on this, the copy in the explanation screen is generated according to the displayed text data, including:
[0114] determining, according to data attributes of the displayed text data, first text data for user interaction, a display style of the first text data, and an interaction mode of the first text data in the displayed text data;
[0115] An interactive text component in the explanation screen is generated according to the first text data, the display style of the first text data, and the interaction mode of the first text data.
[0116] In this embodiment, the displayed text data has data attributes that indicate whether the displayed text data is interactive, the interaction method, and the display method. Interaction methods include click interaction, long press interaction, drag interaction, etc., and display methods include font, size, color, position, etc. Based on this, according to the data attributes of the displayed text data, the first text data used for user interaction, the display style of the first text data, and the interaction method of the first text data are determined in the displayed text data. In addition, the second text data not used for user interaction and the display style of the second text data can also be determined in the displayed text data.
[0117] Next, for the first text data, an interactive copy component is generated in the explanation screen according to the first text data, the display style of the first text data and the interaction method of the first text data. That is, according to the display method and interaction method of the first text data, the text content of the first text data is controlled to be displayed and interacted in the explanation screen, and an interactive copy component is generated according to the text content after display and interaction.
[0118] Next, for the second text data, a non-interactive copy component is generated in the explanation screen according to the second text data and the display style of the second text data. That is, according to the display mode of the second text data, the text content of the second text data is controlled to be displayed in the explanation screen, and a non-interactive copy component is generated according to the displayed text content.
[0119] It can be seen that through this embodiment, the first text data for user interaction, the display style of the first text data and the interaction method of the first text data in the displayed text data can be determined according to the data attributes of the displayed text data, and the interactive copy component in the explanation screen is generated according to the first text data, the display style of the first text data and the interaction method of the first text data, so that the user can interact with the interactive explanation video through the interactive copy component, thereby improving the user's experience of watching the interactive explanation video.
[0120] FIG2 is a schematic diagram of the principle of generating a question explanation video provided by an embodiment of the present disclosure. As shown in FIG2 , first, the stem data of the first question is obtained. Based on the stem data of the first question, the solution data of the first question is obtained and the question type of the first question is determined. Then, based on the question type of the first question, the explanation method requirements and the explanation data requirements corresponding to the first question are determined. Then, the solution data, explanation method requirements, and explanation data requirements of the first question are input into the generative model. The explanation data of the first question is generated by the generative model. The explanation data includes the display text data corresponding to the explanation screen of the first question and the broadcast text data corresponding to the explanation screen of the first question, as well as the sub-animation data corresponding to the explanation screen of the first question. In FIG2 , the explanation data is split to obtain the display text data, the broadcast text data, and the sub-animation data.
[0121] Next, the text for the explanation screen is generated based on the displayed text data. The animation engine generates an animation view for the explanation screen based on the sub-animation data. The text-to-audio model generates audio data corresponding to the explanation screen based on the broadcast text data. The text, animation view, and audio data are assembled to produce an interactive explanation video for the first question. The text includes interactive text components, and / or the animation view includes interactive animation elements.
[0122] In this embodiment, the interactive explanation video of the first question can be realized through H5 (Hypertext Markup Language) data. The server sends the generated interactive explanation video to the user's terminal device, and the terminal device parses and displays the interactive explanation video according to the parsing rules of the H5 data.
[0123] Figure 3a is a schematic diagram of an interactive explanation video provided by an embodiment of the present disclosure. As shown in Figure 3a, taking an explanation screen of the interactive explanation video as an example, the explanation screen shows the problem analysis stage. The screen analyzes the question in the form of a table, and the table is also a specific animation form. In addition, the explanation screen also includes interactive copy components "Previous Step", "Next Step" and "Water Quality" and "Salt Quality". The user turns pages through "Previous Step" and "Next Step", and answers questions in the video through "Water Quality" and "Salt Quality". The explanation screen will give a corresponding response after the user answers, such as displaying "Correct Answer" or "What a pity".
[0124] Similar to Figure 3a, the explanation screen can also display the corresponding content of the problem-solving stage, the problem-reading stage, and the summary stage, and each explanation stage can include interactive copy components and / or interactive animation elements, which are not illustrated in this embodiment.
[0125] FIG3 b is a schematic diagram of an interactive explanation video provided by another embodiment of the present disclosure. As shown in FIG3 b , a question-and-answer component is also displayed in the interactive explanation video. Users can interact with the question-and-answer component through text or voice, and the question-and-answer component can answer specific questions of users.
[0126] FIG3c is a schematic diagram of an interactive explanation video provided by another embodiment of the present disclosure. As shown in FIG3c , a question feedback entrance is also displayed in the interactive explanation video, and users can provide feedback on corresponding questions through the question feedback entrance.
[0127] In summary, through this embodiment, on the one hand, interactive explanation videos of questions can be automatically generated based on the generative model, without the need for manual video recording, which improves the video generation efficiency and reduces labor costs. On the other hand, through the interactive explanation video of the question, the question can be explained to the user in multiple forms such as text, animation and dubbing, thereby improving the question explanation effect. On the other hand, the user can interact with the explanation video through the interactive text components and / or interactive animation elements in the interactive explanation video during the process of explaining the question, and realize at least one of the functions of selecting the explanation screen, turning pages, adjusting animation effects, displaying knowledge point information, etc., so as to facilitate the user to better understand the explanation process of the question. On the other hand, through long-term accumulation, more appropriate explanation method requirements and explanation data requirements can be obtained, and iterative updates of model processing can be realized to generate interactive explanation videos that better meet user needs.
[0128] FIG4 is a schematic diagram of the structure of a device for generating a topic explanation video according to an embodiment of the present disclosure. As shown in FIG4 , the device includes:
[0129] The data acquisition unit 41 is configured to acquire stem data of a first question, acquire solution data for the first question based on the stem data of the first question, and determine the type of the first question;
[0130] A first generating unit 42 is configured to generate explanation data for the first question using a generative model based on the problem-solving data and the problem type; the explanation data includes text data for explaining the first question and animation data for explaining the first question;
[0131] The second generation unit 43 is used to generate an interactive explanation video of the first topic based on the explanation data; the interactive explanation video displays the text corresponding to the text data and the animation view corresponding to the animation data; the text includes an interactive text component, and / or the animation view includes an interactive animation element.
[0132] Optionally, the data acquisition unit 41 is specifically configured to:
[0133] Retrieving the solution data of the first question in a question bank according to the question stem data of the first question;
[0134] or,
[0135] The problem-solving model is used to generate problem-solving data for the first problem based on the stem data of the first problem.
[0136] Optionally, the data acquisition unit 41 is specifically configured to:
[0137] The question type discrimination model is used to determine the question type of the first question according to the question stem data of the first question and / or the solution data of the first question.
[0138] Optionally, the first generating unit 42 is specifically configured to:
[0139] Determining, according to the topic type, explanation method requirements corresponding to the first topic and explanation data requirements corresponding to the first topic;
[0140] The generative model is used to generate explanation data for the first question according to the problem-solving data, the explanation method requirement, and the explanation data requirement.
[0141] Optionally, the first generating unit 42 is further specifically configured to:
[0142] Determining, by the generative model, the explanation content for the first question according to the problem-solving data and the explanation method requirements;
[0143] The generative model generates explanation data for the first question according to the explanation data requirements and the explanation content of the first question; the explanation data is used to represent the explanation content of the first question.
[0144] Optionally, the first generating unit 42 is further specifically configured to:
[0145] Determining, by the generative model, the explanation content for each explanation screen of the first question and the expression method of the explanation content based on the explanation content quantity requirement for the explanation screen in the explanation data requirement and the explanation content of the first question;
[0146] Through the generative model, explanation data for the first question is generated according to the data format requirements for explanation data in the explanation data requirements, the explanation content of the explanation screen of the first question and the expression method of the explanation content; the text data in the explanation data includes the display text data corresponding to the explanation screen of the first question and the broadcast text data corresponding to the explanation screen of the first question; the animation data in the explanation data includes the sub-animation data corresponding to the explanation screen of the first question.
[0147] Optionally, the sub-animation data is written in a domain-specific language; the first generating unit 42 is further specifically configured to:
[0148] Determining, by the generative model, the explanation content to be expressed through animation based on the explanation content of the explanation screen for the first question and the expression method of the explanation content;
[0149] The sub-animation data corresponding to the explanation screen of the first topic is generated through the generative model according to the format requirements for the domain-specific language in the data format requirements and the explanation content expressed through animation.
[0150] Optionally, the text data in the explanation data includes display text data corresponding to the explanation screen of the first question and broadcast text data corresponding to the explanation screen of the first question; the animation data in the explanation data includes sub-animation data corresponding to the explanation screen of the first question; and the second generating unit 43 is specifically configured to:
[0151] generating an animation view in the explanation screen according to the sub-animation data;
[0152] Generating the text in the explanation screen according to the display text data;
[0153] Generating explanation audio data corresponding to the explanation screen according to the broadcast text data;
[0154] The text, the animation view and the explanation audio are assembled to obtain an interactive explanation video of the first topic.
[0155] Optionally, the animation view includes interactive animation elements; the sub-animation data is written in a domain-specific language; and the second generating unit 43 is further specifically configured to:
[0156] Parsing the sub-animation data according to a parsing rule corresponding to the domain-specific language to determine a first animation element for user interaction represented by the sub-animation data, a display mode of the first animation element, and an interaction mode of the first animation element;
[0157] An interactive animation element in the explanation screen is generated according to the first animation element, the display mode of the first animation element, and the interaction mode of the first animation element.
[0158] Optionally, the copy includes an interactive copy component; the second generating unit 43 is further specifically configured to:
[0159] determining, according to data attributes of the display text data, first text data for user interaction, a display style of the first text data, and an interaction mode of the first text data in the display text data;
[0160] An interactive text component in the explanation screen is generated according to the first text data, the display style of the first text data, and the interaction mode of the first text data.
[0161] The apparatus for generating a question explanation video in the embodiment of the present disclosure can implement each process of the above-mentioned embodiment of the method for generating a question explanation video and achieve the same effects and functions, which will not be repeated here.
[0162] An embodiment of the present disclosure also provides an electronic device. FIG5 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. As shown in FIG5 , the electronic device may have relatively large differences due to different configurations or performances, and may include one or more processors 501 and a memory 502. One or more applications or data may be stored in the memory 502. Among them, the memory 502 may be a temporary storage or a persistent storage. The application stored in the memory 502 may include one or more modules (not shown in the figure), each module may include a series of computer executable instructions in the electronic device. Furthermore, the processor 501 may be configured to communicate with the memory 502 and execute a series of computer executable instructions in the memory 502 on the electronic device. The electronic device may also include one or more power supplies 503, one or more wired or wireless network interfaces 504, one or more input or output interfaces 505, one or more keyboards 506, etc.
[0163] In a specific embodiment, the electronic device includes a processor; and a memory configured to store computer-executable instructions, wherein when the computer-executable instructions are executed, the processor implements the following process:
[0164] Obtaining stem data of a first question, obtaining solution data for the first question based on the stem data of the first question, and determining a question type of the first question;
[0165] generating explanation data for the first question using a generative model based on the problem-solving data and the question type; the explanation data including text data for explaining the first question and animation data for explaining the first question;
[0166] An interactive explanation video of the first topic is generated based on the explanation data; the interactive explanation video displays text corresponding to the text data and an animation view corresponding to the animation data; the text includes an interactive text component, and / or the animation view includes an interactive animation element.
[0167] The electronic device in the embodiment of the present disclosure can implement each process of the above-mentioned embodiment of the method for generating a topic explanation video and achieve the same effects and functions, which will not be repeated here.
[0168] Another embodiment of the present disclosure further provides a computer-readable storage medium for storing computer-executable instructions. When the computer-executable instructions are executed by a processor, the following process is implemented:
[0169] Obtaining stem data of a first question, obtaining solution data for the first question based on the stem data of the first question, and determining a question type of the first question;
[0170] generating explanation data for the first question using a generative model based on the problem-solving data and the question type; the explanation data including text data for explaining the first question and animation data for explaining the first question;
[0171] An interactive explanation video of the first topic is generated based on the explanation data; the interactive explanation video displays text corresponding to the text data and an animation view corresponding to the animation data; the text includes an interactive text component, and / or the animation view includes an interactive animation element.
[0172] The storage medium in the embodiment of the present disclosure can implement each process of the above-mentioned embodiment of the method for generating a title explanation video and achieve the same effects and functions, which will not be repeated here.
[0173] In various embodiments of the present disclosure, the computer-readable storage medium includes a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0174] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures such as diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0175] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.
[0176] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0177] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing the embodiments of the present disclosure, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0178] Those skilled in the art will appreciate that one or more embodiments of the present disclosure may be provided as a method, system, or computer program product. Therefore, one or more embodiments of the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, one or more embodiments of the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0179] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0180] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0181] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0182] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0183] One or more embodiments of the present disclosure may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. One or more embodiments of the present disclosure may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0184] The various embodiments of this disclosure are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so their description is relatively simple. For relevant portions, refer to the description of the method embodiments.
[0185] The foregoing is merely an embodiment of the present disclosure and is not intended to limit the present disclosure. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present disclosure are intended to be included within the scope of the claims of the present disclosure.
Claims
1. A method for generating a video for explaining questions, characterized in that, Including: Obtain the stem data of the first question. According to the stem data of the first question, obtain the solution data of the first question and determine the question type of the first question. According to the solution data and the question type, generate the explanation data of the first question through a generative model; the explanation data includes text data for explaining the first question and animation data for explaining the first question. Generate an interactive explanation video of the first question according to the explanation data; the interactive explanation video displays the copywriting corresponding to the text data and the animation view corresponding to the animation data; the copywriting includes an interactive copywriting component, and / or the animation view includes an interactive animation element.
2. The method according to claim 1, wherein The obtaining the solution data of the first question according to the stem data of the first question includes: Retrieve the solution data of the first question in the question bank according to the stem data of the first question. Or, Use a problem-solving model to generate the solution data of the first question according to the stem data of the first question.
3. The method according to claim 1, characterized in that, The determining the question type of the first question includes: Use a question type discrimination model to determine the question type of the first question according to the stem data of the first question and / or the solution data of the first question.
4. The method according to claim 1, wherein The generating the explanation data of the first question through a generative model according to the solution data and the question type includes: According to the question type, determine the explanation method requirements corresponding to the first question and the explanation data requirements corresponding to the first question. Through the generative model, generate the explanation data of the first question according to the solution data, the explanation method requirements, and the explanation data requirements.
5. The method according to claim 4, wherein The generating the explanation data of the first question through the generative model according to the solution data, the explanation method requirements, and the explanation data requirements includes: Through the generative model, determine the explanation content of the first question according to the solution data and the explanation method requirements. Through the generative model, generate the explanation data of the first question according to the explanation data requirements and the explanation content of the first question; the explanation data is used to represent the explanation content of the first question.
6. The method according to claim 5, wherein The generating the explanation data of the first question through the generative model according to the explanation data requirements and the explanation content of the first question includes: Through the generative model, determine the explanation content of each explanation screen of the first question and the expression method of the explanation content according to the explanation content amount requirement for the explanation screen in the explanation data requirements and the explanation content of the first question. Through the generative model, according to the data format requirements for the explanatory data in the explanatory data requirements, the explanatory content of the explanatory screen of the first question, and the expression mode of the explanatory content, generate the explanatory data of the first question; the text data in the explanatory data includes the display text data corresponding to the explanatory screen of the first question and the broadcast text data corresponding to the explanatory screen of the first question; the animation data in the explanatory data includes the sub-animation data corresponding to the explanatory screen of the first question.
7. The method according to claim 6, wherein The sub-animation data is written in a domain-specific language; the process of generating the explanatory data of the first question through the generative model according to the data format requirements for the explanatory data in the explanatory data requirements, the explanatory content of the explanatory screen of the first question, and the expression mode of the explanatory content includes: Through the generative model, determine the explanatory content to be expressed through animation according to the explanatory content of the explanatory screen of the first question and the expression mode of the explanatory content; Through the generative model, generate the sub-animation data corresponding to the explanatory screen of the first question according to the format requirements for the domain-specific language in the data format requirements and the explanatory content to be expressed through animation.
8. The method according to claim 1, characterized in that The text data in the explanatory data includes the display text data corresponding to the explanatory screen of the first question and the broadcast text data corresponding to the explanatory screen of the first question; the animation data in the explanatory data includes the sub-animation data corresponding to the explanatory screen of the first question; the process of generating the interactive explanatory video of the first question according to the explanatory data includes: Generate the animation view in the explanatory screen according to the sub-animation data; Generate the copywriting in the explanatory screen according to the display text data; Generate the explanatory audio data corresponding to the explanatory screen according to the broadcast text data; Assemble the copywriting, the animation view, and the explanatory audio to obtain the interactive explanatory video of the first question.
9. The method according to claim 8, wherein The animation view includes interactive animation elements; the sub-animation data is written in a domain-specific language; the process of generating the animation view in the explanatory screen according to the sub-animation data includes: According to the parsing rules corresponding to the domain-specific language, parse the sub-animation data to determine the first animation element for user interaction represented by the sub-animation data, the display mode of the first animation element, and the interaction mode of the first animation element; Generate the interactive animation elements in the explanatory screen according to the first animation element, the display mode of the first animation element, and the interaction mode of the first animation element.
10. The method according to claim 8, wherein The copywriting includes interactive copywriting components; the process of generating the copywriting in the explanatory screen according to the display text data includes: According to the data attributes of the display text data, determine the first text data for user interaction in the display text data, the display style of the first text data, and the interaction mode of the first text data; Generate an interactive copywriting component in the explanation screen according to the first text data, the display style of the first text data, and the interaction method of the first text data.
11. A video generation device for problem explanation, characterized in that, It includes: A data acquisition unit, configured to acquire the stem data of the first question, and according to the stem data of the first question, acquire the solution data of the first question and determine the question type of the first question; A first generation unit, configured to generate the explanation data of the first question through a generative model according to the solution data and the question type; the explanation data includes text data for explaining the first question and animation data for explaining the first question; A second generation unit, configured to generate an interactive explanation video of the first question according to the explanation data; the interactive explanation video displays the copywriting corresponding to the text data and the animation view corresponding to the animation data; the copywriting includes an interactive copywriting component, and / or the animation view includes an interactive animation element. It includes:
12. An electronic device, characterized in that, A processor; And A memory configured to store computer-executable instructions, the computer-executable instructions, when executed, cause the processor to implement the steps of the method according to any one of claims 1-10 above. The computer-readable storage medium is used to store computer-executable instructions, and the computer-executable instructions, when executed by the processor, implement the steps of the method according to any one of claims 1-10 above.
13. A computer-readable storage medium, characterized in that,
Citation Information
Patent Citations
Explanation video generation method and device, and explanation video display method and device
CN112447073A
Question explanation method and system, electronic equipment and computer readable storage medium
CN113127682A
Chinese question solving video generation method and device, electronic equipment and storage medium
CN115114905A
Topic explanation video generation method and device, electronic equipment and storage medium
CN117221656A
Topic explanation video generation method and device, electronic equipment and storage medium
CN117835014A