Method and device for generating video based on visual question solving, and electronic equipment

By obtaining and processing problem-solving copy information and combining intelligent big models to generate video codes, the problem of low efficiency in generating visual problem-solving videos in the existing technology is solved, and the effect of quickly generating 'one question, one video' is achieved.

CN119996763APending Publication Date: 2025-05-13QILIN HESHENG NETWORK TECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510144455.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is less efficient when generating visual problem-solving videos, and cannot achieve the effect of 'one question, one video', and requires users to manually edit the video.

Method used

By obtaining problem-solving copy information, performing speech conversion and parsing processing, generating subtitle files, and using intelligent big models to generate codes for rendering video files, and finally synthesize the visual target video.

Benefits of technology

It realizes the rapid generation of problem-solving videos, realizes the effect of "one question, one video", and improves the efficiency of generating visual problem-solving videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996763A_ABST
    Figure CN119996763A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a generation method and device based on a visual question solving video and electronic equipment, and belongs to the technical field of artificial intelligence. Comprising the following steps: obtaining question-solving copywriting information of a target question, and performing voice conversion on the question-solving copywriting information to obtain a corresponding voice file; performing analysis processing on the voice file to obtain a subtitle file corresponding to the problem solving copywriting information; according to the subtitle file and a first intelligent large model, a rendered video file with a time axis is determined, and the first intelligent large model is used for generating a code for rendering the video file; and according to the voice file and the video file, generating a visual target video corresponding to the problem solving copywriting information. According to the invention, the efficiency of generating the visual question solving video can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device and electronic device for generating a visual problem-solving video. Background Art

[0002] Currently, when tutoring students to practice math, physics and other exercises, if there is a corresponding problem-solving video for each question, visually showing the derivation process of the formula, physical meaning, etc., accompanied by voice explanation, it will be of great value in improving learning outcomes.

[0003] In the related technologies, most of the visual explanation videos of the corresponding exercises are generated through manual editing, which is time-consuming and labor-intensive, resulting in low efficiency in generating visual problem-solving videos and failing to achieve the effect of "one video for one question". Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a method, device and electronic device based on the generation of visual problem-solving videos, so as to solve the problem of low efficiency in generating visual problem-solving videos. For any question of teachers or students, corresponding tutoring problem-solving videos can be generated in real time, thereby greatly facilitating tutoring for problem-solving.

[0005] To solve the above technical problems, the embodiments of the present application are implemented as follows: In a first aspect, an embodiment of the present application provides a method for generating a visual problem-solving video, comprising: obtaining problem-solving text information of a target question, performing voice conversion on the problem-solving text information, and obtaining a corresponding voice file; parsing the voice file to obtain a subtitle file corresponding to the problem-solving text information; determining a rendered video file with a timeline based on the subtitle file and a first intelligent big model, the first intelligent big model being used to generate a code for rendering the video file; and generating a visual target video corresponding to the problem-solving text information based on the voice file and the video file.

[0006] In a second aspect, an embodiment of the present application provides a generation device based on a visual problem-solving video, comprising: an acquisition module, used to obtain the problem-solving text information of a target question, perform voice conversion on the problem-solving text information, and obtain a corresponding voice file; a parsing module, used to parse the voice file, and obtain a subtitle file corresponding to the problem-solving text information; a rendering module, used to determine a rendered video file with a timeline based on the subtitle file and a first intelligent big model, the first intelligent big model being used to generate a code for rendering the video file; a visualization module, used to generate a visualized target video corresponding to the problem-solving text information based on the voice file and the video file.

[0007] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory electrically connected to the processor, the memory storing a computer program, and the processor being used to call and execute the computer program from the memory to implement the above-mentioned method for generating a visual problem-solving video.

[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium for storing a computer program, wherein the computer program can be executed by a processor to implement the above-mentioned method for generating a visual problem-solving video.

[0009] In a fifth aspect, an embodiment of the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the above-mentioned method for generating a visual problem-solving video.

[0010] Adopt the technical scheme of the embodiment of the present application, obtain the solution text information of the target question, perform voice conversion on the solution text information, and obtain the corresponding voice file; perform parsing on the voice file to obtain the subtitle file corresponding to the solution text information; determine the rendered video file with a timeline according to the subtitle file and the first intelligent large model, and the first intelligent large model is used to generate the code for rendering the video file; generate the visualized target video corresponding to the solution text information according to the voice file and the video file. It can be seen that by performing voice conversion on the solution text information of the target question to obtain the voice file, the voice file is parsed to obtain the corresponding subtitle file. Since the subtitle file has a timeline, by inputting the subtitle file into the first intelligent large model and outputting the code for rendering the video file, the video file with a timeline can be automatically generated based on the subtitle file and the first intelligent large model. The video file and the voice file are synthesized into a visualized target video corresponding to the target question, achieving the effect of "one question and one video", and there is no need for the user to edit the video of the solution text information corresponding to the target question in advance, which can improve the efficiency of generating a visualized solution video. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate one or more embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in one or more embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0012] Figure 1 It is a flowchart of a method for generating a visual problem-solving video provided by an embodiment of the present application; Figure 2 It is a schematic diagram of a discrete coordinate system of a math problem provided in an embodiment of the present application.

[0013] Figure 3 This is a screenshot of the video rendered by the Manim code provided in the embodiment of the present application.

[0014] Figure 4 It is a schematic diagram of the process structure of automatically generating a video by a computer or target software provided in an embodiment of the present application.

[0015] Figure 5 It is a scene flow diagram of a method for generating a visual problem-solving video provided by another embodiment of the present application; Figure 6 It is a structural schematic diagram of a generation device based on a visual problem-solving video provided in an embodiment of the present application; Figure 7 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0016] The embodiments of the present application provide a method, device and electronic device for generating a visual problem-solving video, so as to solve the problem of low efficiency in generating a visual problem-solving video.

[0017] In order to enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this application.

[0018] The method for generating a visual problem-solving video provided in the embodiment of the present application can be executed by an electronic device, or by software installed in the electronic device. Specifically, the electronic device can be a terminal device or a server device. Among them, the terminal device can include a smart phone, a laptop computer, a smart wearable device, a vehicle terminal, etc., and the server device can include an independent physical server, a server cluster composed of multiple servers, or a cloud server capable of cloud computing.

[0019] In the following, in combination with the accompanying drawings, a method for generating a visual problem-solving video provided by an embodiment of the present application is described in detail through specific embodiments and application scenarios.

[0020] Figure 1 A flow chart of a method for generating a visual problem-solving video according to an embodiment of the present application is shown, and the method comprises the following steps: S102, obtaining the solution text information of the target question, performing voice conversion on the solution text information, and obtaining a corresponding voice file.

[0021] The target topic includes: the topic text entered by the user or the topic picture taken by the user.

[0022] The problem-solving copy information includes: by answering the target question, copy information is generated to describe the problem-solving steps corresponding to the target question.

[0023] Voice conversion includes: converting the problem-solving text information into corresponding voice files through voice conversion software such as Text To Speech (TTS) software.

[0024] S104, parsing the voice file to obtain a subtitle file corresponding to the solution text information.

[0025] The analysis process includes: analyzing the voice file by voice recognition software to generate the analyzed voice file, that is, the subtitle file.

[0026] Subtitle files are specifically text format (Subripper Text, SRT) subtitle files, such as SRT subtitle files. SRT subtitle files are files with timelines. Specifically, SRT subtitle files are a timeline-based subtitle format that contains subtitle text and the time period for each line of subtitle text. Each subtitle entry consists of four parts: subtitle number, timestamp, subtitle text, and blank line.

[0027] Since the first intelligent large model can automatically generate text information but cannot determine the exact timeline, the subtitle file is obtained by processing the problem-solving text information, which can determine the accurate time information for the subsequent generation of the target video.

[0028] S106, determining a rendered video file with a timeline according to the subtitle file and the first intelligent large model.

[0029] Among them, the first intelligent large model is used to generate code for rendering video files.

[0030] The first intelligent big model includes: a "large parameter" model trained through large-scale data and powerful computing power, which usually has a high degree of versatility and generalization ability, such as a basic big model, etc., which is used to call the corresponding required functions through prompt words and generate code for rendering video files. The first intelligent big model is generally a general big model that does not need to be trained again. Specifically, the prompt words include: a guiding text generated for the first intelligent big model, triggering the first intelligent big model to generate content in a specific direction or perform a special task, such as: a simple question, a description, a command or any form of text input. The required functions include: generating content in a specific direction or performing special tasks, etc. For example: the prompt word is "generate code for coordinate axes", and the first intelligent big model can generate code for rendering coordinate axes.

[0031] By calling the first intelligent big model, the subtitle file is input into the first intelligent big model. Since the subtitle file has a timeline, the code generated by the first intelligent big model according to the subtitle file also has a timeline. Then, according to the subtitle file and the generated code, the rendered video file also has a timeline.

[0032] As an example, the problem-solving text information corresponding to the subtitle file is the detailed solution process of a math problem, so the subtitle file is the detailed solution process of the math problem with time information. The subtitle file is input into the first intelligent big model, and the first intelligent big model generates a code for rendering the video corresponding to the subtitle file. According to the code and the corresponding subtitle file, a video file with timeline information is generated. Each frame corresponding to the video file can include a specific display of the detailed solution process, such as the changes of the corresponding coordinate axis or coordinate points can be displayed in the video.

[0033] S108, generating a visual target video corresponding to the problem-solving copy information according to the voice file and the video file.

[0034] Since the voice file includes the voice of the problem-solving copy information, the video file includes the video of the problem-solving process corresponding to the problem-solving copy information generated according to the subtitle file, and the subtitle file includes the subtitles of the problem-solving copy information, the voice file and the video file are combined and processed to obtain a visual problem-solving detailed process with voice, subtitles and video corresponding to the problem-solving copy information.

[0035] Adopt the technical scheme of the embodiment of the present application, obtain the solution text information of the target question, perform voice conversion on the solution text information, and obtain the corresponding voice file; perform parsing on the voice file to obtain the subtitle file corresponding to the solution text information; determine the rendered video file with a timeline according to the subtitle file and the first intelligent large model, and the first intelligent large model is used to generate the code for rendering the video file; generate the visualized target video corresponding to the solution text information according to the voice file and the video file. It can be seen that by performing voice conversion on the solution text information of the target question to obtain the voice file, the voice file is parsed to obtain the corresponding subtitle file. Since the subtitle file has a timeline, by inputting the subtitle file into the first intelligent large model and outputting the code for rendering the video file, the video file with a timeline can be automatically generated based on the subtitle file and the first intelligent large model. The video file and the voice file are synthesized into a visualized target video corresponding to the target question, achieving the effect of "one question and one video", and there is no need for the user to edit the video of the solution text information corresponding to the target question in advance, which can improve the efficiency of generating a visualized solution video.

[0036] In one embodiment, to obtain the solution text information of the target question (ie, S102), the following steps A1-A3 may be performed: Step A1, obtain the target question, answer the target question through the first intelligent big model, and obtain the solution steps corresponding to the target question.

[0037] Obtaining the target topic includes obtaining the specific text content of the target topic and obtaining the information of the topic image of the target topic.

[0038] The problem-solving steps include: the calculation process of obtaining the answer to the target question. For example: the user enters the text of the target question, or takes and uploads a picture of the target question. The target question is input into the first intelligent big model, and the first intelligent big model answers the target question to obtain the problem-solving steps corresponding to the target question. The specific problem-solving process using the first intelligent big model includes: training the first intelligent big model to generate the problem-solving steps corresponding to the target question; the problem-solving steps of each target question can also be stored according to artificial pre-setting in advance for the convenience of calling the first intelligent big model; or the corresponding problem-solving steps are obtained by reasoning according to the inference rule library, such as generating the problem-solving code through the ability of generating dynamic language Python code in the first intelligent big model, so as to generate the corresponding problem-solving steps based on the inference rule library, mathematical templates, etc., without specific limitation.

[0039] It should be noted that after obtaining the target question, the corresponding problem-solving steps can also be obtained by manually answering the target question in real time.

[0040] As an example, Figure 2 As shown in the figure, it is a schematic diagram of a discrete coordinate system for a math problem, in which the target problem has Figure 2 The coordinate system in the scatter plot also includes: Based on the best fit line in the scatter plot, predict the y value when x = 2. The options include: A: 7, B: 3, C: 1, D: 0.

[0041] The steps to solve the problem include: 1. Observation Figure 2 : First, we need to carefully observe the given scatter plot and the best fit line. The best fit line is a straight line that is as close as possible to the data points in the graph.

[0042] 2. Find the position of x=2: Find the position of 2 on the x-axis.

[0043] 3. Find the corresponding y value: Move vertically up or down from the x=2 position until it intersects the best fit line.

[0044] 4. Read the y value: observe the corresponding value of the intersection on the y-axis.

[0045] 5. According to Figure 2 As shown, when x=2, the intersection of the best fit line and the y-axis is approximately at y=3.

[0046] Therefore, the answer is B:3.

[0047] Summary: This question mainly tests the understanding of scatter plots and best fit lines, as well as the ability to make numerical predictions based on diagrams. The key to solving the problem is to accurately find the point on the best fit line corresponding to the x value and read its y value.

[0048] Step A2, determining the second intelligent large model to be trained according to preset core indicators, wherein the preset core indicators include: one or more of the quality of the generated problem-solving copy information, the accuracy of the generated problem-solving copy information, and the speed of generating the problem-solving copy information.

[0049] The second intelligent big model includes a general big model that meets user needs, such as an artificial intelligence big model. The general big model is a model that performs well in a variety of tasks and applications, is not limited to a specific field, and can be generalized in different fields. The selection of the second intelligent big model includes: first, the selection of the general big model base corresponding to the second intelligent big model, such as various open source models with big model functions. Secondly, the parameter scale of the big model base is selected, such as selecting a suitable parameter scale according to the characteristics and difficulty of the target questions to be answered. Taking the big model with natural language processing function in multi-task scenarios as an example, there are parameters such as 1B, 3B, 7B, 14B, and 72B. The larger the parameter, the stronger the big model with natural language processing function in multi-task scenarios, but the slower the speed and the more resources consumed. Conversely, if the parameter is not large enough, the big model with natural language processing function in multi-task scenarios is insufficient, and the understanding and generation of the question may be biased, resulting in errors.

[0050] It can be seen that the second intelligent big model to be trained can be determined by determining the preset core indicators, that is, one or more of the quality of the problem-solving copy information generated by the general big model as the base, the accuracy of the problem-solving copy information generated, and the speed of generating the problem-solving copy information (how many words generated per minute) are used as assessment indicators, and the general big models with different parameters are tested respectively, and the general big model with the smallest parameter scale that meets the user's preset core indicators is selected, and the general big model is used as the base of the second intelligent big model to be trained for training the second intelligent big model to be trained.

[0051] Step A3, train the second intelligent big model to be trained, input the target question and problem-solving steps into the trained second intelligent big model, and output the problem-solving copy information through the trained second intelligent big model.

[0052] After training the second intelligent big model to be trained, a trained second intelligent big model is obtained, and the target question and problem-solving steps obtained in step A1 are input into the trained second intelligent big model, so that the problem-solving copy information can be output.

[0053] Following the example of step A1 above, according to the above problem-solving steps, the corresponding problem-solving copy information obtained includes: "Let's take a look at this question. The question is, in the given scatter plot, when x is equal to 2, this best fit line can help us predict what y is approximately? First, let's understand what a scatter plot and a best fit line are.

[0054] Scatter plot: You can think of it as a map with dots on it to represent the relationship between different things. For example, we can use a scatter plot to represent the relationship between a person's height and weight. Each dot represents a person's height and weight.

[0055] Best fit line: Imagine that we want to use a straight line to describe the overall trend of all the points in the scatter plot. This straight line is the best fit line. It is like a rope that passes through all the points as much as possible, but it does not necessarily pass through every point.

[0056] Now, let’s get back to this question.

[0057] 1. Find the position where x=2: Find the number 2 on the horizontal axis (x-axis).

[0058] 2. Draw a vertical line: Draw a straight line from x=2 upward until the line touches the best fit line.

[0059] 3. Find the intersection point: The point where the straight line just drawn intersects with the best fit line is the point we found.

[0060] 4. Take a look at the y-axis: Draw a horizontal line from the intersection to the left, and this horizontal line will pass through the vertical axis (y-axis). See what number this horizontal line corresponds to on the y-axis, which is our predicted y value.

[0061] Why can we do this? Because the best fit line can help us predict. When we know an x ​​value, we can use this line to roughly estimate the corresponding y value.

[0062] Now, let’s count together. When x=2, what is y approximately? 1, 2, 3, yes, 3! So, the answer to this question is B: 3.

[0063] To summarize: We observe the scatter plot, find the best fit line, and then use the best fit line to predict the corresponding y value based on the given x value.

[0064] Classmates, have you learned everything? If you still don’t understand, you can always ask me! " In this embodiment, problem-solving steps are obtained through the target question, and the second intelligent big model to be trained is accurately selected according to the core indicators preset by the user. By training the second intelligent big model to be trained, the trained second intelligent big model is obtained. The trained second intelligent big model can output the problem-solving copy information corresponding to the target question. Compared with the general big model, the language style of the trained second intelligent big model will be more in line with the style of teachers tutoring students, and can describe the problem-solving steps in a more step-by-step and persuasive manner, so that learners can have a clearer understanding of the target question.

[0065] In one embodiment, the training of the second intelligent big model (i.e., step A3) may perform the following steps B1-B3: Step B1, obtain multiple sample target questions, and sample problem-solving steps and sample answers corresponding to the multiple sample target questions, input each sample target question and the sample problem-solving steps and sample answers corresponding to the sample target questions into the second intelligent large model to be trained, and obtain a detailed explanation of the sample problem-solving.

[0066] The sample target topic includes: the specific text content of the sample target topic, or the information of the topic image of the sample target topic.

[0067] The sample problem-solving steps include: the calculation process to obtain the answer to the sample target question.

[0068] Sample problem solving solutions are provided to help you understand the detailed steps of the problem solving process, and can also be detailed solutions to the target questions.

[0069] Specifically, we ask people (such as teachers, annotation engineers, etc.) to write scripts for sample target questions. The scripts are sample solutions to sample target questions obtained based on the sample solution steps. People write scripts that can teach students step by step and in a gentle and persuasive manner according to their understanding of tutoring students. Multiple versions of scripts can be written for the same question, and the quality of these scripts is manually ranked according to whether they are easy to understand, concise, and in line with the syllabus, so that we can get sample solution solutions in order.

[0070] Among them, a sample target question can correspond to a sample solution step and multiple sample solution details. However, among the multiple sample solution details, a most accurate script can be determined, and the most accurate script is determined as the sample solution copy information. The sample solution copy information includes: the final detailed solution process corresponding to the target question, which is used to specifically describe the sample solution steps.

[0071] It should be noted that the ranking of multiple sample problem-solving explanations can be used as feedback information for subsequent training of the second intelligent large model, and each sentence in the multiple sample problem-solving explanations can be ranked according to whether they are easy to understand, concise, and in line with the teaching syllabus; the multiple sample problem-solving explanations as a whole can also be ranked according to whether they are easy to understand, concise, and in line with the teaching syllabus.

[0072] As an example, multiple sample target questions, sample problem-solving steps and sample answers are used as training data sets. For example, 1,000 questions are obtained as sample target questions, and sample problem-solving steps and sample answers corresponding to the 1,000 questions are obtained. Personnel are asked to write multiple scripts for the sample target questions based on the sample problem-solving steps and sample answers as sample problem-solving explanations.

[0073] Step B2, based on the detailed solution of the sample problem, fine-tune the second intelligent large model to be trained through fine-tuning technology to obtain the second intelligent large model to be trained after fine-tuning.

[0074] By fine-tuning the Fine-tuning technology, the second intelligent large model to be trained can learn more general feature representations, so that it can perform better on new data.

[0075] Specifically, according to step A1, multiple sample solutions corresponding to the target questions can be obtained. By fine-tuning the second intelligent big model to be trained, the fine-tuned second intelligent big model to be trained can be obtained. The fine-tuned second intelligent big model to be trained can accurately output the sample solutions.

[0076] Step B3, using the reward model to perform alignment training on the second intelligent large model to be trained after fine-tuning, to obtain the trained second intelligent large model.

[0077] The reward model includes a model obtained by training by obtaining a plurality of sample problem-solving detailed solutions in a sequential order, and is used for alignment training of the second intelligent large model to be trained after fine-tuning.

[0078] Aligned training RLHF, which refers to reinforcement learning based on human feedback information, is the most popular training method for neural networks. The feedback information can include the order of the sample solution details in step B1.

[0079] Specifically, multiple sample solutions for the same target question are obtained, the sample solutions are sorted, and sample solutions with order are obtained. The sample solutions with order are used as a training set to train the reward model to be trained, and the trained reward model is obtained. The reward model obtained by training is used to fine-tune the parameters of the second intelligent large model to be trained after fine-tuning, so that the sample solutions output by the second intelligent large model to be trained are more accurate, and the accurate sample solutions are used as sample solution copy information. For example, when the second intelligent large model can output script A, script B, and script C, according to user feedback information, the order of the scripts is that the quality of script A is greater than that of script B and greater than that of script C, and the reward model obtained by training can obtain the order. The second intelligent large model to be trained after fine-tuning outputs script B. The reward model can perform alignment training on the second intelligent large model to be trained after fine-tuning, and inform the second intelligent large model to be trained after fine-tuning that the output result is inaccurate. Through fine-tuning learning again, the result output by the second intelligent large model to be trained is script A, and script A is used as the sample solution copy information. Therefore, the second intelligent large model to be trained after alignment training through the reward model can output more accurate sample solution copy information as the trained second intelligent large model. The feedback information can include the ranking of multiple sample solution details, and can also include the latest feedback information of the user on the output results of the second intelligent large model.

[0080] In this embodiment, the trained second intelligent big model is obtained by fine-tuning the second intelligent big model to be trained and using the reward model for alignment training. The trained second intelligent big model can output more accurate sample problem-solving copy information.

[0081] In one embodiment, the voice file is parsed to obtain a subtitle file corresponding to the solution text information (ie, S104), and the following steps C1-C2 may be performed: Step C1, by parsing the voice file, a plurality of split subtitle files after parsing are obtained.

[0082] Among them, before parsing the solution copy information, the first intelligent big model is used to check the solution copy information. After confirming that it is consistent with the problem-solving steps and the answer is correct, the solution copy information is parsed and processed. The parsing process includes: using TTS software to convert the solution copy information into a voice file, parsing the voice file through voice recognition software, and generating a parsed voice file. The parsed voice file is a text file with time sequence.

[0083] The parsed voice file is split and processed to obtain multiple split subtitle files with time information.

[0084] Step C2, determining the subtitle file with the timestamp information format corresponding to the solution text information based on the multiple split subtitle files.

[0085] Since the split subtitle file includes multiple text files with time information, the split subtitle files are merged into a subtitle file with a timestamp information format according to a preset time sequence, and the preset time sequence can be a time sequence. The split subtitle files obtained after splitting the voice file are generated according to the time sequence as corresponding subtitle files with a timestamp information format. As an example, the complete voice file includes 6 minutes, and the starting point 0 minute 0 second to 3 minute 20 seconds is used as the first split subtitle file, and the 3 minute 21 second and the end point 6 minute 0 second are the second split subtitle file. The first split subtitle file and the second split subtitle file are sorted according to time and combined into a subtitle file with a timestamp information format.

[0086] In this embodiment, by splitting the parsed voice file to obtain multiple split subtitle files including time information, and merging the multiple split subtitle files into a subtitle file with timestamp information, the problem that the first intelligent large model can only recognize the text but cannot generate an accurate video timeline can be solved, and the accurate video frame sequence can be provided for the final generated target video.

[0087] In one embodiment, according to the subtitle file and the first intelligent big model, the rendered video file with a timeline is determined (ie, S106), and the following steps D1-D2 may be performed: Step D1, calling the first intelligent big model, inputting the subtitle file into the first intelligent big model, and generating a code for rendering the video file through the first intelligent big model.

[0088] Since the subtitle file has a timeline, it can play the role of a storyboard. For example, following the example in step C2 above, the first split subtitle file is used as the first storyboard, and the second split subtitle file is used as the second storyboard.

[0089] By calling the first intelligent large model and using the prompt word in the first intelligent large model, the prompt word can be a split subtitle file of the subtitle file, and the code required for rendering the video file corresponding to the subtitle file is generated. Following the example in the above step D1, the code required for rendering the video file is Manim code, which is a Python library for creating mathematical animations, mainly used to make dynamic demonstrations of mathematical formulas and concepts. The Manim code may include a first Manim code and a second Manim code, the first Manim code is used to generate a video corresponding to the first split subtitle file, and the second Manim code is used to generate a video corresponding to the second split subtitle file.

[0090] Step D2, according to the code, determine one or more video screenshots, and generate a video file with a timeline based on the video screenshots.

[0091] Since the subtitle file has timestamp information, according to the timestamp information and the code corresponding to each split subtitle file generated by the first intelligent large model, in the environment of executing the code, a video screenshot corresponding to each split subtitle file of the subtitle file is generated, or the split subtitle file can also correspond to a video with the same timestamp information, and a video file with a timeline is generated according to multiple video screenshots corresponding to multiple split subtitle files in chronological order. It can be seen that since the subtitle file has a timestamp information format, the video of each split subtitle file can be merged into a video file with a timeline.

[0092] It should be noted that the obtained complete video file with a timeline is input into a large model with video understanding capabilities, and the obtained video file is proofread to check whether the generated video file with a timeline is coherent and complete, and to check whether there are errors and contradictions between the video file and the problem-solving steps and the problem-solving copy information. The large model with video understanding capabilities can be the first intelligent large model or a large model with this capability after training. If there are errors, the generation process of the subtitle file and the code generation process are adjusted accordingly. Among them, the integrity and correctness of the video file can also be checked manually.

[0093] In this embodiment, by calling the first intelligent big model, based on the subtitle file, a code for rendering the video file is generated, and by executing the code, a video file including a timeline is generated. A video file with timeline information can be automatically generated based on the first intelligent big model and the subtitle file.

[0094] In one embodiment, after the code for rendering the video file is generated by the first intelligent big model (ie, step D1), the following steps E1-E2 may be performed: Step E1, inputting the code into the sandbox environment to run the code, and determining the running result of the code.

[0095] A sandbox environment refers to an isolated testing environment that allows running software, programs, or code in a closed setting without affecting external systems or environments.

[0096] By testing the code in a sandbox environment, you can determine the results of the code.

[0097] Step E2: when the running result is an error result, the error result is fed back to the first intelligent big model, and the code is modified by the first intelligent big model to obtain the correct result.

[0098] When the running result is an error, the error result is fed back to the first intelligent big model, and the code is modified by the first intelligent big model. Until no error is reported, the execution result of the code without error is returned to the first intelligent big model, and the first intelligent big model is checked again. If there is an error, the code is modified by the first intelligent big model, and this is repeated until there is no problem. Among them, the execution result can be a video, and the video screen can be various mathematical curves drawn.

[0099] For example, Figure 3 As shown in FIG. 1 , a screenshot of the video rendered by the Manim code is shown. According to the screenshot of the video, when X=2, the y value corresponding to the best fit line can be seen from the figure.

[0100] In this embodiment, by trial running the code in a sandbox environment, the accuracy of the code can be ensured, and the generated code can be tested, so that the accuracy of the target video generated subsequently is higher.

[0101] In one embodiment, based on the voice file and the video file, a visualized target video corresponding to the problem-solving copy information is generated (ie, S108), and the following step F may be performed: Step F, based on the voice file and the video file, combined with the preset template of the target video, generates a visual target video corresponding to the solution copy information.

[0102] The preset template refers to a generated template corresponding to a target video set on a target software, such as a pre-set cover, fixed ending, and background music.

[0103] Since the video file is generated based on the subtitle file, it can be seen that the video file has a timeline. Then, since the voice file also has a timeline, through video synthesis tools, such as multimedia video processing tools (Fast Forward Mpeg, ffmpeg) and other open source video processing tools, the subtitle file, voice file and video file are synthesized according to the timeline information to generate a video file with the same timeline information, and the video file with the same timeline information is merged with the preset template to generate a visual target video corresponding to the solution copy information.

[0104] In this embodiment, by merging the generated voice file and video file with the preset template to generate a visual target video corresponding to the solution text information, the generated target video can be made more complete. The target video can include voice, subtitles, animation, and content in the preset template such as cover, generating a complete visual target video that is easier to understand and watch.

[0105] In one embodiment, Figure 4As shown, it is a schematic diagram of the process structure of automatically generating videos by a computer or target software. When the user enters the target question in the target software, the problem-solving steps are obtained, and the button in the target software can be "Please give me an explanation video" to automatically generate the explanation video corresponding to the target question. Specifically, the problem-solving steps corresponding to the target question are obtained, and the problem-solving copy information is generated by calling the second intelligent large model, and the problem-solving copy information is converted into a voice file through TTS, and the voice file is converted into a subtitle file in SRT format using voice recognition software. The first intelligent large model is called to generate a Manim code animation script based on the subtitle file, and a video file including a timeline is obtained by optimizing the script corresponding to the video and the proofreading script, and the video file and the voice file are synthesized and processed, and a preset synthesis template is used to obtain a visualized target video, and finally the target video is output to explain the corresponding problem-solving steps for obtaining the answer to the target question.

[0106] Figure 5 is a scene flow diagram of a method for generating a visual problem-solving video according to another embodiment of the present application, such as Figure 5 As shown, the method comprises the following steps: S501, obtaining the target topic uploaded by the user to the application software.

[0107] The target topic includes text input by the user into the application software, or a picture of the topic taken and uploaded.

[0108] S502, solving the target question through the first intelligent big model to obtain the problem-solving steps.

[0109] S503, input the problem-solving steps into the second intelligent big model, and use the second intelligent big model to generate problem-solving copy information.

[0110] S504, calling the first intelligent big model to check the problem-solving copy information and determine whether the problem-solving copy information is consistent with the problem-solving steps.

[0111] If there is any inconsistency, the correct solution information can be determined by modifying the parameters of the second intelligent large model.

[0112] In the specific implementation, this step can also be omitted without affecting the implementation of this patent.

[0113] S505, using TTS software to synthesize the problem-solving text information into a voice file.

[0114] S506, using speech recognition software to analyze the synthesized speech file, and generating an SRT subtitle file with a time axis according to the speech file.

[0115] S507, input the SRT subtitle file into the first intelligent large model, and generate the Manim code required for rendering the video corresponding to the SRT subtitle file through the prompt words of the first intelligent large model.

[0116] S508, by trial running the Manim code in the sandbox environment, the execution result of the Manim code is returned to the first intelligent large model.

[0117] S509, if the execution result is wrong, the Manim code is modified through the first intelligent large model. If the execution result is correct, the Manim code is executed and the corresponding video is obtained as a video file.

[0118] S510, using a large model with video understanding capabilities to check the integrity of the video file to obtain a correct video file.

[0119] S511, merging the correct video file, voice file and preset template to obtain a visualized target video corresponding to the problem-solving copy information.

[0120] The specific process from S501 to S511 has been described in detail in the above embodiment and will not be repeated here.

[0121] Adopt the technical scheme of the embodiment of the present application, obtain the solution text information of the target question, perform voice conversion on the solution text information, and obtain the corresponding voice file; perform parsing on the voice file to obtain the subtitle file corresponding to the solution text information; determine the rendered video file with a timeline according to the subtitle file and the first intelligent large model, and the first intelligent large model is used to generate the code for rendering the video file; generate the visualized target video corresponding to the solution text information according to the voice file and the video file. It can be seen that by performing voice conversion on the solution text information of the target question to obtain the voice file, the voice file is parsed to obtain the corresponding subtitle file. Since the subtitle file has a timeline, by inputting the subtitle file into the first intelligent large model and outputting the code for rendering the video file, the video file with a timeline can be automatically generated based on the subtitle file and the first intelligent large model. The video file and the voice file are synthesized into a visualized target video corresponding to the target question, achieving the effect of "one question and one video", and there is no need for the user to edit the video of the solution text information corresponding to the target question in advance, which can improve the efficiency of generating a visualized solution video.

[0122] In summary, specific embodiments of the present subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recorded in the claims can be performed in a different order and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing can be advantageous.

[0123] The above is a method for generating a visual problem-solving video provided in an embodiment of the present application. Based on the same idea, an embodiment of the present application also provides a device for generating a visual problem-solving video.

[0124] Figure 6 is a schematic diagram of the structure of a device for generating a visual problem-solving video according to an embodiment of the present invention. Figure 6 As shown, the generation device based on the visual problem-solving video includes: an acquisition module 61, a parsing module 62, a rendering module 63, and a visualization module 64: The acquisition module 61 is used to acquire the solution text information of the target question, perform voice conversion on the solution text information, and obtain the corresponding voice file; The parsing module 62 is used to parse the voice file to obtain a subtitle file corresponding to the solution text information; A rendering module 63, used to determine a rendered video file with a timeline according to the subtitle file and the first intelligent big model, the first intelligent big model is used to generate a code for rendering the video file; The visualization module 64 is used to generate a visualized target video corresponding to the problem-solving text information based on the voice file and the video file.

[0125] In one embodiment, the acquisition module 61 includes: An acquisition unit, used to acquire a target question, answer the target question through the first intelligent big model, and obtain the solution steps corresponding to the target question; A determination unit, used to determine the second intelligent large model to be trained according to preset core indicators, where the preset core indicators include: one or more of the quality of generated problem-solving copy information, the accuracy of generated problem-solving copy information, and the speed of generating problem-solving copy information; The training unit is used to train the second intelligent big model to be trained, and input the target question and problem-solving steps into the trained second intelligent big model, and output the problem-solving copy information through the trained second intelligent big model.

[0126] In one embodiment, the training unit is specifically used to obtain multiple sample target questions, sample problem-solving steps and corresponding sample problem-solving explanations, input the multiple sample target questions and multiple sample problem-solving steps into a second intelligent big model to be trained, and output multiple sample problem-solving copy information; based on the sample problem-solving explanations and sample problem-solving copy information, fine-tune the second intelligent big model to be trained through fine-tuning technology to obtain the second intelligent big model to be trained after fine-tuning; use the reward model to perform alignment training on the second intelligent big model to be trained to obtain the trained second intelligent big model.

[0127] In one embodiment, the parsing module 62 is used to parse the voice file to obtain multiple parsed subtitle files; based on the multiple split subtitle files, determine the subtitle file with the timestamp information format corresponding to the solution text information.

[0128] In one embodiment, the rendering module 63 is specifically used to call the first intelligent big model, input the subtitle file into the first intelligent big model, generate a code for rendering the video file through the first intelligent big model; according to the code, determine one or more video screenshots, and generate a video file with a timeline based on the video screenshots.

[0129] In one embodiment, the device also includes a feedback module, which is used to determine the running result of the code by inputting the code into a sandbox environment for running; when the running result is an error result, the error result is fed back to the first intelligent big model, and the code is modified by the first intelligent big model to obtain the correct result.

[0130] In one embodiment, the visualization module 64 is specifically used to generate a visualized target video corresponding to the problem-solving text information based on the voice file and the video file in combination with a preset template of the target video.

[0131] Adopt the technical scheme of the embodiment of the present application, obtain the solution text information of the target question, perform voice conversion on the solution text information, and obtain the corresponding voice file; perform parsing on the voice file to obtain the subtitle file corresponding to the solution text information; determine the rendered video file with a timeline according to the subtitle file and the first intelligent large model, and the first intelligent large model is used to generate the code for rendering the video file; generate the visualized target video corresponding to the solution text information according to the voice file and the video file. It can be seen that by performing voice conversion on the solution text information of the target question to obtain the voice file, the voice file is parsed to obtain the corresponding subtitle file. Since the subtitle file has a timeline, by inputting the subtitle file into the first intelligent large model and outputting the code for rendering the video file, the video file with a timeline can be automatically generated based on the subtitle file and the first intelligent large model. The video file and the voice file are synthesized into a visualized target video corresponding to the target question, achieving the effect of "one question and one video", and there is no need for the user to edit the video of the solution text information corresponding to the target question in advance, which can improve the efficiency of generating a visualized solution video.

[0132] Those skilled in the art should understand that Figure 6 The generation device based on visual problem-solving video can be used to implement the generation method based on visual problem-solving video described above. The detailed description should be similar to the description of the method part above. To avoid tediousness, it will not be repeated here.

[0133] Based on the same technical concept, an embodiment of the present application further provides an electronic device, which is used to execute the above-mentioned method for generating a visual problem-solving video. Figure 7 A schematic diagram of the structure of an electronic device for implementing various embodiments of the present application. The electronic device may have relatively large differences due to different configurations or performances, and may include a processor (processor) 710, a communication interface (Communications Interface) 720, a memory (memory) 730 and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 may call a computer program stored in the memory 730 and executable on the processor 710 to perform the following steps: Obtaining the solution text information of the target question, performing voice conversion on the solution text information, and obtaining the corresponding voice file; Parse the voice file to obtain the subtitle file corresponding to the solution text information; Determine a rendered video file with a timeline according to the subtitle file and the first intelligent big model, the first intelligent big model is used to generate a code for rendering the video file; Based on the voice file and video file, a visual target video corresponding to the problem-solving copy information is generated.

[0134] Adopt the technical scheme of the embodiment of the present application, obtain the solution text information of the target question, perform voice conversion on the solution text information, and obtain the corresponding voice file; perform parsing on the voice file to obtain the subtitle file corresponding to the solution text information; determine the rendered video file with a timeline according to the subtitle file and the first intelligent large model, and the first intelligent large model is used to generate the code for rendering the video file; generate the visualized target video corresponding to the solution text information according to the voice file and the video file. It can be seen that by performing voice conversion on the solution text information of the target question to obtain the voice file, the voice file is parsed to obtain the corresponding subtitle file. Since the subtitle file has a timeline, by inputting the subtitle file into the first intelligent large model and outputting the code for rendering the video file, the video file with a timeline can be automatically generated based on the subtitle file and the first intelligent large model. The video file and the voice file are synthesized into a visualized target video corresponding to the target question, achieving the effect of "one question and one video", and there is no need for the user to edit the video of the solution text information corresponding to the target question in advance, which can improve the efficiency of generating a visualized solution video.

[0135] The specific execution steps can refer to the various steps of the above-mentioned embodiment of the method for generating a visual problem-solving video, and the same technical effect can be achieved. To avoid repetition, they will not be repeated here.

[0136] It should be noted that the electronic devices in the embodiments of the present application include: servers, terminals, or other devices except terminals.

[0137] The above electronic device structure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. For example, the input unit may include a graphics processing unit (GPU) and a microphone, and the display unit may be configured with a display panel in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit includes a touch panel and at least one of other input devices. The touch panel is also called a touch screen. Other input devices may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.

[0138] The memory can be used to store software programs and various data. The memory may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, an application program or instructions required for at least one function (such as a sound playback function, an image playback function, etc.), etc. In addition, the memory may include a volatile memory or a non-volatile memory, or the memory may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct memory bus random access memory (DRRAM).

[0139] The processor may include one or more processing units; optionally, the processor integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to the operating system, user interface, and application programs, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It is understandable that the modem processor may not be integrated into the processor.

[0140] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned embodiment of the method for generating a visual problem-solving video is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0141] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0142] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned embodiment of the method for generating a visual problem-solving video, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0143] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0144] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the processor is used to run the program or instructions to implement the various processes of the above-mentioned product recommendation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0145] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0146] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0147] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A method for generating a visual problem-solving video, characterized in that: The method comprises: Obtaining solution text information of a target question, performing voice conversion on the solution text information, and obtaining a corresponding voice file; Parsing the voice file to obtain a subtitle file corresponding to the problem-solving text information; Determine a rendered video file with a timeline according to the subtitle file and the first intelligent big model, wherein the first intelligent big model is used to generate a code for rendering the video file; A visual target video corresponding to the problem-solving text information is generated based on the voice file and the video file.

2. The method according to claim 1, characterized in that The step of obtaining the solution text information of the target question includes: Obtain the target question, and answer the target question by using the first intelligent big model to obtain the problem-solving steps corresponding to the target question; Determine the second intelligent large model to be trained according to the preset core indicators, wherein the preset core indicators include: one or more of the quality of generating the problem-solving copy information, the accuracy of generating the problem-solving copy information, and the speed of generating the problem-solving copy information; The second intelligent big model to be trained is trained, and the target question and the problem-solving steps are input into the trained second intelligent big model, and the problem-solving copy information is output through the trained second intelligent big model.

3. The method according to claim 2, characterized in that The training of the second intelligent large model includes: Obtain a plurality of sample target questions, and a plurality of sample problem-solving steps and sample answers corresponding to the sample target questions, and input each of the sample target questions and the sample problem-solving steps and sample answers corresponding to the sample target questions into the second intelligent big model to be trained to obtain a detailed solution to the sample problem; Based on the sample problem-solving detailed explanation, fine-tune the second intelligent big model to be trained by fine-tuning technology to obtain the second intelligent big model to be trained after fine-tuning; The reward model is used to perform alignment training on the second intelligent large model to be trained after fine-tuning to obtain the trained second intelligent large model.

4. The method according to claim 1, characterized in that The parsing process of the voice file to obtain a subtitle file corresponding to the problem-solving text information includes: By parsing the voice file, a plurality of split subtitle files after parsing are obtained; According to the split subtitle file, the subtitle file having a timestamp information format corresponding to the problem-solving text information is determined.

5. The method according to claim 1, characterized in that The step of determining a rendered video file with a timeline according to the subtitle file and the first intelligent model includes: Calling the first intelligent big model, inputting the subtitle file into the first intelligent big model, and generating a code for rendering the video file through the first intelligent big model; According to the code, one or more video screenshots are determined, and based on the video screenshots, the video file with the timeline is generated.

6. The method according to claim 5, characterized in that After the code for rendering the video file is generated by the first intelligent big model, the method includes: By inputting the code into a sandbox environment for execution, determining the execution result of the code; When the running result is an error result, the error result is fed back to the first intelligent big model, and the code is modified by the first intelligent big model to obtain a correct result.

7. The method according to claim 1, characterized in that The step of generating a visualized target video corresponding to the problem-solving text information according to the voice file and the video file includes: According to the voice file and the video file, combined with the preset template of the target video, the visualized target video corresponding to the problem-solving copy information is generated.

8. A device for generating a visual problem-solving video, characterized in that: The device comprises: An acquisition module is used to acquire the solution text information of the target question, perform voice conversion on the solution text information, and obtain a corresponding voice file; A parsing module, used for parsing the voice file to obtain a subtitle file corresponding to the solution text information; A rendering module, used to determine a rendered video file with a timeline according to the subtitle file and the first intelligent big model, wherein the first intelligent big model is used to generate a code for rendering the video file; A visualization module is used to generate a visualized target video corresponding to the problem-solving text information based on the voice file and the video file.

9. An electronic device, characterized in that: It includes a processor and a memory electrically connected to the processor, the memory stores a computer program, and the processor is used to call and execute the computer program from the memory to implement a method for generating a visual problem-solving video as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The storage medium is used to store a computer program, and the computer program can be executed by a processor to implement a method for generating a visual problem-solving video as described in any one of claims 1 to 7.