Explanation video generation method, device, computer equipment and storage medium
By generating and integrating videos with target file content and commentary text, and automatically creating commentary videos, the problem of low manual recording efficiency in the existing technology is solved, and efficient generation of commentary videos is achieved.
Patent Information
- Application Number
- CN202111269644.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-10-29
AI Technical Summary
In the prior art, generating commentary videos requires a lot of manual recording, resulting in inefficiency.
Generate the first video and the commentary text based on the content of the target file to generate the second video, and merge it, automatically create the commentary video, reducing the manual recording steps.
It realizes that the commentary video can be automatically generated without manual recording, saving manpower and time, and improving generation efficiency.
Smart Images

Figure CN114339074B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computer technology, and in particular, to a method, device, computer device, and storage medium for generating an explanatory video. Background Art
[0002] With the rapid development of computer technology, video, as a medium for information expression, is increasingly widely used in people's daily lives. In scenarios such as news broadcasts, weather forecasts, online teaching, and sports commentaries, videos can be produced to display and explain relevant information.
[0003] In the related art, a real person explains the information, and the process of explaining the information is recorded to obtain an explanatory video corresponding to the information. However, since manual video production requires a large amount of manpower and time, the efficiency of producing explanatory videos is low. Summary of the Invention
[0004] Embodiments of the present application provide a method, device, computer device, and storage medium for generating an explanatory video, which can improve the efficiency of generating explanatory videos. The technical solutions are as follows:
[0005] On the one hand, a method for generating an explanatory video is provided. The method includes:
[0006] In response to an instruction for generating an explanatory video of a target file, a first video is generated based on the content of the target file, and the pictures in the first video are used to display the content in the target file;
[0007] Based on the explanatory text corresponding to the target file, a second video is generated, and the second video includes the explanatory audio and explanatory pictures corresponding to the explanatory text;
[0008] The first video and the second video are fused to obtain an explanatory video corresponding to the target file.
[0009] Optionally, the method further includes:
[0010] In response to a content switching operation on the first content, the content displayed in the target file is switched from the first content to the second content; or,
[0011] When the content display duration of the first content reaches the target duration corresponding to the first content, the content displayed in the target file is switched from the first content to the second content.
[0012] Optionally, the step of forming the first video by the generated multiple display images includes:
[0013] Based on the content display duration corresponding to each of the display images, a plurality of the display images are combined to form the first video, and the content display duration corresponding to the display image is the display duration of the content in the display image in the target file.
[0014] Optionally, generating the display images including each content in sequence based on a plurality of contents in the target file includes:
[0015] When reading the target file through a browser, taking screenshots of a plurality of contents in the target file in sequence through the browser to obtain the display images including each content respectively.
[0016] Optionally, the explanatory video generation instruction carries the storage address of the target file, and in response to the explanatory video generation instruction for the target file, generating the first video based on the content of the target file includes:
[0017] In response to the explanatory video generation instruction, obtaining the target file based on the storage address;
[0018] Generating the first video based on the content of the target file.
[0019] Optionally, in response to the explanatory video generation instruction for the target file, generating the first video based on the content of the target file includes:
[0020] Sending the explanatory video generation instruction to the browser through an explanatory video generation application;
[0021] Through the browser, in response to the explanatory video generation instruction, generating the display images including each content in sequence based on a plurality of contents in the target file, and sending the generated plurality of display images to the explanatory video generation application;
[0022] Receiving the plurality of display images through the explanatory video generation application, and combining the plurality of display images to form the first video.
[0023] On the other hand, an explanatory video generation device is provided, and the device includes:
[0024] A first video generation module, configured to generate a first video based on the content of the target file in response to the explanatory video generation instruction for the target file, and the pictures in the first video are used to display the content in the target file;
[0025] A second video generation module, configured to generate a second video based on the explanatory text corresponding to the target file, and the second video includes the explanatory audio and explanatory pictures corresponding to the explanatory text;
[0026] A third video generation module, configured to fuse the first video and the second video to obtain an explanatory video corresponding to the target file.
[0027] Optionally, the apparatus further includes:
[0028] An instruction generation module, configured to generate explanatory video generation instructions for each of the target files respectively in response to a batch selection operation on multiple target files;
[0029] A parallel execution module, configured to execute the steps of generating an explanatory video corresponding to each target file in parallel in response to the explanatory video generation instructions for each target file.
[0030] Optionally, the second video generation module includes:
[0031] An audio generation unit, configured to generate the explanatory audio based on the explanatory text;
[0032] An explanatory image generation unit, configured to generate a set of explanatory images, the set of explanatory images including multiple explanatory images, where each explanatory image includes the same explanatory object, and there are at least two explanatory images in the set of explanatory images with different lip shapes of the explanatory object;
[0033] A first generation unit, configured to generate the second video based on the explanatory audio and the set of explanatory images.
[0034] Optionally, the audio generation unit is configured to perform at least one of the following:
[0035] The explanatory video generation instruction carries a model identifier, and calls the audio conversion model indicated by the model identifier to convert the explanatory text into the explanatory audio;
[0036] The explanatory video generation instruction carries an explanatory speed, and converts the explanatory text into an explanatory audio with the explanatory speed.
[0037] Optionally, the explanatory image generation unit is configured to perform at least one of the following:
[0038] The explanatory video generation instruction carries an object identifier, and generates the set of explanatory images based on the explanatory object indicated by the object identifier, so that the multiple explanatory images include the explanatory object;
[0039] The explanatory video generation instruction carries an action identifier, and generates the set of explanatory images based on the dynamic action indicated by the action identifier, so that the action combinations of the explanatory object in the multiple explanatory images form the dynamic action;
[0040] The explanation video generation instruction carries a position parameter, and based on the position indicated by the position parameter, the set of explanation images is generated so that the explained object in the multiple explanation images is located at the position.
[0041] Optionally, the target file includes multiple data segments, the explanation audio includes at least one explanation audio segment corresponding to the data segments, the set of explanation images includes an explanation image subset corresponding to each data segment, and the first generation unit is configured to:
[0042] Generate the explanation audio based on at least one explanation audio segment corresponding to each data segment in the arrangement order of each data segment;
[0043] Generate the explanation video based on the explanation image subset corresponding to each data segment in the arrangement order of each data segment;
[0044] Generate the second video based on the explanation audio and the explanation video, wherein the start playback time point of the explanation audio segment corresponding to the same data segment is the same as the start playback time point of the corresponding explanation image subset.
[0045] Optionally, the target file includes a first data segment and a second data segment, the explanation text includes at least one explanation text segment, wherein the first data segment has a corresponding explanation text segment, and the second data segment does not have a corresponding explanation text segment; the explanation image generation unit is configured to:
[0046] Generate an explanation image subset corresponding to the first data segment, and in the explanation image subset corresponding to the first data segment, there are at least two explanation images with different mouth shapes of the explained object;
[0047] Generate an explanation image subset corresponding to the second data segment, and in the explanation image subset corresponding to the second data segment, the mouth shapes of the explained object in different explanation images are the same.
[0048] Optionally, the explanation audio includes at least one explanation audio segment corresponding to the first data segment, and the explanation image generation unit is configured to:
[0049] Generate a first image subset based on the first duration of the explanation audio segment corresponding to the first data segment, the first image subset includes multiple explanation images and the playback duration corresponding to each explanation image, the mouth shapes of the explained object in different explanation images are different, and the sum of the playback durations corresponding to the multiple explanation images is equal to the first duration;
[0050] When the display duration of the content in the first data segment is greater than the first duration, determine a first duration difference between the display duration of the content and the first duration;
[0051] Based on the first duration difference, generate a second subset of images, where the second subset of images includes at least one of the explanatory images and the playback duration corresponding to each explanatory image. The mouth shapes of the explanatory objects in different explanatory images are the same, and the sum of the playback durations corresponding to at least one of the explanatory images is equal to the first duration difference;
[0052] Combine the first subset of images and the second subset of images to form the subset of explanatory images corresponding to the first data segment.
[0053] Optionally, the first video generation module includes:
[0054] A display image generation unit configured to sequentially generate display images including each content based on multiple contents in the target file;
[0055] A second generation unit configured to form the first video from the multiple generated display images.
[0056] Optionally, the display image generation unit is configured to:
[0057] Read the target file, and when the content displayed in the target file is the first content, generate a display image including the first content;
[0058] When the content displayed in the target file switches from the first content to the second content, continue to generate a display image including the second content until the reading of the target file ends, obtaining multiple display images, where the second content is the next content after the first content.
[0059] Optionally, the device further includes:
[0060] A content switching module configured to, in response to a content switching operation on the first content, switch the content displayed in the target file from the first content to the second content; or,
[0061] The content switching module is configured to, when the display duration of the first content reaches the target duration corresponding to the first content, switch the content displayed in the target file from the first content to the second content.
[0062] Optionally, when the second content is the last content in the target file, the device further includes:
[0063] A file reading module, configured to stop reading the target file when the content display duration of the second content has reached the target duration corresponding to the second content, and the duration from the time point when reading the target file starts to the current time point has reached the playback duration of the second video;
[0064] The file reading module is further configured to continue reading the target file that is displaying the second content when the content display duration of the second content has not reached the target duration corresponding to the second content, or the duration from the time point when reading the target file starts to the current time point has not reached the playback duration of the second video.
[0065] Optionally, the second generating unit is configured to:
[0066] Based on the content display duration corresponding to each display image, form the first video from multiple display images, where the content display duration corresponding to the display image is the display duration of the content in the display image in the target file.
[0067] Optionally, the display image generating unit is configured to:
[0068] When reading the target file through a browser, sequentially take screenshots of multiple contents in the target file through the browser to obtain display images including each content respectively.
[0069] Optionally, the commentary video generation instruction carries the storage address of the target file, and the first video generation module includes:
[0070] A file acquisition unit, configured to, in response to the commentary video generation instruction, acquire the target file based on the storage address;
[0071] A third generating unit, configured to generate the first video based on the content of the target file.
[0072] Optionally, the first video generation module includes:
[0073] An instruction sending unit, configured to send the commentary video generation instruction to the browser through a commentary video generation application;
[0074] A display image generating unit, configured to, through the browser, in response to the commentary video generation instruction, sequentially generate display images including each content based on multiple contents in the target file, and send the generated multiple display images to the commentary video generation application;
[0075] A second generation unit, configured to generate an application through the commentary video, receive a plurality of the display images, and form the plurality of display images into the first video.
[0076] On the other hand, a computer device is provided, which includes a processor and a memory. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the commentary video generation method as described in the above aspect.
[0077] On the other hand, a computer-readable storage medium is provided. At least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor to implement the operations performed by the commentary video generation method as described in the above aspect.
[0078] On the other hand, a computer program product or a computer program is provided. The computer program product or the computer program includes computer program code. The computer program code is stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the computer device implements the operations performed by the commentary video generation method as described in the above aspect.
[0079] The method, device, computer device, and storage medium provided in the embodiments of the present application automatically generate a first video for displaying the content of a target file based on the content of the target file, automatically generate a second video for commenting on the content of the target file based on the commentary text corresponding to the target file, and fuse the first video and the second video, so as to obtain a commentary video that can simultaneously display and comment on the content of the target file. Therefore, only by providing the target file, the commentary video corresponding to the target file can be automatically generated, without manually recording the commentary video, saving manpower and time, and improving the efficiency of generating the commentary video. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0081] Figure 1 is a schematic diagram of an implementation environment provided in an embodiment of the present application.
[0082] Figure 2 is a flowchart of a commentary video generation method provided in an embodiment of the present application.
[0083] Figure 3 It is a flowchart of another method for generating an explanatory video provided by an embodiment of the present application.
[0084] Figure 4 It is a schematic diagram of an explanatory image provided by an embodiment of the present application.
[0085] Figure 5 It is a flowchart of a method for generating explanatory audio and explanatory images provided by an embodiment of the present application.
[0086] Figure 6 It is a schematic diagram of the picture of an explanatory video provided by an embodiment of the present application.
[0087] Figure 7 It is a flowchart of yet another method for generating an explanatory video provided by an embodiment of the present application.
[0088] Figure 8 It is a schematic diagram of the structure of an explanatory video generation device provided by an embodiment of the present application.
[0089] Figure 9 It is a schematic diagram of the structure of another explanatory video generation device provided by an embodiment of the present application.
[0090] Figure 10 It is a schematic diagram of the structure of a terminal provided by an embodiment of the present application.
[0091] Figure 11 It is a schematic diagram of the structure of a server provided by an embodiment of the present application. Detailed implementation manners
[0092] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0093] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the present application, the first video may be referred to as the second video, and similarly, the second video may be referred to as the first video.
[0094] Among them, "at least one" means one or more than one. For example, at least one video can be one video, two videos, three videos, etc., which is any integer greater than or equal to one. "Multiple" means two or more than two. For example, multiple videos can be two videos, three videos, etc., which is any integer greater than or equal to two. "Each" means every one in at least one. For example, each video means every one of multiple videos. If there are three videos in multiple videos, then each video means every one of the three videos.
[0095] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0096] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.
[0097] Computer Vision (CV) is a science that studies how to make machines "see". Further speaking, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition, tracking, and measurement on targets, and further perform graphic processing to make the images processed by the computer more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D (3-Dimension) technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0098] The method for generating an explanatory video provided in the embodiments of the present application will be described below based on artificial intelligence technology and computer vision technology.
[0099] The method for generating an explanatory video provided in the embodiments of the present application can be used in a computer device. Optionally, the computer device is a terminal or a server. Optionally, the server is an independent physical server, or the server is a server cluster or a distributed system composed of multiple physical servers, or the server is a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the terminal is a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto.
[0100] In a possible implementation manner, the computer program involved in the embodiments of the present application can be deployed on one computer device for execution, or on multiple computer devices located at one location for execution, or on multiple computer devices distributed at multiple locations and interconnected through a communication network. The multiple computer devices distributed at multiple locations and interconnected through a communication network can form a blockchain system. In a possible implementation manner, the computer device for generating an explanatory video in the embodiments of the present application is a node in the blockchain system, and the node can store the generated explanatory video in the blockchain. After that, the node or other devices in the blockchain can obtain the explanatory video in the blockchain.
[0101] Figure 1 is a schematic diagram of an implementation environment provided in the embodiments of the present application. Refer to Figure 1 , this implementation environment includes computer device 101 and computer device 102. Computer device 101 and computer device 102 can be directly or indirectly connected through a wired or wireless communication method. Computer device 101 is used to generate a corresponding explanatory video based on a target file, computer device 102 is used to provide the target file to computer device 101, and computer device 101 can also provide the generated explanatory video to computer device 102 subsequently.
[0102] The method for generating an explanatory video provided in the embodiments of the present application can be applied to any scenario for generating an explanatory video.
[0103] For example, in the field of education, the target file is a courseware PPT (PowerPoint, slides), and the explanatory video is a teaching video. After the computer device obtains the courseware PPT, based on the content of the courseware PPT, a first video is generated. The picture of the first video is the content in the courseware PPT. Based on the teaching text of the courseware PPT, a second video is generated. The second video includes teaching voice and teaching pictures. The teaching pictures include a virtual teacher who is teaching. The computer device fuses the first video and the second video to obtain the teaching video corresponding to the courseware PPT. The teaching video includes both the content in the courseware PPT and the virtual teacher who is explaining the content in the courseware PPT. Therefore, the method provided in the embodiments of the present application can automatically generate a teaching video without the teacher recording the teaching video personally, improving the efficiency and convenience of generating the teaching video.
[0104] In addition, the explanatory video generation method provided in the embodiments of the present application can also be applied to generate news broadcast videos, product introduction videos, scenic spot introduction videos, etc. The embodiments of the present application do not limit the application scenarios of the explanatory video generation method.
[0105] Figure 2 It is a flowchart of an explanatory video generation method provided in the embodiments of the present application. The execution subject of the embodiments of the present application is a computer device. Refer to Figure 2 , and the method includes:
[0106] 201. The computer device responds to an explanatory video generation instruction for a target file and generates a first video based on the content of the target file.
[0107] The explanatory video generation instruction for the target file is used to request the generation of an explanatory video corresponding to the target file. The explanatory video is used to display and explain the content of the target file. The target file is a file of any type, and the content of the target file can also be content of any type.
[0108] The computer device responds to the explanatory video generation instruction and generates a first video corresponding to the target file based on the content of the target file. Among them, the picture in the first video is used to display the content in the target file.
[0109] 202. The computer device generates a second video based on the explanatory text corresponding to the target file.
[0110] Since the first video only includes pictures for displaying the content in the target file and cannot explain the target file. There is an explanatory text corresponding to the target file. Therefore, the computer device will also generate a second video corresponding to the target file based on the explanatory text corresponding to the target file.
[0111] Among them, the explanatory text corresponding to the target file is used to explain the content of the target file. The second video includes the explanatory audio and explanatory images corresponding to the explanatory text, and the second video is used to explain the content of the target file.
[0112] 203. The computer device fuses the first video and the second video to obtain an explanatory video corresponding to the target file.
[0113] The images in the first video are used to display the content in the target file. The second video includes explanatory audio and explanatory images for explaining the content of the target file. In order to obtain an explanatory video that can simultaneously display and explain the content of the target file, the computer device fuses the first video and the second video to obtain an explanatory video corresponding to the target file. The images in the explanatory video are used to display and explain the content in the target file, and the explanatory video also includes explanatory audio for explaining the content in the target file.
[0114] The method provided by the embodiments of the present application automatically generates a first video for displaying the content of the target file based on the content of the target file, automatically generates a second video for explaining the content of the target file based on the explanatory text corresponding to the target file, and fuses the first video and the second video to obtain an explanatory video that can simultaneously display and explain the content of the target file. Therefore, only by providing the target file, the explanatory video corresponding to the target file can be automatically generated without manually recording the explanatory video, saving manpower and time and improving the efficiency of generating the explanatory video.
[0115] Figure 3 It is a flowchart of a method for generating an explanatory video provided by an embodiment of the present application. The execution subject of the embodiment of the present application is a computer device. Refer to Figure 3 and the method includes:
[0116] 301. In response to an explanatory video generation instruction for a target file, the computer device sequentially generates display images including each content based on multiple contents in the target file.
[0117] In response to an explanatory video generation instruction for a target file, the computer device sequentially determines multiple contents in the target file, and based on each content in the target file, sequentially generates display images including each content. The display images are used to display the content in the target file, and the display images generated by the computer device correspond one-to-one to the contents in the target file.
[0118] Among them, the target file can be any type of file. For example, the target file is a web page file, a PPT file, or a Word (text) file, etc. The content of the target file can be any type of file. For example, in the field of education, the content of the target file is the lecture content; in the field of media dissemination, the content of the target file is the content of news broadcasts; in the field of e-commerce, the content of the target file is the product introduction; in the field of tourism, the content of the target file is the scenic spot introduction, etc.
[0119] Among them, the target file includes multiple contents. One content in the target file refers to the content in a picture shown in the target file. Multiple different pictures will be shown in the target file, so the target file includes the content corresponding to each picture.
[0120] In a possible implementation manner, when the computer device reads the target file through the browser, through the browser, it sequentially takes screenshots of multiple contents in the target file, and respectively obtains display images including each content.
[0121] During the process of the browser of the computer device reading the target file, multiple different contents will be sequentially shown in the target file. Therefore, the computer device takes screenshots of each content shown in the target file through the browser. Optionally, the browser is a headless browser (Puppeteer). A headless browser is a browser without a user interface. During the process of reading the target file, the headless browser executes the logic of displaying the content in the target file in the background, but does not display the content in the target file in the foreground. The headless browser executes the logic of taking screenshots of multiple contents in the target file in the background, and the screenshot process is not displayed in the foreground either.
[0122] In a possible implementation manner, the computer device reads the target file. When the content shown in the target file is the first content, it generates a display image including the first content. When the content shown in the target file switches from the first content to the second content, it continues to generate a display image including the second content until the target file reading ends, and multiple display images are obtained. Among them, the multiple contents in the target file are arranged in order. Whenever one content is shown, the currently shown content is switched to the next content of this content, and the second content is the next content of the first content.
[0123] During the process of a computer device reading a target file, different contents will be sequentially displayed in the target file. For the currently displayed content, the computer device will generate a display image including this content. Whenever the computer device detects that the content displayed in the target file has changed, it will continue to generate a display image including the changed content until the reading of the target file ends, that is, the last content in the target file has been displayed. The computer device will obtain multiple display images, and each of these multiple display images includes different contents in the target file. In the embodiments of this application, during the process of reading the target file, as the content displayed in the target file is continuously switched, the computer device automatically generates display images including each content in sequence according to each displayed content. The information contained in the multiple generated display images is the same as the information contained in the target file, thereby realizing the conversion of the target file into multiple display images.
[0124] In a possible implementation manner, the computer device switches the content displayed in the target file, including the following two methods:
[0125] (1) In response to a content switching operation on the first content, switch the content displayed in the target file from the first content to the second content.
[0126] When the content displayed in the target file is the first content, if it is necessary to switch the displayed content, perform the content switching operation on the first content. The computer device responds to this content switching operation, determines the next content of the first content, that is, the second content, and switches the content displayed in the target file from the first content to the second content.
[0127] Optionally, the computer device reads the target file through a headless browser. Since this headless browser has no user interface and the user cannot perform the content switching operation on the first content, the headless browser calls the user simulation interface to simulate the user's execution of the content switching operation on the first content. This user simulation interface is used to simulate the user's operations. Optionally, the first content corresponds to a content switching button, and this content switching operation is a click operation on this content switching button. Then the computer device calls this user simulation interface through the headless browser to simulate the click operation on this content switching button. For example, this user simulation interface is function vdcUserClick(selector), where selector represents a selector used to select which button to simulate clicking. The identifier of this content switching button is start, then the browser's call instruction to this user simulation interface is "vdcUserClick("#start")", indicating a request to simulate the click operation on the content switching button with the identifier start.
[0128] (2) When the content display duration of the first content reaches the target duration corresponding to the first content, switch the content displayed in the target file from the first content to the second content.
[0129] When the content displayed in the target file is the first content, the computer device determines the target duration corresponding to the first content. The target file includes the target duration, which represents the duration for which the first content should be displayed, and the target duration can be preset. During the process that the content displayed in the target file is the first content, the computer device monitors the content display duration of the first content in real time. If the content display duration of the first content reaches the target duration, then switch the content displayed in the target file from the first content to the second content, that is, switch to the next content. If the content display duration of the first content does not reach the target duration, then the content displayed in the target file remains the first content and no switching is performed.
[0130] In the embodiments of the present application, two methods, namely, according to the content switching operation and according to the display duration, are provided to determine the timing of switching the content displayed in the target file, which improves the flexibility of switching the content displayed in the target file.
[0131] In a possible implementation manner, when the second content is the last content in the target file, the computer device stops reading the target file when the content display duration of the second content has reached the target duration corresponding to the second content and the duration from the time point of starting to read the target file to the current time point has reached the playback duration of the second video; when the content display duration of the second content has not reached the target duration corresponding to the second content, or the duration from the time point of starting to read the target file to the current time point has not reached the playback duration of the second video, continue to read the target file that is currently displaying the second content.
[0132] First, since the second content is the last content in the target file, stop reading the target file after the second content is displayed. The target file includes the target duration corresponding to the second content, which represents the duration for which the second content should be displayed. Therefore, when stopping reading the target file, the content display duration of the second content should be not less than the target duration to ensure that the second content has been displayed.
[0133] In addition, the duration from the start of reading the target file to the current time point is equivalent to the total display duration for displaying multiple contents in the target file. Subsequently, the computer device will form the first video from multiple display images, and the playback duration corresponding to the first video is equal to the total display duration corresponding to the multiple contents. The pictures in the first video are used to display the contents of the target file, and the second video is used to explain the contents of the target file. To ensure that the playback duration corresponding to the first video is not less than that of the second video, the total display duration corresponding to the multiple contents in the target file should not be less than the playback duration corresponding to the second video.
[0134] Therefore, in the case where the second content is the last content in the target file, stopping reading the target file needs to meet two conditions. One is that the content display duration of the second content has reached the target duration corresponding to the second content, and the other is that the duration from the time point of starting to read the target file to the current time point has reached the playback duration corresponding to the second video. If either of the two conditions is not met, the computer device continues to read the target file that is currently displaying the second content.
[0135] In a possible implementation manner, the computer device generates an instruction for generating an explanatory video for the target file in response to an operation of generating an explanatory video for the target file. In another possible implementation manner, the instruction for generating the explanatory video is sent by the terminal to the computer device. After receiving the instruction for generating the explanatory video, the computer device uses the method for generating an explanatory video provided in the embodiments of the present application to generate the explanatory video corresponding to the target file and returns the explanatory video to the terminal.
[0136] 302. The computer device forms the first video from the generated multiple display images.
[0137] After the computer device sequentially generates display images including each content, it forms the first video from the multiple display images according to the generation order of the multiple display images. The pictures in the first video are used to display the contents of the target file. Since the first video is only composed of display images, the first video only includes pictures and does not include audio.
[0138] In a possible implementation manner, the computer device forms the first video from the multiple display images based on the content display duration corresponding to each display image, where the content display duration corresponding to the display image is the display duration of the content in the display image in the target file.
[0139] After generating a display image including the content in the target file based on the content in the target file, the computer device further determines the display duration of the content in the target file, determines the display duration of the content as the content display duration of the content, and determines the content display duration as the content display duration corresponding to the display image. Therefore, the computer device obtains the display image including each content and the content display duration corresponding to each display image. Then, the computer device obtains the content display duration corresponding to each display image, and based on the content display duration corresponding to each display image, forms the multiple display images into a first video, so that the playback duration corresponding to each display image in the first video is equal to the content display duration corresponding to the display image.
[0140] For example, among the multiple display images, there are a first display image and a second display image. The content included in the second display image is the next content of the content included in the first display image, and the second display image is the next display image of the first display image. The content display duration corresponding to the first display image is 5 seconds, and the content display duration corresponding to the second display image is 10 seconds. Then, in the first video formed by the multiple display images, the playback duration of the first display image is 5 seconds. After the first display image finishes playing, the second display image is played in the first video, and the playback duration of the second display image is 10 seconds.
[0141] In a possible implementation manner, the computer device forms the multiple display images into the first video through the FFmpeg program. The FFmpeg program is used to record, convert, and stream audio or video, etc. Through the FFmpeg program, the computer device can convert data in multiple image formats into data in video format.
[0142] It should be noted that the above steps 301 - 302 take the example of first generating the display image and then forming the first video to illustrate the process of the computer device generating the first video based on the content of the target file in response to the instruction for generating the explanatory video of the target file. In addition, the computer device can also directly record the screen of the content of the target file to obtain the first video including the content in the target file. Or the computer device first generates video segments including each content based on the multiple contents in the target file in sequence, and combines the generated multiple video segments into the first video. The embodiments of the present application do not limit the manner of generating the first video based on the content in the target file.
[0143] In a possible implementation manner, the instruction for generating the explanatory video carries the storage address of the target file. The computer device responds to the instruction for generating the explanatory video, obtains the target file based on the storage address, and generates the first video based on the content of the target file.
[0144] The computer device obtains the target file in the storage space indicated by the storage address. Optionally, the storage address is a URL (Uniform Resource Locator), and the computer device obtains the target file through a browser based on the URL.
[0145] In another possible implementation, the computer device sends an explanatory video generation instruction to the browser through an explanatory video generation application; through the browser, in response to the explanatory video generation instruction, based on multiple contents in the target file, display images including each content are sequentially generated, and the generated multiple display images are sent to the explanatory video generation application; through the explanatory video generation application, the multiple display images are received and the multiple display images are composed into a first video.
[0146] An explanatory video generation application runs on the computer device. The explanatory video generation application is used to generate an explanatory video corresponding to any application. A browser also runs on the computer device, and the browser has a function of generating images, such as generating images by taking screenshots, etc. Through the explanatory video generation application, an explanatory video generation instruction for the target file is generated and sent to the browser. Optionally, the URL corresponding to the target file is carried in the explanatory video generation instruction, and the browser obtains the target file based on the URL.
[0147] 303. The computer device generates an explanatory audio based on the explanatory text corresponding to the target file.
[0148] An explanatory text corresponds to the target file. The explanatory text is used to explain the content in the target file. The computer device generates an explanatory audio corresponding to the target file based on the explanatory text. The explanatory audio is used to explain the content in the target file. Optionally, the explanatory text is also carried in the explanatory video generation instruction, and the computer device obtains the explanatory text from the explanatory video generation instruction. Optionally, the explanatory text corresponding to the target file is stored in the computer device, and the computer device directly obtains the locally stored explanatory text.
[0149] In a possible implementation, the computer device generates an explanatory audio based on the explanatory text, including at least one of the following:
[0150] (1) The explanatory video generation instruction carries a model identifier, and the computer device calls the audio conversion model indicated by the model identifier to convert the explanatory text into an explanatory audio.
[0151] The computer device stores multiple audio conversion models. Each audio conversion model is used to convert any text into audio, and different audio conversion models have different parameters such as intonation, tone, timbre, and language, so as to generate audio with different characteristics. For example, the audio converted by the audio conversion model can belong to the voice of a lively cartoon character, a deep male voice, a gentle female voice, etc. Optionally, the audio conversion model in the embodiments of the present application is a TTS (Text To Speech) model or other models, etc.
[0152] The computer device obtains the model identifier carried in the commentary video generation instruction, determines the audio conversion model indicated by the model identifier among the multiple audio conversion models, inputs the commentary text into the audio conversion model, and the audio conversion model converts the commentary text, thereby outputting the commentary audio corresponding to the commentary text.
[0153] (2) The commentary video generation instruction carries the commentary speed, and the computer device converts the commentary text into commentary audio with the commentary speed.
[0154] The commentary speed is used to indicate the speed of the commentary audio. For example, the commentary speed is 200 words per minute. Or the commentary speed is used to indicate that the speed of the commentary audio is fast, medium, or slow, then the computer device converts the commentary text into commentary audio with the commentary speed according to the indication of the commentary speed.
[0155] 304. The computer device generates a set of commentary images.
[0156] The set of commentary images includes multiple commentary images, where each commentary image includes the same commentary object, and there are at least two commentary images in the multiple sets of commentary images with different mouth shapes of the commentary object. Optionally, the commentary object is a virtual object, such as a 2D virtual object or a 3D virtual object generated by AI technology, etc. The virtual object can be in any form, such as a human form or a cartoon form, etc. The embodiments of the present application do not make any limitations in this regard. Figure 4 is a schematic diagram of a commentary image provided by an embodiment of the present application, as Figure 4 shown, the commentary image includes a commentary object 401.
[0157] In a possible implementation manner, the computer device generates a set of commentary images, including at least one of the following:
[0158] (1) The commentary video generation instruction carries an object identifier, and the computer device generates a set of commentary images based on the commentary object indicated by the object identifier, so that multiple commentary images include the commentary object.
[0159] The computer device stores multiple commentary objects, each with a different image. For example, the commentary objects include a male in human form, a female in human form, or a cartoon form, etc. The computer device obtains the object identifier carried in the commentary video generation instruction, determines the commentary object indicated by the object identifier among the multiple commentary objects, and generates multiple commentary images including the commentary object based on the commentary object, forming the commentary image set.
[0160] (2) The commentary video generation instruction carries an action identifier. The computer device generates a commentary image set based on the dynamic action indicated by the action identifier, so that the action combinations of the commentary objects in the multiple commentary images form the dynamic action.
[0161] The action identifier carried in the commentary video generation instruction is used to indicate a dynamic action. For example, the dynamic action includes greeting, making gestures, etc. The commentary video requested to be generated by the commentary video generation instruction includes a commentary object performing the dynamic action. If the computer device can drive the commentary object to perform a specific action, the computer device determines the dynamic action indicated by the action identifier and generates multiple commentary images based on the dynamic action, so that the action combinations of the commentary objects in the multiple commentary images form the dynamic action, that is, by connecting the actions of the commentary objects in each commentary image continuously, the dynamic action can be formed. The computer device forms the generated multiple commentary images into the commentary image set.
[0162] (3) The commentary video generation instruction carries a position parameter. The computer device generates a commentary image set based on the position indicated by the position parameter, so that the commentary objects in the multiple commentary images are located at the position.
[0163] The position parameter carried in the commentary video generation instruction is used to indicate a position. The commentary video requested to be generated by the commentary video generation instruction includes a commentary object located at the position. If the computer device can control the position of the commentary object, the computer device determines the position indicated by the position parameter and generates multiple commentary images based on the position, so that the commentary objects in the multiple commentary images are located at the position. The computer device forms the generated multiple commentary images into the commentary image set.
[0164] In the embodiments of the present application, through the object identifier, action identifier, and position parameter carried in the commentary video generation instruction, the multiple commentary images generated by the computer device can include a specified commentary object, a commentary object performing a specified action, or a commentary object located at a specified position, thereby improving the flexibility and diversity of controlling the commentary objects in the commentary images, and further improving the flexibility and diversity of the generated second video.
[0165] In a possible implementation, the target file includes a first data segment and a second data segment, and the explanatory text includes at least one explanatory text segment. Among them, the first data segment has a corresponding explanatory text segment, and the second data segment does not have a corresponding explanatory text segment. Then the computer device generates an explanatory image set, including: generating a sub-set of explanatory images corresponding to the first data segment, in the sub-set of explanatory images corresponding to the first data segment, there are at least two explanatory images with different mouth shapes of the explanatory object; generating a sub-set of explanatory images corresponding to the second data segment, in the sub-set of explanatory images corresponding to the second data segment, the mouth shapes of the explanatory objects in different explanatory images are the same. The computer device constitutes the explanatory image set corresponding to the target file with the sub-set of explanatory images corresponding to the first data segment and the sub-set of explanatory images corresponding to the second data segment.
[0166] Since the first data segment has a corresponding explanatory text segment, it can be considered that the content of the first data segment needs to be explained. The sub-set of explanatory images corresponding to the first data segment is used to constitute the explanatory picture for explaining the content of the first data segment. Then, the explanatory object that is explaining the content of the first data segment should be included in the explanatory picture. In order to simulate the scene where the explanatory object is explaining, in the sub-set of explanatory images corresponding to the first data segment, there should be at least two explanatory images with different mouth shapes of the explanatory object, so as to achieve the effect that the explanatory object is speaking, and it is more in line with the real explanatory scene.
[0167] Since the second data segment does not have a corresponding explanatory text segment, it can be considered that the content of the second data segment does not need to be explained. In order to ensure the smoothness of the picture in the finally generated explanatory video, the explanatory object in the explanatory video should not appear and disappear suddenly, but should always exist in the explanatory video. Therefore, the second data segment corresponds to a sub-set of explanatory images, and the explanatory images in the sub-set of explanatory images also include the explanatory object, but the mouth shapes of the explanatory objects in different explanatory images are the same. For example, the mouth shapes of the explanatory objects are all in a closed state, so as to achieve the effect that the explanatory object is not speaking.
[0168] In a possible implementation, the target file includes a first data segment and a second data segment, and the explanatory text includes at least one explanatory text segment. Among them, the first data segment has a corresponding explanatory text segment, and the second data segment does not have a corresponding explanatory text segment. Then the computer device generates explanatory audio based on the explanatory text, including: generating an explanatory audio segment corresponding to the first data segment based on the explanatory text segment corresponding to the first data segment. In addition, the computer device can also generate a blank audio segment corresponding to the second data segment, and there is no sound in the blank audio segment corresponding to the second data segment.
[0169] In a possible implementation, the commentary audio includes commentary audio segments corresponding to at least one first data segment. Then, the computer device generates a subset of commentary images corresponding to the first data segment, including the following steps 3041 - 3043.
[0170] 3041. The computer device generates a first subset of images based on the first duration of the commentary audio segment corresponding to the first data segment. The first subset of images includes multiple commentary images and the playback duration corresponding to each commentary image. The mouth shapes of the commentary objects in different commentary images are different, and the sum of the playback durations corresponding to the multiple commentary images is equal to the first duration.
[0171] The commentary screen formed by the first subset of images corresponding to the first data segment is synchronized with the commentary audio segment corresponding to the first data segment in time. The computer device determines the first duration of the commentary audio segment corresponding to the first data segment. The first duration represents the playback duration of the commentary audio segment. First, based on the first duration, a first subset of images is generated so that the sum of the playback durations corresponding to the multiple commentary videos in the first subset of images is equal to the first duration, thereby ensuring that the time for playing the commentary audio segment is the same as the time for playing the commentary images in the first subset of images, so as to realize the synchronous playback of the commentary audio segment corresponding to the first data segment and the corresponding multiple commentary images.
[0172] Among them, the first subset of images also includes the playback duration corresponding to each commentary image, and the playback durations corresponding to each commentary image can be the same or different. When the playback durations corresponding to each commentary object are the same, the playback duration is equivalent to the frame rate corresponding to the multiple commentary images.
[0173] Among them, in the first subset of images, the mouth shapes of the commentary objects in different commentary images are different, so as to simulate the effect of the commentary object speaking. Optionally, the computer device controls the mouth shapes of the commentary objects in the multiple commentary images based on the commentary text segment corresponding to the first data segment, so that the combination of the mouth shapes of the commentary objects in the multiple commentary images forms the mouth shapes of reading the commentary text segment, so as to simulate the effect of the commentary object reading the commentary text segment.
[0174] 3042. When the content display duration of the first data segment is greater than the first duration, the computer device determines the first duration difference between the content display duration and the first duration. Based on the first duration difference, a second subset of images is generated. The second subset of images includes at least one commentary image and the playback duration corresponding to each commentary image. The mouth shapes of the commentary objects in different commentary images are the same, and the sum of the playback durations corresponding to at least one commentary image is equal to the first duration difference.
[0175] The display duration of the content of the first data segment refers to the display duration of the first data segment in the target file, and also the display duration of the first data segment in the finally generated commentary video. The commentary screen formed by the set of commentary images corresponding to the first data segment should be synchronized with the display screen formed by the content of the first data segment in terms of time. Among them, in step 3041 above, the computer device has generated a first subset of images, and the playback durations of the multiple commentary images in the first subset of images are equal to the first duration. If the display duration of the content of the first data segment is greater than the first duration, it means that the content of the first data segment has not finished playing, while the commentary screen corresponding to the first data segment has finished playing, resulting in the inability to synchronize the display screen corresponding to the first data segment with the commentary screen, and causing the commentary object to disappear before the display screen corresponding to the first data segment has finished playing. To avoid this problem, the computer device determines the first duration difference between the content display duration and the first duration, and generates a second subset of images whose sum of playback durations is equal to the first duration difference to fill in the commentary screen, so as to ensure the synchronization of the display screen corresponding to the first data segment with the commentary screen.
[0176] Since the commentary audio segment has finished playing and there is no need to provide further commentary for the commentary object, in the second subset of images, the mouth shapes of the commentary object in different commentary images are the same. For example, the mouth shapes of the commentary object are all in a closed state, thus achieving the effect that the commentary object is not speaking.
[0177] 3043. The computer device combines the first subset of images and the second subset of images to form a subset of commentary images corresponding to the first data segment.
[0178] After the computer device obtains the first subset of images and the second subset of images, it combines the first subset of images and the second subset of images to form a subset of commentary images corresponding to the first data segment. That is to say, the subset of commentary images includes the commentary images in the first subset of images and the commentary images in the second subset of images. Among them, the commentary images in the first subset of images and the second subset of images are arranged in sequence, so the commentary images in the subset of commentary images are also arranged in sequence, and the multiple commentary images in the first subset of images are located in front of the multiple commentary images in the second subset of images.
[0179] 305. The computer device generates a second video based on the commentary audio and the set of commentary images.
[0180] After the computer device generates the commentary audio and the set of commentary images, it generates a second video corresponding to the target file based on the commentary audio and the set of commentary images. The second video includes the commentary audio and the commentary screen corresponding to the commentary text, and the second video is used to provide commentary on the content of the target file.
[0181] In a possible implementation, the target file includes multiple data segments, the explanatory audio includes explanatory audio segments corresponding to at least one data segment, and the set of explanatory images includes subsets of explanatory images corresponding to each data segment. The computer device generates a second video based on the explanatory audio and the set of explanatory images, including: generating explanatory audio based on the explanatory audio segments corresponding to at least one data segment in the arrangement order of each data segment; generating explanatory images based on the subsets of explanatory images corresponding to each data segment in the arrangement order of each data segment; generating a second video based on the explanatory audio and the explanatory images, wherein the start playback time point of the explanatory audio segment corresponding to the same data segment is the same as the start playback time point of the subset of explanatory images corresponding thereto.
[0182] When each data segment among the multiple data segments corresponds to an explanatory audio segment, the explanatory audio segments corresponding to the multiple data segments are synthesized into explanatory audio in the arrangement order of each data segment. When there is a data segment among the multiple data segments that does not have a corresponding explanatory audio segment, a blank audio segment corresponding to the data segment is generated according to the display duration corresponding to the data segment, and there is no sound in the blank audio segment. Then the computer device synthesizes the explanatory audio segments corresponding to the multiple data segments and the blank audio segments into explanatory audio in the arrangement order of each data segment. Therefore, the playback duration corresponding to the explanatory audio generated by the computer device is equal to the sum of the display durations corresponding to each data segment in the target file.
[0183] The display duration corresponding to each data segment is equal to the sum of the playback durations of multiple explanatory images in the subset of explanatory images corresponding to the data segment. Therefore, the playback duration corresponding to the explanatory images generated by the computer device is equal to the sum of the display durations corresponding to each data segment in the target file, and the playback durations of the explanatory audio and the explanatory images are equal. The computer device generates a second video based on the explanatory audio and the explanatory images. The playback duration corresponding to the second video is equal to the sum of the display durations corresponding to each data segment in the target file, or in other words, the playback duration is equal to the reading duration of the target file.
[0184] Moreover, since each data segment and the explanatory audio segment are in one-to-one correspondence, and each data segment and the subset of explanatory images are also in one-to-one correspondence, the explanatory audio segment and the subset of explanatory images are also in one-to-one correspondence. The explanatory audio segment corresponding to the same data segment and the corresponding subset of explanatory images are played synchronously, that is, the start playback time points in the second video are the same.
[0185] In a possible implementation, a computer device is connected to a background server, which is used to synthesize an explanation object. The computer device generates an explanation audio and an explanation image set corresponding to a target file by calling multiple interfaces provided by the background server related to generating an explanation video. Among them, the multiple interfaces include a start interface, a generation interface, a callback interface, a video description interface, an explanation object hiding interface, a user simulation interface, and an end interface. The start interface is function vdcStart(), which is used to notify the background server that the process of synthesizing an explanation object starts.
[0186] The generation interface is function vdcShow(timestamp,id,param,callback,description), which is used to notify the background server to generate an explanation image and an explanation audio. The timestamp (timestamp) parameter in the generation interface is used to transfer the duration between the time point when the start interface is called and the current time point. The type of the timestamp parameter is an integer, and the unit is milliseconds, which is used to synchronize the screen corresponding to the target file and the screen corresponding to the explanation object. The id (identifier) parameter in the generation interface is used to indicate a call, and the id parameter is carried during the callback. The param parameter in the generation interface is used for extension, and the param parameter is a string in json (JavaScript Object Notation, JS, object notation) format. Among them, the param parameter includes various parameters as shown in Table 1.
[0187] Table 1
[0188]
[0189] The callback interface is used to callback and notify that the synthesis of the explanation object is completed. It is of string type. The callback interface includes receiving a string type parameter, which is used to receive the above id parameter and is used to associate an action of synthesizing an explanation object. The video description interface is used to describe the generated explanation image set. The explanation object hiding interface is used to hide the explanation object in the explanation image. The user simulation interface is used to simulate the operations of a user. The end interface is used to notify to stop generating the explanation image. Among them, after calling the start interface, the computer device will also receive a return value sent by the background server. The return value is of numerical type, and the definition of the return value is shown in Table 2 below.
[0190] Table 2
[0191] Return value Description 0 Call succeeded -1 Call failed -2 The call is in progress. Please wait for the callback or call again later
[0192] Figure 5It is a flowchart for generating explanatory audio and explanatory images provided by an embodiment of the present application. Taking the target file as a web page file as an example, the web page file includes data segment 1 and data segment 2. 501. The computer device waits for the web page file to finish loading. 502. The computer device calls the start interface to notify the background server to prepare to execute the process of generating explanatory images and explanatory audio, so that the background server is ready. 503. The computer device calls the generation interface to request the background server to generate the explanatory images and explanatory audio corresponding to data segment 1. 504. The background server generates multiple explanatory images and explanatory audio corresponding to data segment 1. 505. The background server notifies the computer device that the generation of the explanatory images and explanatory audio corresponding to data segment 1 is completed. 506. The computer device calls the generation interface to request the background server to generate the explanatory images and explanatory audio corresponding to data segment 2. 507. The background server generates multiple explanatory images and explanatory audio corresponding to data segment 2. 508. The background server notifies the computer device that the generation of the explanatory images and explanatory audio corresponding to data segment 2 is completed. 509. The computer device calls the end interface to notify the background server to end the process of generating explanatory images and explanatory audio.
[0193] It should be noted that the above steps 303-305 illustrate the process of the computer device generating the second video based on the explanatory text corresponding to the target file, taking the example of first generating the explanatory audio and the set of explanatory images and then forming the second video. In another embodiment, the explanatory video of the second video may not include the explanatory object. For example, the explanatory video includes descriptive text of the target file, etc.
[0194] 306. The computer device fuses the first video and the second video to obtain the explanatory video corresponding to the target file.
[0195] The picture in the first video is used to display the content in the target file. The second video includes the explanatory audio and the explanatory picture for explaining the content of the target file. In order to obtain the explanatory video that simultaneously displays and explains the content of the target file, the computer device fuses the first video and the second video to obtain the explanatory video corresponding to the target file.
[0196] In a possible implementation manner, the explanatory video is composed of a first layer and a second layer. The second layer is located above the first layer. The picture in the first video is located in the first layer, and the picture in the second video is located in the second layer. The pictures in the two layers are superimposed to obtain the picture in the explanatory video.
[0197] In a possible implementation manner, the picture in the second video includes the explanatory object and a green screen. The green screen is the background of the explanatory object in the second video. Then the computer device first deletes the green screen in the second video, and fuses the second video after deleting the green screen with the first video to obtain the explanatory video.
[0198] Figure 6 It is a schematic diagram of the screen of an explanatory video provided by an embodiment of the present application. Figure 6 It shows the screens 601, 602, and 603 in the explanatory video, where the same explanatory object 604 is included in each of the screens 601, 602, and 603, and the actions of the explanatory object 604 are different in each screen. The content other than the explanatory object 604 in each screen is the content in the target file.
[0199] It should be noted that, taking the order of steps 301 - step 306 as an example, the embodiment of the present application illustrates the process of a computer device generating an explanatory video. However, the embodiment of the present application does not limit the specific order of generating the first video and the second video. For example, the computer device simultaneously executes the above steps 301, 303, and step 304 to obtain multiple display images, explanatory audio, and an explanatory image set. Then, the computer device simultaneously executes the above steps 302 and 305 to generate the first video and the second video, and then executes step 306 to fuse the first video and the second video into an explanatory video.
[0200] It should be noted that the embodiment of the present application only takes the example of generating an explanatory video corresponding to one target file for illustration. In another embodiment, the computer device responds to a batch selection operation on multiple target files, respectively generates an explanatory video generation instruction for each target file, and in response to the explanatory video generation instruction for each target file, executes the steps of generating an explanatory video corresponding to each target file in parallel. The computer device batch generates explanatory video generation instructions for multiple target files. For each explanatory video generation instruction corresponding to a target file, the computer device executes the above steps 301 - 306 to generate an explanatory video corresponding to each target file in parallel.
[0201] In the related art, it is necessary for a real person to record an explanatory video. However, in the embodiment of the present application, the computer device can batch generate explanatory videos without manual participation in the generation process of the explanatory videos, improving the efficiency of generating explanatory videos.
[0202] The method provided by the embodiment of the present application automatically generates a first video for displaying the content of the target file based on the content of the target file, automatically generates a second video for explaining the content of the target file based on the explanatory text corresponding to the target file, and fuses the first video and the second video to obtain an explanatory video that can simultaneously display and explain the content of the target file. Therefore, only by providing the target file, the explanatory video corresponding to the target file can be automatically generated without manual recording of the explanatory video, saving manpower and time and improving the efficiency of generating explanatory videos.
[0203] Moreover, during the process of reading the target file, as the content displayed in the target file is continuously switched, the computer device automatically generates display images including each piece of content in sequence according to each piece of content displayed. The information contained in the multiple generated display images is the same as the information contained in the target file, thus realizing the conversion of the target file into multiple display images.
[0204] Moreover, two methods, namely, according to the content switching operation and according to the display duration, are provided to determine the timing of switching the content displayed in the target file, improving the flexibility of switching the content displayed in the target file.
[0205] Moreover, through the object identifier, action identifier, and position parameter carried in the commentary video generation instruction, the multiple commentary images generated by the computer device include the specified commentary object, the commentary object performing the specified action, or the commentary object located at the specified position, improving the flexibility and diversity of controlling the commentary object, and further improving the flexibility and diversity of the generated second video.
[0206] Moreover, when the first data segment has a corresponding commentary text segment, the content of the first data segment needs to be explained. In order to simulate the scenario where the commentary object is explaining, in the subset of commentary images corresponding to the first data segment, there are at least two commentary images with different mouth shapes of the commentary object, thus achieving the effect that the commentary object is speaking.
[0207] Figure 7 It is a flowchart of another method for generating a commentary video provided by an embodiment of the present application. As Figure 7 shown, the file provider is a device that provides a web page file, and the background server is used to generate a commentary video. The background server provides a headless browser, a commentary video generation application, and an AI (Artificial Intelligence) synthesis service. The AI synthesis service is used to synthesize a commentary object. The method includes the following steps:
[0208] 701. The commentary video generation application sends the URL corresponding to the web page file to the headless browser, requesting the headless browser to load the URL.
[0209] 702. The headless browser receives the URL, loads the URL, and obtains the web page file provided by the file provider.
[0210] 703. The file provider calls the start interface to notify the headless browser to start the process of generating a commentary video.
[0211] 704. During the process of a headless browser reading a web page file, by taking screenshots of the content in the web page file, display images and the content display duration corresponding to each display image are generated in sequence. Among them, the headless browser generates the content display duration corresponding to the display image according to the screenshot frequency. Whenever the content rendered on the web page changes, the headless browser takes a screenshot. Figure 1 times.
[0212] 705. The file provider controls the headless browser to display the content in data segment 1 of the web page file. After the content in data segment 1 is displayed, the page is in a waiting state until a callback notification that the explanatory image and explanatory audio are generated is received.
[0213] 706. The file provider calls the generation interface and sends a generation request to the headless browser to request the generation of an explanatory image subset and an explanatory audio segment corresponding to data segment 1. The generation request also carries an explanatory text segment corresponding to data segment 1, which is used to generate the explanatory audio segment. For example, the explanatory text segment is "A fast horse travels 240 li every day".
[0214] 707 - 708. Through the headless browser and the explanatory video generation application, the generation request is forwarded to the AI synthesis service.
[0215] 709. The AI synthesis service generates an explanatory audio segment corresponding to data segment 1 and multiple explanatory images corresponding to data segment 1 based on the explanatory text segment. Among the multiple explanatory images generated by the AI synthesis service, there is an explanatory object. The mouth shapes of the explanatory object in the multiple explanatory images generated based on the explanatory text segment are different, so as to achieve the effect of the explanatory object reading the explanatory text segment. After the multiple explanatory images are generated, if data segment 1 has not been fully displayed, the AI synthesis service continues to generate multiple explanatory images including the explanatory object, but the explanatory object in the multiple explanatory images is in a waiting state. The waiting state means that the mouth shape of the explanatory object is in a closed state and the actions of the explanatory object are all consistent.
[0216] 710 - 712. After the explanatory image subset and the explanatory audio segment corresponding to data segment 1 are generated, the AI synthesis service sends a generation completed notification to the file provider through the explanatory video generation application and the headless browser.
[0217] When the file provider receives this notification, it controls the headless browser to stop displaying data segment 1 and start displaying data segment 2, and the file provider continues to call the generation interface to request the generation of an explanatory image subset and an explanatory audio segment corresponding to data segment 2. This process is the same as the process in steps 705 - 712 above.
[0218] 713. When all n data segments in the web page file are displayed, the file provider calls the end interface to notify the headless browser to end the process of generating display images.
[0219] 714. The headless browser stops taking screenshots, and multiple display images corresponding to the web page file are generated.
[0220] 715. The headless browser notifies the commentary video generation application to end the process of generating commentary images and commentary audio.
[0221] 716. The commentary video generation application notifies the AI synthesis service to end the process of generating commentary images and commentary audio. The AI synthesis service generates a second video based on the commentary audio and the set of commentary images.
[0222] 717. The commentary video generation application composes the multiple display videos into a first video.
[0223] 718. The commentary video generation application fuses the first video and the second video into a commentary video.
[0224] Figure 8 It is a schematic structural diagram of a commentary video generation device provided by an embodiment of the present application. Refer to Figure 8 , the device includes:
[0225] A first video generation module 801, configured to generate a first video based on the content of the target file in response to a commentary video generation instruction for the target file, and the pictures in the first video are used to display the content of the target file;
[0226] A second video generation module 802, configured to generate a second video based on the commentary text corresponding to the target file, and the second video includes commentary audio and commentary pictures corresponding to the commentary text;
[0227] A third video generation module 803, configured to fuse the first video and the second video to obtain a commentary video corresponding to the target file.
[0228] The commentary video generation device provided by the embodiment of the present application automatically generates a first video for displaying the content of the target file based on the content of the target file, automatically generates a second video for explaining the content of the target file based on the commentary text corresponding to the target file, and fuses the first video and the second video, so that a commentary video that can simultaneously display and explain the content of the target file can be obtained. Therefore, only by providing the target file, the commentary video corresponding to the target file can be automatically generated, without manually recording the commentary video, saving manpower and time, and improving the efficiency of generating the commentary video.
[0229] Optionally, refer to Figure 9 , the device further includes:
[0230] An instruction generation module 804, configured to generate an explanatory video generation instruction for each target file respectively in response to a batch selection operation on multiple target files;
[0231] A parallel execution module 805, configured to execute the steps of generating an explanatory video corresponding to each target file in parallel in response to the explanatory video generation instruction for each target file.
[0232] Optionally, referring to Figure 9 , a second video generation module 802, including:
[0233] An audio generation unit 812, configured to generate explanatory audio based on the explanatory text;
[0234] An explanatory image generation unit 822, configured to generate a set of explanatory images, the set of explanatory images including multiple explanatory images, where each explanatory image includes the same explanatory object, and there are at least two explanatory images in the multiple sets of explanatory images with different lip movements of the explanatory object;
[0235] A first generation unit 832, configured to generate a second video based on the explanatory audio and the set of explanatory images.
[0236] Optionally, referring to Figure 9 , the audio generation unit 812, configured to perform at least one of the following:
[0237] The explanatory video generation instruction carries a model identifier, and calls an audio conversion model indicated by the model identifier to convert the explanatory text into explanatory audio;
[0238] The explanatory video generation instruction carries an explanatory speed, and converts the explanatory text into explanatory audio with the explanatory speed.
[0239] Optionally, referring to Figure 9 , the explanatory image generation unit 822, configured to perform at least one of the following:
[0240] The explanatory video generation instruction carries an object identifier, and generates a set of explanatory images based on the explanatory object indicated by the object identifier, so that the multiple explanatory images include the explanatory object;
[0241] The explanatory video generation instruction carries an action identifier, and generates a set of explanatory images based on the dynamic action indicated by the action identifier, so that the action combinations of the explanatory objects in the multiple explanatory images form a dynamic action;
[0242] The explanatory video generation instruction carries a position parameter, and generates a set of explanatory images based on the position indicated by the position parameter, so that the explanatory object in the multiple explanatory images is located at the position.
[0243] Optionally, referring to Figure 9, the target file includes multiple data segments, the commentary audio includes commentary audio segments corresponding to at least one data segment, the commentary image set includes commentary image subsets corresponding to each data segment, and the first generation unit 832 is configured to:
[0244] Generate commentary audio based on the commentary audio segments corresponding to at least one data segment in the arrangement order of each data segment;
[0245] Generate a commentary screen based on the commentary image subsets corresponding to each data segment in the arrangement order of each data segment;
[0246] Generate a second video based on the commentary audio and the commentary screen, wherein the start playback time point of the commentary audio segment corresponding to the same data segment is the same as the start playback time point of the corresponding commentary image subset.
[0247] Optionally, refer to Figure 9 , the target file includes a first data segment and a second data segment, the commentary text includes at least one commentary text segment, wherein the first data segment has a corresponding commentary text segment and the second data segment does not have a corresponding commentary text segment; the commentary image generation unit 822 is configured to:
[0248] Generate a commentary image subset corresponding to the first data segment, and in the commentary image subset corresponding to the first data segment, there are at least two commentary images with different mouth shapes of the commentary object;
[0249] Generate a commentary image subset corresponding to the second data segment, and in the commentary image subset corresponding to the second data segment, the mouth shapes of the commentary objects in different commentary images are the same.
[0250] Optionally, refer to Figure 9 , the commentary audio includes commentary audio segments corresponding to at least one first data segment, and the commentary image generation unit 822 is configured to:
[0251] Generate a first image subset based on the first duration of the commentary audio segment corresponding to the first data segment, the first image subset includes multiple commentary images and the playback duration corresponding to each commentary image, the mouth shapes of the commentary objects in different commentary images are different, and the sum of the playback durations corresponding to the multiple commentary images is equal to the first duration;
[0252] When the content display duration of the first data segment is greater than the first duration, determine the first duration difference between the content display duration and the first duration;
[0253] Generate a second image subset based on the first duration difference, the second image subset includes at least one commentary image and the playback duration corresponding to each commentary image, the mouth shapes of the commentary objects in different commentary images are the same, and the sum of the playback durations corresponding to the at least one commentary image is equal to the first duration difference;
[0254] Construct the explanatory image subset corresponding to the first data segment from the first image subset and the second image subset.
[0255] Optionally, refer to Figure 9 , the first video generation module 801 includes:
[0256] A display image generation unit 811, configured to sequentially generate display images including each content based on multiple contents in the target file;
[0257] A second generation unit 821, configured to form the generated multiple display images into a first video.
[0258] Optionally, refer to Figure 9 , the display image generation unit 811 is configured to:
[0259] Read the target file, and when the content displayed in the target file is the first content, generate a display image including the first content;
[0260] When the content displayed in the target file is switched from the first content to the second content, continue to generate a display image including the second content until the target file reading ends, obtaining multiple display images, where the second content is the next content of the first content.
[0261] Optionally, refer to Figure 9 , the apparatus further includes:
[0262] A content switching module 806, configured to, in response to a content switching operation on the first content, switch the content displayed in the target file from the first content to the second content; or,
[0263] A content switching module 806, configured to, when the content display duration of the first content reaches the target duration corresponding to the first content, switch the content displayed in the target file from the first content to the second content.
[0264] Optionally, refer to Figure 9 , where the second content is the last content in the target file, the apparatus further includes:
[0265] A file reading module 807, configured to stop reading the target file when the content display duration of the second content has reached the target duration corresponding to the second content and the duration between the time point when the target file starts to be read and the current time point has reached the playing duration corresponding to the second video;
[0266] The file reading module 807 is further configured to continue reading the target file for the second content being displayed if the content display duration of the second content does not reach the target duration corresponding to the second content, or if the duration from the time point when reading the target file starts to the current time point does not reach the playing duration corresponding to the second video.
[0267] Optionally, referring to Figure 9 , the second generation unit 821 is configured to:
[0268] Based on the content display duration corresponding to each display image, form a first video from multiple display images, where the content display duration corresponding to a display image is the display duration of the content in the display image in the target file.
[0269] Optionally, referring to Figure 9 , the display image generation unit 811 is configured to:
[0270] When reading the target file through a browser, take screenshots of multiple contents in the target file in sequence through the browser to obtain display images including each content respectively.
[0271] Optionally, referring to Figure 9 , the commentary video generation instruction carries the storage address of the target file, and the first video generation module 801 includes:
[0272] The file acquisition unit 831 is configured to, in response to the commentary video generation instruction, acquire the target file based on the storage address;
[0273] The third generation unit 841 is configured to generate a first video based on the content of the target file.
[0274] Optionally, referring to Figure 9 , the first video generation module 801 includes:
[0275] The instruction sending unit 851 is configured to send a commentary video generation instruction to the browser through the commentary video generation application;
[0276] The display image generation unit 811 is configured to, through the browser, in response to the commentary video generation instruction, sequentially generate display images including each content based on multiple contents in the target file, and send the generated multiple display images to the commentary video generation application;
[0277] The second generation unit 821 is configured to receive multiple display images through the commentary video generation application and form a first video from the multiple display images.
[0278] It should be noted that: when the explanation video generation device provided in the above embodiment generates an explanation video, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be assigned to different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the explanation video generation device provided in the above embodiment and the embodiment of the explanation video generation method belong to the same concept. For the specific implementation process, please refer to the method embodiment and will not be elaborated here.
[0279] An embodiment of the present application also provides a computer device, which includes a processor and a memory. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the operations performed in the explanation video generation method of the above embodiment.
[0280] Optionally, the computer device is provided as a terminal. Figure 10 The structural schematic diagram of a terminal 1000 provided by an exemplary embodiment of the present application is shown.
[0281] The terminal 1000 includes: a processor 1001 and a memory 1002.
[0282] The processor 1001 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1001 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1001 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 1001 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for the rendering and drawing of the content to be displayed on the display screen. In some embodiments, the processor 1001 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process the computational operations related to machine learning.
[0283] The memory 1002 may include one or more computer-readable storage media, which may be non-transitory. The memory 1002 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1002 is used to store at least one computer program, and the at least one computer program is used to be possessed by the processor 1001 to implement the explanation video generation method provided in the method embodiments of this application.
[0284] In some embodiments, the terminal 1000 may further optionally include: a peripheral device interface 1003 and at least one peripheral device. The processor 1001, the memory 1002, and the peripheral device interface 1003 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1003 through a bus, signal lines, or a circuit board. Optionally, the peripheral device includes: a radio frequency circuit 1004.
[0285] The peripheral device interface 1003 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 1001 and the memory 1002. In some embodiments, the processor 1001, the memory 1002, and the peripheral device interface 1003 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1001, the memory 1002, and the peripheral device interface 1003 may be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0286] The radio frequency circuit 1004 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1004 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1004 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1004 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 1004 may communicate with other devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: a metropolitan area network, generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1004 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0287] Those skilled in the art can understand that Figure 10 the structure shown in Figure 10 does not constitute a limitation on the terminal 1000, and it may include more or fewer components than those shown in the figure, or combine certain components, or adopt a different component arrangement.
[0288] Optionally, the computer device is provided as a server. Figure 11 FIG. Figure 11 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 1100 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 1101 and one or more memories 1102. Among them, at least one computer program is stored in the memory 1102, and the at least one computer program is loaded and executed by the processor 1101 to implement the methods provided by the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. The server may also include other components for implementing the functions of the device, which will not be elaborated here.
[0289] An embodiment of the present application also provides a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is loaded and executed by a processor to implement the operations performed by the explanatory video generation method in the above embodiment.
[0290] An embodiment of the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer program code. The computer program code is stored in a computer-readable storage medium. The processor of the computer device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the computer device implements the operations performed by the explanatory video generation method in the above embodiment. In some embodiments, the computer program involved in the embodiments of the present application may be deployed to be executed on one computer device, or on multiple computer devices located at one place, or, on multiple computer devices distributed at multiple places and interconnected by a communication network. The multiple computer devices distributed at multiple places and interconnected by a communication network may form a blockchain system.
[0291] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a disk, or an optical disc, etc.
[0292] The above are only alternative embodiments of the embodiments of the present application and are not intended to limit the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the embodiments of the present application shall be included within the protection scope of the present application.
Claims
1. A method for generating an explanatory video, characterized in that, The method includes: In response to an instruction for generating an explanatory video of a target file, generating a first video based on the content of the target file, where the pictures in the first video are used to display the content in the target file; Generating explanatory audio based on the explanatory text corresponding to the target file, where the target file includes multiple first data segments, the explanatory text includes explanatory text segments corresponding to the multiple first data segments, and the explanatory audio includes explanatory audio segments corresponding to the multiple first data segments; For any first data segment, generating a first subset of images based on the first duration of the explanatory audio segment corresponding to the first data segment, where the first subset of images includes multiple explanatory images and the playback duration corresponding to each explanatory image, the mouth shapes of the explanatory objects in different explanatory images are different, and the sum of the playback durations corresponding to the multiple explanatory images is equal to the first duration; When the content display duration of the first data segment is greater than the first duration, determining a first duration difference between the content display duration and the first duration, where the content display duration is the display duration of the first data segment in the first video; Generating a second subset of images based on the first duration difference, where the second subset of images includes at least one of the explanatory images and the playback duration corresponding to each explanatory image, the mouth shapes of the explanatory objects in different explanatory images are the same, and the sum of the playback durations corresponding to at least one of the explanatory images is equal to the first duration difference; Combining the first subset of images and the second subset of images to form an explanatory subset of images corresponding to the first data segment, where the multiple explanatory images in the first subset of images in the explanatory subset of images are located in front of the multiple explanatory images in the second subset of images; Generating a second video based on the explanatory audio and the explanatory subsets of images corresponding to the multiple first data segments, where the second video includes the explanatory audio and explanatory pictures corresponding to the explanatory text; Fusing the first video and the second video to obtain an explanatory video corresponding to the target file.
2. The method according to claim 1, wherein The method further includes: In response to a batch selection operation on multiple target files, respectively generating instructions for generating explanatory videos for each target file; In response to the instructions for generating explanatory videos for each target file, executing the steps of generating explanatory videos corresponding to each target file in parallel.
3. The method according to claim 1, characterized in that, The generating of the explanatory audio based on the explanatory text corresponding to the target file includes at least one of the following: The explanatory video generation instruction carries a model identifier, and an audio conversion model indicated by the model identifier is called to convert the explanatory text into the explanatory audio; The explanatory video generation instruction carries an explanatory speed, and the explanatory text is converted into explanatory audio with the explanatory speed.
4. The method according to claim 1, wherein The generation process of the explanatory subset of images corresponding to any first data segment further includes at least one of the following: The explanation video generation instruction carries an object identifier, and based on the explanation object indicated by the object identifier, generates the subset of explanation images, so that multiple explanation images in the subset of explanation images include the explanation object; The explanation video generation instruction carries an action identifier, and based on the dynamic action indicated by the action identifier, generates the subset of explanation images, so that the action combination of the explanation object in multiple explanation images in the subset of explanation images constitutes the dynamic action; The explanation video generation instruction carries a position parameter, and based on the position indicated by the position parameter, generates the subset of explanation images, so that the explanation object in multiple explanation images in the subset of explanation images is located at the position; 5. The method according to claim 1, characterized in that, The target file further includes a second data segment, and the second data segment does not have a corresponding explanation text segment; the method further includes: Generating a subset of explanation images corresponding to the second data segment, and in the subset of explanation images corresponding to the second data segment, the mouth shapes of the explanation objects in different explanation images are the same; 6. The method according to any one of claims 1-5, characterized in that, Generating a first video based on the content of the target file includes: Sequentially generating display images including each content based on multiple contents in the target file; Constructing the generated multiple display images into the first video; 7. The method according to claim 6, characterized in that Sequentially generating display images including each content based on multiple contents in the target file, including: Reading the target file, and when the content displayed in the target file is the first content, generating a display image including the first content; When the content displayed in the target file is switched from the first content to the second content, continuing to generate a display image including the second content until the reading of the target file ends, obtaining multiple display images, where the second content is the next content of the first content; 8. The method according to claim 7, characterized in that When the second content is the last content in the target file, the method further includes: When the content display duration of the second content has reached the target duration corresponding to the second content, and the duration from the time point when reading the target file starts to the current time point has reached the playback duration corresponding to the second video, stopping reading the target file; When the content display duration of the second content has not reached the target duration corresponding to the second content, or the duration from the time point when reading the target file starts to the current time point has not reached the playback duration corresponding to the second video, continuing to read the target file that is displaying the second content; 9. An explanatory video generation device, characterized in that, The device includes: A first video generation module, configured to respond to an explanation video generation instruction for a target file, and generate a first video based on the content of the target file, where the picture in the first video is used to display the content in the target file; A second video generation module, configured to generate a commentary audio based on the commentary text corresponding to the target file, where the target file includes a plurality of first data segments, the commentary text includes commentary text segments corresponding to the plurality of first data segments, and the commentary audio includes commentary audio segments corresponding to the plurality of first data segments; for any first data segment, generate a first image subset based on the first duration of the commentary audio segment corresponding to the first data segment, where the first image subset includes a plurality of commentary images and the playback duration corresponding to each commentary image, the mouth shapes of the commentary objects in different commentary images are different, and the sum of the playback durations corresponding to the plurality of commentary images is equal to the first duration; in the case where the content display duration of the first data segment is greater than the first duration, determine a first duration difference between the content display duration and the first duration, where the content display duration is the display duration of the first data segment in the first video; generate a second image subset based on the first duration difference, where the second image subset includes at least one of the commentary images and the playback duration corresponding to each commentary image, the mouth shapes of the commentary objects in different commentary images are the same, and the sum of the playback durations corresponding to at least one of the commentary images is equal to the first duration difference; form the first image subset and the second image subset into the commentary image subset corresponding to the first data segment, where the plurality of commentary images in the first image subset in the commentary image subset are in front of the plurality of commentary images in the second image subset; generate a second video based on the commentary audio and the commentary image subsets corresponding to the plurality of first data segments, where the second video includes the commentary audio and commentary pictures corresponding to the commentary text; A third video generation module, configured to fuse the first video and the second video to obtain a commentary video corresponding to the target file.
10. The device according to claim 9, characterized in that The apparatus further includes: An instruction generation module, configured to generate commentary video generation instructions for each of the target files respectively in response to a batch selection operation on a plurality of the target files; A parallel execution module, configured to execute the steps of generating a commentary video corresponding to each of the target files in parallel in response to the commentary video generation instructions for each of the target files.
11. The device according to claim 9, characterized in that, The second video generation module is configured to perform at least one of the following: The commentary video generation instruction carries a model identifier, and calls the audio conversion model indicated by the model identifier to convert the commentary text into the commentary audio; The commentary video generation instruction carries a commentary speed, and converts the commentary text into a commentary audio with the commentary speed.
12. The device according to claim 9, wherein, The second video generation module is further configured to perform at least one of the following: The commentary video generation instruction carries an object identifier, and generates the commentary image subset based on the commentary object indicated by the object identifier, so that the plurality of commentary images in the commentary image subset include the commentary object; The explanation video generation instruction carries an action identifier, and based on the dynamic action indicated by the action identifier, generates the subset of explanation images, so that the action combination of the explanation object in the multiple explanation images in the subset of explanation images constitutes the dynamic action; The explanation video generation instruction carries a position parameter, and based on the position indicated by the position parameter, generates the subset of explanation images, so that the explanation object in the multiple explanation images in the subset of explanation images is located at the position.
13. The device according to claim 9, wherein The target file further includes a second data segment, and the second data segment does not have a corresponding explanation text segment; the second video generation module is further configured to: Generate a subset of explanation images corresponding to the second data segment, and in the subset of explanation images corresponding to the second data segment, the mouth shapes of the explanation object in different explanation images are the same.
14. The device according to any one of claims 9 - 13, characterized in that, The first video generation module includes: A display image generation unit, configured to sequentially generate display images including each content based on multiple contents in the target file; A second generation unit, configured to form the first video from the multiple generated display images.
15. The device according to claim 14, wherein The display image generation unit is configured to: Read the target file, and when the content displayed in the target file is the first content, generate a display image including the first content; When the content displayed in the target file is switched from the first content to the second content, continue to generate a display image including the second content until the reading of the target file ends, obtaining multiple display images, where the second content is the next content of the first content.
16. The device according to claim 15, characterized in that, When the second content is the last content in the target file, the device further includes: A file reading module, configured to stop reading the target file when the content display duration of the second content has reached the target duration corresponding to the second content and the duration between the time point when the target file starts to be read and the current time point has reached the playback duration corresponding to the second video; The file reading module is further configured to continue reading the target file that is currently displaying the second content when the content display duration of the second content has not reached the target duration corresponding to the second content, or the duration between the time point when the target file starts to be read and the current time point has not reached the playback duration corresponding to the second video.
17. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the explanation video generation method according to any one of claims 1 to 8.
18. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor to implement the operations performed by the explanation video generation method according to any one of claims 1 to 8.
19. A computer program product comprising a computer program, characterized in that, The computer program is loaded and executed by a processor to implement the operations performed by the explanation video generation method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method and system for applying virtual character to automatic video production, and storage medium
CN113259778A