Animation video generation method, system, computer device and storage medium

Through the animation video generation method, animation videos are generated using the play text and pending pictures, which solves the problem of low efficiency in demonstration pictures and achieves fast and efficient animation video generation and voice explanation synchronization effects.

CN114399571BActive Publication Date: 2025-06-10CHINA PING AN LIFE INSURANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210058098.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-19
Publication Date
2025-06-10
Estimated Expiration
2042-01-19

AI Technical Summary

Technical Problem

Explanation based on demonstration pictures is low efficiency, the existing technology requires a lot of manpower and time for post-processing, and the different processor methods are not common.

Method used

Through an animation video generation method, the animation video voice is generated using the play text, the candidate display object is obtained from the pending pictures, the video display object is determined based on the play text text, and the object output time is determined based on the video display object and the animation video voice, a coordinate reference table is generated, and the animation video is finally generated based on the coordinate reference table and the animation video voice.

Benefits of technology

It improves the interpretation efficiency based on demonstration pictures, can quickly generate animated videos for explanations, reduces manpower and time costs, and realizes the synchronization effect of animation videos and voice explanations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114399571B_ABST
    Figure CN114399571B_ABST
Patent Text Reader

Abstract

The present application discloses an animation video generation method, system, computer device, and storage medium. The animation video generation method includes: generating animation video voice according to the script text; obtaining candidate display objects from the pictures to be processed; determining video display objects from the candidate display objects according to the script text; determining the object output time according to the video display objects and the animation video voice; generating a coordinate reference table according to the preset animation video display mode, video display objects, and object output time; and generating an animation video according to the coordinate reference table and the animation video voice. The animation video generation method can improve the explanation efficiency based on the demonstration pictures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular, to an animated video generation method, system, computer device, and storage medium. Background Art

[0002] For presentation pictures used in PPT (PowerPoint, a graphical presentation software), since most of the presentation pictures are static pictures when explaining PPT, audiences are prone to distraction and inattentiveness during the lecture. Currently, some presentation effects can be manually added to the presentation pictures in the later stage to assist the explanation. However, this method requires a large amount of manpower and time, and the explanation and processing methods of presentation pictures by different processors are not universal. Currently, the explanation efficiency based on presentation pictures is relatively low. Summary of the Invention

[0003] In view of this, embodiments of this application provide an animated video generation method, system, computer device, and storage medium to solve the problem of relatively low explanation efficiency based on presentation pictures.

[0004] In a first aspect, embodiments of this application provide an animated video generation method, and the method includes:

[0005] Generate animated video voice according to the script text;

[0006] Obtain candidate display objects from the pictures to be processed;

[0007] Determine video display objects from the candidate display objects according to the script text;

[0008] Determine the object output time according to the video display objects and the animated video voice;

[0009] Generate a coordinate reference table according to a preset animated video display method, the video display objects, and the object output time;

[0010] Generate an animated video according to the coordinate reference table and the animated video voice.

[0011] In the above aspect and any possible implementation manner, a further implementation manner is provided. The candidate display objects include candidate text information objects and candidate image information objects. The obtaining of candidate display objects from the pictures to be processed includes:

[0012] Use an OCR recognition model to recognize the pictures to be processed to obtain the candidate text information objects;

[0013] Use a saliency segmentation algorithm to segment the pictures to be processed to obtain the region of interest;

[0014] Perform binarization processing on the region of interest to obtain a binarized region of interest;

[0015] Obtain the candidate image information object according to the binarized region of interest.

[0016] In the above aspect and any possible implementation manner, a further implementation manner is provided. The candidate display object includes a candidate text information object and a candidate image information object. Determining a video display object from the candidate display objects according to the script text includes:

[0017] Obtain the text meaning of the candidate image information object;

[0018] Based on semantic association, according to the text meaning of the candidate image information object, associate the candidate image information object with the candidate text information object;

[0019] Perform string matching between the script text and the candidate text information object, and use the candidate text information object that matches the script text and the candidate image information object associated with the candidate text information object that matches the script text as the video display object.

[0020] In the above aspect and any possible implementation manner, a further implementation manner is provided. Determining the object output time according to the video display object and the animated video voice includes:

[0021] Determine the voice playing duration of the animated video voice according to the voice playing rate input by the user;

[0022] Determine the display order of the video display object according to the animated video voice;

[0023] Determine the object output time according to the display order of the video display object and the voice playing duration of the animated video voice.

[0024] In the above aspect and any possible implementation manner, a further implementation manner is provided. Generating an animated video according to the coordinate reference table and the animated video voice includes:

[0025] Determine the animated picture corresponding to the number of frames according to the coordinate reference table;

[0026] Determine the animated video voice corresponding to the playing of the animated picture according to the number of frames;

[0027] Generate the animated video according to the animated picture corresponding to the number of frames and the animated video voice corresponding to the playing of the animated picture.

[0028] For the aspects and any possible implementation manners as described above, a further implementation manner is provided. The video display object includes a display text information object and a display image information object. Determining an animated picture corresponding to a frame number according to the coordinate reference table includes:

[0029] Reset the picture to be processed to an initial display picture, where the initial display picture hides the display text information object and displays the display image information object;

[0030] According to the coordinate reference table, using the initial display picture as the starting picture, generate the animated picture frame by frame according to the frame number.

[0031] For the aspects and any possible implementation manners as described above, a further implementation manner is provided. After generating the animated video according to the coordinate reference table and the animated video voice, the method further includes:

[0032] Obtain the drawing area input by the user on the picture to be processed;

[0033] Determine the annotation coordinate parameters according to the drawing area;

[0034] Modify the coordinate reference table according to the annotation coordinate parameters and the preset annotation type, and generate an animated video with annotation explanations according to the animated video voice and the modified coordinate reference table.

[0035] In a second aspect, an animated video generation system provided by an embodiment of the present application includes:

[0036] A first generation module, configured to generate an animated video voice according to the script text;

[0037] An acquisition module, configured to acquire candidate display objects from the picture to be processed;

[0038] A first determination module, configured to determine a video display object from the candidate display objects according to the script text;

[0039] A second determination module, configured to determine the object output time according to the video display object and the animated video voice;

[0040] A second generation module, configured to generate a coordinate reference table according to a preset animated video display manner, the video display object, and the object output time;

[0041] A third generation module, configured to generate an animated video according to the coordinate reference table and the animated video voice.

[0042] Further, the candidate display objects include candidate text information objects and candidate image information objects.

[0043] Further, the obtaining module is specifically configured to:

[0044] Use an OCR recognition model to recognize the to-be-processed picture, and obtain the candidate text information object;

[0045] Use a saliency segmentation algorithm to segment the to-be-processed picture, and obtain the region of interest;

[0046] Perform binarization processing on the region of interest to obtain a binarized region of interest;

[0047] Obtain the candidate image information object according to the binarized region of interest.

[0048] Further, the first determination module is specifically configured to:

[0049] Obtain the text meaning of the candidate image information object;

[0050] Based on semantic association, according to the text meaning of the candidate image information object, associate the candidate image information object with the candidate text information object;

[0051] Perform string matching on the script text and the candidate text information object, and use the candidate text information object that matches the script text and the candidate image information object associated with the candidate text information object that matches the script text as the video display object.

[0052] Further, the second determination module is specifically configured to:

[0053] Determine the speech playing duration of the animation video voice according to the speech playing rate input by the user;

[0054] Determine the display order of the video display object according to the animation video voice;

[0055] Determine the object output time according to the display order of the video display object and the speech playing duration of the animation video voice.

[0056] Further, the third generation module is specifically configured to:

[0057] Determine the animation picture corresponding to the number of frames according to the coordinate reference table;

[0058] Determine the animation video voice corresponding to the playing of the animation picture according to the number of frames;

[0059] Generate the animation video according to the animation picture corresponding to the number of frames and the animation video voice corresponding to the playing of the animation picture.

[0060] Further, the video display object includes a display text information object and a display image information object.

[0061] Further, the third generation module is further specifically configured to:

[0062] Reset the to-be-processed picture to an initial display picture, where the initial display picture hides the display text information object and displays the display image information object;

[0063] According to the coordinate reference table, using the initial display picture as the starting picture, generate the animation pictures frame by frame according to the number of frames.

[0064] Further, the animation video generation system is further specifically configured to:

[0065] Obtain the drawing area input by the user on the to-be-processed image;

[0066] Determine the annotation coordinate parameters according to the drawing area;

[0067] Modify the coordinate reference table according to the annotation coordinate parameters and the preset annotation type, and generate an animated video with annotation explanations according to the animated video voice and the modified coordinate reference table.

[0068] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, it executes the steps of the animated video generation method as described in the first aspect.

[0069] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of the animated video generation method as described in the first aspect are implemented.

[0070] In the embodiments of the present application, first, animated video voice is generated according to the script text so that there is voice-assisted explanation during the playback of the animated video; then, candidate display objects are obtained from the to-be-processed pictures, and potential objects to be displayed in the animated video can be initially screened out; then, video display objects are determined from the candidate display objects according to the script text, which can be combined with the script text to quickly determine the objects that the user wants to focus on displaying in the animated video; then, the object output time is determined according to the video display objects and the animated video voice, and the objects can be displayed in sequence and with emphasis during the playback of the animated video; then, a coordinate reference table is generated according to the preset animated video display method, the video display objects, and the object output time. Through this coordinate reference table, the change of the video display objects over time can be determined, so as to accurately achieve the animated display effect; finally, an animated video is generated according to the coordinate reference table and the animated video voice, and when the video display changes over time to present an animated display effect, the voice explanation can be synchronized. The animated video in the embodiments of the present application can quickly generate an animated video for explanation based on the demonstration pictures, which can improve the explanation efficiency based on the demonstration pictures. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 is a flowchart of a method for generating an animated video in an embodiment of the present application;

[0072] Figure 2 is a schematic block diagram of an animated video generation system corresponding one-to-one to the method for generating an animated video in an embodiment of the present application;

[0073] Figure 3 is a schematic diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0074] In order to better understand the technical solutions of the present application, the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0075] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0076] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms of "a", "the", and "said" used in the embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0077] It should be understood that the term "and / or" used herein is merely a description of the same field of related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this text generally represents an "or" relationship between the related objects before and after.

[0078] It should be understood that although terms such as first, second, and third may be used in the embodiments of this application to describe preset ranges, etc., these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from each other. For example, without departing from the scope of the embodiments of this application, the first preset range can also be referred to as the second preset range, and similarly, the second preset range can also be referred to as the first preset range.

[0079] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".

[0080] Figure 1 is a flowchart of an animation video generation method in the embodiments of this application. This animation video generation method can be applied to a scenario where an explanation is given based on demonstration pictures (pictures for demonstration and explanation purposes). The explainer can use this animation video generation method to quickly generate an animation video for explanation through the demonstration pictures. As Figure 1 shown, this animation video generation method includes the following steps:

[0081] S10: Generate the voice of the animation video according to the script text.

[0082] Among them, the script text refers to the script for the video explanation. This script text determines the overall idea and key points of the video explanation.

[0083] In one embodiment, the TTS (Text to Speech) technology can be used to convert the script text into the voice of the animation video for the narration of the video. It can be understood that the script text can be preset by the user. The user can preset the script text according to the needs of the explanation idea and key points. In this way, when generating the animation video using the demonstration pictures (such as PPT pictures), the content that needs to be highlighted in the demonstration pictures can be determined through this script text.

[0084] S20: Obtain candidate display objects from the pictures to be processed.

[0085] Among them, the picture to be processed can specifically refer to a presentation picture such as a PPT picture. The candidate display object refers to an object that may potentially be used for animation video display, and the object can include text-type objects and image-type objects. Understandably, the picture to be processed can include various text-type objects and image-type objects, and some of these objects are the key contents of the video explanation.

[0086] In one embodiment, the content of the picture to be processed including various objects can be screened first to obtain objects that may potentially be used for animation video display.

[0087] S30: Determine the video display object from the candidate display objects according to the script text.

[0088] Among them, the video display object refers to an object that needs to be highlighted and displayed in the animation video, and these objects come from the picture to be processed.

[0089] In one embodiment, the video display object can be determined through the script text. For different script texts, it is supported to determine different video display objects from the candidate display objects. In this way, when the user's explanation idea and key points of explanation change, the object that needs to be highlighted and displayed can be re-determined from the candidate display objects by adjusting the script text.

[0090] S40: Determine the object output time according to the video display object and the animation video voice.

[0091] Among them, the object output time refers to the specific time when the video display object is played in the animation video.

[0092] In one embodiment, on the premise that the animation video voice is determined, the total duration of the animation video voice playback can be obtained. Then, under the known total duration, the object output time can be determined according to the explanation order of the video display object in the script text, so that the video display object can be displayed in sequence and with emphasis during the animation video playback. Understandably, determining the object output time also helps to determine the start and end times of the animation playback effect of the video display object in the animation video, and can more accurately control the animation playback duration of the video display object.

[0093] S50: Generate a coordinate reference table according to the preset animation video display method, video display object, and object output time.

[0094] Among them, the coordinate reference table records the coordinate conditions of the video display object changing with time in the animated video. It can be understood that the animated video is obtained from multiple pictures at a certain frame rate. When the video display object has an animation playback effect, the change of the corresponding frame picture mainly lies in the change of the video display object. Therefore, by using the animated video display method (preset animation playback effect), the change rule of the video display object can be determined, and then combined with the object output time, the pictures of the corresponding frames when the video display object has an animation effect can be generated.

[0095] S60: Generate an animated video according to the coordinate reference table and the animated video voice.

[0096] In one embodiment, with the coordinate reference table, there is the change rule of the video display object in the animated video over time. Based on the existing image to be processed and combined with the animated video voice, the animated video can be quickly generated. It should be noted that this process does not require a large amount of memory resources, and the animated video can be generated based on the coordinate reference table and the image to be processed.

[0097] In the embodiment of the present application, first, the animated video voice is generated according to the script text so that there is voice assistance for explanation when the animated video is played; then the candidate display objects are obtained from the pictures to be processed, and the potential objects to be displayed in the animated video can be initially screened out; then the video display object is determined from the candidate display objects according to the script text, which can be combined with the script text to quickly determine the object that the user wants to focus on displaying in the animated video; then the object output time is determined according to the video display object and the animated video voice, and the objects can be displayed in sequence and with emphasis when the animated video is played; then the coordinate reference table is generated according to the preset animated video display method, the video display object and the object output time. Through this coordinate reference table, the change situation of the video display object over time can be determined, so as to accurately achieve the animation display effect; finally, the animated video is generated according to the coordinate reference table and the animated video voice, and when the video display changes over time to present an animation display effect, the voice explanation can be synchronized. The animated video in the embodiment of the present application can quickly generate an animated video for explanation based on the demonstration pictures, which can improve the explanation efficiency based on the demonstration pictures. Further, the candidate display objects include candidate text information objects and candidate image information objects.

[0098] Further, the candidate display objects include candidate text information objects and candidate image information objects.

[0099] It can be understood that the key parts on the demonstration pictures include text-type objects and image-type objects. In this embodiment, the objects are mainly divided into these two parts.

[0100] Further, in step S20, that is, when obtaining the candidate display objects from the pictures to be processed, the following steps are included:

[0101] S21: Identify the image to be processed using an OCR recognition model to obtain a candidate text information object.

[0102] In one embodiment, OCR (Optical Character Recognition) technology can be used to process the image to be processed, and a candidate text information object that may potentially be used for display can be obtained from the image to be processed. For text-based objects, OCR recognition can be used to simply and quickly extract them.

[0103] S22: Segment the image to be processed using a saliency segmentation algorithm to obtain a region of interest.

[0104] Among them, the saliency segmentation algorithm is an image segmentation algorithm used to focus on foreground and background details.

[0105] In one embodiment, the image to be processed can be segmented from the perspective of foreground and background details of the image to obtain a region of interest. Among them, the region of interest may refer to some iconic icons, and these iconic images generally have a large difference from the surrounding colors and line distributions. These iconic icons can be associated with the text, such as a title icon used to summarize the text.

[0106] S23: Perform binarization processing on the region of interest to obtain a binarized region of interest.

[0107] In one embodiment, binarization processing can also be performed on the region of interest, which can more precisely delimit the range of iconic images, such as removing the interference of some other graphics or text.

[0108] S24: Obtain a candidate image information object based on the binarized region of interest.

[0109] In one embodiment, considering the display processing for facilitating the animation playback effect, after using the binarization operation to remove the interference of other graphics or text on the region of interest, the binarized region of interest can be expanded into a frame region with a larger area range to obtain a candidate image information object. In this way, the iconic images in the original region of interest can be retained, and it is convenient for the display processing of the animation playback effect.

[0110] In steps S21 - S24, when obtaining candidate display objects from the image to be processed, the candidate text information object and the candidate image information object can be accurately extracted separately. Among them, when extracting the candidate image information object, the interference of other graphics or text can also be removed, and the image quality of the obtained candidate image information object is better.

[0111] Further, in step S30, that is, when determining the video display object from the candidate display objects according to the script text, the following steps are included:

[0112] S31: Obtain the text meaning of the candidate image information object.

[0113] Understandably, there is generally an association between the candidate image information object and the candidate text information object. For example, for a title icon (one of the candidate image information objects) used to summarize text, the associated candidate text information object is related to the title content.

[0114] In one embodiment, the user can pre-enter the text meaning of the candidate image information object. After the system receives this text meaning, it can determine the relationship between the candidate image information and the candidate text information object based on this text meaning.

[0115] S32: Based on semantic association, according to the text meaning of the candidate image information object, associate the candidate image information object with the candidate text information object.

[0116] In one embodiment, a semantic analysis model can be specifically used to analyze the semantic association between the text meaning and the candidate text information object. For example, if the keyword of the text meaning of the candidate image information object is "medical health", then the semantic analysis model will search for the associated candidate text information object according to this keyword of "medical health". Understandably, the associated candidate text information object and the associated candidate image information object often appear simultaneously or have an animation display effect within the same time period during video playback. In this application, the originally fragmented candidate text information object and candidate image information object are associated, so that the problem of disjointed text and image animation display can be avoided during the animation video display.

[0117] S33: Perform string matching between the script text and the candidate text information object, and use the candidate text information object that matches the script text and the candidate image information object associated with the candidate text information object that matches the script text as the video display object.

[0118] Understandably, the candidate text information object is a preliminarily extracted object, and it is also necessary to combine the script text to determine whether the candidate text information should be highlighted with an animation effect in the animation video.

[0119] In one embodiment, the script text and the candidate text information object can be subjected to string matching. If there is a string in the candidate text information that matches the script text, it can be considered that the matching string part is the key object of this video explanation. Synchronously, the candidate image information object associated with the candidate text information object that matches the script text should also be synchronously determined as the key object of the video explanation, so as to determine the final video display object. It should be understood that this video display object does not refer to the only object shown in the animated video, but should be considered as the key object shown in the animated video. The animated video may include other content in the picture to be processed that is not prominently shown as the key point.

[0120] In steps S31 - S33, a specific implementation manner for determining the video display object is provided. The candidate text information object and the candidate image information object can be associated through semantic association, and the video display object can be determined through string matching, which can accurately determine the video display object for this animated video explanation.

[0121] Furthermore, in step S40, that is, in determining the object output time according to the video display object and the animated video voice, the following steps are included:

[0122] S41: Determine the voice playback duration of the animated video voice according to the voice playback rate input by the user.

[0123] It can be understood that the voice playback duration is controlled by the voice playback rate. In this embodiment, the user can control the voice playback rate accordingly according to the actual voice playback duration requirement. In the case where the lengths of the script texts are different, the voice playback duration can be controlled within the expected duration by adjusting the voice playback rate.

[0124] S42: Determine the display order of the video display objects according to the animated video voice.

[0125] It can be understood that there are multiple video display objects, and the animated display times of each video display object in the animated video are different. In one embodiment, the playback order of the animated video voice can be used as a reference for the animated display time of the video display object. Specifically, when the animated video voice plays to a specific script line sentence, the video display object corresponding to the semantics of this script line sentence will be shown in the animated video. In this embodiment, the explanation order of the script lines in the animated video voice can be used to determine the display order of the video display objects.

[0126] S43: Determine the object output time according to the display order of the video display objects and the voice playback duration of the animated video voice.

[0127] In one embodiment, after determining the voice playback duration of the animated video voice, the time corresponding to the display of the video display object can be allocated according to the display order of the video display objects, that is, the object output time. Specifically, when the current video display object is ranked according to the display order, the voice playback duration of the animated video voice is used as the playback start time point of the current video display object. When the next video display object is ranked according to the display order, the voice playback duration of the animated video voice is used as the playback end time point of the current video display object. The object output time can be determined from the playback start time point and the playback end time point. This object output time helps to establish a coordinate reference table in terms of time (corresponding to the number of frames of the video), so as to generate an animated video more simply and quickly.

[0128] In steps S41 - S43, a specific implementation manner for determining the object output time is provided. The display order of the video display objects can be determined by using the playback order included in the animated video voice, and further, the display time of the video display objects on the animated video can be determined one by one based on the voice playback duration. This implementation manner can quickly determine the object output time.

[0129] Furthermore, in step S60, that is, generating an animated video according to the coordinate reference table and the animated video voice, includes the following steps:

[0130] S61: Determine the animated pictures corresponding to the number of frames according to the coordinate reference table.

[0131] Among them, the number of frames can adopt a preset fixed number of frames.

[0132] It can be understood that an animated video is obtained from multiple frames of animated pictures. In the non - static pictures of the animated video, each frame of the animated picture is different. During the display of the video display object, although each frame of the animated picture is different, the main change between the animated pictures is the change of the video display object. In the embodiments of the present application, by using the coordinate situation recording the change of the video display object in the animated video over time, the animated pictures corresponding to the number of frames can be generated quickly and simply according to the preset fixed number of frames.

[0133] S62: Determine the animated video voice corresponding to the playback of the animated pictures according to the number of frames.

[0134] In one embodiment, the number of frames determines the number of animated pictures required for displaying the animation effect and the display quality of the animation effect. For example, when a higher number of frames is used for the animated video, more animated pictures are required, and the display quality of the animation effect is also higher. In one embodiment, the total number of animated pictures to be played can be determined according to the number of frames, and the corresponding animated video voice during playback can be determined. This helps to accurately display the video display object and the animation effect during the playback of the animated video voice.

[0135] S63: Generate an animated video based on the animated pictures corresponding to the number of frames and the animated video voice corresponding to the playback of the animated pictures.

[0136] In one embodiment, an animated video can be obtained by combining the animated pictures and the animated video voice according to the playback time. The animated video includes a speech explanation related to the script text. When the animated video is played and explained, the corresponding video display object will display the animation effect synchronously with the animated video voice. The audience can more easily notice the key content in the animated video.

[0137] In steps S61 - S63, the animated pictures corresponding to the number of frames can be quickly and simply generated according to the preset fixed number of frames, and combined with the animated video voice to generate an animated video.

[0138] Furthermore, the video display object includes an object for displaying text information and an object for displaying image information.

[0139] It can be understood that the video display object includes an object of text type and an object of image type, which are respectively called an object for displaying text information and an object for displaying image information.

[0140] Furthermore, in step S61, that is, to determine the animated pictures corresponding to the number of frames according to the coordinate reference table, the following steps are included:

[0141] S611: Reset the picture to be processed to the initial display picture, where the initial display picture hides the object for displaying text information and displays the object for displaying image information.

[0142] It can be understood that the coordinate reference table is obtained according to the display mode of the animated video, and the coordinate information recorded therein can reflect the display mode of the animated video. In the embodiments of the present application, before generating the animated pictures, the picture to be processed can be first reset to the initial display picture. Specifically, the initial display picture can be the animated picture corresponding to the number of frames at the starting time point of the playback of the video display object. The initial display picture can be initially set to hide the object for displaying text information and display the object for displaying image information. For example, the object for displaying image information as the title is displayed in the initial display picture, and the remaining objects for displaying text information associated with the title will be displayed in the animated pictures corresponding to the subsequent frames.

[0143] S612: According to the coordinate reference table, starting from the initial display picture, generate animated pictures frame by frame according to the number of frames.

[0144] In one embodiment, after setting the initial display picture, animated pictures can be generated frame by frame based on the coordinates in the coordinate reference table according to the initial display picture. Further, if the real-time computing power of the playback carrier of the animated video itself is strong enough, the coordinate reference table can be used to pre-generate the animated pictures needed next during the playback of the animated video. In this way, the animated pictures do not have to occupy storage all the time. After the animated pictures are generated and played, they can be deleted, and only the coordinate reference table needs to be retained.

[0145] In steps S611 - S612, animated pictures can be generated frame by frame based on the coordinate reference table on the basis of setting the initial display picture. In addition, the present application also utilizes the characteristic that the change of animated pictures between adjacent frames is small, so that it will be more convenient to generate animated pictures according to the number of frames, and the computing power resources occupied will be smaller.

[0146] Further, after step S60, that is, after generating the animated video according to the coordinate reference table and the animated video voice, the following steps are also included:

[0147] S71: Obtain the drawing area input by the user on the image to be processed.

[0148] In one embodiment, the user (explainer) can also pre-draw the area for annotation and explanation on the image to be processed, and this area is the drawing area.

[0149] S72: Determine the annotation coordinate parameters according to the drawing area.

[0150] Among them, the annotation coordinate parameters refer to the coordinates of the pixels corresponding to the drawing area in the image to be processed.

[0151] In one embodiment, using the pixel distribution of the image to be processed, the annotation coordinate parameters corresponding to the drawing area can be determined by comparing the pixel differences before and after the input drawing area.

[0152] S73: Modify the coordinate reference table according to the annotation coordinate parameters and the preset annotation type, and generate an animated video with annotation explanation according to the animated video voice and the modified coordinate reference table.

[0153] Among them, the preset annotation type can specifically be the type of annotation behavior and annotation action, including the annotation behavior specifically performed in the drawing area. Modify the coordinate reference table according to the annotation coordinate parameters (initial) and the preset annotation type (including annotation behaviors and actions that change with the number of frames). In this way, animated pictures with annotation explanation can be generated frame by frame according to the modified coordinate reference table, and an animated video with annotation explanation can be generated in combination with the animated video voice.

[0154] In steps S71 - S73, the coordinate reference table can be modified based on the drawing area input on the image to be processed and the preset annotation types, so as to generate an animated video with annotation explanations. In this way, when playing the animated video, contents of annotation explanations such as drawing red lines and drawing circles will be displayed on the animated video.

[0155] Furthermore, the animated video can be generated according to actual playback requirements, specifically achieved by correspondingly modifying the coordinate reference table, and the steps for generating the animated video with annotation explanations (annotation requirements) in steps S71 - S73 can be referred to.

[0156] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0157] Figure 2 is a schematic block diagram of an animated video generation system corresponding one - to - one to the animated video generation method in the embodiments of the present application. As Figure 2 shown, the animated video generation system includes a first generation module 10, an acquisition module 20, a first determination module 30, a second determination module 40, a second generation module 50, and a third generation module 60.

[0158] The first generation module 10 is used to generate the voice of the animated video according to the script text.

[0159] The acquisition module 20 is used to acquire candidate display objects from the image to be processed.

[0160] The first determination module 30 is used to determine the video display objects from the candidate display objects according to the script text.

[0161] The second determination module 40 is used to determine the object output time according to the video display objects and the voice of the animated video.

[0162] The second generation module 50 is used to generate a coordinate reference table according to the preset animated video display mode, video display objects, and object output time.

[0163] The third generation module 60 is used to generate an animated video according to the coordinate reference table and the voice of the animated video.

[0164] Optionally, the candidate display objects include candidate text information objects and candidate image information objects.

[0165] Optionally, the acquisition module 20 is specifically used for:

[0166] Using an OCR recognition model to recognize the image to be processed to obtain candidate text information objects;

[0167] Segment the image to be processed using a saliency segmentation algorithm to obtain the region of interest;

[0168] Perform binarization processing on the region of interest to obtain a binarized region of interest;

[0169] Obtain candidate image information objects based on the binarized region of interest.

[0170] Optionally, the first determination module 30 is specifically configured to:

[0171] Obtain the text meaning of the candidate image information object;

[0172] Based on semantic association, according to the text meaning of the candidate image information object, associate the candidate image information object with the candidate text information object;

[0173] Perform string matching on the script text and the candidate text information object, and use the candidate text information object that matches the script text and the candidate image information object associated with the candidate text information object that matches the script text as the video display object.

[0174] Optionally, the second determination module 40 is specifically configured to:

[0175] Determine the speech playback duration of the animated video voice according to the speech playback rate input by the user;

[0176] Determine the display order of the video display object according to the animated video voice;

[0177] Determine the object output time according to the display order of the video display object and the speech playback duration of the animated video voice.

[0178] Optionally, the third generation module 60 is specifically configured to:

[0179] Determine the animated picture corresponding to the frame number according to the coordinate reference table;

[0180] Determine the animated video voice corresponding to the playback of the animated picture according to the frame number;

[0181] Generate an animated video according to the animated picture corresponding to the frame number and the animated video voice corresponding to the playback of the animated picture.

[0182] Optionally, the video display object includes a display text information object and a display image information object.

[0183] Optionally, the third generation module 60 is also specifically configured to:

[0184] Reset the image to be processed to the initial display picture, where the initial display picture hides the display text information object and displays the display image information object;

[0185] According to the coordinate reference table, using the initial display picture as the starting picture, generate animated pictures frame by frame according to the number of frames.

[0186] Optionally, the animated video generation system is further specifically configured to:

[0187] Obtain the drawing area input by the user on the image to be processed;

[0188] Determine the annotation coordinate parameters according to the drawing area;

[0189] Modify the coordinate reference table according to the annotation coordinate parameters and the preset annotation type, and generate an animated video with annotation explanations according to the animated video voice and the modified coordinate reference table.

[0190] In the embodiments of the present application, first, generate the animated video voice according to the script text so that there is voice assistance for explanation when the animated video is played; then obtain the candidate display objects from the image to be processed, and potentially screen out the objects to be displayed in the animated video; then determine the video display objects from the candidate display objects according to the script text, which can be combined with the script text to quickly determine the objects that the user wants to focus on displaying in the animated video; then determine the object output time according to the video display objects and the animated video voice, which can display the objects in sequence and with emphasis when the animated video is played; then generate the coordinate reference table according to the preset animated video display method, video display objects, and object output time. Through this coordinate reference table, the change situation of the video display objects over time can be determined, so as to accurately achieve the animated display effect; finally, generate the animated video according to the coordinate reference table and the animated video voice, which can synchronously cooperate with the voice explanation when the video display presents the animated display effect over time. The animated video in the embodiments of the present application can quickly generate an animated video for explanation based on the demonstration picture, which can improve the explanation efficiency based on the demonstration picture.

[0191] Figure 3 It is a schematic diagram of a computer device in the embodiments of the present application.

[0192] As Figure 3 shown, the computer device 110 includes a processor 111, a memory 112, and computer-readable instructions 113 stored in the memory 112 and executable on the processor 111. When the processor 111 executes the computer-readable instructions 113, each step of the animated video generation method is implemented.

[0193] Exemplarily, the computer-readable instructions 113 may be divided into one or more modules / units, which are stored in the memory 112 and executed by the processor 111 to complete this application. One or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer-readable instructions 113 in the computer device 110.

[0194] The computer device 110 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device may include, but is not limited to, a processor 111 and a memory 112. Those skilled in the art can understand that Figure 3 This is only an example of the computer device 110 and does not constitute a limitation on the computer device 110. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the computer device may also include input / output devices, network access devices, a bus, etc.

[0195] The so-called processor 111 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0196] The memory 112 may be an internal storage unit of the computer device 110, such as the hard disk or memory of the computer device 110. The memory 112 may also be an external storage device of the computer device 110, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 110. Further, the memory 112 may also include both the internal storage unit and the external storage device of the computer device 110. The memory 112 is used to store computer-readable instructions and other programs and data required by the computer device. The memory 112 may also be used to temporarily store data that has been output or will be output.

[0197] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0198] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0199] In the embodiments of the present application, the server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0200] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the system is divided into different functional units or modules to complete all or part of the functions described above.

[0201] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0202] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present application, it can also be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium. When the computer-readable instructions are executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer-readable instructions include computer-readable instruction codes, and the computer-readable instruction codes can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium can include: any entity or device that can carry the computer-readable instruction code, recording media, USB flash drives, mobile hard disks, magnetic disks, optical discs, computer memories, read-only memories (ROMs), random access memories (RAMs), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0203] The present application also provides a computer-readable storage medium. The computer-readable storage medium stores computer-readable instructions. When the computer-readable instructions are executed by a processor, an animation video generation method is implemented.

[0204] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for generating an animated video, characterized in that, it includes: generating the voice of the animated video according to the script text; obtaining candidate display objects from the pictures to be processed; determining video display objects from the candidate display objects according to the script text, where the candidate display objects include candidate text information objects and candidate image information objects, including: obtaining the text meaning of the candidate image information object; based on semantic association, according to the text meaning of the candidate image information object, associating the candidate image information object with the candidate text information object; performing string matching on the script text and the candidate text information object, and taking the candidate text information object that matches the script text and the candidate image information object associated with the candidate text information object that matches the script text as the video display object; determining the object output time according to the video display object and the voice of the animated video; generating a coordinate reference table according to the preset animated video display method, the video display object and the object output time; generating an animated video according to the coordinate reference table and the voice of the animated video.

2. The method according to claim 1, characterized in that, the candidate display objects include candidate text information objects and candidate image information objects, and the obtaining candidate display objects from the pictures to be processed includes: using an OCR recognition model to recognize the pictures to be processed to obtain the candidate text information objects; using a saliency segmentation algorithm to segment the pictures to be processed to obtain the region of interest; performing binarization processing on the region of interest to obtain a binarized region of interest; obtaining the candidate image information object according to the binarized region of interest.

3. The method according to claim 1, characterized in that, the determining the object output time according to the video display object and the voice of the animated video includes: determining the voice playing duration of the voice of the animated video according to the voice playing rate input by the user; determining the display order of the video display objects according to the voice of the animated video; determining the object output time according to the display order of the video display objects and the voice playing duration of the voice of the animated video.

4. The method according to claim 1, characterized in that, the generating an animated video according to the coordinate reference table and the voice of the animated video includes: determining the animated pictures corresponding to the number of frames according to the coordinate reference table; determining the voice of the animated video corresponding to the playing of the animated pictures according to the number of frames; generating the animated video according to the animated pictures corresponding to the number of frames and the voice of the animated video corresponding to the playing of the animated pictures.

5. The method according to claim 4, characterized in that, the video display objects include display text information objects and display image information objects, and the determining the animated pictures corresponding to the number of frames according to the coordinate reference table includes: resetting the pictures to be processed to the initial display pictures, where the initial display pictures hide the display text information objects and display the display image information objects; Based on the coordinate reference table, using the initial display picture as the starting picture, generate the animation pictures frame by frame according to the number of frames.

6. The method according to any one of claims 1-5, wherein, after generating the animation video according to the coordinate reference table and the animation video voice, the method further includes: Obtain the drawing area input by the user on the image to be processed; Determine the annotation coordinate parameters according to the drawing area; Modify the coordinate reference table according to the annotation coordinate parameters and the preset annotation type, and generate an animation video with annotation explanations according to the animation video voice and the modified coordinate reference table.

7. An animation video generation system, wherein, comprising: A first generation module for generating animation video voice according to the script text; An acquisition module for acquiring candidate display objects from the image to be processed; A first determination module for determining video display objects from the candidate display objects according to the script text, the candidate display objects including candidate text information objects and candidate image information objects, including: Obtain the text meaning of the candidate image information object; Based on semantic association, according to the text meaning of the candidate image information object, associate the candidate image information object and the candidate text information object; Perform string matching on the script text and the candidate text information object, and use the candidate text information object that matches the script text and the candidate image information object associated with the candidate text information object that matches the script text as the video display object; A second determination module for determining the object output time according to the video display object and the animation video voice; A second generation module for generating a coordinate reference table according to the preset animation video display method, the video display object and the object output time; A third generation module for generating an animation video according to the coordinate reference table and the animation video voice.

8. A computer device, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein, when the processor executes the computer-readable instructions, it executes the steps of the animation video generation method according to any one of claims 1-6.

9. A computer-readable storage medium storing computer-readable instructions, wherein, when the computer-readable instructions are executed by a processor, the steps of the animation video generation method according to any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Method and system for automatically generating demonstration video, equipment and storage medium

    CN111538851A

  • Method and apparatus for animation

    US20060221084A1