Text animation video generation method and device, electronic equipment and storage medium

By generating multi-level semantic labels and text timestamp sequences, the problems of low generation efficiency and poor visual effects of text animation videos are solved, and the text group content matches the audio beats are achieved, and the special effects display effect of text animation videos is improved.

CN120499476APending Publication Date: 2025-08-15BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510585939.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, generating text animation videos requires manual segmentation, layout and adding dynamic special effects, resulting in low generation efficiency and poor visual effects.

Method used

Generate multi-level semantic labels by obtaining the text semantics of the video content text, generate text timestamp sequences based on the multi-level semantic labels and target audio, and generate text animation special effects based on the text timestamp sequence to achieve matching the text group content and the audio beat.

Benefits of technology

It improves the generation efficiency and visual effects of text animation videos, and realizes the audio-visual effect of audio-visual synchronization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499476A_ABST
    Figure CN120499476A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a text animation video generation method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a video content text, generating a multi-level semantic tag according to the text semantics of the video content text, and enabling the multi-level semantic tag to be used for representing at least two text groups forming the video content text, obtaining semantic information of each text group; and generating a text timestamp sequence according to the multi-level semantic tag and the target audio, generating a text animation special effect corresponding to each text group based on the text timestamp sequence, and generating a text animation video based on each text animation special effect. Therefore, the display opportunity of each text group content in the text animation video is matched with the audio beat of the target audio, audio sticking point display of the text content is realized, the special effect display effect of the text animation video is improved, and meanwhile, the generation efficiency of the text animation video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of video processing technology, and in particular to a method, device, electronic device, and storage medium for generating text animation videos. Background Art

[0002] As a media resource with better visual effects and richer information expression methods, video has better visual expressiveness than static text. Currently, users achieve better information dissemination effects by turning text into dynamic video resources and publishing and delivering them on media platforms.

[0003] In the existing technology, in order to achieve better information display effects, relevant technical solutions further add background music and font animation to the text video generated based on text to generate a text animation video. The text animation video has an audio-visual effect of synchronized audio and video, that is, the so-called "card points" of video content and audio content, so that the text animation video has better audio and video expressiveness.

[0004] However, in the existing technology, in order to generate the above-mentioned text animation video with audio and video synchronized audio and video effects, users need to manually segment, layout and add dynamic special effects to produce the video, resulting in low efficiency and poor visual effects of text animation video generation. Summary of the Invention

[0005] The embodiments of the present disclosure provide a method, device, electronic device, and storage medium for generating text animation videos to overcome the problems of low efficiency and poor visual effects in generating text animation videos.

[0006] In a first aspect, an embodiment of the present disclosure provides a method for generating a text animation video, comprising:

[0007] The video content text is obtained, and multi-level semantic tags are generated based on the text semantics of the video content text, wherein the multi-level semantic tags are used to characterize at least two text groups constituting the video content text and the semantic information of the text groups; a text timestamp sequence is generated based on the multi-level semantic tags and target audio, wherein the text timestamp sequence is used to characterize the display timing corresponding to the text groups; based on the text timestamp sequence, text animation effects corresponding to the text groups are generated, and a text animation video is generated based on the text animation effects.

[0008] In a second aspect, an embodiment of the present disclosure provides a device for generating a text animation video, comprising:

[0009] A first generation module is configured to obtain video content text and generate multi-level semantic tags based on the text semantics of the video content text, wherein the multi-level semantic tags are used to represent at least two text groups constituting the video content text and semantic information of the text groups;

[0010] A second generating module is configured to generate a text timestamp sequence based on the multi-level semantic tags and the target audio, wherein the text timestamp sequence is used to represent a presentation timing corresponding to the text group;

[0011] The special effects module is used to generate text animation special effects corresponding to the text group based on the text timestamp sequence, and generate a text animation video based on the text animation special effects.

[0012] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;

[0013] The memory stores computer-executable instructions;

[0014] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the text animation video generation method described in the first aspect and various possible designs of the first aspect.

[0015] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer execution instructions are stored. When a processor executes the computer execution instructions, the text animation video generation method described in the first aspect and various possible designs of the first aspect is implemented.

[0016] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the text animation video generation method described in the first aspect and various possible designs of the first aspect.

[0017] The text animation video generation method, device, electronic device and storage medium provided in this embodiment obtain video content text and generate multi-level semantic tags based on the text semantics of the video content text. The multi-level semantic tags are used to represent at least two text groups constituting the video content text, as well as the semantic information of the text groups; based on the multi-level semantic tags and target audio, a text timestamp sequence is generated, and the text timestamp sequence is used to represent the display timing corresponding to the text groups; based on the text timestamp sequence, text animation special effects corresponding to the text groups are generated, and a text animation video is generated based on the text animation special effects. By generating multi-level semantic tags based on the text semantics of the video content text, semantic-based sentence segmentation grouping of the video content text is achieved. Then, based on the multi-level semantic tags and the target audio, a text timestamp sequence is generated that matches the audio beat of the target audio and represents the text group timestamps corresponding to the text group. Then, based on the text timestamp sequence, text animation special effects corresponding to each text group are generated. The text animation special effects are then combined to generate a text animation video, so that the display timing of each text group content in the text animation video matches the audio beat of the target audio, realizing the audio card point display of the text content, improving the special effects display effect of the text animation video, and at the same time improving the generation efficiency of the text animation video. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0019] Figure 1 A diagram of an application scenario of the text animation video generation method provided in an embodiment of the present disclosure;

[0020] Figure 2 Schematic diagram of the process of generating text animation video provided by the embodiment of the present disclosure Figure 1 ;

[0021] Figure 3 for Figure 2 A flowchart of a specific implementation method of step S102 in the embodiment shown;

[0022] Figure 4 for Figure 3 A flowchart of a specific implementation method of step S1022 in the embodiment shown;

[0023] Figure 5 A schematic diagram of a text timestamp sequence generation process provided by an embodiment of the present disclosure;

[0024] Figure 6 for Figure 5 A flowchart of a specific implementation method of step S1022-3 in the embodiment shown;

[0025] Figure 7 Schematic diagram of the process of generating text animation video provided by the embodiment of the present disclosure Figure 2 ;

[0026] Figure 8 for Figure 7 A flowchart of a specific implementation method of step S202 in the embodiment shown;

[0027] Figure 9 for Figure 7 A flowchart of a specific implementation method of step S205 in the embodiment shown;

[0028] Figure 10 A schematic diagram of a processing link for generating text animation videos provided by an embodiment of the present disclosure;

[0029] Figure 11 A structural block diagram of a text animation video generating device provided by an embodiment of the present disclosure;

[0030] Figure 12 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure;

[0031] Figure 13 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0033] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0034] The following explains the application scenarios of the embodiments of the present disclosure:

[0035] The text animation video generation method provided by the embodiment of the present disclosure can be applied to applications (APP, Application) with video generation functions, such as video editing applications, AI smart assistant applications, etc. More specifically, it can be applied to application scenarios in which text animation videos are generated based on text. The execution subject of this embodiment can be a terminal device that runs the above-mentioned application with video generation function, or a server that deploys the server corresponding to the above-mentioned application, or other electronic devices that perform similar functions. Among them, when the execution subject is a terminal device, the terminal device executes the method provided by this embodiment by running the above-mentioned application; when the execution subject is a server, the server of the above-mentioned application with video generation function can be partially or completely run on the server, and execute the method provided by this embodiment on the server side, and the terminal device runs the client of the application, and the communication between the server and the terminal device is based on the server-client, so that the terminal device can obtain the execution result of the method provided by this embodiment and display it as needed.

[0036] Among them, in some embodiments, the terminal device or server can implement the text animation video generation method provided by the embodiment of the present disclosure by running various computer executable instructions or computer programs. For example, computer executable instructions can be program-level commands, machine instructions or software instructions. The computer program can be a native program or software module in the operating system; it can be a local application, that is, a program that needs to be installed in the operating system to run, or it can be a small program embedded in any APP, that is, a program that runs based on a browser environment. In summary, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form, and the specific implementation form can be configured as needed. Furthermore, in the process of implementing the text animation video generation method provided by the embodiment of the present disclosure, the terminal device can execute the method by running a computer executable instruction or computer program set locally, or it can execute the method by calling a computer executable instruction or computer program set in an external server. In some embodiments, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud storage, cloud communications, cloud databases, cloud computing, cloud functions, network services, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Among them, cloud services can be interactive processing services for terminal devices to call.

[0037] Figure 1This is an application scenario diagram of the text animation video generation method provided by the embodiment of the present disclosure, referring to Figure 1 As shown in , taking a terminal device as an example, a target application with a video generation function is running in the terminal device. The user triggers the text animation video generation function through the interactive interface of the target application, and the text animation video generation interface is displayed. In one possible implementation, a text input box is configured in the text animation video generation interface. The user generates the corresponding text animation video by entering the video content text into the text input box. In another possible implementation, a pre-generated text document can be loaded to obtain the video content text recorded in the text document, and then the corresponding text animation video is generated.

[0038] The so-called text animation video is an animated video that displays text in the video content text through animation effects such as "flashing" and "flying in and out." For example, as shown in the figure, the content of the video content text includes, for example, "This is a true story..." In the text animation video generated based on the above video content text, the above video content text is dynamically displayed through "flashing". For example, in the first frame, the text "this is" is displayed, in the 10th frame, the text "a" is displayed, and in the 20th frame, the text "true story" is displayed. Between the above videos, there can also be video frames showing the text changing process, thereby achieving a "dynamic" visual effect. At the same time, the above-mentioned text animation video is accompanied by background music while displaying the text content. In order to achieve the audio-visual effect of synchronized audio and video, it is necessary to match the timing of displaying the text content in the text animation video with the timing of the audio beat. For example, when playing the first beat of the background music, the text animation video displays the text "this is", when playing the second beat of the background music, the text animation video displays the text "one", and when playing the third beat of the background music, the text animation video displays the text "true story", thereby achieving the audio-visual effect of synchronized audio and video.

[0039] In the prior art, users are required to manually segment, layout, and add dynamic special effects to produce text animation videos with audio and video synchronization. This results in low efficiency and poor visual effects. The present disclosure provides a method for generating text animation videos to address these issues.

[0040] refer to Figure 2 , Figure 2 Schematic diagram of the process of generating text animation video provided by the embodiment of the present disclosure Figure 1. The method of this embodiment can be applied in a terminal device or a server. In the case where the terminal device executes the method provided by this embodiment, in one possible implementation, the terminal device can implement the text animation video generation method provided by this embodiment by executing the program code deployed locally. In another possible implementation, the server can be used to deploy a functional service based on the text animation video generation method provided by this embodiment, and the terminal device can implement the text animation video generation method provided by this embodiment by accessing the above server and calling the corresponding functional service. For example, the text animation video generation method provided by this embodiment includes:

[0041] Step S101: obtaining video content text, and generating multi-level semantic tags based on the text semantics of the video content text, wherein the multi-level semantic tags are used to represent at least two text groups constituting the video content text, and semantic information of the text groups.

[0042] Step S102: Generate a text timestamp sequence based on the multi-level semantic tags and the target audio. The text timestamp sequence is used to represent the presentation timing corresponding to the text group. The presentation timing corresponding to the text group matches the audio beat of the target audio.

[0043] refer to Figure 1 The application scenario diagram shown in the figure is as follows. In this embodiment, the text animation video generation method provided is introduced with the server as the execution subject. For example, the server is deployed with a server corresponding to an application with video generation function, and the terminal device runs the client of the application. The server and the terminal device obtain the video content text sent by the terminal device based on the server-client communication. The video content text is the initial text used to generate the text animation video. Figure 1As shown, for example, it can be "dialogue text", "speech text", "product promotion text", etc. The specific content of the video content text can be set as needed and is not limited here. Afterwards, the server parses the video content text and divides the video content text according to the text semantics of the video content text to generate multi-level semantic tags, wherein the multi-level semantic tags are used to represent at least two text groups constituting the video content text, as well as the semantic information of each text group. For example, the multi-level semantic tags include [(tag_1, info_1), (tag_2, info_2), ...]. Among them, tag_1 refers to the first text group of the video content text, which includes the characters "this" and "is". More specifically, tag_1 can be a set of the above characters, such as tag_1 = ("this", "is"), or a set of character identifiers representing the above characters, such as the Unicode encoding corresponding to "this" and "is". It can also be a character offset or offset interval of the video content text, for example, tag_1 = (1, 2) represents the first and second characters in the video content text. Info_1 represents the semantic information corresponding to tag_1. For example, it can be a semantic feature vector or a semantic feature vector matrix. For another example, info_1 can also be a part-of-speech type determined based on the semantics of tag_1. For example, tag_1 = 1 represents tag_1 as a verb, while tag_1 = 2 represents tag_1 as a noun, and so on. By combining (tag_1, info_1), we can obtain a text group that constitutes the video content text, that is, a sentence segment of the video content text, and the semantic information corresponding to this text group (the sentence segment). The content and meaning of other multi-level semantic tags such as (tag_2, info_2) are similar and will not be repeated here.

[0044] Furthermore, the process of generating multi-level semantic tags based on the textual semantics of the video content text can be implemented through a pre-trained language processing model. That is, after the video content text is input into the language processing model, the semantic understanding and reasoning capabilities of the language processing model are used to segment the video content text into multiple text groups, thereby generating multi-level semantic tags describing the text groups and the semantic information corresponding to the text groups. Alternatively, the video content text can be segmented according to a preset segmentation logic, for example, based on different parts of speech types such as subject, predicate, and object, thereby generating multi-level semantic tags describing the text groups and the part of speech types corresponding to the text groups. The specific configuration can be as needed and is not limited here.

[0045] Furthermore, after the multi-level semantic tags are generated, it is equivalent to completing the parsing and segmentation process of the video content text. After that, it is necessary to determine the display timing corresponding to each text group (sentence segmentation) based on the multi-level semantic tags and the target audio, so that the display timing of the text group matches the audio beat of the target audio, and achieve the audio-visual effect of "card point" audio and video synchronization. Specifically, based on the multi-level semantic tags and the target audio, a text timestamp sequence is generated to characterize the display timing corresponding to each text group. The text timestamp sequence may include the appearance time corresponding to each text group, and may also include the display time period corresponding to each text group in the text timestamp sequence (i.e., the display start time and the display end time). At the same time, the display timing of the text group described by the text timestamp sequence is matched with the audio beat of the target audio; that is, when the audio beat of the target audio is played, the text content corresponding to the text group is synchronously displayed, thereby achieving the audio-visual synchronization effect of "card point".

[0046] In one possible implementation, Figure 3 As shown, the specific implementation of step S102 includes:

[0047] Step S1021: performing beat recognition on the target audio to generate beat information, where the beat information is used to indicate the beat timestamps of the audio beats constituting the target audio and the beat types of the audio beats;

[0048] Step S1022: Generate a text timestamp sequence based on the beat information and multi-level semantic tags.

[0049] Exemplarily, first, after the server obtains the target audio, it performs beat recognition on the target audio to obtain a beat timestamp indicating the audio beat constituting the target audio, and beat information of the beat type of the audio beat. In one possible implementation, the beat information records the beat timestamp corresponding to each beat, and the beat type corresponding to the beat. The beat type includes, for example, a single beat, a negative beat, etc. More specifically, taking a single beat as an example, it can also be refined to include beat types such as two beats and three beats. The beat type can be pre-set, and its specific implementation method can be determined as needed, which is not limited here. Afterwards, based on the beat information and multi-level semantic labels, the beat timestamp is time-aligned with the text group described by the text timestamp sequence to generate a text timestamp sequence. When the text content corresponding to the text group is displayed based on the text timestamp sequence, the display timing of the text content corresponding to the text group is consistent with the audio beat of the target audio.

[0050] Furthermore, in a possible implementation, as Figure 4 As shown, the specific implementation of step S1022 includes:

[0051] Step S1022 - 1 : Obtain a beat timestamp sequence according to the beat information, where the beat timestamp sequence includes at least two beat timestamps.

[0052] Step S1022 - 2 : Obtain at least two text groups constituting the video content text according to the multi-level semantic tags.

[0053] Step S1022 - 3 : aligning the beat timestamps in the beat timestamp sequence with at least two text groups to obtain a text group timestamp corresponding to each text group.

[0054] Step S1022 - 4 : Generate a text timestamp sequence according to the text group timestamp corresponding to each text group.

[0055] For example, Figure 5 A schematic diagram of a text timestamp sequence generation process provided by an embodiment of the present disclosure is shown below in conjunction with Figure 5 The above process is further introduced as follows: Figure 5 As shown, after obtaining the target audio, the target audio is parsed to generate beat information. This beat information includes beat timestamps such as time points t0_1, t1_2, t0_3, and t1_4, i.e., beat timestamps. t0 represents the first type of beat, t1 represents the second type of beat, and 1, 2, 3, and 4 represent timestamps, respectively. Next, based on these beat timestamps, the text groups indicated by the multi-level semantic tags are aligned, such as text group T1, text group T2, and text group T3. This generates a text timestamp sequence, which records the presentation timing corresponding to each text group. For example, the presentation timing corresponding to text group T1 is t0_1, the presentation timing corresponding to text group T2 is t1_1, the presentation timing corresponding to text group T3 is t0_2, and so on.

[0056] In this embodiment, by performing beat recognition on the target audio, generating beat information, and then generating a text timestamp sequence based on the beat information, the beat alignment of each text group in the text animation video with the audio beat of the target audio is achieved, thereby achieving the audio-visual effect of "stuck-point" audio and video synchronization.

[0057] Furthermore, for each text group, a different number of beat timestamps may be corresponding, thereby achieving the effect of mapping different types of text groups to different beat durations. Specifically, Figure 6 As shown, in a possible implementation, the specific implementation of step S1022-3 includes:

[0058] Step S1022 - 3A: Determine the text display type corresponding to each text group based on the semantic information of each text group. The text display type includes at least an emphasized display type and a non-emphasized display type.

[0059] Step S1022 - 3B: Determine the number of beat timestamps corresponding to each text group according to the text presentation type corresponding to each text group.

[0060] Step S1022 - 3C: Obtain the text group timestamp corresponding to each text group according to the number of beat timestamps corresponding to each text group and the beat timestamp sequence.

[0061] Exemplarily, based on multi-level semantic tags, semantic information of each text group can be obtained. Then, based on the content semantics of the text group represented by the semantic information, the text display type corresponding to each text group is determined, wherein the text display type includes at least an emphasized display type and a non-emphasized display type. In short, it includes at least text groups that need to be emphasized and text groups that do not need to be emphasized. For example, based on the semantic information of the text group, the text group containing keywords is determined to be the text group that needs to be emphasized, and the text display type of such text group is the emphasized display type; wherein the keyword can be a character with specific content or specific pronunciation. Afterwards, according to different text display types, the number of beats corresponding to each text group is determined, wherein the number of beat timestamps is the number of audio beats, and one audio beat corresponds to one beat timestamp in the beat timestamp sequence. For example, for a text group of non-emphasized display type, the number of corresponding beat timestamps is 1, that is, one text group of non-emphasized display type corresponds to the display duration of one audio beat (that is, corresponds to one beat timestamp); and for a text group of emphasized display type, the number of corresponding beat timestamps is 2, that is, one text group of emphasized display type corresponds to the display duration of two audio beats (that is, corresponds to two beat timestamps); thereby achieving the purpose of text groups of different text display types corresponding to different numbers of beat timestamps, that is, achieving the technical effect of different text contents in the video content text corresponding to different display durations and beat combinations. Afterwards, a text timestamp sequence is constructed through the text group timestamps generated by the above steps, and a text animation video is generated based on the text timestamp sequence. This can achieve the technical effect of emphasizing key information such as theme characters and rhyming characters when displaying the text group, thereby improving the information display effect of the text animation video.

[0062] Furthermore, exemplarily, the target audio may be acquired randomly, or may be recommended based on the video content text by an audio recommendation system, or may be manually selected by a user. In the above cases, the playback duration of the target audio may not match the text length of the video content text. For example, if the playback duration of the target audio is shorter and the text length of the video content text is longer, it is impossible to assign a corresponding audio beat to each text group. In this case, it is necessary to adjust the beat timestamp sequence or the text group indicated by the multi-level semantic label to solve the problem of length mismatch between the audio content and the video content (i.e., text content).

[0063] In a possible implementation, before executing step S1022-3, the process further includes:

[0064] Step S1022 - 0 : Process the beat timestamp sequence or the text group to obtain a beat timestamp sequence and at least two text groups with matching lengths.

[0065] For example, in one possible implementation, when the number of beat timestamps in the beat timestamp sequence is less than the number of text groups represented by the multi-level semantic tags (i.e., the length-matched beat timestamp sequence and at least two text groups do not match), the beat timestamp sequence can be extended. Specifically, after adding the audio duration of the target audio to several beat timestamps at the head of the beat timestamp sequence, an extended beat timestamp is generated and sequentially arranged to the end of the beat timestamp sequence, thereby delaying the beat timestamp sequence. This implementation step is equivalent to replaying the target audio after the target audio is played, thereby delaying the playback duration of the target audio to match the length of the video content text. In another possible implementation, the text groups can be processed, for example, by merging the short text groups of two vectors into a long text group, thereby reducing the number of text groups, wherein a short text group refers to a text group with a smaller number of characters, and a long text group refers to a text group with a larger number of characters. It is understood that in another possible implementation, the above two solutions can be performed simultaneously, that is, the beat timestamp sequence is extended while the text groups are merged, thereby obtaining a beat timestamp sequence of matching lengths and at least two text groups. The above solutions for processing the beat timestamp sequence or the text groups can be implemented as needed and are not further described here.

[0066] Step S103: Generate text animation effects corresponding to the text group based on the text timestamp sequence, and generate a text animation video based on the text animation effects.

[0067] Exemplarily, after obtaining a text timestamp sequence, corresponding text animation special effects are generated based on the display timing of each text group described by the text timestamp sequence, wherein the difference between the display timing (display time) of two adjacent text groups, that is, the special effect duration of the text animation special effect corresponding to the previous text group in the two adjacent text groups. After generating corresponding text animation special effects for each text group according to this principle, they are combined to obtain a text animation video.

[0068] In a possible implementation, the specific implementation of step S103 includes:

[0069] Step S1031: reconstructing the text timestamp sequence and the text groups corresponding to the text timestamp sequence according to the multi-level semantic tags and the target audio to obtain at least two reconstructed text groups and corresponding reconstructed text timestamp sequences, wherein at least one reconstructed text group includes at least two text groups.

[0070] Step S1032: generating text animation effects corresponding to each reconstructed text group based on the reconstructed text group and the corresponding reconstructed text timestamp sequence.

[0071] For example, after obtaining the text timestamp sequence, in one possible implementation, the text animation special effects corresponding to each text group can be directly generated based on the text timestamp sequence, that is, the text timestamp corresponding to each text group. In another possible implementation, each text group can be further reorganized to make it more compatible with the target audio and achieve better audio and video synchronization effects. Specifically, first, based on the semantic information of each text group represented by the multi-level semantic tags and the beat information of the target audio, the text timestamp sequence and the text groups corresponding to the text timestamp sequence are reconstructed. Specifically, adjacent text groups with similar or identical semantic information can be reconstructed into a reconstructed text group, and corresponding reconstructed text timestamps are generated for the reconstructed text group. After the reconstruction process for the text group is completed, a reconstructed text timestamp sequence is obtained; or, a text group is split into at least two reconstructed text groups, each of which contains one or more characters, and corresponding reconstructed text timestamps are generated for the reconstructed text group. After the reconstruction process for the text group is completed, a reconstructed text timestamp sequence is obtained. Afterwards, based on the reconstructed text group and the corresponding reconstructed text timestamp sequence, the text animation effects corresponding to each reconstructed text group are generated. This process is similar to the process of generating the corresponding text animation effects based on the text group, which has been introduced in detail before and will not be repeated here.

[0072] In the steps of this embodiment, the text timestamp sequence and the text group corresponding to the text timestamp sequence are reconstructed according to the multi-level semantic tags and the target audio, thereby further improving the adaptability of the reconstructed text group and the target audio, and improving the "card point" effect and the visual expressiveness of the finally generated text animation special effects.

[0073] In this embodiment, by obtaining video content text and generating multi-level semantic tags based on the text semantics of the video content text, the multi-level semantic tags are used to characterize at least two text groups constituting the video content text, as well as the semantic information of each text group; based on the multi-level semantic tags and the target audio, a text timestamp sequence is generated, the text timestamp sequence is used to characterize the display timing corresponding to each text group; based on the text timestamp sequence, text animation special effects corresponding to each text group are generated, and a text animation video is generated based on each text animation special effect. By generating multi-level semantic tags based on the text semantics of the video content text, the video content text is segmented and grouped based on semantics; then, based on the multi-level semantic tags and the target audio, a text timestamp sequence is generated that matches the audio beat of the target audio and characterizes the text group timestamps corresponding to the text group; then, based on the text timestamp sequence, text animation special effects corresponding to each text group are generated; and then, the text animation special effects are combined to generate a text animation video, so that the display timing of each text group content in the text animation video matches the audio beat of the target audio, achieving audio card display of the text content, improving the special effect display effect of the text animation video, and improving the generation efficiency of the text animation video.

[0074] refer to Figure 7 , Figure 7 Schematic diagram of the process of generating text animation video provided by the embodiment of the present disclosure Figure 2 In this embodiment Figure 2 On the basis of the embodiment shown, step S103 is further refined and a step of obtaining target audio is added. The text animation video generation method includes:

[0075] Step S201: Obtain video content text, and generate multi-level semantic tags based on the text semantics of the video content text. The multi-level semantic tags are used to represent at least two text groups constituting the video content text, and semantic information of each text group.

[0076] Step S202: Generate retrieval features based on multi-level semantic tags.

[0077] Step S203: calling the audio recommendation service to process the retrieval features and obtain target audio that matches the multi-level semantic tags.

[0078] For example, in this embodiment, after obtaining multi-level semantic tags, retrieval features are first constructed based on the multi-level semantic tags. In one possible implementation method, secondary extraction can be performed based on the semantic information of each text group represented by the multi-level semantic tags to obtain special effects representing the theme and summary of the video content text, that is, retrieval features. Alternatively, the semantic information of each text group represented by the multi-level semantic tags can be statistically analyzed, and those with the same semantic information (or similarity greater than a threshold) are classified into the same category. Afterwards, the semantic information within a category with the most semantic information is determined as the retrieval feature. Among them, the retrieval feature can be a category identifier, feature vector, etc. that represents the theme or summary of the video content text. Afterwards, the audio recommendation service is called, and a search is performed based on the retrieval feature to obtain recommended audio that matches the retrieval feature, that is, the target audio that matches the multi-level semantic tags.

[0079] Furthermore, in a possible implementation, as Figure 8 As shown, the specific implementation of step S202 includes:

[0080] Step S2021: Obtain content keywords corresponding to the video content text based on the multi-level semantic tags, and / or obtain the average text length of each text group.

[0081] Step S2022: Generate retrieval features based on content keywords and / or average text length.

[0082] For example, in one possible implementation, content keywords corresponding to the video content text are first obtained based on multi-level semantic tags. For example, after reasoning and summarizing the semantic information of each text group, a content keyword representing the content theme of the video content text is generated. Content keywords may be, for example, "product introduction," "novel story," "life record," etc. This step can be implemented by calling a language model. Subsequently, retrieval features are generated based on the content keyword for subsequent target audio retrieval. Through the steps of this embodiment, the retrieved target audio can be matched with the content keyword corresponding to the video content text, thereby improving the matching degree between the target audio and the video content of the final generated text animation video.

[0083] In another possible implementation, based on multi-level semantic tags, the average text length of the text group is counted, for example, each text group consists of several characters, wherein the shorter the text length, the higher the corresponding text display frequency should be, and accordingly, the faster the audio beat of the target audio matching the above text group. Afterwards, based on the average text length, a retrieval feature is generated to perform subsequent target audio retrieval. Through the steps of this embodiment, the audio beat of the retrieved target audio can be matched with the text length of the text group, thereby improving the matching degree between the target audio and the group display frequency of the text group in the finally generated text animation video, and improving the effect of audio and video synchronization.

[0084] Step S204: Generate a text timestamp sequence based on the multi-level semantic tags and the target audio. The text timestamp sequence is used to represent the presentation timing corresponding to each text group. The presentation timing corresponding to the text group matches the audio beat of the target audio.

[0085] Step S205: obtaining stylized information, where the stylized information is used to characterize the special effects style of the text animation special effects.

[0086] Step S206: generating text animation effects corresponding to each text group based on the stylized information and the text timestamp sequence, wherein the text animation effects have a target effect style that matches the stylized information.

[0087] For example, after determining the target audio, a text timestamp sequence can be constructed based on the multi-level semantic tags and the target audio. The construction method of the text timestamp sequence is described in Figure 2 The embodiment shown has been introduced in detail and will not be repeated here. Afterwards, before generating the text animation special effects corresponding to the text group, stylized information is obtained. The stylized information is information used to characterize the special effects style of the text animation special effects. Since the text animation video only contains content display for text, the stylized information is also information that characterizes the style of the text special effects. Specifically, the stylized information may include font outline information, font color information, font action information and other related information for representing font special effects. Examples will not be given here one by one. Afterwards, based on the stylized information and the text timestamp sequence, the characters corresponding to each text group are processed to generate the corresponding text animation special effects. The text animation special effects corresponding to each text group have a target special effects style that matches the stylized information.

[0088] Among them, in one possible implementation, Figure 9 As shown, the specific implementation of step S205 includes:

[0089] Step S2051: Obtaining content keywords corresponding to the video content text based on the multi-level semantic tags;

[0090] Step S2052: Obtain stylized information based on content keywords.

[0091] For example, in one possible implementation, stylized information can be generated based on user operations or obtained based on preset configuration information, that is, the stylized information is pre-generated fixed information. In another possible implementation, the stylized information can also be dynamically generated based on the video content text. That is, when the content of the video content text is different, different stylized information will be obtained, thereby generating text animation videos with different target special effects styles. Specifically, first, based on multi-level semantic tags, the content keywords corresponding to the video content text are obtained. For example, after reasoning and summarizing the semantic information of each text group, a content keyword representing the content theme of the video content text is generated. The content keyword can be, for example, "product introduction", "novel story", "life record", etc. This step can be achieved by calling a language model. Afterwards, based on the content keyword and the preset mapping relationship, the stylized information that matches the content keyword is determined, that is, the special effects style of the text animation special effect. For example, when the content keyword is "product introduction", the corresponding stylized information Info_1 represents the business-style text animation effects; when the content keyword is "life record", the corresponding stylized information Info_2 represents the cartoon-style text animation effects.

[0092] In this embodiment, the content keywords corresponding to the video content text are determined through multi-level semantic tags, and then the stylized information matching the content keywords is obtained based on the content keywords, so that the text animation effects corresponding to the text group match the content of the video content text, and the final generated text animation video has a better visual performance effect.

[0093] Step S207: Generate a text animation video based on the text animation special effects and the target audio.

[0094] For example, after obtaining the text animation effects corresponding to each text group, the text animation effects corresponding to each text group are packaged, and corresponding image materials are added to each text animation effect and rendered. Then, the audio channel data corresponding to the target audio is merged to generate a text animation video.

[0095] In a possible implementation, this embodiment further includes:

[0096] Step S208: Generate reading timbre information according to the multi-level semantic tags, and generate text speech corresponding to the video content text according to the reading timbre information and the video content text.

[0097] Accordingly, in the case of including the above step S206, the specific implementation of step S207 includes:

[0098] Step S207A: Generate a text animation video based on the text animation special effects, text voice and target audio.

[0099] Exemplarily, on the other hand, based on the generation of text animation special effects, in this embodiment, it is also possible to further generate reading timbre information based on multi-level semantic tags, wherein the reading timbre information represents the voice timbre and / or reading frequency of the human voice, such as children's timbre, girl's timbre, fast reading, slow reading, etc. Specifically, for example, the corresponding content keywords are determined based on the multi-level semantic tags, and then the corresponding reading timbre information is obtained based on the content keywords. Afterwards, based on the reading timbre information, the video content text is converted into text to speech (Text To Speech, TTS), and the text speech corresponding to the video content text, that is, the reading language, can be generated. The specific process of converting text to speech based on TTS technology is an existing technology and will not be repeated here. Afterwards, the text animation special effects, text speech and target audio are rendered and merged to obtain a text animation video with human voice reading speech.

[0100] In this embodiment, multi-level semantic tags are further used to generate reading timbre information, and corresponding text speech is generated based on the reading timbre information. Then, the text speech is combined with the text speech to generate a text animation video with a human voice reading speech. Because the text speech is determined based on the multi-level semantic tags generated in the previous step, the reading rhythm and human voice timbre of the text speech match the multi-level semantic tags, that is, the semantic information of the text groups represented by the multi-level semantic tags, as well as the grouping method of each text group. This improves the compatibility of the text speech with the video content text, and enhances the audio and video synchronization and audio and video compatibility of the final generated text animation video.

[0101] Figure 10 A schematic diagram of a processing link for generating text animation video provided by an embodiment of the present disclosure is shown below. Figure 10 The above process is further introduced as follows: Figure 10As shown, illustratively, the server first obtains the video content text, which can be user-entered or intelligently generated. It then performs semantic understanding on the video content text and, based on the results of semantic understanding, roughly groups the video content text to generate multi-level semantic tags. Next, based on these multi-level semantic tags, it invokes an audio recommendation service to obtain target audio, then invokes an audio card service to obtain the target audio's tempo information. Finally, combining this tempo information with the multi-level semantic tags, it performs text card processing to obtain a text timestamp sequence. Then, based on the text timestamp sequence, the audio information corresponding to the target audio and the multi-level semantic labels, the text group and the corresponding text timestamp sequence are reconstructed, that is, the text is regrouped, to generate at least two reconstructed text groups and corresponding reconstructed text timestamp sequences that are better adapted to the target audio. Then, through user selection or intelligent style recommendation, matching stylized information is obtained, and based on the stylized information, text animation effects are generated. Then, by accessing the material resource library, the corresponding texture resources are obtained, the text animation effects are packaged and rendered, and combined with the target audio to generate the final text animation video.

[0102] In this embodiment, the implementation of step S201 and step S204 is the same as that of the present disclosure. Figure 2 The implementation methods of step S101 and step S102 in the illustrated embodiment are the same and will not be described in detail here.

[0103] Corresponding to the text animation video generation method of the above embodiment, Figure 11 This is a block diagram of the structure of the text animation video generation device provided in an embodiment of the present disclosure. The methods described in the above embodiments can be executed by this text animation video generation device, which can be implemented using software and / or hardware and integrated into electronic devices with certain data processing capabilities. These electronic devices may include, but are not limited to, mobile terminals with big data processing capabilities, as well as fixed terminals with big data processing capabilities, such as desktop computers and supercomputers.

[0104] For ease of explanation, only the parts related to the embodiments of the present disclosure are shown. Figure 11 , the text animation video generating device 3 includes:

[0105] A first generating module 31 is configured to obtain video content text and generate multi-level semantic tags based on the textual semantics of the video content text. The multi-level semantic tags are used to represent at least two text groups constituting the video content text and semantic information of the text groups.

[0106] A second generating module 32 is configured to generate a text timestamp sequence based on the multi-level semantic tags and the target audio, wherein the text timestamp sequence is used to represent the presentation timing corresponding to the text group;

[0107] The special effects module 33 is configured to generate text animation special effects corresponding to the text group based on the text timestamp sequence, and generate a text animation video based on the text animation special effects.

[0108] According to one or more embodiments of the present disclosure, the second generation module 32 is specifically used to: perform beat recognition on the target audio and generate beat information, where the beat information is used to indicate the beat timestamps of the audio beats constituting the target audio and the beat types of the audio beats; and generate a text timestamp sequence based on the beat information and multi-level semantic labels.

[0109] According to one or more embodiments of the present disclosure, when the second generation module 32 generates a text timestamp sequence based on beat information and multi-level semantic labels, it is specifically used to: obtain a beat timestamp sequence based on the beat information, wherein the beat timestamp sequence includes at least two beat timestamps; obtain at least two text groups constituting the video content text based on the multi-level semantic labels; obtain text group timestamps corresponding to the text groups by aligning the beat timestamps in the beat timestamp sequence and at least two text groups; and generate a text timestamp sequence based on the text group timestamps corresponding to the text groups.

[0110] According to one or more embodiments of the present disclosure, when the second generation module 32 obtains the text group timestamp corresponding to the text group by aligning the beat timestamps and at least two text groups in the beat timestamp sequence, it is specifically used to: determine the text display type corresponding to the text group based on the semantic information of the text group, the text display type at least including the emphasized display type and the non-emphasized display type; determine the number of beat timestamps corresponding to the text group based on the text display type corresponding to the text group; obtain the text group timestamp corresponding to the text group based on the number of beat timestamps corresponding to the text group and the beat timestamp sequence.

[0111] According to one or more embodiments of the present disclosure, before obtaining the text group timestamps corresponding to the text groups by aligning the beat timestamps and at least two text groups in the beat timestamp sequence, the second generation module 32 is further used to: process the beat timestamp sequence or text group to obtain a beat timestamp sequence and at least two text groups with matching lengths.

[0112] According to one or more embodiments of the present disclosure, the second generation module 32 is further configured to: generate retrieval features based on multi-level semantic tags; and call an audio recommendation service to process the retrieval features to obtain target audio that matches the multi-level semantic tags.

[0113] According to one or more embodiments of the present disclosure, when the second generation module 32 generates retrieval features based on multi-level semantic tags, it is specifically used to: obtain content keywords corresponding to the video content text based on the multi-level semantic tags, and / or obtain the average text length of the text group; generate retrieval features based on the content keywords and / or the average text length.

[0114] According to one or more embodiments of the present disclosure, when the special effects module 33 generates text animation special effects corresponding to a text group based on a text timestamp sequence, it is specifically used to: obtain stylized information, where the stylized information is used to characterize the special effects style of the text animation special effects; and generate text animation special effects corresponding to the text group based on the stylized information and the text timestamp sequence, wherein the text animation special effects have a target special effects style that matches the stylized information.

[0115] According to one or more embodiments of the present disclosure, when acquiring stylized information, the special effects module 33 is specifically configured to: obtain content keywords corresponding to the video content text according to multi-level semantic tags; and obtain stylized information according to the content keywords.

[0116] According to one or more embodiments of the present disclosure, when the special effects module 33 generates text animation special effects corresponding to a text group based on a text timestamp sequence, it is specifically used to: reconstruct the text timestamp sequence and the text group corresponding to the text timestamp sequence according to multi-level semantic tags and target audio, and obtain at least two reconstructed text groups and corresponding reconstructed text timestamp sequences, at least one reconstructed text group includes at least two text groups; based on the reconstructed text group and the corresponding reconstructed text timestamp sequence, generate text animation special effects corresponding to the reconstructed text group.

[0117] According to one or more embodiments of the present disclosure, the special effects module 33 is further used to: generate reading timbre information based on multi-level semantic tags, the reading timbre information representing the human voice timbre and / or reading frequency; generate text speech corresponding to the video content text based on the reading timbre information and the video content text; when the special effects module 33 generates a text animation video based on the text animation special effects, it is specifically used to: generate a text animation video based on the text animation special effects and text speech.

[0118] The first generation module 31, the second generation module 32 and the special effect module 33 are connected in sequence. The text animation video generation device 3 provided in this embodiment can implement the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, which will not be repeated in this embodiment.

[0119] Figure 12 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure is shown in FIG. Figure 12 As shown, the electronic device 4 includes:

[0120] A processor 41, and a memory 42 communicatively connected to the processor 41;

[0121] Memory 42 stores computer-executable instructions;

[0122] The processor 41 executes the computer execution instructions stored in the memory 42 to implement the following Figure 2-Figure 10 The text animation video generation method in the illustrated embodiment.

[0123] Optionally, the processor 41 and the memory 42 are connected via a bus 43 .

[0124] For related instructions, please refer to Figure 2-Figure 10 The relevant descriptions and effects corresponding to the steps in the corresponding embodiments can be understood, and no further details are given here.

[0125] The present invention provides a computer-readable storage medium that stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the computer-executable instructions are used to implement the present invention. Figure 2-Figure 10 The text animation video generation method provided in any one of the corresponding embodiments.

[0126] The present invention provides a computer program product, including a computer program, which implements the present invention when executed by a processor. Figure 2-Figure 10 The text animation video generation method provided in any one of the corresponding embodiments.

[0127] In order to implement the above embodiment, the embodiment of the present disclosure further provides an electronic device.

[0128] refer to Figure 13 , which shows a schematic structural diagram of an electronic device 900 suitable for implementing an embodiment of the present disclosure. The electronic device 900 may be a terminal device or a server. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 13 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0129] like Figure 13As shown, the electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the electronic device 900 are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0130] Typically, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device 900 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 13 The electronic device 900 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0131] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0132] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0133] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0134] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.

[0135] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0136] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0137] The units or modules involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit or module does not, in some cases, limit the unit itself.

[0138] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0139] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0140] In a first aspect, according to one or more embodiments of the present disclosure, a method for generating a text animation video is provided, comprising:

[0141] The video content text is obtained, and multi-level semantic tags are generated based on the text semantics of the video content text, wherein the multi-level semantic tags are used to characterize at least two text groups constituting the video content text, and semantic information of each of the text groups; a text timestamp sequence is generated based on the multi-level semantic tags and target audio, wherein the text timestamp sequence is used to characterize the display timing corresponding to each of the text groups, and the display timing corresponding to the text groups matches the audio beat of the target audio; based on the text timestamp sequence, text animation effects corresponding to each of the text groups are generated, and a text animation video is generated based on each of the text animation effects.

[0142] According to one or more embodiments of the present disclosure, generating a text timestamp sequence based on the multi-level semantic tags and the target audio includes: performing beat recognition on the target audio to generate beat information, wherein the beat information is used to indicate the beat timestamps of the audio beats constituting the target audio, and the beat types of the audio beats; generating the text timestamp sequence based on the beat information and the multi-level semantic tags.

[0143] According to one or more embodiments of the present disclosure, generating the text timestamp sequence based on the beat information and the multi-level semantic labels includes: obtaining a beat timestamp sequence based on the beat information, wherein the beat timestamp sequence includes at least two beat timestamps; obtaining at least two text groups constituting the video content text based on the multi-level semantic labels; obtaining text group timestamps corresponding to each of the text groups by aligning the beat timestamps in the beat timestamp sequence and the at least two text groups; and generating the text timestamp sequence based on the text group timestamps corresponding to each of the text groups.

[0144] According to one or more embodiments of the present disclosure, the method of aligning the beat timestamps in the beat timestamp sequence and the at least two text groups to obtain the text group timestamps corresponding to each text group includes: determining the text display type corresponding to each text group based on the semantic information of each text group, the text display type including at least an emphasized display type and a non-emphasized display type; determining the number of beat timestamps corresponding to each text group based on the text display type corresponding to each text group; and aligning the corresponding number of beat timestamps in the beat timestamp sequence based on the number of beat timestamps corresponding to each text group to obtain the text group timestamps corresponding to each text group.

[0145] According to one or more embodiments of the present disclosure, before aligning the beat timestamps in the beat timestamp sequence and the at least two text groups to obtain the text group timestamps corresponding to each text group, the method further includes: processing the beat timestamp sequence or the text group to obtain a beat timestamp sequence and the at least two text groups with matching lengths.

[0146] According to one or more embodiments of the present disclosure, the method further includes: generating retrieval features based on the multi-level semantic tags; and calling an audio recommendation service to process the retrieval features to obtain target audio that matches the multi-level semantic tags.

[0147] According to one or more embodiments of the present disclosure, generating retrieval features based on the multi-level semantic tags includes: obtaining content keywords corresponding to the video content text based on the multi-level semantic tags, and / or obtaining the average text length of each of the text groups; generating the retrieval features based on the content keywords and / or the average text length.

[0148] According to one or more embodiments of the present disclosure, generating text animation special effects corresponding to each of the text groups based on the text timestamp sequence includes: obtaining stylization information, wherein the stylization information is used to characterize the special effect style of the text animation special effect; generating text animation special effects corresponding to each of the text groups based on the stylization information and the text timestamp sequence, wherein the text animation special effect has a target special effect style that matches the stylization information.

[0149] According to one or more embodiments of the present disclosure, the acquiring of stylized information includes: obtaining content keywords corresponding to the video content text according to the multi-level semantic tags; and obtaining the stylized information according to the content keywords.

[0150] According to one or more embodiments of the present disclosure, based on the text timestamp sequence, text animation special effects corresponding to each of the text groups are generated, including: reconstructing the text timestamp sequence and the text groups corresponding to the text timestamp sequence according to the multi-level semantic tags and the target audio, to obtain at least two reconstructed text groups and corresponding reconstructed text timestamp sequences, at least one of the reconstructed text groups including at least two of the text groups; based on the reconstructed text groups and the corresponding reconstructed text timestamp sequences, generating text animation special effects corresponding to each of the reconstructed text groups.

[0151] According to one or more embodiments of the present disclosure, the method further includes: generating reading timbre information based on the multi-level semantic tags, the reading timbre information representing the human voice timbre and / or reading frequency; generating text speech corresponding to the video content text based on the reading timbre information and the video content text; generating a text animation video based on each of the text animation special effects, including: generating a text animation video based on each of the text animation special effects and the text speech.

[0152] In a second aspect, according to one or more embodiments of the present disclosure, a text animation video generating apparatus is provided, comprising:

[0153] A first generation module is configured to obtain video content text and generate multi-level semantic tags based on the textual semantics of the video content text, wherein the multi-level semantic tags are used to represent at least two text groups constituting the video content text and semantic information of each of the text groups;

[0154] A second generating module is configured to generate a text timestamp sequence based on the multi-level semantic tags and the target audio, wherein the text timestamp sequence is used to represent a presentation timing corresponding to each of the text groups;

[0155] The special effects module is used to generate text animation special effects corresponding to each of the text groups based on the text timestamp sequence, and generate text animation videos based on each of the text animation special effects.

[0156] According to one or more embodiments of the present disclosure, the second generation module is specifically used to: perform beat recognition on the target audio to generate beat information, where the beat information is used to indicate the beat timestamps of the audio beats constituting the target audio and the beat types of the audio beats; and generate the text timestamp sequence based on the beat information and the multi-level semantic labels.

[0157] According to one or more embodiments of the present disclosure, when the second generation module generates the text timestamp sequence based on the beat information and the multi-level semantic labels, it is specifically used to: obtain a beat timestamp sequence based on the beat information, wherein the beat timestamp sequence includes at least two beat timestamps; obtain at least two text groups constituting the video content text based on the multi-level semantic labels; obtain text group timestamps corresponding to each of the text groups by aligning the beat timestamps in the beat timestamp sequence and the at least two text groups; and generate the text timestamp sequence based on the text group timestamps corresponding to each of the text groups.

[0158] According to one or more embodiments of the present disclosure, when the second generation module obtains the text group timestamp corresponding to each text group by aligning the beat timestamps in the beat timestamp sequence and the at least two text groups, the second generation module is specifically used to: determine the text display type corresponding to each text group according to the semantic information of each text group, and the text display type at least includes an emphasized display type and a non-emphasized display type; determine the number of beat timestamps corresponding to each text group according to the text display type corresponding to each text group; align the corresponding number of beat timestamps in the beat timestamp sequence according to the number of beat timestamps corresponding to each text group to obtain the text group timestamp corresponding to each text group.

[0159] According to one or more embodiments of the present disclosure, before obtaining the text group timestamps corresponding to each text group by aligning the beat timestamps in the beat timestamp sequence and the at least two text groups, the second generation module is further used to: process the beat timestamp sequence or the text group to obtain a beat timestamp sequence and the at least two text groups with matching lengths.

[0160] According to one or more embodiments of the present disclosure, the second generation module is further used to: generate retrieval features based on the multi-level semantic tags; call the audio recommendation service to process the retrieval features to obtain target audio that matches the multi-level semantic tags.

[0161] According to one or more embodiments of the present disclosure, when the second generation module generates retrieval features based on the multi-level semantic tags, it is specifically used to: obtain content keywords corresponding to the video content text based on the multi-level semantic tags, and / or obtain the average text length of each of the text groups; generate the retrieval features based on the content keywords and / or the average text length.

[0162] According to one or more embodiments of the present disclosure, when the special effects module generates text animation special effects corresponding to each of the text groups based on the text timestamp sequence, it is specifically used to: obtain stylization information, where the stylization information is used to characterize the special effects style of the text animation special effects; and generate text animation special effects corresponding to each of the text groups based on the stylization information and the text timestamp sequence, wherein the text animation special effects have a target special effects style that matches the stylization information.

[0163] According to one or more embodiments of the present disclosure, when acquiring stylized information, the special effects module is specifically configured to: obtain content keywords corresponding to the video content text according to the multi-level semantic tags; and obtain the stylized information according to the content keywords.

[0164] According to one or more embodiments of the present disclosure, when the special effects module generates text animation special effects corresponding to each of the text groups based on the text timestamp sequence, it is specifically used to: reconstruct the text timestamp sequence and the text groups corresponding to the text timestamp sequence according to the multi-level semantic tags and the target audio, and obtain at least two reconstructed text groups and corresponding reconstructed text timestamp sequences, at least one of the reconstructed text groups includes at least two of the text groups; based on the reconstructed text groups and the corresponding reconstructed text timestamp sequences, generate text animation special effects corresponding to each of the reconstructed text groups.

[0165] According to one or more embodiments of the present disclosure, the special effects module is further used to: generate reading timbre information based on the multi-level semantic tags, wherein the reading timbre information represents the human voice timbre and / or reading frequency; generate text speech corresponding to the video content text based on the reading timbre information and the video content text; when the special effects module 33 generates a text animation video based on each of the text animation special effects, it is specifically used to: generate a text animation video based on each of the text animation special effects and the text speech.

[0166] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, comprising: at least one processor and a memory;

[0167] The memory stores computer-executable instructions;

[0168] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the text animation video generation method described in the first aspect and various possible designs of the first aspect.

[0169] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the text animation video generation method described in the first aspect and various possible designs of the first aspect is implemented.

[0170] In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the text animation video generation method as described in the first aspect and various possible designs of the first aspect.

[0171] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0172] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0173] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for generating text animation video, characterized in that: include: Acquire video content text, and generate a multi-level semantic tag based on the text semantics of the video content text, wherein the multi-level semantic tag is used to represent at least two text groups constituting the video content text and semantic information of the text groups; Generating a text timestamp sequence based on the multi-level semantic tags and the target audio, wherein the text timestamp sequence is used to represent a presentation timing corresponding to the text group, and the presentation timing corresponding to the text group matches the audio beat of the target audio; Based on the text timestamp sequence, a text animation special effect corresponding to the text group is generated, and a text animation video is generated based on the text animation special effect.

2. The method according to claim 1, characterized in that Generating a text timestamp sequence according to the multi-level semantic tags and the target audio includes: Performing beat recognition on the target audio to generate beat information, where the beat information is used to indicate a beat timestamp of an audio beat constituting the target audio and a beat type of the audio beat; The text timestamp sequence is generated according to the beat information and the multi-level semantic labels.

3. The method according to claim 2, characterized in that Generating the text timestamp sequence according to the beat information and the multi-level semantic labels includes: Obtaining a beat timestamp sequence according to the beat information, wherein the beat timestamp sequence includes at least two of the beat timestamps; Obtaining at least two text groups constituting the video content text according to the multi-level semantic tags; Obtaining text group timestamps corresponding to the text groups by aligning the beat timestamps in the beat timestamp sequence with the at least two text groups; The text timestamp sequence is generated according to the text group timestamps corresponding to the text groups.

4. The method according to claim 3, characterized in that The step of aligning the beat timestamps in the beat timestamp sequence with the at least two text groups to obtain text group timestamps corresponding to the text groups includes: determining, based on semantic information of the text group, a text presentation type corresponding to the text group, the text presentation type including at least an emphasized presentation type and a non-emphasized presentation type; determining the number of beat timestamps corresponding to the text group according to the text presentation type corresponding to the text group; The text group timestamp corresponding to the text group is obtained according to the number of beat timestamps corresponding to the text group and the beat timestamp sequence.

5. The method according to claim 3, characterized in that Before obtaining the text group timestamps corresponding to the text groups by aligning the beat timestamps in the beat timestamp sequence with the at least two text groups, the method further includes: The beat timestamp sequence or the text group is processed to obtain a beat timestamp sequence and the at least two text groups with matching lengths.

6. The method according to claim 1, characterized in that The method further comprises: generating retrieval features according to the multi-level semantic labels; An audio recommendation service is called to process the retrieval features to obtain target audio that matches the multi-level semantic tags.

7. The method according to claim 6, characterized in that Generating retrieval features according to the multi-level semantic tags includes: Obtaining content keywords corresponding to the video content text according to the multi-level semantic tags, and / or obtaining an average text length of the text group; The search feature is generated according to the content keywords and / or the average text length.

8. The method according to claim 1, characterized in that Generating a text animation effect corresponding to the text group based on the text timestamp sequence includes: Acquire stylized information, where the stylized information is used to characterize the special effects style of the text animation special effects; Based on the stylization information and the text timestamp sequence, a text animation special effect corresponding to the text group is generated, wherein the text animation special effect has a target special effect style that matches the stylization information.

9. The method according to claim 8, characterized in that The obtaining of stylized information includes: Obtaining content keywords corresponding to the video content text according to the multi-level semantic tags; The stylized information is obtained according to the content keywords.

10. The method according to claim 1, characterized in that Generating a text animation effect corresponding to the text group based on the text timestamp sequence, including: reconstructing the text timestamp sequence and the text group corresponding to the text timestamp sequence according to the multi-level semantic tags and the target audio to obtain at least two reconstructed text groups and corresponding reconstructed text timestamp sequences, at least one of the reconstructed text groups including at least two of the text groups; Based on the reconstructed text group and the corresponding reconstructed text timestamp sequence, a text animation effect corresponding to the reconstructed text group is generated.

11. The method according to claim 1, wherein The method further comprises: generating reading timbre information according to the multi-level semantic tags, wherein the reading timbre information represents the human voice timbre and / or reading frequency; Generating a text speech corresponding to the video content text according to the reading timbre information and the video content text; Generating a text animation video based on the text animation special effects includes: A text animation video is generated based on the text animation special effects and the text voice.

12. A text animation video generating device, characterized in that: include: A first generation module is configured to obtain video content text and generate multi-level semantic tags based on the text semantics of the video content text, wherein the multi-level semantic tags are used to represent at least two text groups constituting the video content text and semantic information of the text groups; A second generating module is configured to generate a text timestamp sequence based on the multi-level semantic tags and the target audio, wherein the text timestamp sequence is used to represent a presentation timing corresponding to the text group; The special effects module is used to generate text animation special effects corresponding to the text group based on the text timestamp sequence, and generate a text animation video based on the text animation special effects.

13. An electronic device, characterized in that: include: processor and memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the text animation video generation method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When the processor executes the computer-executable instructions, the text animation video generation method according to any one of claims 1 to 11 is implemented.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the text animation video generating method according to any one of claims 1 to 11 is implemented.