Animation generation method and device, equipment and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT TECH SHANGHAI
- Filing Date
- 2024-11-20
- Publication Date
- 2026-05-22
AI Technical Summary
Existing deep learning models suffer from high randomness and difficulty in human control when generating animations corresponding to audio, resulting in low accuracy and quality of the generated animations, especially in the expression of head movements, which are prone to conflict.
By extracting semantic and accent keywords from the audio, target animation clips are matched from a pre-defined animation library. Taking into account semantic and accent features, the animation is generated in the target matching order to avoid conflicts between animation clips and ensure smooth splicing.
It improves the accuracy and quality of animation, making the generated animation fully reflect the audio, reliable in quality, and smoother in splicing between animation clips.
Smart Images

Figure CN122072986A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular to an animation generation method, apparatus, device, and storage medium. Background Technology
[0002] With the development of computer vision technology, deep learning models are used to generate animations that match audio in order to improve efficiency.
[0003] In related technologies, when generating animations based on deep learning models, audio is input into the deep learning model, and the model outputs an animation corresponding to the semantic information of the audio. Since the same semantic meaning can be expressed by multiple actions, and different semantic meanings can correspond to similar actions, and because deep learning models are generative models, the animation generation process is highly unpredictable and difficult to control, resulting in low accuracy of the final animation. Summary of the Invention
[0004] This application provides an animation generation method, apparatus, device, and storage medium, which improves the accuracy and quality of animation generated from audio. The technical solution provided by this application is as follows.
[0005] According to one aspect of the embodiments of this application, an animation generation method is provided, the method comprising:
[0006] Obtain multiple keywords at multiple time points in the audio, including semantic keywords and stress keywords. Each semantic keyword corresponds to a semantic category, and each stress keyword is related to the stress feature at the corresponding time point.
[0007] Based on the target matching order, at least two target animation segments that match the audio are sequentially determined from the animation library. The matching process of the keyword with a later matching order refers to the target animation segment corresponding to the keyword with a earlier matching order. The target matching order is associated with the semantic category of the multiple keywords and semantic keywords.
[0008] An animation of the audio is generated based on the at least two target animation clips;
[0009] The animation library includes animation clips corresponding to at least two semantic categories and animation clips corresponding to accent features.
[0010] According to another aspect of the embodiments of this application, an animation generation apparatus is provided, the apparatus comprising:
[0011] The acquisition module is used to acquire multiple keywords at multiple time points in the audio. The multiple keywords include semantic keywords and stress keywords. Each semantic keyword corresponds to a semantic category, and each stress keyword is related to the stress feature at the corresponding time point.
[0012] The determination module is used to sequentially determine at least two target animation segments that match the audio from the animation library based on the target matching order, wherein the matching process of the keyword with a later matching order refers to the target animation segment corresponding to the keyword with a earlier matching order, and the target matching order is associated with the semantic category of the plurality of keywords and semantic keywords;
[0013] A generation module is used to generate an animation of the audio based on the at least two target animation clips;
[0014] The animation library includes animation clips corresponding to at least two semantic categories and animation clips corresponding to accent features.
[0015] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory being used to store a computer program, the computer program being loaded and executed by the processor to implement the animation generation method in the embodiments of this application.
[0016] On the other hand, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium, the computer program being loaded and executed by a processor to implement the animation generation method in the embodiments of this application.
[0017] On the other hand, a computer program product is provided, the computer program product including a computer program stored in a computer-readable storage medium, a processor of a computer device reading the computer program from the computer-readable storage medium, the processor executing the computer program, causing the computer device to perform the animation generation method described in any of the above implementations.
[0018] In this embodiment, semantic keywords and accent keywords are extracted from the audio. Based on these keywords, target animation segments matching the audio are selected from an animation library. Then, an animation of the audio is generated based on these target animation segments. This approach considers both the semantic and accent features of the audio during animation generation, ensuring a comprehensive and accurate representation of the audio. Furthermore, the method matches animation segments from a pre-defined animation library, making the quality of the final animation controllable and reliable. Additionally, when matching animation segments from the library, the matching process for keywords with later matching sequences references the target animation segments corresponding to keywords with earlier matching sequences. This avoids conflicts between at least two identified target animation segments, resulting in a more diverse range of animation segments and smoother transitions between them, thus improving overall animation quality. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the implementation environment of a solution provided in one embodiment of this application;
[0020] Figure 2 This is a schematic diagram of a computer system provided in one embodiment of this application;
[0021] Figure 3 This is a flowchart of an animation generation method provided in one embodiment of this application;
[0022] Figure 4 This is a schematic diagram of a triple provided in one embodiment of this application;
[0023] Figure 5 This is a flowchart of an animation generation method provided in one embodiment of this application;
[0024] Figure 6 This is a functional schematic diagram of a speech recognition module provided in one embodiment of this application;
[0025] Figure 7 This is a schematic diagram illustrating the correspondence between keywords and semantic categories provided in one embodiment of this application;
[0026] Figure 8 This is a schematic diagram of an accent feature provided in one embodiment of this application;
[0027] Figure 9 This is a schematic diagram illustrating the priority provided in one embodiment of this application;
[0028] Figure 10 This is a schematic diagram illustrating a script example and a motion capture requirement example provided in one embodiment of this application;
[0029] Figure 11 This is a schematic diagram of time point annotation for an animation clip provided in one embodiment of this application;
[0030] Figure 12 This is a probability correspondence diagram between semantic categories and accent features and head movements provided in one embodiment of this application;
[0031] Figure 13 This is a schematic diagram of a fused animation clip provided in one embodiment of this application;
[0032] Figure 14 This is a schematic diagram of a spliced animation clip provided in one embodiment of this application;
[0033] Figure 15 This is a schematic diagram of a spliced animation clip provided in one embodiment of this application;
[0034] Figure 16 This is a schematic diagram of a spliced animation clip provided in one embodiment of this application;
[0035] Figure 17 This is a flowchart of an animation generation method provided in one embodiment of this application;
[0036] Figure 18 This is a schematic diagram of the interface of a plugin provided in one embodiment of this application;
[0037] Figure 19 This is a comparison diagram of results provided by one embodiment of this application;
[0038] Figure 20 This is a block diagram of an animation generation apparatus provided in one embodiment of this application;
[0039] Figure 21 This is a structural block diagram of a computer device provided in one embodiment of this application. Detailed Implementation
[0040] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0041] In related technologies, when generating animations from audio, the audio is typically input into a deep learning model, which then outputs an animation based on the audio's semantic information. However, since the same semantic meaning can be expressed by multiple actions, and different semantic meanings can correspond to the same or similar actions, and because deep learning models are generative models, the animation generation process is highly unpredictable and difficult to control, resulting in low accuracy. Taking the generation of head animations from audio as an example, keywords at multiple time points in the audio can have multiple semantic meanings. Therefore, multiple head actions can be generated for keywords at multiple time points. However, since the same semantic meaning can be expressed by multiple head actions, and different semantic meanings can be expressed by the same head action, and the deep learning model only generates the animation for the head action at each time point based on its semantic meaning, conflicts may occur between multiple head actions at multiple time points in the animation. For example, multiple consecutive time points might show head shaking, obviously resulting in poor animation quality and low accuracy.
[0042] Please refer to Figure 1 The diagram illustrates an animation generation method provided in one embodiment of this application. The method includes at least one of the following steps.
[0043] 1. Keyword Acquisition: Acquire multiple keywords at multiple time points in audio 101, including semantic keyword 102 and accent keyword 103. Each semantic keyword corresponds to a semantic category, and each accent keyword is related to the accent features at the corresponding time point.
[0044] 2. Determine the target animation segments. The animation library 104 includes multiple animation segments 105, including animation segments corresponding to semantic categories and animation segments corresponding to accent features. Based on the semantic keywords 102, semantic categories, accent keywords 103, and target matching order 106 determined above, at least two target animation segments 107 that match the audio are sequentially determined from the animation library 104. The matching process for keywords with later matching order references the target animation segments corresponding to keywords with earlier matching order. This matching of animation segments from a pre-set animation library makes the generated animation controllable, ensuring quality. Furthermore, referencing the matching process for keywords with later matching order to the target animation segments corresponding to keywords with earlier matching order avoids conflicts between the determined at least two target animation segments, resulting in a more diverse range of animation segments in the generated animation and smoother splicing between them, further ensuring the quality of the generated animation and improving its accuracy.
[0045] 3. Generate animation. Based on the identified at least two target animation segments 107, generate animation 108 for audio 101, which means splicing at least two target animation segments 106 according to the timing of their corresponding keywords to obtain animation 108 for audio 101.
[0046] Please refer to Figure 2 The diagram illustrates a computer system provided in one embodiment of this application. The computer system includes at least one of the following: a terminal device 10 and a server 20.
[0047] Terminal device 10 can be an electronic device such as a mobile phone, tablet computer, multimedia playback device, PC (Personal Computer), wearable device, in-vehicle terminal device, VR (Virtual Reality) device, AR (Augmented Reality) device, MR (Mixed Reality) device, etc. Terminal device 10 can run a client with a target application, which can be an application that generates animation for audio. This application embodiment does not limit the implementation form of the target application; for example, it can be an application that requires downloading and installation, a small program that does not require installation, a web application, etc.
[0048] In some embodiments, terminal device 10 acquires audio and sends it to server 20. Server 20 acquires multiple keywords at multiple time points in the audio, generates an animation based on these keywords, and returns the animation to terminal device 10 for display. Alternatively, terminal device 10 acquires multiple keywords at multiple time points in the audio, generates an animation based on these keywords, and displays the animation.
[0049] In this embodiment, server 20 provides background services for the target application. For example, server 20 may be used to generate animation for audio. Server 20 can be a single server, a server cluster consisting of multiple servers, or a cloud computing service center. Terminal device 10 can communicate with server 20 via a network, such as a wireless or wired network.
[0050] Please refer to Figure 3 The diagram illustrates a flowchart of an animation generation method provided in one embodiment of this application. The execution entity for each step of this method can be a computer device, such as... Figure 2 The terminal device 10 in the computer system shown can also be a server 20. Taking the server as the executing entity of each step of the method as an example, the method may include at least one of the following steps 310 to 330.
[0051] Step 310: Obtain multiple keywords at multiple time points in the audio. The multiple keywords include semantic keywords and stress keywords. Each semantic keyword corresponds to a semantic category, and each stress keyword is related to the stress feature at the corresponding time point.
[0052] In this embodiment of the application, each keyword includes one or more characters, that is, each keyword is a character or word. The audio is audio including speech, in order to facilitate the identification of semantic keywords and semantic categories therein.
[0053] In this embodiment, the multiple keywords include one or more semantic keywords and one or more accent keywords. Each accent keyword corresponds to an accent feature at a specific time point in the audio, and the accent feature is related to at least one of the following: pronunciation duration, intensity, and pitch at that time point. Each accent keyword is associated with the accent feature at its corresponding time point, meaning that words in the audio whose accent features meet preset requirements are output as accent keywords.
[0054] Step 320: Based on the target matching order, at least two target animation segments that match the audio are sequentially determined from the animation library. The matching process of the keyword with a later matching order refers to the target animation segment corresponding to the keyword with a earlier matching order. The target matching order is associated with the semantic categories of multiple keywords and semantic keywords.
[0055] In this embodiment, the animation clips in the animation library can be animations of the movement of objects (such as cars, trains, etc.) or animations of the movement of virtual objects (such as animals, plants, people, etc.). Furthermore, they can be animations of the movement of a specific part of a virtual object. In this embodiment, the animation clips in the animation library that depict the head movement of a virtual object are used as an example. Each animation clip includes a head movement, direction, and amplitude of the virtual object; that is, the virtual object in each animation clip moves based on a head movement, direction, and amplitude.
[0056] The head movements include nodding, tilting the head, shaking the head, extending the head forward, tilting the head and then nodding, and lateral movements. Movement directions include left, right, towards the center, from left to right, and from right to left. Movement amplitudes include small and medium amplitudes. A small amplitude means the rotation angle does not exceed a first angle, while a medium amplitude means the rotation angle exceeds the first angle but does not exceed a second angle, where the second angle is greater than the first angle. Multiple head movements, movement directions, and movement amplitudes can be freely combined to obtain at least two triplets, each containing head movement, movement direction, and movement amplitude. Each animation segment corresponds to one triplet. For example, see [link to example]. Figure 4 , Figure 4 This is a schematic diagram of a triple provided in one embodiment of this application. Each triple includes a head movement, a movement direction, and a movement amplitude.
[0057] In this embodiment, the animation library includes animation clips corresponding to at least two semantic categories and animation clips corresponding to accent features. The head movements in the animation clips corresponding to each semantic category can represent the semantics of that category. For example, for an affirmative semantic category, a nodding movement in its corresponding animation clip can represent an affirmative semantics. For a negative semantic category, a head shaking movement in its corresponding animation clip can represent a negative semantics.
[0058] Specifically, the process involves identifying target animation segments that match the audio from the animation library, which means identifying animation segments corresponding to at least some of the keywords from the animation library. Among these keywords, the animation segments corresponding to semantic keywords are the animation segments corresponding to their semantic categories, and the animation segments corresponding to accent keywords are the animation segments corresponding to accent features.
[0059] In this application embodiment, a semantic category can correspond to one or more animation segments, and an accent feature can correspond to one or more animation segments. Therefore, when determining the target animation segment that matches each keyword from the animation library, the matched target animation segments are also referenced to select the most suitable target animation segment from the multiple animation segments corresponding to the semantic category or accent feature of the keyword.
[0060] Step 330: Generate an audio animation based on at least two target animation clips.
[0061] In this embodiment of the application, each target animation segment corresponds to a keyword. At least two target animation segments are spliced together according to the time sequence of the corresponding keywords to obtain the audio animation.
[0062] In this embodiment, semantic keywords and accent keywords are extracted from the audio. Based on these keywords, target animation segments matching the audio are selected from an animation library. Then, an animation of the audio is generated based on these target animation segments. This approach considers both the semantic and accent features of the audio during animation generation, ensuring a comprehensive and accurate representation of the audio. Furthermore, the method matches animation segments from a pre-defined animation library, making the quality of the final animation controllable and reliable. Additionally, when matching animation segments from the library, the matching process for keywords with later matching sequences references the target animation segments corresponding to keywords with earlier matching sequences. This avoids conflicts between at least two identified target animation segments, resulting in a more diverse range of animation segments and smoother transitions between them, thus improving overall animation quality.
[0063] Please refer to Figure 5 The diagram illustrates a flowchart of an animation generation method provided in one embodiment of this application. The execution entity for each step of this method can be a computer device, such as... Figure 2 The terminal device 10 in the computer system shown can also be a server 20. Taking the server as the executing entity of each step of the method as an example, the method includes at least one of the following steps 510 to 540.
[0064] Step 510: Obtain multiple keywords at multiple time points in the audio. The multiple keywords include semantic keywords and stress keywords. Each semantic keyword corresponds to a semantic category, and each stress keyword is related to the stress feature at the corresponding time point.
[0065] In this application embodiment, the semantic keywords and accent keywords of the audio are obtained respectively. The process of obtaining multiple keywords at multiple time points of the audio includes the following steps (1)-(2).
[0066] (1) Obtain the text of the audio and input the text into the large language model. The large language model is used to output at least one semantic keyword and the semantic category of each semantic keyword based on the semantic prompt information. The semantic prompt information is used to prompt the large language model to obtain the semantics of the text.
[0067] In this embodiment, obtaining the text from the audio is equivalent to performing speech recognition on the audio to obtain the text. Optionally, speech recognition is performed on the audio using a speech recognition module. The speech recognition module is used to perform speech recognition on the audio to obtain the text from the audio. Further, the speech recognition module is also used to identify the language of the audio and the pronunciation time of each word in the text, including the start and end times of pronunciation. Optionally, the speech recognition module uses a deep learning speech recognition model to perform speech recognition, and the speech recognition model can recognize the speech of audio in multiple languages.
[0068] For example, see Figure 6 , Figure 6 This is a functional diagram of a speech recognition module provided in one embodiment of this application. The speech recognition module can not only recognize text, but also recognize the language type and the pronunciation time of each word.
[0069] In this embodiment, semantic prompts are used to provide the large language model with background and initial guidance for the task, enabling the large language model to understand the task requirements and generate results that meet the requirements. In some embodiments, the animation library includes only a few specified semantic categories. When identifying semantic keywords and semantic categories based on the large language model, only semantic keywords belonging to these semantic categories are identified. Accordingly, the semantic prompts can specify these semantic categories to instruct the large language model to identify semantic keywords in the text that belong to these semantic categories.
[0070] In this embodiment, the semantic prompt information can be: You are a semantic analyst, and you want to analyze and list the words with special semantics in the text. You will select words from the text that have the following semantic categories (affirmative, negative, intensity) and list them. Furthermore, the semantic prompt information can also include a specific example so that the large language model can refer to the example to generate results that meet the requirements.
[0071] In this embodiment, the large language model can select different strategies for semantic analysis based on the language of the audio. For example, for one language, the large language model can analyze the semantics of each character in the text. For characters with semantic meaning, the large language model will return the name of the semantic category, while characters without semantic meaning will return "free". For another language, the large language model can directly use text matching to achieve semantic analysis, that is, pre-define semantic categories for multiple characters, match the text with the pre-defined multiple characters, and then find the semantically meaningful keywords in the text.
[0072] For example, see Figure 7 , Figure 7 This is a schematic diagram illustrating the correspondence between keywords and semantic categories provided in one embodiment of this application. Each semantic category can have multiple different keywords. "We," "Me," etc., represent semantic categories.
[0073] In some embodiments, each animation clip in the animation library is labeled with a semantic category. A large language model is used to identify semantic keywords in the text that belong to the semantic categories found in the animation library. Since a label is also a semantic keyword, the large language model can detect and classify keywords in the text that are the same as or similar to the labels included in the animation library, thus obtaining keywords belonging to the semantic categories found in the animation library. For example, if the animation library contains a label for the semantic category "very," and the text contains synonymous keywords such as "extremely," then that keyword is classified into the semantic category of "very."
[0074] In this embodiment of the application, the semantic keywords and semantic categories of each semantic keyword in the text are obtained by using a large language model, which can improve the accuracy and efficiency of obtaining semantic keywords and semantic categories.
[0075] In some embodiments, the large language model is obtained by fine-tuning a general large language model, that is, the large language model is also trained. The training process of the large language model includes the following steps: obtaining at least two training samples and semantic prompts, each training sample including a text sample and the semantic categories of at least two keywords in the text sample; adjusting the model parameters of the large language model based on the at least two training samples and semantic prompts to obtain a trained large language model.
[0076] In some embodiments, during the acquisition of training samples, multiple semantic category labels can be preset. These labels are also the labels included in the animation library. The labels are input into a large language model, which generates a large amount of text containing these semantic keywords. Then, the semantic and non-semantic parts of the text are separated, and the semantic category labels corresponding to the semantic keywords are labeled, resulting in a large number of positive samples. Furthermore, incorrect semantic category labels can be used to label the semantic keywords, resulting in a large number of negative samples. This ensures that the training samples include both positive and negative samples, improving the generalization ability of the trained large language model.
[0077] The process involves iteratively adjusting the model parameters of the large language model based on at least two training samples and semantic prompts. In each iteration, a text sample is input into the large language model, which outputs a predicted keyword and its semantic category. The model parameters are adjusted based on the differences between the keywords in the text sample and the predicted keywords, as well as the differences between the semantic categories of the keywords and the predicted keywords, until an iteration stopping condition is met. This stopping condition can be reaching a preset number of iterations or the difference converging; no specific limitation is made here.
[0078] In this embodiment, the large speech model is fine-tuned based on the text sample and the semantic categories of at least two keywords in the text sample, so that the large language model can better identify semantic keywords and semantic categories in the text, thereby improving the accuracy of semantic keywords and semantic categories obtained by the large language model.
[0079] In this embodiment, the example of obtaining semantic keywords and semantic categories in audio through a large language model is used for illustration. In other embodiments, semantic keywords and semantic categories in audio can also be obtained manually, which is not specifically limited here.
[0080] (2) Determine the stress features of each word in the audio, and output the words with the maximum stress features in the audio as stress keywords.
[0081] In some embodiments, the stress feature is related to factors such as pronunciation duration, intensity, and pitch. The process of determining the stress feature of each word in the audio includes the following steps: weighted summation of at least two of the pronunciation duration, intensity, pitch, and pitch difference of the word in the audio to obtain the stress feature of the word. The pitch difference refers to the difference between the maximum pitch and the minimum pitch of the word.
[0082] In some embodiments, the computer device uses an accent detection module to obtain accented keywords in the audio. Optionally, the accent detection module also considers the language of the audio when determining accented features. For example, if there are pauses in the audio, and in some languages these pauses can be represented by symbols such as periods, commas, question marks, and semicolons, meaning there are no words at the pauses, then the accented feature of the symbols representing pauses in the text can be directly assigned a value of zero.
[0083] In this embodiment, sound intensity is measured in decibels (dB), and pitch is measured in Hz. The weights of sound duration, sound intensity, pitch, and pitch difference can be set as needed and are not specifically limited here. In some embodiments, to improve computational efficiency, the sound duration, sound intensity, pitch, and pitch difference are normalized before being weighted and summed. The normalization process is used to limit the values to a preset range, such as [0, 1].
[0084] In this embodiment, the stress feature is obtained by weighted summation of the above four items, i.e., stress feature = weight 1 × pronunciation duration + weight 2 × intensity + weight 3 × pitch + weight 4 × pitch range. See also Figure 8 , Figure 8 This is a schematic diagram of the accent feature provided in one embodiment of this application. The accent feature is obtained by weighted summation of the normalized values of pronunciation time, intensity, pitch and pitch difference.
[0085] In this embodiment, since the pronunciation duration, intensity, pitch, and pitch difference of a word at a certain point in the audio can reflect the rhythm of that point in the audio, and the rhythmic prominence in the audio is the point where the virtual object in the animation is most likely to perform an action, the stress features can be determined based on these factors, and then the maximum value of the stress features can be determined, thus the rhythmic prominence in the audio can be determined, which means that the determined stress keywords are accurate, and the animation generated based on the stress keywords can accurately match the audio, thereby improving the quality of the animation.
[0086] In this embodiment of the application, words with maximum stress features in the audio are used as stress keywords. Since a maximum value represents the maximum value of a local area in the audio, the maximum value can represent the point where the rhythm of the local area in the audio is most prominent. Therefore, using words with maximum values as stress keywords makes it easier to find the animation clip that most accurately reflects the audio rhythm of the local area based on the maximum value. In other words, the stress keywords determined by this scheme are accurate.
[0087] In this embodiment, the determination of accent keywords is illustrated using a maximum value as an example. In other embodiments, the words in the audio can be sorted at multiple time points based on the accent features of each word. The larger the accent feature, the higher the ranking. The word ranked first in a preset position is taken as the accent keyword. In this way, the accent keywords selected are the words with the most prominent rhythm in the audio, and the animation clips determined based on the accent keywords are more in line with the audio rhythm.
[0088] It should be noted that the sequence numbers of steps (1)-(2) above are only for the convenience of explanation and are not used to restrict the execution order of the two. For example, step (1) can be executed before or after step (2), and the two can also be executed synchronously to improve efficiency.
[0089] Step 520: Based on the priority of the semantic category of each semantic keyword and the priority of the stress feature of each stressed keyword, determine at least two target keywords from multiple keywords.
[0090] In this embodiment, each semantic category corresponds to a priority, and the stress feature also corresponds to a priority; different stressed keywords correspond to the same priority. Optionally, the priority of the stress feature lies among the various priorities of the semantic categories. For example, the priority of the stress feature is 3, and the priorities of various semantic categories are 1, 2, 4, 5, etc. The smaller the number, the higher the priority. For example, see... Figure 9 , Figure 9 This is a schematic diagram illustrating the priority provided in one embodiment of this application. Different semantic categories can correspond to the same priority, meaning different semantic categories can have the same importance. Since a keyword can be both a semantic keyword and an accent keyword, the higher of the semantic category priority and the accent feature priority can be used as the keyword's priority.
[0091] In some embodiments, the target keyword is a high-priority keyword. The process of determining at least two target keywords from multiple keywords based on the priority of the semantic category of each semantic keyword and the priority of the accent feature of each accent keyword includes the following steps: for any keyword, if the interval between the keyword and the temporally adjacent keyword is not less than a first duration, the keyword is output as the target keyword; if the interval between the keyword and the temporally adjacent keyword is less than the first duration, the keyword with the higher priority among the temporally adjacent keywords is output as the target keyword.
[0092] The first duration can be set as needed, and no specific limit is set here.
[0093] In this embodiment, when the interval between two keywords is short, only the keyword with higher priority is retained. This avoids the problem of difficulty in splicing animation clips of the two keywords due to the short interval. Furthermore, retaining the keyword with higher priority allows the determined animation clips to more accurately represent the audio, improving the accuracy of the final generated animation.
[0094] In other embodiments, only high-priority keywords are retained. The process of determining target keywords may also include the following implementation: Multiple keywords are sorted based on the priority of their respective semantic categories and the priority of the stress features of each stressed keyword; the keyword ranked first in a preset position is output as the target keyword. A higher priority for any keyword indicates greater importance and influence on the audio. Therefore, after determining such keywords as target keywords, animation is generated based on the corresponding animation clips, enabling the final animation to more richly and accurately express the audio, thus improving the accuracy of the generated animation.
[0095] Step 530: Match target animation clips from the animation library sequentially for each target keyword according to the target matching order of at least two target keywords, and obtain the target animation clips matched for each target keyword. The target matching order is obtained based on the priority of the semantic category or the priority of the accent features of each of the at least two target keywords.
[0096] In the embodiments of this application, the target matching order is obtained based on the priority of the semantic category of each of the at least two target keywords or the priority of the accent feature. That is, based on the priority of the semantic category of the semantic keyword and the priority of the accent feature of the accent keyword, the at least two target keywords are sorted to obtain the target matching order of the at least two target keywords. The higher the priority, the earlier the ranking.
[0097] In this embodiment of the application, since different semantic categories can correspond to the same priority, there may be keywords with the same priority among the at least two target keywords. Therefore, determining the target matching order also includes the following implementation: if there are at least two first keywords among the at least two target keywords, the duration priority of each first keyword is determined based on the interval duration between each first keyword and the target endpoint of the audio. The target endpoint refers to the endpoint with the shorter interval duration between the two endpoints of the audio and the first keyword. The at least two first keywords refer to target keywords with the same priority. The duration priority of each first keyword is negatively correlated with its corresponding interval duration. Based on the duration priority of each of the at least two first keywords, the at least two first keywords are arranged in the priority order to obtain the target matching order of the at least two target keywords. The priority order refers to the order in which the target keywords other than the at least two first keywords have been arranged according to priority.
[0098] Among them, the time priority of each first keyword is negatively correlated with its corresponding interval length. That is, the longer the interval length, the lower the priority, and the shorter the interval length, the higher the priority.
[0099] In this embodiment, the target endpoint is described as the endpoint closest to the target keyword. In other embodiments, the target endpoint is directly the start point or the end point of the audio; no specific limitation is made here.
[0100] The priority order is determined by sorting the target keywords other than the primary keywords based on their respective priorities. Then, based on the duration priority of each of the primary keywords, the primary keywords are arranged into a priority order. This means that the primary keywords are first sorted according to their respective priority durations to obtain a first order, and then, based on the same priority of these primary keywords, the two primary keywords in this first order are arranged into a priority order to obtain the target matching order.
[0101] In this embodiment, when two target keywords have the same priority, the keyword closer to the audio endpoint is given a higher priority. Since people pay more attention to animation segments closer to the audio endpoint when watching audio animations, they have higher requirements for these segments. Therefore, giving the keyword at that location a higher priority allows for earlier matching of the target animation segment, reducing the impact of other already matched target animation segments on matching the segment at that location. This facilitates matching a more accurate animation segment for the target keyword, thereby generating a more accurate animation and improving the quality of the animation.
[0102] In some embodiments, at least two first keywords may contain keywords with the same duration priority. The process of arranging at least two first keywords into a priority order based on their respective duration priorities to obtain the target matching order of at least two target keywords includes the following steps: if at least two first keywords exist among at least two target keywords, and at least two second keywords exist among at least two first keywords, and at least two second keywords are both stressed keywords, the stress priority of each second keyword is determined based on the stress feature of each second keyword, where at least two second keywords refer to first keywords with the same duration priority, and the stress priority of each second keyword is positively correlated with its respective stress feature; and the at least two first keywords are arranged into a priority order based on their respective duration priorities and stress priorities to obtain the target matching order of at least two target keywords.
[0103] Among them, the stress priority of each second keyword is positively correlated with its respective stress feature. That is, the larger the stress feature, the higher the stress priority, and the smaller the stress feature, the lower the stress priority.
[0104] Specifically, based on the duration priority of each of the at least two first keywords and the accent priority of each of the at least two second keywords, the at least two first keywords are arranged into a priority order. That is, based on the priority duration of each of the at least two first keywords, the first keywords other than the at least two second keywords are sorted first to obtain a second order. Then, based on the duration priority and accent priority of each of the at least two second keywords, the at least two second keywords are arranged into the second order to obtain a first order. Finally, based on the same priority of the at least two first keywords, the at least two first keywords of the first order are arranged into a priority order to obtain the target matching order.
[0105] In this embodiment, when two target keywords have the same duration priority, the keyword with a larger accent feature is assigned a higher priority. Since a larger accent feature indicates a stronger rhythmic characteristic of the target keyword and a greater impact on the audio, assigning it a higher priority allows for earlier matching of the target animation segment. This reduces the influence of other already matched target animation segments on matching the target keyword, facilitating the matching of more accurate animation segments and improving the quality of the animation.
[0106] In some embodiments, if at least two second keywords are semantic keywords, the time priority order among these at least two second keywords can be randomly sorted.
[0107] In this embodiment, each animation segment includes a virtual object's head movement, direction, and amplitude. Each animation segment is pre-recorded, and the process includes the following steps: A script is written for each animation segment, including semantic keywords and accent keywords, specifying head movements, directions, and amplitudes for each semantic keyword and accent keyword. Actors perform based on the script; that is, actors perform corresponding movements while reciting the semantic keywords and accent keywords in the script, based on the head movements, directions, and amplitudes set for them. For other lines, a baseline animation (idel) is performed, where the head movements are limited to a preset range. Motion capture tools such as VICON are used to capture the actors, obtaining animation data corresponding to the script. This animation data is then mapped onto a virtual object to control the virtual object to perform corresponding movements, resulting in the corresponding animation segment.
[0108] For example, see Figure 10 , Figure 10 This is a schematic diagram illustrating an example script and motion capture requirements provided in one embodiment of this application. Semantic keywords and accented keywords in the script are highlighted. Furthermore, certain requirements are imposed on the actors' movements to ensure that the captured movements are as close as possible to everyday human actions and to capture as diverse movements as possible.
[0109] Since the script includes multiple semantic keywords and multiple accent keywords, the animation consists of multiple animation segments corresponding to each keyword. Therefore, the resulting animation is segmented into multiple animation segments, each containing only one type of head movement with moderate speed and amplitude, exhibiting only medium and small amplitude movements, without large amplitude movements. Furthermore, to ensure smooth splicing of subsequent animation segments, a baseline animation of a certain duration, such as no more than 0.5 seconds, can be retained in the animation before and after the head movement in each segment. Optionally, the segmentation can be performed using a motion capture tool like Motion Builder.
[0110] In some embodiments, to facilitate the segmentation of animation into animation segments, the captured animation data can be imported into a rendering engine such as UNREAL ENGINE for processing. Since these engines have standard virtual object templates, it is also convenient to generate multiple animation segments with uniform virtual objects. Furthermore, the amplitude, speed, and length of the animation segments can be manually adjusted. Moreover, the engine can align the first frame of all segmented animation segments with the first frame of the base animation, that is, unify the local rotation angles of the head, neck, and neck1 bones in the first frame of the animation segment, aligning the animation segments with the base animation to the same metric dimension, facilitating the subsequent splicing of multiple animation segments based on the base animation. The head posture of the virtual object can be represented by three rotation angles, which refer to the rotation angles of the three bones of the head around three mutually perpendicular axes in the three-dimensional coordinate system, namely the pitch angle, yaw angle, and roll angle of the virtual object, corresponding to the head-up posture, head-down posture, and head-turning posture, respectively.
[0111] In some embodiments, animation segments are further annotated with time points, including the animation start point (action_start), the center point (stroke), and the animation end point (action_end). The animation start point is the time when the action in the animation segment begins, the animation end point is the time when the action in the animation segment ends, and the center point refers to the time when the action with the largest amplitude of movement is performed, such as the time point corresponding to the lowest point of a head-down movement. Alternatively, the center point refers to the time point slightly before the time when the action with the largest amplitude of movement is performed. Further, the animation start point and animation end point are scaled according to a preset scaling factor, using the center point as the center, to obtain the center start point (play_start) and center end point (play_end). For example, if the preset scaling factor is 0.5, then the center start point is the time point at the center position between the animation start point and the center point, and the center end point is the time point at the center position between the center point and the animation end point. The segment between the center start point and the center end point is the central segment of the animation segment, the segment between the animation start point and the center start point is the first transition segment of the animation segment, and the segment between the center end point and the animation end point is the second transition segment of the animation segment. The central segment refers to the segment that cannot be modified during the splicing of animation clips, while the transition segment refers to the segment that can be modified during the splicing of animation clips to facilitate the connection between them. For example, see... Figure 11 , Figure 11 This is a schematic diagram of time point annotation for an animation clip provided in one embodiment of this application.
[0112] In this embodiment of the application, after the animation segments are processed, they are exported from the rendering engine. Optionally, the animation data of each animation segment is read from the rendering engine using the FBX SDK tool, and an animation library is formed based on the animation data of each animation segment.
[0113] In this embodiment, the matching process of keywords with a later matching order refers to the target animation segment corresponding to the keywords with a earlier matching order. Each animation segment in the animation library includes a head action, action direction, and action amplitude of a virtual object. The animation library includes at least one animation segment corresponding to each semantic category and at least one animation segment corresponding to an accent feature. The process of matching target animation segments from the animation library sequentially for each target keyword according to the target matching order of at least two target keywords to obtain the target animation segment matched by each target keyword includes the following steps (1)-(3).
[0114] (1) For each target keyword, determine at least two first animation segments that match the target keyword from the animation library. The semantic category of the first animation segment is the same as the semantic category of the target keyword, or the first animation segment is an animation segment corresponding to the accent feature of the target keyword.
[0115] Wherein, if the target keyword is a semantic keyword, the first animation segment is an animation segment in the animation library corresponding to the semantic category of that semantic keyword. If the target keyword is an accent keyword, the first animation segment is an animation segment in the animation library corresponding to the accent feature. If the target keyword is both a semantic keyword and a target keyword, then at least two first animation segments include the animation segment in the animation library corresponding to the semantic category of the target keyword and the animation segment corresponding to the accent feature.
[0116] In this embodiment, a semantic category can correspond to at least one head action. A semantic category may have the same or different probabilities for different head actions; head actions whose probabilities meet a preset requirement are designated as the head actions corresponding to that semantic category. Optionally, the probabilities of these head actions are normalized, and head actions whose normalized probabilities meet a preset requirement are designated as the head actions corresponding to that semantic category. Furthermore, animation clips including these head actions are determined as the animation clips corresponding to that semantic category. For example, the preset requirement is that the normalized probability is greater than a preset probability.
[0117] In this embodiment, the accent feature can correspond to any head action; that is, any head action in the animation library can represent the accent keyword. To facilitate data processing, only a small number of animation segments are selected for the accent keyword each time. Therefore, the probability of each head action relative to the accent feature can be pre-set. When determining the first animation segment for the accent keyword, the animation segments containing one or more head actions are output as the first animation segment for the accent keyword, based on the probability of each head action.
[0118] For example, see Figure 12 , Figure 12 This is a probability correspondence diagram between semantic categories and stress features and head actions provided in one embodiment of this application. Each semantic category corresponds to at least one head action, with different probabilities for different head actions. Stress features correspond to multiple head actions, with different probabilities for different head actions.
[0119] In this embodiment, a semantic category corresponds to at least one head action, and each head action can correspond to at least two animation segments with different action directions and amplitudes. Animation segments with the same action direction and amplitude can further include at least two animation segments with different parameters such as action speed. That is, one semantic category corresponds to at least two first animation segments, and one semantic keyword can correspond to at least two first animation segments. Similarly, an accent keyword corresponds to at least one head action, and each head action can correspond to at least two animation segments with different action directions and amplitudes. Animation segments with the same action direction and amplitude can further include at least two animation segments with different parameters such as action speed. That is, one accent keyword can correspond to at least two first animation segments.
[0120] (2) Determine at least one second animation segment from at least two first animation segments. The second animation segment has different head movements and movement directions from the first target animation segment. The first target animation segment refers to the target animation segment whose target keyword has been determined and is adjacent to the target keyword in the time sequence.
[0121] In this embodiment of the application, among at least two first animation segments matched from the target keyword, a second animation segment is selected that is different from the head movements and movement directions of the adjacent determined target animation segments. This makes the determined adjacent target animation segments as different as possible, which increases the diversity of animation segments in the animation. This can bring a better viewing experience to the audience when the animation is played, that is, improve the quality of the animation.
[0122] In other embodiments, if the second animation segment is not included in at least two first animation segments, at least one third animation segment is determined from the at least two first animation segments, wherein the direction of motion of the virtual object in the third animation segment is different from that in the first target animation segment; and a target animation segment matching the target keyword is determined from the at least one third animation segment.
[0123] In this embodiment, if no second animation segment with different head movements and directions can be found, a third animation segment with different head movements is selected. This avoids the situation where no target animation segment can be selected for the target keyword. Furthermore, selecting a third animation segment with different head movements increases the diversity of animation segments in the animation, ensuring the quality of the animation. The specific process of determining the target animation segment matching the target keyword from at least one third animation segment is the same as the specific process of determining the target animation segment matching the target keyword from at least one second animation segment, and will not be repeated here.
[0124] Furthermore, if a third animation segment cannot be determined from at least two first animation segments, then both of the at least two first animation segments are considered as second animation segments.
[0125] (3) Identify the target animation segment that matches the target keyword from at least one second animation segment.
[0126] In this embodiment of the application, each animation segment is divided into animation start point, center start point, center point, center end point and animation end point from left to right according to the timeline, and the segment between the center start point and the center end point is the center segment of the animation segment; the process of determining the target animation segment that matches the target keyword from at least one second animation segment includes the following steps: outputting the second animation segment that meets the preset conditions from at least one second animation segment as the target animation segment that matches the target keyword, and the preset conditions include at least one of the following.
[0127] The first item is that the duration of the center segment of the second animation segment is less than the interval duration. The interval duration refers to the time interval between the center segments of the two second target animation segments when the center points of the two second target animation segments are aligned with the time points of their respective target keywords. The two second target animation segments refer to two target animation segments whose target keywords are determined and are adjacent to the time sequence of the target keywords.
[0128] The interval between the center segments of the two second target animation segments is also the interval between the last frame of the center segment of the first second target animation segment and the first frame of the center segment of the second target animation segment that comes later in time.
[0129] In this embodiment of the application, the duration of the center segment of the second animation segment is less than the interval duration between the center segments of the two second target animation segments, so that the center segment of the second animation segment can be inserted between the two determined adjacent target animation segments, thus avoiding splicing conflicts between animation segments.
[0130] The second point is that the duration of the central segment of the second animation clip is less than the duration of the corresponding target keyword in the audio.
[0131] The duration of the target keyword in the audio, i.e. the pronunciation duration of the target keyword, is determined based on the start and end times of the pronunciation of the target keyword in the audio.
[0132] In this embodiment of the application, the duration of the central segment of the second animation clip is less than the duration of the corresponding target keyword in the audio, so that the central segment can correspond to the audio in terms of timing. This avoids the problem of mismatch between the auditory and visual aspects of the animation due to the duration of the second animation clip exceeding the duration of the target keyword, thus ensuring the quality of the animation.
[0133] Thirdly, the switching speed between the center segment of the second animation segment and the center segment of each second target animation segment is less than the preset speed.
[0134] The two second target animation segments are the target animation segments adjacent to the left and right of the first target animation segment. Optionally, for the second target animation segment adjacent to the left, the switching speed refers to the speed at which the head switches from the last frame of the central segment of the second target animation segment to the first frame of the central segment. This switching speed is the quotient of the head pose difference between these two frames and the interval between the two frames. The head pose can be represented by three rotation angles, which are the rotation angles of the head around three mutually perpendicular axes in the three-dimensional coordinate system, namely pitch, yaw, and roll, corresponding to head tilt, head shake, and head turn, respectively. The head pose difference refers to the difference between these three angles between the two frames. Correspondingly, the switching speed is the angle that can be deflected per unit time. Since each of the three rotation angles can correspond to a switching speed, the switching speed between two frames can be the average of the switching speeds of these three rotation angles. Similarly, for the second target animation segment adjacent to the right, the switching speed refers to the speed at which the head switches from the last frame of the central segment of the second target animation segment to the first frame of the central segment. This switching speed is the quotient of the head pose difference between these two frames and the interval between the two frames.
[0135] In this embodiment, the switching speed between the center segment of the second animation segment and the center segment of the adjacent target animation segment is less than a preset speed, so that the switching between the two center segments is smooth and avoids abrupt switching, thus ensuring the visual effect of the animation and thus ensuring the quality of the animation.
[0136] In some embodiments, the second animation segment also satisfies at least one of the following conditions: If the target keyword is the first target keyword in the time sequence among at least two target keywords, then the head movement of the virtual object in the second animation segment is not a lateral movement. If the head movement of the virtual object in the second target animation segment is a head shake and the target keyword is an accented keyword, then the head movement of the virtual object in the second animation segment is not a lateral movement. If the head movement of the virtual object in the third target animation segment is a head raise followed by a nod, then the head movement of the virtual object in the second animation segment is not a head raise, and the third target animation segment is a target animation segment for which the interval between the current target keyword and the target keyword has been determined, with the interval duration not exceeding the fourth duration. If the head movement of the virtual object in the third target animation segment is a head raise, then the head movement of the virtual object in the second animation segment is not a head raise followed by a nod. If the head movement of the virtual object in the third target animation segment is a lateral movement, a head tilt, a head raise, or a forward stretch, then the head movements corresponding to the second animation segment and the third target animation segment are different, in order to reduce the repetition of these head movements.
[0137] In this embodiment of the application, the second animation segment is made to meet the above conditions, so that the determined second animation segment is more consistent with other determined target animation segments, thereby improving the viewing experience of the animation.
[0138] In some embodiments, when determining the target animation segment, the animation segment is also matched with the audio features of the target keyword. That is, for each target keyword, based on the audio features of the target keyword, a target animation segment whose motion amplitude of the virtual object matches the audio features of the target keyword is determined from at least one second animation segment. The audio features include at least one of intensity and pitch. In this embodiment, matching the motion amplitude with the audio features makes the determined target animation segment more closely match the target keyword, thereby making the audio and animation more closely match, ensuring the quality of the animation.
[0139] The animation clips in the animation library correspond to two types of motion amplitude, each corresponding to a different audio feature. If the audio feature belongs to the first feature range, the target keyword corresponds to the smaller amplitude of the two motion amplitudes; if the audio feature belongs to the second feature range, the target keyword corresponds to the medium amplitude of the two motion amplitudes. The maximum value of the first feature range is less than the minimum value of the second feature range. For example, if the audio feature refers to sound intensity, the first feature range refers to the first decibel range, and the second feature range refers to the second decibel range.
[0140] It should be noted that the target animation segment can be determined directly from the second animation segment based on audio features, or the second animation segment can be determined based on audio features, that is, audio features can be used as one of the aforementioned preset conditions to determine the second animation segment.
[0141] In this embodiment of the application, since there are multiple conditions and restrictions when determining the target animation segment for the target keyword, optionally, if the target animation segment cannot be determined for the target keyword based on the above conditions, the target keyword is skipped and no more target animation segments are matched for the target keyword.
[0142] In this embodiment, steps 520-530 above achieve the process of sequentially determining at least two target animation segments that match the audio from the animation library based on the target matching order. In this embodiment, after obtaining the semantic keywords and accent keywords in the audio, the semantic keywords and accent keywords are further filtered based on priority to obtain target keywords. Then, the target animation segments are determined based on the target keywords. This avoids the problem of too many animation segments that are difficult to splice and conflict due to matching animation segments for each keyword. Moreover, filtering keywords based on priority ensures that the selected keywords are those that are highly important and influential to the audio. Based on this, the animation segments are determined and the animation is generated, so that the final generated animation can accurately express the audio, has a high degree of fit with the audio, and improves the quality of the animation.
[0143] Step 540: Generate an audio animation based on at least two target animation clips.
[0144] In some embodiments, each animation segment includes, from left to right, an animation start point, a center start point, a center point, a center end point, and an animation end point, and the segment between the center start point and the center end point is the central segment of the animation segment; the process of generating audio animation based on at least two target animation segments includes the following steps: based on the timing of the center points of at least two target animation segments and the time points of their respective corresponding target keywords, splicing at least two target animation segments together to obtain audio animation.
[0145] In this embodiment, at least two target animation segments are spliced together based on the time sequence of the center point of the target animation segment and the corresponding target key point time point, so that the target animation segments in the generated animation are played at the time points of their respective target keywords, so that the visual information of the animation and the auditory information of the audio match, ensuring the quality of the animation.
[0146] In some embodiments, the process of splicing at least two target animation segments based on the timing of the center point of at least two target animation segments and the timing of their respective target keywords to obtain the audio animation includes the following steps: the number of at least two target animation segments is M, and for the target animation segment whose timing is at the i-th position among the at least two target animation segments, the following steps are taken, where i is an integer greater than 0 and less than or equal to M.
[0147] (1) If i = 1, the baseline animation and the first target animation segment are merged to obtain the merged animation segment. The baseline animation refers to the animation in which the range of head movements is limited to a preset range.
[0148] In this embodiment, the reference animation refers to head movements whose amplitude is limited to a preset range, that is, simulating the slight head and neck movements of a person when they are quiet and not speaking. Optionally, the reference animation includes multiple repeated reference animation segments, in which the amplitude of head movements in each reference animation segment is limited to a preset range.
[0149] In this embodiment, the starting part of the animation is generally a base animation. The base animation and the first target animation segment are merged to obtain a merged animation segment, that is, the first target animation segment is spliced onto the base animation. Optionally, before merging, the center point of the first target animation segment is aligned with the time point of the corresponding target keyword in the audio. More specifically, the center point of the first target animation segment is aligned with the starting time point of the target keyword in the audio. The animation corresponding to the time point before the first target animation segment in the audio is the base animation. The number of base animation segments included in the base animation is determined based on the duration of the time period corresponding to the base animation.
[0150] In some embodiments, to ensure smooth splicing between the base animation and the first target animation segment, the process of merging the base animation and the first target animation segment to obtain a merged animation segment includes the following steps: frame mixing of the base animation and a first transition segment of the first target animation segment to obtain a first mixed segment; sequentially splicing the base animation, the first mixed segment, the center segment of the first target animation segment, and the second transition segment to obtain the merged animation segment. The first transition segment refers to the segment between the animation start point and the center start point, and the second transition segment refers to the segment between the center end point and the animation end point. For example, see... Figure 13 , Figure 13 This is a schematic diagram of a fused animation clip provided in one embodiment of this application.
[0151] Frame blending technology generates a new animation frame by comparing pixel differences between two animation frames. In this embodiment, a reference animation with the same number of animation frames as the first transition segment is obtained, and this reference animation is aligned one-to-one with the animation frames in the first transition segment. For any two aligned animation frames, the head motion parameters in the reference animation frames are weighted and summed with the head motion parameters in the first transition segment to obtain a blended frame. Multiple blended frames are then sequentially stitched together to obtain the first blended segment. The head motion parameters include pitch angle, yaw angle, and roll angle, etc.
[0152] Optionally, in order to further improve the smoothness of the splicing between the base animation and the first target animation segment, as the interval between the animation frame in the first transition segment and the starting animation frame of the first transition segment is longer, the weight of the animation frame is greater during the frame blending process, and the weight of the animation frame in the base animation is smaller, so that the motion features in the first target animation segment in the first blend segment gradually become stronger, and the motion features in the base animation gradually become weaker.
[0153] The first blending segment is used to transition from the base animation to the first target animation segment; this first blending segment can also be called a BlendIn segment. If there are some animation frames before the start point of the first target animation segment, the first transition segment refers to the segment between the start point and the center start point of the target animation segment. Similarly, if there are some animation frames after the end point of the first target animation segment, the second transition segment refers to the segment between the center end point and the end point of the target animation segment.
[0154] In this embodiment, the starting part of the animation is a base animation. For the first target animation segment, the first transition segment is frame-blended based on the base animation to obtain a first blended segment, which realizes the fusion of the base animation and the first transition segment. Then, the base animation and the first target animation segment are spliced based on the first blended segment. In other words, the splicing transition between the base animation and the first target animation segment is realized through the first blended segment, making the splicing between the two smoother, the visual effect better, and thus improving the quality of the animation.
[0155] (2) If i is greater than 1 and less than M, based on the baseline animation, splice the i-th target animation segment onto the generated fused animation segment.
[0156] In some embodiments, the process of stitching the i-th target animation segment onto the generated fused animation segment based on the baseline animation includes any of the following cases.
[0157] The first method involves the following steps: If the interval between the i-th target animation segment and the already generated merged animation segment is longer than the second duration, the second transition segment of the last animation segment in the already generated merged animation segment is frame-blended with the base animation to obtain the second blended segment. The first transition segment of the i-th target animation segment is then frame-blended with the base animation to obtain the third blended segment. The second transition segment of the last animation segment in the already generated merged animation segment is removed. The merged animation segment after removal, the second blended segment, the base animation, the third blended segment, the center segment of the i-th target animation segment, and the second transition segment are then sequentially spliced together to attach the i-th target animation segment to the already generated merged animation segment. For example, see [link to example]. Figure 14 , Figure 14 This is a schematic diagram of a spliced animation clip provided in one embodiment of this application.
[0158] Specifically, before splicing the i-th target animation segment onto the already generated merged animation segment, the center point of the i-th target animation segment is aligned with the time point of the corresponding target keyword in the audio. Furthermore, the center point of the i-th target animation segment is aligned with the start time point of the target keyword in the audio. The interval between the i-th target animation segment and the already generated merged animation segment is also the interval between the last frame of the already generated merged animation segment and the first frame of the i-th target animation segment.
[0159] The second blending segment is used to transition from the generated blended animation clip to the baseline animation; this second blending segment can also be called a Blendout segment. The third blending segment is used to transition from the baseline animation to the i-th target animation clip; this third blending segment can also be called a BlendIn segment.
[0160] In this embodiment, when the duration between the i-th target animation segment and the generated fused animation segment is relatively long, a reference animation is filled in the interval between the two, and the reference animation and the second transition segment of the last animation segment on the generated fused animation segment are mixed. This facilitates the splicing of the generated fused animation segment and the reference animation based on the obtained second mixed segment. Furthermore, the reference animation and the first transition segment of the i-th target animation segment are also mixed, which facilitates the splicing of the reference animation and the i-th target animation segment based on the obtained third mixed segment. This makes the splicing between the i-th target animation segment and the generated fused animation segment smoother and improves the quality of the animation.
[0161] The second method involves extending the duration of the second transition segment of the last animation segment if the interval between the i-th target animation segment and the generated merged animation segment is less than the second duration but greater than the third duration. Then, interpolation is performed between the last frame of the extended second transition segment and the first frame of the center segment of the i-th target animation segment to obtain a first interpolated segment. The extended merged animation segment, the first interpolated segment, the center segment of the i-th target animation segment, and the second transition segment are then sequentially spliced together to attach the i-th target animation segment to the generated merged animation segment. For example, see [link to example]. Figure 15 , Figure 15 This is a schematic diagram of a spliced animation clip provided in one embodiment of this application.
[0162] In some embodiments, the duration of the second transition segment of the last animation clip can be lengthened by reducing the playback speed of the second transition segment, such as adjusting the playback speed to 0.5x. Interpolation can be implemented using spherical linear interpolation.
[0163] In this embodiment, the second duration is greater than the third duration. When the interval between the i-th target animation segment and the generated merged animation segment is neither long enough nor short enough, the interval is insufficient to allow for a transition to the base animation followed by a transition from the base animation to the i-th target animation segment. In this case, the second transition segment of the last animation segment is lengthened to fill the gap between them. Furthermore, interpolation is performed between the last frame and the first frame of the center segment of the i-th target animation segment to fill the gap between the two animation frames. A first interpolated segment is obtained based on this interpolation, enabling a smooth transition between the i-th target animation segment and the generated merged animation segment. This results in a smoother splicing between the i-th target animation segment and the generated merged animation segment, improving the animation quality.
[0164] The third method involves interpolating between the i-th target animation segment and the already generated merged animation segment if the interval is less than the third duration. This interpolation occurs between the last frame of the center segment of the last animation segment and the first frame of the center segment of the i-th target animation segment, resulting in a second interpolated segment. The second transition segment of the last animation segment on the already generated merged animation segment is then removed. The merged animation segment with the removed segments, the second interpolated segment, the center segment of the i-th target animation segment, and the second transition segment are then sequentially spliced together to attach the i-th target animation segment to the already generated merged animation segment. For example, see [link to example]. Figure 16 , Figure 16 This is a schematic diagram of a spliced animation clip provided in one embodiment of this application.
[0165] In this embodiment, when the duration between the i-th target animation segment and the generated fused animation segment is short, the second transition segment of the last animation segment on the generated fused animation segment is removed. The fused animation segment after removal, the second interpolated segment, the center segment of the i-th target animation segment, and the second transition segment are then sequentially spliced together. That is, the second transition segment of the last animation segment and the first transition segment of the i-th target animation segment are removed from the generated fused animation segment. Interpolation is then performed directly between the center end point of the last animation segment and the center start point of the i-th target animation segment to splice the i-th target animation segment onto the generated fused animation segment. This splicing of the two center segments through interpolation makes the splicing smoother and eliminates the need to consider transition segments and baseline animation, resulting in higher splicing efficiency.
[0166] (3) If i = M, based on the baseline animation, splice the Mth target animation segment onto the generated fused animation segment to obtain the audio animation.
[0167] The process of splicing the Mth target animation segment onto the generated blended animation segment is the same as the process of splicing the ith target animation segment onto the generated blended animation segment, and will not be repeated here.
[0168] In this embodiment, the first target animation segment is merged with the base animation to obtain the first merged animation segment. Then, each subsequent target animation segment is spliced onto the previously merged animation segment based on the base animation. This achieves the splicing of at least two target animation segments in sequence, thereby making multiple animation segments in the animation correspond to multiple time points in the audio and ensuring the quality of the generated animation.
[0169] For example, see Figure 17 , Figure 17 This is a flowchart of an animation generation method provided in one embodiment of this application. The method involves first obtaining multiple keywords from the audio, then selecting target keywords from these keywords, matching semantic keywords within the target keywords with animation segments corresponding to semantic categories, and matching accent keywords within the target keywords with animation segments corresponding to accent features, resulting in at least two target animation segments. These at least two target animation segments are then concatenated to generate an animation of the audio.
[0170] The solutions provided in this application can be applied to online dialogue scenarios, or embedded into plugins of some applications to generate animations offline. For example, see... Figure 18 , Figure 18This is a schematic diagram of the interface of a plugin provided in one embodiment of this application. The plugin loads a character model (virtual object) via the "Binding" option; loads an audio file via "Load Audio"; and sets the parameters for "Head Dynamics" in "Animation Adjustment" to control the scaling of the generated animation's movement. Finally, triggering the "Generate Head Animation" button allows for offline output of the corresponding character model's head animation data in a dialogue scene. The generated animation can also be previewed and edited. Clearly, this process can significantly improve the efficiency of animators when creating animations and save costs.
[0171] In the embodiments of this application, the generated animation can be played in various animation playback software, and no specific limitation is made here. See also Figure 19 , Figure 19 This is a comparison diagram of results provided in one embodiment of this application. Figure 19 In the example (a), multiple animation frames are obtained from an animation clip using other methods. Figure 19 In (b), multiple animation frames in an animation clip obtained through the solution provided in the embodiments of this application correspond to the text "don't know" in the same audio. By comparison, it can be seen that other solutions failed to generate a head shaking motion corresponding to the text "don't know" when encountering it, while the solution provided in the embodiments of this application generated a high-quality head shaking motion animation. Obviously, the head motion of this animation is closer to the head motion in daily conversation. It can be seen that the animation generated by the solution provided in the embodiments of this application has controllability and semantic saliency, and the effect is better.
[0172] In this embodiment, semantic keywords and accent keywords are extracted from the audio. Based on these keywords, target animation segments matching the audio are selected from an animation library. Then, an animation is generated based on these target animation segments. This approach considers both the semantic and accent features of the audio during animation generation, ensuring comprehensiveness and resulting in an accurate and comprehensive representation of the audio. Furthermore, the method matches animation segments from a pre-defined animation library, making the quality of the final animation controllable and reliable. Additionally, when matching animation segments from the library, the matching process for keywords with later matching sequences references the target animation segments corresponding to keywords with earlier matching sequences. This avoids conflicts between the identified at least two target animation segments, and the animation generated based on these segments includes more diverse segments with smoother transitions, thus improving overall animation quality.
[0173] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0174] Please refer to Figure 20 This diagram illustrates a block diagram of an animation generation apparatus according to an embodiment of this application. The apparatus has the function of implementing the above-described animation generation method; this function can be implemented in hardware or by hardware executing corresponding software. The apparatus can be the computer device described above, or it can be installed within a computer device. For example... Figure 20 As shown, the device may include:
[0175] The acquisition module 2010 is used to acquire multiple keywords at multiple time points in the audio. The multiple keywords include semantic keywords and stress keywords. Each semantic keyword corresponds to a semantic category, and each stress keyword is related to the stress feature at the corresponding time point.
[0176] The determination module 2020 is used to sequentially determine at least two target animation segments that match the audio from the animation library based on the target matching order. The matching process of the keyword with a later matching order refers to the target animation segment corresponding to the keyword with a earlier matching order. The target matching order is associated with the semantic category of multiple keywords and semantic keywords.
[0177] Generation module 2030 is used to generate an audio animation based on at least two target animation clips;
[0178] The animation library includes animation clips corresponding to at least two semantic categories and animation clips corresponding to accent features.
[0179] In some embodiments, the determining module 2020 is configured to:
[0180] Based on the priority of the semantic category of each semantic keyword and the priority of the stress feature of each stressed keyword, at least two target keywords are determined from multiple keywords;
[0181] According to the target matching order of at least two target keywords, target animation clips are matched sequentially from the animation library for each target keyword to obtain the target animation clip matched for each target keyword. The target matching order is obtained based on the priority of the semantic category or the priority of the accent features of each of the at least two target keywords.
[0182] In some embodiments, the determining module 2020 is configured to:
[0183] For any keyword, if the interval between the keyword and the keyword with an adjacent time sequence is not less than the first time sequence, the keyword will be output as the target keyword.
[0184] If the interval between a keyword and a keyword adjacent to it in time is less than the first time interval, the keyword with the higher priority among the keywords adjacent to it in time will be output as the target keyword.
[0185] In some embodiments, the determining module 2020 is further configured to:
[0186] If there are at least two first keywords among at least two target keywords, the duration priority of each first keyword is determined based on the interval between each first keyword and the target endpoint of the audio. The target endpoint refers to the endpoint of the two endpoints of the audio with the shorter interval between it and the first keyword. At least two first keywords refer to target keywords with the same priority. The duration priority of each first keyword is negatively correlated with its corresponding interval duration.
[0187] Based on the duration priority of each of the at least two primary keywords, the at least two primary keywords are arranged in a priority order to obtain the target matching order of the at least two target keywords. The priority order refers to the order in which the target keywords other than the at least two primary keywords have been arranged according to their priority.
[0188] In some embodiments, the determining module 2020 is configured to:
[0189] If at least two target keywords contain at least two first keywords, and at least two first keywords contain at least two second keywords, and at least two second keywords are stressed keywords, the stress priority of each second keyword is determined based on the stress characteristics of each second keyword. At least two second keywords refer to first keywords with the same duration priority. The stress priority of each second keyword is positively correlated with its respective stress characteristics.
[0190] Based on the duration priority of at least two primary keywords and the accent priority of at least two secondary keywords, the at least two primary keywords are arranged in a priority order to obtain the target matching order of at least two target keywords.
[0191] In some embodiments, each animation segment in the animation library includes a head movement, movement direction and movement amplitude of a virtual object, and the animation library includes at least one animation segment corresponding to each semantic category and at least one animation segment corresponding to an accent feature.
[0192] Determine module 2020, used for:
[0193] For each target keyword, at least two first animation segments that match the target keyword are determined from the animation library. The semantic category of the first animation segment is the same as the semantic category of the target keyword, or the first animation segment is the animation segment corresponding to the accent feature of the target keyword.
[0194] Determine at least one second animation segment from at least two first animation segments. The head movements and movement directions of the second animation segment are different from those of the first target animation segment. The first target animation segment refers to the target animation segment whose target keyword is determined and whose time sequence is adjacent to the target keyword.
[0195] Identify the target animation segment that matches the target keyword from at least one second animation segment.
[0196] In some embodiments, each animation segment includes, from left to right, an animation start point, a center start point, a center point, a center end point, and an animation end point along a timeline, and the segment between the center start point and the center end point is the central segment of the animation segment; the determining module 2020 is used for:
[0197] Output at least one of the second animation clips that meets the preset conditions as the target animation clip that matches the target keyword. The preset conditions include at least one of the following:
[0198] The duration of the central segment of the second animation segment is less than the interval duration. The interval duration refers to the time interval between the central segments of the two second target animation segments when the center points of the two second target animation segments are aligned with the time points of their respective target keywords. The two second target animation segments refer to two target animation segments whose target keywords are determined and are adjacent to the time sequence of the target keywords.
[0199] The duration of the central segment of the second animation clip is less than the duration of the corresponding target keyword in the audio;
[0200] The switching speed between the center segment of the second animation clip and the center segment of each second target animation clip is less than the preset speed.
[0201] In some embodiments, the determining module 2020 is further configured to:
[0202] If at least two first animation segments do not include a second animation segment, determine at least one third animation segment from the at least two first animation segments, wherein the direction of motion of the virtual object in the third animation segment is different from that in the first target animation segment;
[0203] Identify the target animation segment that matches the target keyword from at least one third animation segment.
[0204] In some embodiments, the determining module 2020 is configured to:
[0205] For each target keyword, based on the audio features of the target keyword, a target animation segment is determined from at least one second animation segment whose motion amplitude of the virtual object matches the audio features of the target keyword, the audio features including at least one of intensity and pitch.
[0206] In some embodiments, each animation segment includes, from left to right, an animation start point, a center start point, a center point, a center end point, and an animation end point, and the segment between the center start point and the center end point is the central segment of the animation segment;
[0207] Module 2030 is generated for:
[0208] Based on the timing of the center points of at least two target animation clips and the time points of their respective target keywords, at least two target animation clips are spliced together to obtain an audio animation.
[0209] In some embodiments, the generation module 2030 is configured to:
[0210] The number of at least two target animation segments is M. For the target animation segment whose timing is at the i-th position among the at least two target animation segments, the following steps are taken, where i is an integer greater than 0 and less than or equal to M:
[0211] If i=1, the baseline animation and the first target animation segment are merged to obtain the merged animation segment. The baseline animation refers to the animation in which the range of head movements is limited to a preset range.
[0212] If i is greater than 1 and less than M, based on the baseline animation, the i-th target animation segment is spliced onto the already generated fused animation segment;
[0213] If i = M, based on the baseline animation, the Mth target animation segment is spliced onto the already generated merged animation segment to obtain the audio animation.
[0214] In some embodiments, the generation module 2030 is configured to:
[0215] The reference animation is frame-blended with the first transition segment of the first target animation segment to obtain the first blended segment. The reference animation, the first blended segment, the center segment of the first target animation segment, and the second transition segment are sequentially spliced together to obtain the merged animation segment. The first transition segment refers to the segment between the animation start point and the center start point, and the second transition segment refers to the segment between the center end point and the animation end point.
[0216] In some embodiments, the generation module 2030 is configured to perform any of the following:
[0217] If the interval between the i-th target animation segment and the generated merged animation segment is longer than the second duration, the second transition segment of the last animation segment in the generated merged animation segment is frame-blended with the base animation to obtain the second blended segment. The first transition segment of the i-th target animation segment is frame-blended with the base animation to obtain the third blended segment. The second transition segment of the last animation segment in the generated merged animation segment is removed. The merged animation segment after removal, the second blended segment, the base animation, the third blended segment, the center segment of the i-th target animation segment, and the second transition segment are sequentially spliced together to splice the i-th target animation segment onto the generated merged animation segment.
[0218] If the interval between the i-th target animation segment and the generated merged animation segment is less than the second duration but greater than the third duration, the duration of the second transition segment of the last animation segment is lengthened. Interpolation is performed between the last frame of the lengthened second transition segment and the first frame of the center segment of the i-th target animation segment to obtain the first interpolated segment. The lengthened merged animation segment, the first interpolated segment, the center segment of the i-th target animation segment, and the second transition segment are sequentially spliced together to splice the i-th target animation segment onto the generated merged animation segment.
[0219] If the interval between the i-th target animation segment and the generated merged animation segment is less than the third duration, interpolation is performed between the last frame of the center segment of the last animation segment and the first frame of the center segment of the i-th target animation segment to obtain a second interpolated segment. The second transition segment of the last animation segment on the generated merged animation segment is removed. The merged animation segment after removal, the second interpolated segment, the center segment of the i-th target animation segment and the second transition segment are sequentially spliced together to splice the i-th target animation segment onto the generated merged animation segment.
[0220] In some embodiments, the acquisition module 2010 is used for:
[0221] The text of the audio is obtained and input into a large language model. The large language model is used to output at least one semantic keyword and the semantic category of each semantic keyword based on semantic prompts. The semantic prompts are used to prompt the large language model to obtain the semantics of the text.
[0222] In some embodiments, the apparatus further includes a training module for:
[0223] Obtain at least two training samples and semantic prompts, where each training sample includes a text sample and the semantic categories of at least two keywords in the text sample.
[0224] Based on at least two training samples and semantic prompts, the model parameters of the large language model are adjusted to obtain a well-trained large language model.
[0225] In some embodiments, the acquisition module 2010 is used for:
[0226] Determine the stress features of each word in the audio, and output the words with the maximum stress feature as stress keywords.
[0227] In some embodiments, the acquisition module 2010 is used for:
[0228] The stress feature of a word is obtained by weighted summing of at least two of the pronunciation duration, intensity, pitch, and pitch difference in the audio. The pitch difference refers to the difference between the maximum and minimum pitch of the word.
[0229] In this embodiment, semantic keywords and accent keywords are extracted from the audio. Based on these keywords, target animation segments matching the audio are selected from an animation library. Then, an animation is generated based on these target animation segments. This approach considers both the semantic and accent features of the audio during animation generation, ensuring comprehensiveness and resulting in an accurate and comprehensive representation of the audio. Furthermore, the method matches animation segments from a pre-defined animation library, making the quality of the final animation controllable and reliable. Additionally, when matching animation segments from the library, the matching process for keywords with later matching sequences references the target animation segments corresponding to keywords with earlier matching sequences. This avoids conflicts between the identified at least two target animation segments, and the animation generated based on these segments includes more diverse segments with smoother transitions, thus improving overall animation quality.
[0230] Please refer to Figure 21 This diagram illustrates a structural block diagram of a computer device 2100 provided in one embodiment of this application. The computer device 2100 may be a terminal device 10 or a server 20 in an implementation environment, and is used to implement the animation generation method provided in the above embodiments. Specifically:
[0231] Typically, computer device 2100 includes a processor 2110 and a memory 2120.
[0232] Processor 2110 may include one or more processing cores, such as a quad-core processor or a 21-core processor. Processor 2110 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). Processor 2110 may also include a main processor and a coprocessor. The main processor, also known as the central processing unit, is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 2110 may include a GPU, which is responsible for executing the method steps provided in this application. In some embodiments, processor 2110 may also include an AI processor, which is used to handle computational operations related to machine learning.
[0233] The memory 2120 may include one or more computer-readable storage media, which may be non-transitory. The memory 2120 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 2120 are used to store a computer program configured to be executed by one or more processors (such as a GPU) to implement the above-described animation generation method.
[0234] Those skilled in the art will understand that Figure 21 The structure shown does not constitute a limitation on computer device 2100 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0235] This application also provides a computer-readable storage medium storing a computer program, which is loaded and executed by a processor to implement the animation generation method of any of the above implementations.
[0236] This application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the animation generation method of any of the above implementations.
[0237] In some embodiments, the computer program product involved in the present application can be deployed and executed on a computer device, or on multiple computer devices located in one location, or on multiple computer devices distributed in multiple locations and interconnected through a communication network. Multiple computer devices distributed in multiple locations and interconnected through a communication network can form a blockchain system.
[0238] It should be noted that the data collection and processing in this application should strictly comply with the requirements of relevant national laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0239] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.
[0240] All the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here. The above are only optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An animation generation method, characterized in that, The method includes: Obtain multiple keywords at multiple time points in the audio, including semantic keywords and stress keywords. Each semantic keyword corresponds to a semantic category, and each stress keyword is related to the stress feature at the corresponding time point. Based on the target matching order, at least two target animation segments that match the audio are sequentially determined from the animation library. The matching process of the keyword with a later matching order refers to the target animation segment corresponding to the keyword with a earlier matching order. The target matching order is associated with the semantic category of the multiple keywords and semantic keywords. An animation of the audio is generated based on the at least two target animation clips; The animation library includes animation clips corresponding to at least two semantic categories and animation clips corresponding to accent features.
2. The method according to claim 1, characterized in that, The step of determining at least two target animation clips that match the audio from the animation library based on the target matching order includes: Based on the priority of the semantic category of each of the semantic keywords and the priority of the stress feature of each of the stressed keywords, at least two target keywords are determined from the plurality of keywords; According to the target matching order of the at least two target keywords, target animation segments are matched sequentially from the animation library for each target keyword to obtain the target animation segments matched for each target keyword. The target matching order is obtained based on the priority of the semantic category or the priority of the accent feature of each of the at least two target keywords.
3. The method according to claim 2, characterized in that, The determination of at least two target keywords from the plurality of keywords based on the priority of the semantic category of each semantic keyword and the priority of the stress feature of each stress keyword includes: For any keyword, if the interval between the keyword and a keyword with an adjacent time sequence is not less than a first time sequence, the keyword is output as the target keyword. If the interval between the keyword and the keyword with adjacent time sequence is less than a first time sequence, the keyword with higher priority among the keyword and the keyword with adjacent time sequence is output as the target keyword.
4. The method according to claim 2, characterized in that, The method further includes: If there are at least two first keywords among the at least two target keywords, the duration priority of each first keyword is determined based on the interval between each first keyword and the target endpoint of the audio. The target endpoint refers to the endpoint of the two endpoints of the audio that has a shorter interval with the first keyword. The at least two first keywords refer to target keywords with the same priority. The duration priority of each first keyword is negatively correlated with its corresponding interval. Based on the duration priority of each of the at least two first keywords, the at least two first keywords are arranged in a priority order to obtain the target matching order of the at least two target keywords. The priority order refers to the order in which the target keywords other than the at least two first keywords are arranged according to priority.
5. The method according to claim 4, characterized in that, The step of arranging the at least two first keywords into a priority order based on their respective duration priorities to obtain the target matching order of the at least two target keywords includes: If there are at least two first keywords among the at least two target keywords, and at least two second keywords among the at least two first keywords, and both of the at least two second keywords are stressed keywords, the stress priority of each second keyword is determined based on the stress feature of each second keyword. The at least two second keywords refer to first keywords with the same duration priority, and the stress priority of each second keyword is positively correlated with its respective stress feature. Based on the duration priority of each of the at least two first keywords and the accent priority of each of the at least two second keywords, the at least two first keywords are arranged in the priority order to obtain the target matching order of the at least two target keywords.
6. The method according to any one of claims 2-5, characterized in that, Each animation segment in the animation library includes a head movement, movement direction and movement amplitude of a virtual object. The animation library includes at least one animation segment corresponding to each semantic category and at least one animation segment corresponding to an accent feature. The step of matching target animation clips from the animation library sequentially for each target keyword according to the target matching order of the at least two target keywords, to obtain the target animation clip matched for each target keyword, includes: For each target keyword, at least two first animation segments matching the target keyword are determined from the animation library, wherein the semantic category of the first animation segment is the same as the semantic category of the target keyword or the first animation segment is an animation segment corresponding to the accent feature of the target keyword; At least one second animation segment is determined from the at least two first animation segments, wherein the head movements and movement directions of the second animation segment are different from those of the first target animation segment, and the first target animation segment refers to the target animation segment whose target keyword is determined and is adjacent to the target keyword in the time sequence. The target animation segment matching the target keyword is determined from the at least one second animation segment.
7. The method according to claim 6, characterized in that, Each animation segment, arranged from left to right along the timeline, includes an animation start point, a center start point, a center point, a center end point, and an animation end point. The segment between the center start point and the center end point is the central segment of the animation segment. Determining the target animation segment matching the target keyword from the at least one second animation segment includes: The second animation clip that meets the preset conditions is output as the target animation clip that matches the target keyword. The preset conditions include at least one of the following: The duration of the central segment of the second animation segment is less than the interval duration. The interval duration refers to the interval duration between the central segments of the two second target animation segments when the center points of the two second target animation segments are aligned with the time points of their respective target keywords. The two second target animation segments refer to two target animation segments whose target keywords are determined and are adjacent to the time sequence of the target keyword. The duration of the central segment of the second animation clip is less than the duration of the corresponding target keyword in the audio; The switching speed between the center segment of the second animation segment and the center segment of each second target animation segment is less than the preset speed.
8. The method according to claim 6, characterized in that, The method further includes: If the second animation segment is not included in the at least two first animation segments, at least one third animation segment is determined from the at least two first animation segments, wherein the direction of motion of the virtual object in the third animation segment is different from that in the first target animation segment; The target animation segment matching the target keyword is determined from the at least one third animation segment.
9. The method according to claim 6, characterized in that, Determining the target animation segment matching the target keyword from the at least one second animation segment includes: For each target keyword, based on the audio features of the target keyword, a target animation segment from the at least one second animation segment is determined whose motion amplitude of the virtual object matches the audio features of the target keyword, the audio features including at least one of intensity and pitch.
10. The method according to any one of claims 2-9, characterized in that, Each animation segment, arranged from left to right along the timeline, includes an animation start point, a center start point, a center point, a center end point, and an animation end point. The segment between the center start point and the center end point is the central segment of the animation segment. The step of generating the animation based on the at least two target animation clips includes: Based on the timing of the center points of the at least two target animation clips and the time points of their respective target keywords, the at least two target animation clips are spliced together to obtain the animation of the audio.
11. The method according to claim 10, characterized in that, The step of splicing the at least two target animation clips together based on the time sequence of their center points and the time points of their corresponding target keywords to obtain the audio animation includes: The number of the at least two target animation segments is M. For the target animation segment whose timing is at the i-th position among the at least two target animation segments, the following steps are taken, where i is an integer greater than 0 and less than or equal to M: If i=1, the baseline animation and the first target animation segment are merged to obtain a merged animation segment. The baseline animation refers to the animation in which the range of head movements is limited to a preset range. If i is greater than 1 and less than M, based on the baseline animation, the i-th target animation segment is spliced onto the generated fused animation segment; If i = M, based on the baseline animation, the Mth target animation segment is spliced onto the generated fused animation segment to obtain the animation of the audio.
12. The method according to claim 11, characterized in that, The process of fusing the baseline animation and the first target animation segment to obtain the fused animation segment includes: The reference animation and the first transition segment of the first target animation segment are frame-blended to obtain a first blended segment. The reference animation, the first blended segment, the center segment and the second transition segment of the first target animation segment are sequentially spliced together to obtain the merged animation segment. The first transition segment refers to the segment between the animation start point and the center start point, and the second transition segment refers to the segment between the center end point and the animation end point.
13. The method according to claim 11, characterized in that, The step of stitching the i-th target animation segment onto the generated merged animation segment based on the baseline animation includes any of the following: If the interval between the i-th target animation segment and the generated fused animation segment is longer than the second duration, the second transition segment of the last animation segment in the generated fused animation segment is frame-mixed with the reference animation to obtain a second mixed segment. The first transition segment of the i-th target animation segment is frame-mixed with the reference animation to obtain a third mixed segment. The second transition segment of the last animation segment in the generated fused animation segment is removed. The fused animation segment after removal, the second mixed segment, the reference animation, the third mixed segment, the center segment and the second transition segment of the i-th target animation segment are sequentially spliced together to splice the i-th target animation segment onto the generated fused animation segment. If the interval between the i-th target animation segment and the generated fused animation segment is less than the second duration and greater than the third duration, the duration of the second transition segment of the last animation segment is lengthened. Interpolation is performed between the last frame of the lengthened second transition segment and the first frame of the center segment of the i-th target animation segment to obtain a first interpolated segment. The lengthened fused animation segment, the first interpolated segment, the center segment of the i-th target animation segment, and the second transition segment are sequentially spliced together to splice the i-th target animation segment onto the generated fused animation segment. If the interval between the i-th target animation segment and the generated fused animation segment is less than the third duration, interpolation is performed between the last frame of the center segment of the last animation segment and the first frame of the center segment of the i-th target animation segment to obtain a second interpolated segment. The second transition segment of the last animation segment on the generated fused animation segment is removed. The fused animation segment after removal, the second interpolated segment, the center segment of the i-th target animation segment, and the second transition segment are sequentially spliced together to splice the i-th target animation segment onto the generated fused animation segment.
14. The method according to any one of claims 1-13, characterized in that, The acquisition of multiple keywords at multiple time points in the audio includes: The text of the audio is obtained and input into a large language model. The large language model is used to output at least one semantic keyword and the semantic category of each semantic keyword based on semantic prompt information. The semantic prompt information is used to prompt the large language model to obtain the semantics of the text.
15. The method according to claim 14, characterized in that, The training process of the large language model includes: Obtain at least two training samples and the semantic prompt information, wherein each training sample includes a text sample and the semantic categories of at least two keywords in the text sample; Based on the at least two training samples and the semantic prompt information, the model parameters of the large language model are adjusted to obtain a trained large language model.
16. The method according to any one of claims 1-15, characterized in that, The acquisition of multiple keywords at multiple time points in the audio includes: Determine the stress features of each word in the audio, and output the words with the maximum stress features in the audio as the stress keywords.
17. The method according to claim 16, characterized in that, Determining the stress features of each word in the audio includes: The stress feature of the word is obtained by weighted summing of at least two of the pronunciation duration, intensity, pitch and pitch difference of the word in the audio, where the pitch difference refers to the difference between the maximum pitch and the minimum pitch of the word.
18. An animation generation device, characterized in that, The device includes: The acquisition module is used to acquire multiple keywords at multiple time points in the audio. The multiple keywords include semantic keywords and stress keywords. Each semantic keyword corresponds to a semantic category, and each stress keyword is related to the stress feature at the corresponding time point. The determination module is used to sequentially determine at least two target animation segments that match the audio from the animation library based on the target matching order, wherein the matching process of the keyword with a later matching order refers to the target animation segment corresponding to the keyword with a earlier matching order, and the target matching order is associated with the semantic category of the plurality of keywords and semantic keywords; A generation module is used to generate an animation of the audio based on the at least two target animation clips; The animation library includes animation clips corresponding to at least two semantic categories and animation clips corresponding to accent features.
19. A computer device, characterized in that, The computer device includes a processor and a memory, the memory being used to store a computer program, the computer program being loaded by the processor and executed as the animation generation method according to any one of claims 1 to 17.
20. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program for performing the animation generation method according to any one of claims 1 to 17.