Action animation generation method and device, equipment, medium and program product

By acquiring the duration of the text content and the sequence of action frames, and adjusting the action animation to match the text duration, the problem of inconsistent virtual object action playback was solved, improving the expressiveness and interactive experience of virtual objects.

CN121053263APending Publication Date: 2025-12-02TENCENT DIGITAL TIANJIN
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410652388.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-24
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

In existing technologies, when virtual objects narrate dialogue statements, the playback of actions is highly repetitive and the transitions between actions are abrupt, which affects the interactive experience.

Method used

By obtaining the first duration and first action frame sequence corresponding to the text content, the action frame sequence is adjusted to generate an action animation that matches the text content, ensuring that the animation duration is consistent with the text duration and enhancing the contextual fit.

Benefits of technology

It enhances the expressiveness of virtual objects and the efficiency of human-computer interaction, making the motion animation of virtual objects more closely related to text content and improving the realism of interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053263A_ABST
    Figure CN121053263A_ABST
Patent Text Reader

Abstract

The invention discloses an action animation generation method and device, equipment, a medium and a program product, and relates to the technical field of computers. The method comprises the steps of obtaining text content; acquiring a first action frame sequence matched with the text content; and sequence adjustment is performed on the first action frame sequence based on the first duration, a first action animation represented by the virtual object based on the text content is generated, and the animation duration of the first action animation is matched with the first duration. By means of the mode, when the virtual object narrates the text content, the first action animation can be generated by means of the first duration corresponding to the text content and the first action frame sequence corresponding to the text content, the contextual integrating degree between the text content and the first action animation is enhanced, and therefore the expressive force of the virtual object is greatly improved. The method can be applied to various scenes such as cloud technology and artificial intelligence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device, medium, and program product for generating motion animation. Background Technology

[0002] With the development of computer technology, the generation of virtual worlds has received increasing attention. Virtual worlds include various virtual elements such as virtual objects and virtual buildings. The generation of different virtual elements may involve different technologies due to the differences in their implementation effects. Among them, virtual objects in virtual worlds may need to express different postures, such as movement postures and limb movement postures. Therefore, the generation of virtual object movements is very important in the generation technology of virtual worlds.

[0003] In related technologies, when it is necessary to control a virtual object to perform actions during the process of narrating a dialogue statement, the dialogue scenario corresponding to the current dialogue statement is analyzed and a preset action library for the current dialogue scenario is determined. This preset action library stores multiple manually configured preset actions that match the current dialogue scenario. Then, based on the dialogue statement, the action to be performed is randomly selected from the preset action library, so that when the virtual object issues a dialogue statement, an animation of the virtual object performing the action is displayed.

[0004] Actions in the action library typically correspond to a certain playback duration. When the duration of a dialogue statement is less than the action playback duration, the action can only be truncated and played in part. When the duration of a dialogue statement is greater than the action playback duration, the action needs to be played repeatedly until the statement duration is reached. This process suffers from problems such as high repetition rate of action playback and abrupt action transitions, which affect the realism of the dialogue between the user object and the virtual object and impact the user's interactive experience. Summary of the Invention

[0005] This application provides a method, apparatus, device, medium, and program product for generating motion animation. When a virtual object narrates text content, it generates a first motion animation by utilizing a first duration corresponding to the text content and a first action frame sequence corresponding to the text content. This enhances the contextual fit between the text content and the first motion animation, thereby significantly improving the expressiveness of the virtual object. The technical solution is as follows.

[0006] On the one hand, a method for generating motion animation is provided, the method comprising:

[0007] Obtain text content, which is reference information for generating the first motion animation of the virtual object. The text content corresponds to a first duration, which is the duration for the virtual object to narrate the text content.

[0008] Obtain a first action frame sequence that matches the text content, the first action frame sequence being used to generate the action state of the virtual object narrating the text content;

[0009] Based on the first duration, the sequence of the first action frames is adjusted to generate a first action animation of the virtual object based on the text content. The animation duration of the first action animation matches the first duration. The first action animation is the animated performance presented when the virtual object narrates the text content.

[0010] On the other hand, an apparatus for generating motion animation is provided, the apparatus comprising:

[0011] The acquisition module is used to acquire text content, which is reference information for generating the first motion animation of the virtual object. The text content corresponds to a first duration, which is the duration for the virtual object to narrate the text content.

[0012] The acquisition module is also used to acquire a first action frame sequence that matches the text content, the first action frame sequence being used to generate the action state of the virtual object narrating the text content;

[0013] The adjustment module is used to adjust the sequence of the first action frame based on the first duration to generate a first action animation of the virtual object based on the text content. The animation duration of the first action animation matches the first duration, and the first action animation is the animation performance presented by the virtual object when it narrates the text content.

[0014] In an optional embodiment, the adjustment module is further configured to obtain the number of first screen frames corresponding to the first action frame sequence, wherein the number of first screen frames is the number of action screen frames that make up the first action frame sequence; map the first action frame sequence and the number of first screen frames to the latent space corresponding to the encoder, and extract latent feature representations, wherein the latent feature representations are key information representing the action changes of multiple action screen frames from the time dimension and the action change dimension, wherein the multiple action screen frames make up the first action frame sequence; decode the latent feature representations with the first duration as the action generation condition corresponding to the decoder, and generate the first action animation of the virtual object based on the text content, wherein the encoder and the decoder are pre-trained neural network layers involved in the process of generating the first action animation.

[0015] In an optional embodiment, the adjustment module is further configured to obtain the number of second screen frames corresponding to the text content based on the rendering frame rate used when generating the first motion animation and the first duration; extract the quantity feature representation corresponding to the number of second screen frames; concatenate the quantity feature representation and the latent feature representation to obtain a concatenated feature representation; decode the concatenated feature representation and generate the first motion animation of the virtual object based on the text content.

[0016] In an optional embodiment, the adjustment module is further configured to: obtain the number of second frame counts corresponding to the text content based on the rendering frame rate used when generating the first motion animation and the first duration; extract the quantity feature representation corresponding to the number of second frame counts; obtain the attention weights corresponding to the quantity feature representation and the multiple sub-feature representations in the latent feature representation, wherein the attention weights are used to represent the degree of attention paid by the multiple sub-feature representations to the quantity feature representation during the decoding process; perform a weighted summation on the multiple sub-feature representations based on the attention weights to obtain a weighted feature representation; decode the weighted feature representation to generate the first motion animation of the virtual object based on the text content.

[0017] In an optional embodiment, the acquisition module is further configured to acquire multiple action data, the action data including action frame sequences and action text, the action frame sequences describing multiple action screen frames corresponding to the action data, and the action text representing the semantic information represented by the action frame sequences; matching the text content with the action text corresponding to the multiple action data respectively, and determining the first action data matching the text content from the multiple action data, the first action data including the first action frame sequence and the first action text, the first action text having a semantic matching relationship with the text content.

[0018] In an optional embodiment, the acquisition module is further configured to extract the semantic feature representation corresponding to the text content, the semantic feature representation being used to characterize the semantic information of the text content; acquire the text feature representation of the action text corresponding to the plurality of action data respectively, the text feature representation being used to characterize the semantic information of the action text; match the semantic feature representation with the plurality of text feature representations, and determine the first action data that matches the text content from the plurality of action data.

[0019] In an optional embodiment, the acquisition module is further configured to determine the similarity between the semantic feature representation and the plurality of text feature representations, obtain the feature similarity corresponding to the plurality of text feature representations respectively, and the feature similarity is used to characterize the degree of semantic matching between the text content and the action content; and determine the first action data from the plurality of action data based on the plurality of feature similarities.

[0020] In an optional embodiment, the acquisition module is further configured to determine the first action data corresponding to the first text feature representation based on the first text feature representation corresponding to the maximum feature similarity among the plurality of feature similarities; or, if there is a first feature similarity among the plurality of feature similarities that reaches a preset similarity threshold, the first action data is determined based on the first text feature representation corresponding to the first feature similarity, wherein the first text feature representation is the feature representation corresponding to the first action text in the first action data.

[0021] In an optional embodiment, the acquisition module is further configured to generate action description text based on the text content under action restriction conditions, the action description text being used to trigger the first action animation of generating a virtual object, the action description text including restriction information and text content, the restriction information being text information used to describe the action restriction conditions; and to acquire the first action frame sequence matching the action description text, the first action frame sequence having a semantic matching relationship with the text content.

[0022] In an optional embodiment, the apparatus further includes:

[0023] A training module is used to acquire sample text, the sample text corresponding to a sample duration; acquire sample action data based on the sample text, the sample action data including a sample action frame sequence, the sample action frame sequence corresponding to a sample action frame number, the sample action frame number representing the number of action frames in the sample action frame sequence; encode the sample action frame sequence and the sample action frame number through an encoder in a neural network model to obtain a sample latent feature representation; use the sample duration as the action generation condition corresponding to the decoder, decode the sample latent feature representation through the decoder in the neural network model to obtain predicted action data; train the neural network model based on the difference between the sample action data and the predicted action data.

[0024] In an optional embodiment, the training module is further configured to obtain a first prediction loss value between the sample reference data and the predicted action data; obtain a second prediction loss value between the prior distribution and the posterior distribution corresponding to the encoder, wherein the prior distribution is a distribution obtained by the encoder based on prior knowledge, and the posterior distribution is a distribution learned by the encoder based on the latent feature representation of the sample; and train the neural network model with the goal of reducing the first prediction loss value and the second prediction loss value.

[0025] In an optional embodiment, the first motion animation frame is used to present a first object motion of the virtual object;

[0026] The adjustment module is further configured to, when the virtual object performs a second object action within a preset time period after performing a first object action, obtain a second action animation corresponding to the second object action; perform frame interpolation processing on the first action animation and the second action animation to generate an action animation in which the virtual object performs the first object action and the second object action consecutively.

[0027] In an optional embodiment, the adjustment module is further configured to determine the continuation time of the second object action following the first object action performed by the virtual object, wherein the continuation time is the time when the first action animation finishes playing and the second action animation begins playing; obtain a first number of action animation frames before the continuation time from the first action animation, and obtain a second number of action animation frames after the continuation time from the second action animation; and perform frame interpolation processing on the first number of action animation frames and the second number of action animation frames to obtain the action animation.

[0028] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the motion animation generation method as described in any of the embodiments of this application above.

[0029] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, code set, or instruction set is stored therein, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the motion animation generation method as described in any of the embodiments of this application above.

[0030] On the other hand, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the motion animation generation method described in any of the above embodiments.

[0031] The beneficial effects of the technical solutions provided in this application include at least the following:

[0032] Based on the first duration of the virtual object narrating text content, the sequence of the first action frames matching the text content is adjusted to ensure that the generated first action animation better meets the first duration, resulting in the first action animation of the virtual object based on the text content. When the virtual object narrates text content, the first duration corresponding to the text content is used as a generation condition in the animation generation process, and the first action sequence matching the text content is used as auxiliary content in the animation generation process. This ensures that the animation duration of the first action animation generated based on the first action frame sequence matches the first duration, thus strengthening the correlation between the first action animation presented by the virtual object when narrating text content and the text content. This enhances the contextual fit between the text content and the first action animation, significantly improving the expressiveness of the virtual object and also improving human-computer interaction efficiency to a certain extent. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1 This is a structural block diagram of a generation system provided in an exemplary embodiment of this application;

[0035] Figure 2 This is a flowchart of a method for generating motion animation provided in an exemplary embodiment of this application;

[0036] Figure 3 This is a flowchart illustrating the encoding and decoding processes and the resulting first motion animation provided in an exemplary embodiment of this application.

[0037] Figure 4 This is a flowchart of a method for generating motion animation provided in another exemplary embodiment of this application;

[0038] Figure 5This is a flowchart illustrating the training of a neural network model provided in an exemplary embodiment of this application;

[0039] Figure 6 This is a flowchart of a method for generating motion animation provided in another exemplary embodiment of this application;

[0040] Figure 7 This is a schematic diagram of an interface for a first motion animation based on text content display, provided in an exemplary embodiment of this application.

[0041] Figure 8 This is a structural block diagram of a motion animation generation apparatus provided in an exemplary embodiment of this application;

[0042] Figure 9 This is a structural block diagram of a motion animation generation apparatus provided in another exemplary embodiment of this application;

[0043] Figure 10 This is a schematic diagram of the structure of a server provided in an exemplary embodiment of this application. Detailed Implementation

[0044] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0045] First, a brief introduction to the terms used in the embodiments of this application will be given.

[0046] Artificial Intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to have perception, reasoning, and decision-making capabilities. AI technology is a comprehensive discipline involving a wide range of fields, encompassing both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Pre-trained models, also known as large models or foundational models, can be widely applied to downstream tasks in various areas of AI after fine-tuning. AI software technologies mainly include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0047] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP involves natural language, the language people use in daily life, and is closely related to linguistics; it also involves computer science and mathematics. It is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language, the language people use in daily life, and thus it has a close connection with linguistics. Pre-trained models, an important technique for model training in artificial intelligence, evolved from Large Language Models (LLMs) in NLP. After fine-tuning, large language models can be widely applied to downstream tasks. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.

[0048] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning.

[0049] In related technologies, when it's necessary to control virtual objects to perform actions during dialogue, the system analyzes the dialogue scenario corresponding to the current dialogue statement and determines a preset action library for that scenario. This library stores multiple manually configured actions matched to the current dialogue scenario. Then, based on the dialogue statement, an action is randomly selected from this library, resulting in an animation of the virtual object performing that action when it utters a dialogue statement. Actions in the library typically correspond to a certain playback duration. If the dialogue statement's duration is shorter than the action's playback duration, the action is truncated and only a portion is played. If the dialogue statement's duration exceeds the action's playback duration, the action is repeated until the statement's duration is reached. This process suffers from high repetition rates and abrupt action transitions, affecting the realism of the interaction between the user and the virtual object and impacting the user's interactive experience.

[0050] This application provides a method for generating motion animation. When a virtual object narrates text content, it generates a first motion animation using a first duration corresponding to the text content and a first action frame sequence corresponding to the text content. This enhances the contextual fit between the text content and the first motion animation, thereby significantly improving the expressiveness of the virtual object. The motion animation generation method provided in this application can be applied to various human-computer interaction scenarios, such as game interaction scenarios, film and television animation production scenarios, virtual reality (VR) scenarios, augmented reality (AR) scenarios, and medical scenarios. This application does not limit its application to these scenarios.

[0051] In some embodiments, the method for generating motion animation is illustrated by taking the application of the method to a game interaction scene as an example.

[0052] To illustrate, in a game interaction scenario, a game application has a built-in intelligent interaction function. Based on the player's triggering of the intelligent interaction function, a virtual object is displayed. The player can choose the game application's default dialogue options to interact with the virtual object, or they can manually enter text content to have an intelligent dialogue with the virtual object.

[0053] Optionally, the method for generating action animations can be executed by the background program corresponding to the game application, or by the terminal on which the game application is installed. Taking the method of generating action animations executed by the terminal as an example, the terminal considers the dialogue options selected by the player or the text content entered by the player as the text content to be analyzed. The text content is the reference information for generating the first action animation of the virtual object. The text content corresponds to a first duration, which is the duration for the virtual object to narrate the text content, such as 5 seconds. In addition, a first action frame sequence matching the text content can be obtained. This first action frame sequence can roughly represent the action state required by the text content. Then, the first action frame sequence is adjusted based on the first duration to generate the first action animation performed by the virtual object when narrating the text content. During the dialogue interaction between the player and the virtual object, not only can the text content of the virtual object's reply be displayed on the screen, but also the first action animation of the virtual object performing actions based on the text content can be displayed. This complements the display of the first action animation of the virtual object, enriching the dialogue display effect during the dialogue interaction process, enhancing the visual expressiveness, and strengthening the player's interest in the dialogue interaction with the virtual object.

[0054] In some embodiments, the method for generating motion animation is illustrated by applying it to a virtual reality scene.

[0055] In a virtual reality scenario, this includes a completely virtual environment created using computer technology, where users can immerse themselves through access devices (such as head-mounted displays). Virtual objects are pre-configured within the virtual scene, allowing users to interact with them through dialogue, actions, and other means, experiencing a sense of immersion.

[0056] Optionally, the motion animation generation method is executed by a computer device that supports the operation of a virtual environment. This computer device can be a terminal, server, or similar entity. During the execution of the motion animation generation method by the computer device, the user interacts with virtual objects in the virtual environment through voice or text input, conveying certain dialogue statements to the virtual objects. When the computer device generates a corresponding response statement for the virtual object based on the dialogue statements, the response statement is used as text content. This text content serves as reference information for generating the first motion animation of the virtual object, and it corresponds to a first duration, which is the duration for the virtual object to narrate the text content. Furthermore, the computer device also obtains a first sequence of action frames that matches the text content. This first sequence of action frames roughly represents the action state required by the text content. Then, based on the first duration, the first sequence of action frames is adjusted to generate the first motion animation performed by the virtual object when narrating the text content. During the dialogue interaction between the user object and the virtual object, not only can the virtual object be controlled to respond to the user object's dialogue statements through text content, but also a first action animation with consistent animation performance with the text content can be displayed during the virtual object's response. This makes the user object have a more realistic dialogue process with the virtual object in the virtual environment, improves the realism of the virtual environment, and provides the user object with a more immersive experience.

[0057] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the text content, action frame sequences, etc. involved in this application were obtained with full authorization.

[0058] Secondly, the generation system involved in the embodiments of this application will be described. The motion animation generation method provided in the embodiments of this application can be implemented by the terminal alone, by the server, or by the terminal and the server through data interaction. The embodiments of this application do not limit this. Optionally, the method of generating motion animation by interaction between the terminal and the server will be described as an example.

[0059] This is illustrative; please refer to it. Figure 1 The generation system involves a terminal 110 and a server 120, which are connected via a communication network 130.

[0060] In some embodiments, the application installed on the terminal 110 has a built-in dialogue function. During the operation of the application, virtual objects are displayed on the interface of the terminal 110. The dialogue function enables users of the terminal 110 to interact with the virtual objects through dialogue.

[0061] Indicatively, the user inputs a dialogue statement on terminal 110, and terminal 110 invokes the dialogue function to generate text content based on the dialogue statement; or, the user selects from multiple dialogue options provided by the application, and terminal 110 displays text content for responding to the selected dialogue option; or, terminal 110 sends the dialogue statement to server 120 via communication network 130, and server 120 generates text content for responding to the dialogue statement based on the dialogue statement, etc.

[0062] The text content serves as reference information for generating the first motion animation of the virtual object.

[0063] Optionally, in order to improve the vividness of the display interface on the terminal 110 when the virtual object responds to the text content, the first action animation corresponding to the virtual object can be displayed in conjunction with the virtual object's response to the text content. The performance of the first action animation should match the expression of the text content.

[0064] In this context, the text content corresponds to the first duration, which is the duration for the virtual object to narrate the text content. For example, if the virtual object takes approximately 3 seconds to narrate the text content in response to the user object, then 3 seconds will be used as the first duration corresponding to the text content.

[0065] In some embodiments, server 120 generates text content or receives text content sent by terminal 110, and then server 120 analyzes the text content.

[0066] Optionally, server 120 acquires a first action frame sequence that matches the text content.

[0067] The first action frame sequence is used to generate the action state of the virtual object narrating the text content.

[0068] In a schematic way, the first action frame sequence is determined based on the text content. The first action frame sequence is similar to the meaning expressed by the text content and can roughly show the action state of the virtual object when narrating the text content. However, there may be some differences between the first action frame sequence and the text content. Therefore, the first action frame sequence can be adjusted and matched with the text content to generate the action animation.

[0069] In some embodiments, server 120 adjusts the sequence of first action frames based on a first duration to generate a first action animation of a virtual object based on text content.

[0070] Among them, the animation duration of the first action animation matches the first duration, and the first action animation is the animation performance presented when the virtual object narrates the text content.

[0071] In illustrative terms, the first duration, as the duration corresponding to the text content, can constrain the animation duration of the generated first motion animation. Therefore, in the process of generating the first motion animation of the virtual object based on the text content, the first duration is used as a constraint to adjust the sequence of the first motion frame that matches the text content in order to generate the first motion animation of the virtual object based on the text content.

[0072] Optionally, the animation duration of the first action animation being the same as the first duration is considered a match between the animation duration and the first duration, so that the start and end times of the virtual object narrating the text content correspond to the start and end times of the first action animation, respectively; or, the animation duration of the first action animation being within a preset duration range corresponding to the first duration is considered a match between the animation duration and the first duration. For example, if the first duration is 5 seconds and the preset duration range corresponding to the first duration is within a 0.5-second error range, the animation duration of the first action animation being 4.8 seconds is considered a match between the animation duration and the first duration, etc.

[0073] In some embodiments, the server 120 sends the animation data corresponding to the first motion animation to the terminal 110 via the communication network 130, thereby rendering and displaying the first motion animation on the interface of the terminal 110.

[0074] It is worth noting that the aforementioned terminals include, but are not limited to, mobile terminals such as mobile phones, tablets, portable laptops, smart voice interaction devices, smart home appliances, and in-vehicle terminals, as well as desktop computers; the aforementioned servers can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0075] Cloud technology refers to a hosting technology that unifies hardware, applications, networks, and other resources within a wide area network (WAN) or local area network (LAN) to achieve data computation, storage, processing, and sharing. Based on the cloud computing business model, cloud technology encompasses network technology, information technology, integration technology, management platform technology, and application technology. It can form resource pools, providing flexible and convenient on-demand access.

[0076] In some embodiments, the server described above can also be implemented as a node in a blockchain system.

[0077] Based on the above-described terminology and application scenarios, the method for generating motion animations provided in this application will be explained, taking the application of this method to a server as an example. Figure 2 As shown, the method includes the following steps 210 to 230.

[0078] Step 210: Obtain the text content.

[0079] In illustrative terms, text content is composed of text characters, which include at least one of various character types such as Chinese characters, English letters, and punctuation marks.

[0080] Optionally, the text content is the default configuration content.

[0081] To illustrate, the terminal application has multiple response texts and multiple dialogue texts built-in. The dialogue texts are the content used to answer the response texts. During the operation of the application, the terminal displays the dialogue texts for the user to select. Based on the dialogue text selected by the user, the response text used to answer the dialogue text is determined and used as the text content.

[0082] Optionally, the text content is generated based on the dialogue text.

[0083] In illustrative terms, the terminal has an intelligent dialogue function, which is supported by a pre-trained dialogue model. Users input dialogue text through the terminal, and the terminal responds and generates reply text based on the dialogue text using the intelligent dialogue function. Alternatively, the terminal sends the dialogue text to the server and receives the reply text generated by the server, and then uses the reply text as the text content.

[0084] The text content serves as reference information for generating the first motion animation of the virtual object.

[0085] In illustrative terms, a virtual object is an object used to describe text content; a virtual object is implemented as an object displayed in a game application; or a virtual object is implemented as an object displayed in virtual reality technology; or a virtual object is an object designed during the development of a virtual environment, etc.

[0086] Optionally, virtual objects can be presented not only as static states, but also as dynamic states. Dynamic states include at least one of various states that are presented as activities, such as limb movement states (e.g., waving arms, raising legs, etc.), head movement states (e.g., nodding, shaking heads, etc.), and movement states (e.g., running, walking, etc.).

[0087] Considering that if corresponding actions are added during the process of virtual objects narrating text content, that is, if the virtual objects present a dynamic state corresponding to the text content while controlling the virtual objects to narrate text content, the virtual objects will be more human-like and the realism of the interaction between the virtual objects and the user objects will be improved. Therefore, the first action animation of the virtual objects can be generated when the text content is known.

[0088] In other words, the first motion animation is the action that the virtual object needs to perform while narrating the text content. Therefore, the generated first motion animation should match the text content, so that while the virtual object is narrating the text content, the first motion animation that matches the text content is displayed on the terminal interface.

[0089] Among them, the text content corresponds to the first duration, which is the duration for which the virtual object narrates the text content.

[0090] Optionally, the first duration corresponding to the text content is a preset duration. For example, different text contents are preset to correspond to a different duration, and the durations corresponding to different text contents may be the same or different.

[0091] For illustrative purposes, the default duration for the text content "A new store opened at location X today, let's go check it out" is 3 seconds, representing the default duration for the virtual object to narrate the text content, etc.

[0092] Optionally, the first duration corresponding to the text content is determined based on the number of characters in the text content.

[0093] This is an illustrative example of a pre-defined correspondence between the number of characters and the duration. After determining the text content, the first duration corresponding to the text content is determined based on the number of characters in the text content and the pre-defined correspondence. For example, the pre-defined duration is 1 second for characters in the range (1, 5), 2 seconds for characters in the range (6, 10), and 3 seconds for characters in the range (11, 15). Taking the generated text content as an example, if the generated text content is "Of course I like it here, these are all memories," and the current text content has 14 characters, then the first duration corresponding to the text content is determined to be 3 seconds, etc.

[0094] It is worth noting that the above are merely illustrative examples, and the embodiments of this application are not limited thereto.

[0095] Step 220: Obtain the first action frame sequence that matches the text content.

[0096] To illustrate, after obtaining the text content, a first action frame sequence that matches the text content is found. The first action frame sequence is used to generate the action state of the virtual object that narrates the text content.

[0097] Optionally, the text semantics corresponding to the text content are analyzed, and an action frame sequence expressing the current text semantics is selected from multiple action frame sequences based on the text semantics as the first action frame sequence.

[0098] Indicatively, multiple action frame sequences are pre-acquired content, where each action frame sequence includes at least one action frame, which is the content that makes up the action frame.

[0099] Optionally, each action frame sequence corresponds to a text label. When searching for the first action frame sequence from multiple action frame sequences based on text semantics, the text semantics are semantically matched with the multiple text labels respectively, so as to match the action frame sequences with similar or identical semantics from multiple action frame sequences as the first action frame sequence corresponding to the text content.

[0100] Step 230: Adjust the sequence of the first action frame based on the first duration to generate the first action animation of the virtual object based on the text content.

[0101] As an illustration, considering that the performance duration of the first action frame sequence may differ from the first duration, after determining the first action frame sequence that matches the text content, the first action frame sequence can also be adjusted based on the first duration to make the performance duration of the first action frame sequence more consistent with the first duration, thereby generating the first action animation.

[0102] Among them, the animation duration of the first action animation matches the first duration.

[0103] In illustrative terms, the first motion animation includes multiple motion animation frames. The duration of the first motion animation, which is composed of multiple motion animation frames, is the animation duration. Since the first motion animation is obtained under the constraint of the first duration, the animation duration and the first duration are matched.

[0104] Optionally, results with the same animation duration as the first duration are considered to be mutually matched. For example, if the first duration is 5 seconds, and the generated first motion animation also has a duration of 5 seconds, then the animation duration matches the first duration.

[0105] Optionally, results whose animation duration falls within a preset interval corresponding to the first duration are considered as matching results. For example, if the first duration is 5 seconds and the preset interval corresponding to the first duration is (4.5, 5.5), then if the animation duration of the generated first motion animation is 5.2 seconds, then the animation duration matches the first duration, etc.

[0106] The first motion animation is the animated presentation of the virtual object narrating the text content.

[0107] As an illustration, in addition to displaying text content and virtual objects on the terminal interface, the first animation of the virtual object narrating the text content will also be displayed. For example, if the text content is "Hi, how are you?", the first animation will be the virtual object waving and smiling.

[0108] It is worth noting that the above are merely illustrative examples, and the embodiments of this application are not limited thereto.

[0109] In summary, when virtual objects narrate text content, using the first duration corresponding to the text content as the generation condition in the animation generation process, and using the first action sequence matching the text content as auxiliary content in the animation generation process, the animation duration of the first action animation generated based on the first action frame sequence matches the first duration. This makes the first action animation presented by the virtual object when narrating text content more relevant to the text content, and strengthens the contextual fit between the text content and the first action animation. This not only greatly enhances the expressiveness of virtual objects, but also improves the efficiency of human-computer interaction to a certain extent.

[0110] In an optional embodiment, during the process of adjusting the sequence of first action frames based on a first duration to generate a first motion animation, the motion animation generation method is executed using a pre-trained encoder and a pre-trained decoder, with the first duration serving as a constraint condition for the decoder during decoding, thereby obtaining the first motion animation under the constraint of the first duration. (Illustrative example, such as...) Figure 3 As shown above, Figure 2 Step 230 shown can also be implemented as steps 310 to 330.

[0111] Step 310: Obtain the number of first screen frames corresponding to the first action frame sequence.

[0112] Indicatively, the first action frame sequence includes multiple action frame images, each of which is implemented as an image. The multiple action frame images are arranged sequentially according to the image arrangement order to obtain the first action frame sequence. Playing the first action frame sequence is the process of playing multiple action frame images in the image arrangement order.

[0113] Optionally, when adjusting the sequence of the first action frame sequence based on the first duration, the number of first screen frames corresponding to the first action frame sequence is first obtained, where the number of first screen frames is the number of action screen frames that make up the first action frame sequence.

[0114] Indicatively, the number of first frame frames corresponding to the first action frame sequence is 180, meaning that the first action frame sequence consists of 180 action frame frames.

[0115] Optionally, the first action frame sequence corresponds to a second duration, which is the duration of playing the first action frame sequence.

[0116] For illustrative purposes, the second duration corresponding to the first action frame sequence is 3.66 seconds, which means that it takes a total of 3.66 seconds from the first action frame in the first action frame sequence to the last action frame.

[0117] There is a correlation between the number of first screen frames corresponding to the first action frame sequence and the second duration, and this correlation is based on the rendering frame rate.

[0118] Indicatively, based on the difference in rendering frame rate, the first action frame sequence with the first number of frame counts will correspond to different second durations. The second duration is the ratio of the number of action frame counts in the first action frame sequence to the rendering frame rate.

[0119] The rendering frame rate is used to represent the number of frames displayed per unit of time, which is usually in seconds. A frame is the smallest unit that makes up a single still image in a sequence of action frames. If the rendering frame rate is expressed as the number of frames played per second, it is expressed as Frames Per Second, abbreviated as FPS.

[0120] For illustrative purposes, a rendering frame rate of 60 FPS means that 60 frames are displayed per second. If the number of first frames corresponding to the first action frame sequence is 180, it means that the first action frame sequence includes 180 action frames. Based on the rendering frame rate and the number of first frames, the second duration corresponding to the first action frame sequence is determined to be 3 seconds, which means that the duration of playing multiple action frames in the first action frame sequence is 3 seconds.

[0121] Similarly, if the rendering frame rate is known to be 60 FPS and the second duration corresponding to the first action frame sequence is 3 seconds, then the number of first screen frames corresponding to the first action frame sequence can be determined to be 180; that is, the number of first screen frames is the product of the rendering frame rate and the second duration, which will not be elaborated here.

[0122] Step 320: Map the first action frame sequence and the number of first screen frames to the latent space corresponding to the encoder, and extract the latent feature representation.

[0123] Indicatively, after obtaining the first action frame sequence and the number of first screen frames corresponding to the first action frame sequence, considering that the first action frame sequence represents the changes between multiple action screen frames, and the number of first screen frames represents the number of action screen frames under the current rendering frame rate, the first action frame sequence and the number of first screen frames are used together as the input information of the encoder.

[0124] The encoder is a pre-trained neural network layer. The first action frame sequence and the number of first screen frames are input into the encoder so that the encoder maps the first action frame sequence and the number of first screen frames to a low-dimensional latent space. The latent space is an abstract, continuous feature representation space that can capture the key features and changing dimensions of the first action frame sequence and the number of first screen frames input to the encoder.

[0125] Optionally, the encoder is a neural network composed of multiple layers, including at least one of fully connected layers, convolutional layers, and recurrent layers. The first action frame sequence and the number of first screen frames are input into the encoder, thereby gradually extracting features from the input information using the multi-layered neural network within the encoder to remove unnecessary information and retain important features. Finally, a low-dimensional latent feature representation is extracted through a fully connected layer. The dimension of this latent feature representation is much smaller than the dimension of the input data (i.e., the first action frame sequence and the number of first screen frames).

[0126] Among them, latent feature representation is the key information that characterizes the changes in motion from the time dimension and the motion change dimension for multiple motion frame images.

[0127] Indicatively, the time dimension reflects the duration and rate of change of the first action frame sequence. Taking the first action frame sequence as an example, analyzing multiple action frames from the time dimension allows us to analyze the rate of change and duration of the first action from the perspective of the number of first frames. For example, when the encoder analyzes the first action frame sequence and the number of first frames from the time dimension, it analyzes the duration and rate of change of the first action corresponding to the first action frame sequence separately and represents them using different vectors.

[0128] Indicatively, the motion change dimension is used to describe the key changes in the execution of the first action corresponding to the first action frame sequence, such as the start of the first action, important transformations during the execution of the first action, and the state at the end of the first action. These processes can be achieved by analyzing the spatial characteristics of the first action frame sequence, such as analyzing the action effects at different action moments in different first action frame sequences, and then comparing the changes in action effects corresponding to multiple action moments. This achieves the purpose of analyzing multiple action frames in terms of the motion change dimension, such as by representing the motion change dimension using at least one vector.

[0129] Optionally, the latent feature representation is based on key information representing the changes in motion across multiple motion frames from both temporal and motion change dimensions. Therefore, at least one vector corresponding to the temporal dimension and at least one vector corresponding to the motion change dimension are obtained, and the latent feature representation is obtained by combining multiple vectors. For example, at least one vector corresponding to the temporal dimension and the motion change dimension are concatenated, and the concatenated vector is compressed to a low dimension to obtain the latent feature representation.

[0130] In other words, the encoder encodes the first action frame sequence and the number of first screen frames, and outputs a latent feature representation.

[0131] Step 330: Decode the latent feature representation using the first duration as the action generation condition corresponding to the decoder, and generate the first action animation of the virtual object based on the text content.

[0132] The encoder and decoder are neural network layers involved in generating the first motion animation; the decoder is also a pre-trained neural network layer.

[0133] Optionally, the latent feature representation output by the encoder is used as the input data of the decoder. The decoder performs decoding processing on the latent feature representation to restore the low-dimensional latent feature representation to high-dimensional data, thereby generating a first motion animation. The first motion animation is an animated video composed of multiple motion animation frames.

[0134] Indicatively, decoders typically employ neural networks, including at least one of several types such as convolutional layers (CNNs), deconvolutional layers or transpose convolutional layers, dense layers (DFs), activation layers, and batch normalization layers. The neural network in the decoder progressively upsamples the low-dimensional latent feature representations, adding details layer by layer until the original data's dimension and structure are restored, resulting in the generated initial motion animation.

[0135] In some embodiments, during the process of generating the first motion animation through the decoder, decoding conditions are set for the decoder, thereby enabling a targeted decoding process in the decoding process; for example, a first duration is introduced as an action generation condition (i.e., decoding condition), thereby introducing a first duration as a condition variable in addition to the input decoding feature representation, which can work together with the latent feature representation to affect the generation of the first motion animation when the decoder performs decoding processing.

[0136] Optionally, duration is introduced as an action generation condition during the training of the decoder, enabling the decoder to learn effectively step by step when action generation conditions exist. This allows the trained decoder to perform targeted decoding based on latent feature representation and the first duration during the decoding process, and to decode the first action animation.

[0137] Indicatively, the first motion animation is the animation content generated based on the first motion frame sequence. The motion effect it expresses is quite similar to the motion effect of the first motion frame sequence. That is, the overall structure of the generated first motion animation is similar to the overall structure of the first motion frame sequence. However, there are still some differences between the two. These differences are determined by the training and prediction processes of the encoder and decoder.

[0138] In an optional embodiment, the number of second screen frames corresponding to the text content is obtained based on the rendering frame rate of the first motion animation and the first duration.

[0139] The rendering frame rate is used to represent the number of frames displayed per unit of time. The rendering frame rate of the first motion animation is a preset value used to indicate the rendering frame rate used when displaying the first motion animation.

[0140] Optionally, based on the rendering frame rate and the first duration corresponding to the text content, the number of second screen frames corresponding to the text content can be determined. The number of second screen frames is the target number of motion screen frames when generating the first motion animation, that is, the value that the number of screen frames of the first motion animation is expected to match.

[0141] To illustrate, the product of the rendering frame rate and the first duration is calculated to obtain the number of second screen frames corresponding to the text content. For example, if the rendering frame rate used to display the second animation is predetermined to be 60 FPS and the first duration is 4 seconds, then the number of second screen frames corresponding to the text content is 240. That is, it is desired that the number of animation frames generated in the first animation matches 240. For example, it is desired that the number of animation frames in the first animation is 240, or that the number of animation frames in the first animation is within a certain value range of around 240, etc.

[0142] In an optional embodiment, the quantity feature representation corresponding to the number of second frame images is extracted.

[0143] In a schematic manner, during the process of using the first duration as the action generation condition of the decoder and decoding the latent feature representation, the number of second frame segments corresponding to the first duration is used as the input information of the decoder. The feature representation corresponding to the number of second frame segments is extracted by the decoder to obtain the quantitative feature representation corresponding to the number of second frame segments.

[0144] In an optional embodiment, the quantity feature representation and the latent feature representation are spliced ​​together to obtain the spliced ​​feature representation.

[0145] To illustrate, after extracting the quantitative feature representation, the quantitative feature representation is concatenated with the latent feature representation input to the decoder to obtain the concatenated feature representation.

[0146] Optionally, the latent feature representation can be concatenated after the quantitative feature representation to obtain the concatenated feature representation; or, the quantitative feature representation can be concatenated after the latent feature representation to obtain the concatenated feature representation, etc.

[0147] In an optional embodiment, the splicing feature representation is decoded and a first motion animation of the virtual object based on the text content is generated.

[0148] In illustrative terms, the splicing feature representation based on quantitative feature representation can demonstrate a certain duration constraint effect, while the latent feature representation can demonstrate the change effect of multiple action frames in the first action frame sequence. Therefore, by decoding the splicing feature representation obtained after the fusion of the two, the first action animation of the virtual object based on the text content can be generated under the action state constraint of the first action frame sequence.

[0149] It is worth noting that the above are merely illustrative examples, and the embodiments of this application are not limited thereto.

[0150] In an optional embodiment, the number of second screen frames corresponding to the text content is obtained based on the rendering frame rate and first duration of the first motion animation; and the quantitative feature representation corresponding to the number of second screen frames is extracted.

[0151] In a schematic way, the number of second-frames is used as the input information of the decoder. The decoder extracts the feature representation corresponding to the number of second-frames to obtain the quantitative feature representation corresponding to the number of second-frames.

[0152] In an optional embodiment, attention weights corresponding to multiple sub-feature representations in the quantitative feature representation and the latent feature representation are obtained.

[0153] The attention weight is used to represent the degree of attention that multiple sub-feature representations pay to the quantity feature representation in the latent feature representation during the decoding process. These multiple sub-feature representations constitute the latent feature representation.

[0154] Indicatively, a latent feature representation includes multiple sub-feature representations. Sub-feature representations can be understood as feature representations corresponding to different information expressed by the latent feature representation. For example, a latent feature representation may include sub-feature representations that express the magnitude of action changes, as well as sub-feature representations that express the direction of action changes, etc.

[0155] In some embodiments, after determining the multiple sub-feature representations that make up the potential feature representation, the degree of attention each sub-feature representation pays to the quantitative feature representation is determined, thereby obtaining the attention weights corresponding to the multiple sub-feature representations respectively.

[0156] Indicatively, after determining the multiple sub-feature representations that make up the latent feature representation, the similarity between each sub-feature representation and the quantitative feature representation is calculated as an attention weight. For example, the dot product between the sub-feature representation and the quantitative feature representation can be calculated, or a similarity calculation model can be used to calculate the similarity between the two, and then the calculated similarity is used as the attention weight to obtain the attention weights corresponding to the multiple sub-feature representations.

[0157] Optionally, the similarity between each sub-feature representation and the quantitative feature representation is calculated as an attention score, and the attention score is normalized to obtain the attention weight.

[0158] To illustrate, an activation function (such as the softmax function) is used to normalize the attention scores corresponding to the multiple sub-feature representations, resulting in attention weights corresponding to the multiple sub-feature representations.

[0159] In an optional embodiment, multiple sub-feature representations are weighted and summed based on attention weights to obtain a weighted feature representation.

[0160] To illustrate, the attention weights are values ​​obtained after normalization. Therefore, the attention weights corresponding to each sub-feature representation are values ​​between 0 and 1. After obtaining the attention weights corresponding to multiple sub-feature representations, the sub-feature representations are multiplied by their corresponding attention weights to obtain the adjusted feature representations corresponding to the multiple sub-feature representations. The adjusted feature representations are then summed to obtain the weighted feature representation.

[0161] In an optional embodiment, the weighted feature representation is decoded to generate a first motion animation of the virtual object based on the text content representation.

[0162] In illustrative terms, the weighted feature representation reflects the degree of attention paid by different sub-feature representations to the quantitative feature representation. This enables the decoder to leverage the duration constraint effect expressed by the quantitative feature representation and perform decoding based on the weighted feature representation. This achieves the goal of generating the first motion animation of the virtual object based on the text content under the action state constraint of the first action frame sequence.

[0163] In an optional embodiment, the duration feature representation corresponding to the first duration is extracted; the duration feature representation and the latent feature representation are concatenated to obtain the concatenated feature representation; the concatenated feature representation is decoded and a first motion animation of the virtual object based on the text content is generated.

[0164] In an optional embodiment, the duration feature representation corresponding to the first duration is extracted; the attention weights corresponding to multiple sub-feature representations in the duration feature representation and the potential feature representation are obtained respectively; the multiple sub-feature representations are weighted and summed by the attention weights to obtain the weighted feature representation; the weighted feature representation is decoded to generate the first motion animation of the virtual object based on the text content.

[0165] It is worth noting that the above are merely illustrative examples, and the embodiments of this application are not limited thereto.

[0166] In summary, when virtual objects narrate text content, using the first duration corresponding to the text content as the generation condition in the animation generation process, and using the first action sequence matching the text content as auxiliary content in the animation generation process, the animation duration of the first action animation generated based on the first action frame sequence matches the first duration. This makes the first action animation presented by the virtual object when narrating text content more relevant to the text content, and strengthens the contextual fit between the text content and the first action animation. This not only greatly enhances the expressiveness of virtual objects, but also improves the efficiency of human-computer interaction to a certain extent.

[0167] In this embodiment, an encoder and decoder are used to process a first duration to obtain the content of a first motion animation. The encoder can obtain a latent feature representation expressing deep information of multiple motion frame frames based on the input first motion frame sequence. Then, the latent feature representation and the first duration corresponding to the text content are input into the decoder for decoding processing. This allows the decoder to generate the motion animation based on the latent feature representation under the constraint of the first duration, and obtain a first motion animation whose animation duration matches the first duration. Using the first duration as a given condition for the decoder's decoding processing allows the decoder to perform more targeted sequence adjustment and sequence learning on the first motion frame sequence based on the duration limit of the first duration, thereby generating a first motion animation with motion effects similar to the first motion frame sequence. Moreover, this first motion animation can more accurately match the text content, improving the matching degree of the first motion animation presented when the virtual object narrates the text content.

[0168] In an optional embodiment, when obtaining the first action frame sequence based on text content, semantic matching is performed between the text content and multiple pre-obtained action frame sequences to determine the first action frame sequence that matches the text content from the multiple action frame sequences. The process then proceeds by adjusting subsequent sequences based on the first action frame sequence and generating the first motion animation. (Illustrative example, such as...) Figure 4 As shown above, Figure 2 The illustrated embodiment can also be implemented as follows: steps 410 to 440; wherein Figure 2 Step 220 shown can also be implemented as steps 420 to 430.

[0169] Step 410: Obtain the text content.

[0170] Among them, the text content is the reference information for generating the first motion animation of the virtual object, and the text content corresponds to the first duration, which is the duration for the virtual object to narrate the text content.

[0171] For illustrative purposes, the text content is the default configuration content; or, the text content is generated based on the dialogue text, etc.

[0172] Step 420: Obtain multiple action data.

[0173] The motion data includes motion frame sequences and motion text. The motion frame sequences describe multiple motion frames corresponding to the motion data, and the motion text represents the semantic information represented by the motion frame sequences.

[0174] To illustrate, a set of motion data includes not only a sequence of motion frames but also motion text that semantically describes the sequence of motion frames. For example, motion data 1 includes motion frame sequence a, which contains multiple motion frames that, when combined, represent a "waving" motion. Motion data 1 also includes motion text A, which is the text obtained by semantically describing motion frame sequence a, such as "waving".

[0175] Optionally, multiple motion data representing different motion forms can be obtained based on a pre-prepared motion acquisition process; or, multiple motion data can be obtained based on motion capture technology; or, multiple motion data can be obtained based on manual creation by an animator, etc.

[0176] Step 430: Match the text content with the action text corresponding to multiple action data respectively, and determine the first action data that matches the text content from the multiple action data.

[0177] To illustrate, when searching for the first action frame sequence to assist in the generation of motion animation from multiple action data based on text content, semantic matching is performed between the text content and the corresponding action text of each action data based on the action text in the action data, so as to determine the first action data that matches the text content from the multiple action data.

[0178] Among them, the first action data corresponds to the first action frame sequence and the first action text; there is a semantic matching relationship between the first action text and the text content.

[0179] In an optional embodiment, semantic feature representations corresponding to the text content are extracted.

[0180] Among them, semantic features represent the semantic information used to characterize the text content.

[0181] In a schematic way, the text content is input into the feature extraction network, and semantic feature representation is extracted from the feature extraction network. Semantic feature representation is to convert the text content into a numerical form that can reflect its semantic information. Therefore, semantic feature representation is used to capture the deeper meaning, emotion, theme, context and other semantic information in the text content.

[0182] In an optional embodiment, text feature representations of the action texts corresponding to multiple action data are obtained.

[0183] Among them, text features represent the semantic information used to characterize action text.

[0184] As an illustration, after obtaining multiple action data, features can be extracted from the action texts corresponding to each action data, thereby obtaining text feature representations corresponding to each action text; the text feature representations corresponding to different action texts may differ based on semantic information such as their emotion, theme, and deeper meaning.

[0185] In an optional embodiment, semantic feature representations are matched with multiple text feature representations to determine first action data that matches the text content from multiple action data.

[0186] In a schematic manner, the semantic feature representation corresponding to the text content is matched with the text feature representations corresponding to multiple action texts respectively, so as to determine the text feature representation that matches the semantic feature representation from the multiple text feature representations, and the action data to which the action text corresponding to the text feature representation belongs is used as the first action data that matches the text content.

[0187] In some embodiments, the similarity between semantic feature representations and multiple text feature representations is determined to obtain feature similarity corresponding to each of the multiple text feature representations.

[0188] To illustrate, after obtaining the semantic feature representation and multiple text feature representations, the cosine similarity between the semantic feature representation and each text feature representation is calculated as the feature similarity corresponding to the text feature representation. The cosine similarity is used to measure the distance between the semantic feature representation and the text feature representation in the vector space, and can reflect the matching between the text content corresponding to the semantic feature representation and the action text corresponding to the text feature representation.

[0189] Optionally, a higher feature similarity indicates a higher degree of matching between the text content corresponding to the semantic feature representation and the action text corresponding to the text feature representation; conversely, a higher feature similarity indicates a lower degree of matching between the text content corresponding to the semantic feature representation and the action text corresponding to the text feature representation.

[0190] In some embodiments, a first action data is determined from multiple action data based on multiple feature similarities.

[0191] To illustrate, after obtaining the feature similarity corresponding to multiple action texts, the multiple action data are filtered based on the feature similarity to obtain the first action data that matches the text content. In the first action data, there is a semantic matching relationship between the first text feature representation and the semantic feature representation corresponding to the first action text.

[0192] Optionally, based on the first text feature representation corresponding to the largest feature similarity among multiple feature similarities, the first action data corresponding to the first text feature representation is determined.

[0193] In a schematic way, the text feature representation with the highest feature similarity among multiple text feature representations is determined as the first text feature representation. The first text feature representation is the feature representation of the first action text. The first action data to which it belongs is determined based on the first action text.

[0194] Optionally, if a first feature similarity among multiple feature similarities reaches a preset similarity threshold, the first action data is determined based on the first text feature representation corresponding to the first feature similarity.

[0195] In a schematic way, a preset similarity threshold is determined in advance, and multiple feature similarities are compared with the similarity threshold. If there is a first feature similarity that reaches the similarity threshold, the text feature representation corresponding to the first feature similarity is determined as the first text feature representation. The first text feature representation is the feature representation of the first action text in the first action data, and the first action data to which it belongs is determined based on the first action text.

[0196] Optionally, the first text feature representation can be one text feature representation or multiple text feature representations; correspondingly, the first action data can be one action data or multiple action data. When the first action data is multiple action data, an action frame sequence can be randomly selected from the action frame sequences corresponding to the multiple action data as the first action frame sequence, which is not limited here.

[0197] In an optional embodiment, action description text is generated based on the text content under action constraints.

[0198] The action description text is used to trigger the first action animation of the generated virtual object. The action description text includes limiting information and text content.

[0199] In illustrative terms, constraint information is textual information used to describe the restrictions on actions. For example, constraint information could be implemented as "You are an enthusiastic and lively virtual partner. Please match the following sentence with corresponding body language descriptions. These actions must be appropriate for the context of the dialogue, and there should be no physical contact with the main virtual object, and no facial expressions." This constraint information is used to limit the detailed analysis of the text content, so as to extract more suitable text information for obtaining action frame sequences and generate action description text from the known text content.

[0200] Optionally, the text content is appended to the constraint information and sent to the text generation model, which can use its rich knowledge to generate motion description text suitable for generating motion animations.

[0201] To illustrate, the text generation model is implemented as a large language model. Large language models are a type of natural language model, but they are usually larger and more complex than traditional language models. They typically have billions or even hundreds of billions of parameters and are machine learning models trained on massive amounts of data that can understand and generate natural language text. Large language models are usually based on deep learning techniques, especially the Transformer architecture, which can capture complex patterns and dependencies in text. Therefore, large language models are widely used in natural language processing tasks such as text generation, language translation, and question answering.

[0202] In some embodiments, a first action data is matched from multiple action data based on the action description text.

[0203] In a schematic manner, the semantic feature representation of the action description text is extracted; the text feature representation of the action text corresponding to multiple action data is obtained; the semantic feature representation of the action is matched with the multiple text feature representations to determine the first action data that matches the text content from the multiple action data.

[0204] Among them, the first action data includes a semantic matching relationship between the first action text and the action description text.

[0205] Step 440: Adjust the sequence of the first action frame based on the first duration to generate the first action animation of the virtual object based on the text content.

[0206] Among them, the animation duration of the first action animation matches the first duration, and the first action animation is the animation performance presented when the virtual object narrates the text content.

[0207] In an optional embodiment, in addition to adjusting the sequence of the first action frame sequence using the encoder and decoder described above, the first action animation can also be obtained by adjusting the action frame frames in the first action frame sequence.

[0208] Indicatively, at least one action frame in the first action frame sequence is repeatedly played to obtain a first action animation whose animation duration matches the first duration; or, at least one action frame in the first action frame sequence is deleted to obtain a first action animation whose animation duration matches the first duration, etc.

[0209] In some embodiments, the rendering frame rate used in generating the first motion animation is determined; the number of second screen frames corresponding to the text content is determined based on the rendering frame rate and the first duration; the sequence of the first motion frame is adjusted based on the number of second screen frames to obtain the first motion animation of the virtual object based on the text content.

[0210] The rendering frame rate is used to represent the number of frames displayed per unit time; the second frame count is used to represent the number of action frames displayed within the first duration under the rendering frame rate.

[0211] Optionally, based on the number of second screen frames and the number of first screen frames corresponding to the first action frame sequence, the first action frame sequence is adjusted to obtain the first action animation of the virtual object based on the text content.

[0212] Optionally, if the number of first frame frames is less than the number of second frame frames, the number of first frame frames corresponding to the first action frame sequence is increased to the number of second frame frames to obtain the first action animation.

[0213] For illustration purposes, the number of frames in the first screen is 50, and the number of frames in the second screen is 60. At this point, the first action frame sequence cannot yet reach the duration standard of the first action animation (because there are relatively few action animation frames in the first action frame sequence). Therefore, the number of frames in the first screen corresponding to the first action frame sequence can be increased. For example, compare the importance of multiple action animation frames in the first action frame sequence, and repeatedly play at least one of the most important action frames until the number of frames in the second screen is reached and the first action animation is obtained.

[0214] In other words, compared with the first action animation, some action frames of higher importance are repeated in the first action animation.

[0215] Optionally, if the number of first frame frames is greater than the number of second frame frames, the number of first frame frames corresponding to the first action frame sequence is reduced to the number of second frame frames to obtain the first action animation.

[0216] For illustration, the number of frames in the first screen is 60 and the number of frames in the second screen is 50. At this time, the first action frame sequence cannot reach the duration standard of the first action animation (because there are fewer action animation frames in the first action frame sequence). Therefore, the number of frames in the first screen corresponding to the first action frame sequence can be reduced. For example, compare the importance of multiple action animation frames in the first action frame sequence and delete at least one action screen frame with the lowest importance until the number of frames in the second screen is reached and the first action animation is obtained.

[0217] In other words, compared with the first motion animation, some less important motion frames were deleted.

[0218] It is worth noting that the above are merely illustrative examples, and the embodiments of this application are not limited thereto.

[0219] In an optional embodiment, the first motion animation corresponds to the first object action, which represents the action performed by the virtual object in the first motion animation; when the first object action is followed immediately by the second object action, the second motion animation corresponding to the second object action is obtained.

[0220] Optionally, the text content corresponding to the first duration is referred to as the first text content. If the first text content is immediately followed by the second text content, it is considered as the second object action following the first action. Based on this, the second text content and the second action frame sequence matching the second text content are obtained. Then, the second action frame sequence is adjusted based on the duration corresponding to the second text content to generate the second action animation of the virtual object based on the second text content.

[0221] In an optional embodiment, frame interpolation is performed on the first motion animation and the second motion animation to generate motion animation in which the virtual object successively performs the actions of the first object and the second object.

[0222] In illustrative terms, frame interpolation is used to represent the process of adding interpolated intermediate frames between animation frames to smooth the animation effect. Frame interpolation is commonly used in the production of motion animation, where motion animation consists of a series of animation frames, and frame interpolation is used to fill the gaps between these animation frames, thereby making the motion look more natural.

[0223] In some embodiments, the timing of the transition from the first motion animation to the second motion animation is determined.

[0224] Indicatively, the transition moment is the moment when the first animation finishes playing and the second animation begins playing; that is, the transition moment is the end moment of the first animation and the beginning moment of the second animation.

[0225] In some embodiments, a first number of motion animation frames prior to the succession moment are obtained from a first motion animation, and a second number of motion animation frames after the succession moment are obtained from a second motion animation.

[0226] Indicatively, the first motion animation includes multiple motion animation frames, with a first number of motion animation frames obtained before the continuation time; correspondingly, the second motion animation includes multiple motion animation frames, with a second number of motion animation frames obtained after the continuation time. The first number and the second number may be the same or different, and are not limited here.

[0227] For example: obtain 10 motion animation frames from the first motion animation before the continuation time. These 10 motion animation frames are the last 10 motion animation frames corresponding to the first motion animation. Obtain 10 motion animation frames from the second motion animation after the continuation time. These 10 motion animation frames are the first 10 motion animation frames corresponding to the second motion animation.

[0228] In some embodiments, frame interpolation is performed on a first number of motion animation frames and a second number of motion animation frames to obtain motion animation.

[0229] Optionally, at least one of several interpolation methods can be used, such as Linear Interpolation (LERP), Bezier interpolation, and Spherical Linear Interpolation (SLERP). By employing frame interpolation processing, an animation based on the first and second motion animations is obtained, which can display the animation effect of a virtual object successively narrating the first and second text content.

[0230] The LERP algorithm is a basic interpolation technique suitable for handling changes in position. It calculates the specific position of each part of a virtual object within each animation frame. The LERP algorithm is particularly well-suited for handling rotation, ensuring smoothness and consistency during rotation.

[0231] To illustrate, let's take the example of the first number of motion animation frames being the last two motion animation frames of the first motion animation, and the second number of motion animation frames being the first two motion animation frames of the second motion animation; let's call the second to last motion animation frame of the first motion animation frame 1, the last motion animation frame of the first motion animation frame 2, the first motion animation frame of the second motion animation frame 3, and the second motion animation frame of the second motion animation frame 4.

[0232] When performing frame interpolation on the first and second number of motion animation frames, interpolation is performed between motion animation frame 1 and motion animation frame 2, between motion animation frame 2 and motion animation frame 3, and between motion animation frame 3 and motion animation frame 4. For example, taking the interpolation between motion animation frame 2 and motion animation frame 3 as an example, if the virtual arm of the virtual object in motion animation frame 1 points north (rotation angle of 0 degrees) and the virtual object is located at coordinates (0, 0, 0), and the virtual arm of the virtual object in motion animation frame 2 points east (rotation angle of 90 degrees) and the virtual object is located at coordinates (10, 0, 0), that is, the virtual object has a bone rotation angle and a globally unique change, the preset interpolation weight is obtained, and the above rotation angle and coordinate conditions are substituted into the LERP interpolation formula and SLERP interpolation method respectively to calculate the state of the intermediate motion animation frame after interpolation, such as the state of the intermediate motion animation frame: the virtual arm of the virtual object points northeast at 45 degrees, and the coordinates are (5, 0, 0), etc.

[0233] It is worth noting that the above are merely illustrative examples, and the embodiments of this application are not limited thereto.

[0234] In summary, when virtual objects narrate text content, using the first duration corresponding to the text content as the generation condition in the animation generation process, and using the first action sequence matching the text content as auxiliary content in the animation generation process, the animation duration of the first action animation generated based on the first action frame sequence matches the first duration. This makes the first action animation presented by the virtual object when narrating text content more relevant to the text content, and strengthens the contextual fit between the text content and the first action animation. This not only greatly enhances the expressiveness of virtual objects, but also improves the efficiency of human-computer interaction to a certain extent.

[0235] In this embodiment, a method is described for obtaining motion data, including a first motion frame sequence, from multiple motion data based on semantic information of text content, and for adjusting the sequence based on the first motion frame sequence. According to the semantic matching relationship between the motion text and the text content in the motion data, motion data for assisting in generating the first motion animation can be selectively determined from multiple motion data. This allows for precise adjustment of semantically identical or similar first motion frame sequences, ensuring that the first motion animation does not deviate from the semantic information represented by the text content. Combined with the constraint of a first duration corresponding to the text content, the sequence-adjusted first motion animation conforms to both the semantic information corresponding to the text content and the duration constraint of the first duration, thus improving the accuracy of the first motion animation.

[0236] In an optional embodiment, the above-described motion animation generation method is performed using a pre-trained neural network model; wherein the neural network model includes an encoder and a decoder. (Illustrative example, such as...) Figure 5 As shown, the training process of the neural network model is described as follows: the training process of the neural network model can be implemented as follows: steps 510 to 550.

[0237] Step 510: Obtain sample text.

[0238] Among them, the sample text corresponds to the sample duration, which is the duration of the virtual object narrating the text content; the sample duration is used to limit the animation duration of the predicted motion animation generated based on the sample text.

[0239] Optionally, the sample text is text from a pre-collected sample text library, and each sample text is labeled with a sample duration.

[0240] Step 520: Obtain sample action data based on sample text.

[0241] Indicatively, action data that has a semantic match relationship with a sample text is determined from multiple pre-stored action data based on the semantic information of the sample text, and is used as sample action data.

[0242] Optionally, extract the sample feature representation corresponding to the sample text; find multiple action data based on the sample feature representation to obtain sample action data.

[0243] Among them, sample features represent the semantic information used to characterize the sample text.

[0244] To illustrate, the sample text is input into the feature extraction network, and the sample feature representation is extracted from the feature extraction network. The sample feature representation is to convert the sample text into a numerical form that can reflect its semantic information. Therefore, the sample feature representation is used to capture the semantic information such as the deeper meaning, sentiment, theme, and context in the sample text.

[0245] The sample action data includes a sequence of sample action frames, and the number of sample action frames corresponds to the sequence of sample action frames.

[0246] In a schematic way, multiple action data correspond to one action text. The text feature representations of the action texts corresponding to the multiple action data are extracted. The feature similarity between the sample feature representation and the multiple text feature representations is calculated. Based on the similarity calculation results, the action data that best matches the sample text is selected from the multiple action data as the sample action data. There is a semantic matching relationship between the action text in the sample action data and the sample text.

[0247] Indicatively, the sample action frame sequence corresponds to the sample action duration, which describes the duration required to play the sample action frame sequence; based on the sample action duration and the rendering frame rate, the number of sample action frames can be calculated, and the rendering frame rate is a predetermined value.

[0248] Step 530: The sequence of sample action frames and the number of sample action frames are encoded by the encoder in the neural network model to obtain the latent feature representation of the sample.

[0249] As an illustration, the sample action frame sequence includes multiple sample action screen frames; among them, the sample latent feature representation is the key information that characterizes the action changes from the time dimension and the action change dimension for multiple sample action screen frames.

[0250] Step 540: Using the sample duration as the action generation condition corresponding to the decoder, the latent feature representation of the sample is decoded by the decoder in the neural network model to obtain the predicted action data.

[0251] The encoder and decoder are neural network layers that participate in the training of the neural network model; the decoder is set with given conditions, which are the sample duration corresponding to the sample text.

[0252] By using the sample duration as the action generation condition corresponding to the decoder, and by decoding the latent feature representation of the sample through the decoder, predicted action data can be generated.

[0253] Step 550: Train the neural network model based on the difference between the sample action data and the predicted action data.

[0254] In an optional embodiment, a first prediction loss value is obtained between the sample reference data and the predicted action data; the neural network model is trained based on the first prediction loss value.

[0255] In a schematic manner, the loss value between the sample reference data and the predicted action data is calculated to obtain the first predicted loss value. The neural network model is trained with the goal of reducing the first predicted loss value. During the training process, the parameters corresponding to the encoder and decoder are adjusted so that the trained neural network model can analyze the text content.

[0256] In an optional embodiment, a first prediction loss value is obtained between sample reference data and predicted action data; a second prediction loss value is obtained between the prior distribution and the posterior distribution corresponding to the encoder; and a neural network model is trained with the goal of reducing the first prediction loss value and the second prediction loss value.

[0257] The prior distribution is the distribution obtained by the encoder based on prior knowledge, while the posterior distribution is the distribution learned by the encoder based on the latent feature representation of the samples.

[0258] Indicatively, the encoder corresponds to the prior distribution, which is the distribution the encoder obtains based on prior knowledge, usually implemented as a Gaussian distribution. The encoder encodes the input sample action frame sequence and the number of sample action frames into distribution parameters in the latent space, and outputs the mean and variance to form the posterior distribution. That is, the posterior distribution is the distribution of the beat calculated by the encoder, which is a distribution adjusted according to the input sample action frame sequence and the number of sample action frames, and therefore can reflect the information of the sample action frame sequence and the number of sample action frames.

[0259] Optionally, the second prediction loss is implemented as the relative entropy (Kullback-Leibler divergence, KL) divergence. Since KL divergence is usually used to measure the difference between two probability distributions, calculating KL divergence as the second prediction loss can measure the difference between the prior distribution corresponding to the encoder and the posterior distribution obtained after observing the data.

[0260] The neural network model can be trained with the goal of minimizing the second prediction loss to ensure that the encoder can generate a latent representation with sufficient information while retaining sufficient structure; or, the neural network model can be trained with the goal of minimizing the sum of the first and second prediction losses to obtain a trained neural network model for analyzing text content.

[0261] In some embodiments, the trained neural network model includes a trained encoder and a trained decoder. The trained neural network model is applied to the process of analyzing text content to generate a first motion animation corresponding to the text content.

[0262] It is worth noting that the above are merely illustrative examples, and the embodiments of this application are not limited thereto.

[0263] In summary, when virtual objects narrate text content, using the first duration corresponding to the text content as the generation condition in the animation generation process, and using the first action sequence matching the text content as auxiliary content in the animation generation process, the animation duration of the first action animation generated based on the first action frame sequence matches the first duration. This makes the first action animation presented by the virtual object when narrating text content more relevant to the text content, and strengthens the contextual fit between the text content and the first action animation. This not only greatly enhances the expressiveness of virtual objects, but also improves the efficiency of human-computer interaction to a certain extent.

[0264] In an optional embodiment, the above-described method for generating motion animation is referred to as "a method for generating three-dimensional (3D) motion of virtual objects." (Illustrative example, such as...) Figure 6As shown, the above method for generating motion animations is applied to the motion generation scenario of 3D virtual objects, and the implementation process includes the following steps 610 to 640.

[0265] Step 610: Obtain the action description.

[0266] Optionally, the action description can be implemented as the text content mentioned above; or, the action description can be implemented as action description text after processing the text content, etc.

[0267] As an illustration, based on the text content that the virtual object needs to broadcast, an NLP model can be used to generate corresponding action descriptions. For example, an NLP model could be designed with the following prompt: "You are an enthusiastic and lively virtual companion. Please provide corresponding body language descriptions for the following sentence. These actions should be appropriate to the context of the dialogue, without any physical contact with the main virtual object, and without any facial expressions." Sending this prompt along with the text content to the NLP model allows it to leverage the rich knowledge of the NLP model to generate suitable action descriptions.

[0268] Step 620, Action library matching.

[0269] Optionally, a pre-acquired action library stores a large number of 3D actions (action data) with different text descriptions (the aforementioned action text). A contrastive language-image pre-training (CLIP) model is used to encode the text description of each 3D action, thereby obtaining a feature library F. After obtaining the desired action description, the CLIP model is also used to encode this action description to obtain feature f. The cosine similarity between feature f and each feature in feature library F is calculated to obtain the 3D action file corresponding to the feature with the highest similarity as the first action data. The first action data includes a first action frame sequence.

[0270] The following is an illustrative explanation of the content matched by the action library.

[0271] (1) Construction of the action library

[0272] First, a motion library containing various 3D actions needs to be built; these 3D actions may be obtained through motion capture technology or created manually by animators. Each 3D action has one or more associated text descriptions, which may include at least one of the following information: action type, emotion, intensity, etc.

[0273] (2) Application of the CLIP model

[0274] The CLIP model is a multimodal model that can simultaneously understand image content and natural language descriptions. In this application scenario, the language part of the CLIP model is used to process the text descriptions of 3D actions. For example, for each action in the action library, its text description can be input into the CLIP model, and the model will output a high-dimensional feature vector that captures the semantic information of the description; the feature vectors corresponding to multiple actions constitute the feature library F.

[0275] (3) Feature encoding

[0276] To illustrate, when a specific 3D action needs to be retrieved, the action description of that action is first input into the same CLIP model; the model will generate a feature vector f for this new action description, and this feature vector f represents the features of the desired action in the same semantic space.

[0277] (4) Similarity calculation

[0278] To find the action most similar to feature vector f, we can calculate the cosine similarity between feature vector f and each feature vector in the feature library F. Cosine similarity is obtained by calculating the dot product of the two vectors and dividing by their respective magnitudes. This value ranges from -1 to 1, with a larger value indicating a higher similarity.

[0279] (5) Search and sorting

[0280] Based on the calculated cosine similarity, actions in the feature library F can be sorted, and the action with the highest similarity can be selected. Furthermore, a similarity threshold can be set; only actions whose calculated similarity exceeds this threshold are considered a match.

[0281] By following the steps above, a 3D action retrieval system based on natural language description can be realized. This system can quickly and accurately find the specific action required by the user from a large amount of action data.

[0282] Step 630, Variational autoencoder model encoding and decoding.

[0283] A variational autoencoder (VAE) is a generative model that generates new data instances by learning latent feature representations of the input data.

[0284] Optionally, in scenarios involving the generation of motion animation, the VAE model is trained to understand and generate 3D motion sequences. A VAE typically consists of two main parts: an encoder and a decoder. The encoder is responsible for mapping high-dimensional input data to a low-dimensional latent space, where the input data is 3D motion (the first motion frame sequence); while the decoder is responsible for mapping points in this latent space back to the original data space, generating the first motion animation.

[0285] In some embodiments, a VAE model is trained based on all 3D actions in the action library. When generating new actions, the model can choose a different length than the actions encoded, thereby obtaining new actions of the desired length.

[0286] To illustrate, during encoding, input the 3D action selected in the previous step and its original length (i.e., the number of first screen frames mentioned above). During decoding, input the desired target length (i.e., the number of second screen frames mentioned above, which is obtained by multiplying the length t (in seconds) of the audio to be played by the virtual object by the rendering frame rate (e.g., 30 FPS)). This will generate 3D action data of the desired target length.

[0287] The encoding and decoding processes are illustrated below.

[0288] (1) Data preprocessing

[0289] This illustrates how all 3D motion data is collected from the motion library and standardized to a uniform format and size. This includes bone alignment, scaling, and time-series interpolation to ensure data consistency.

[0290] (2) Model Architecture

[0291] The typical VAE architecture is adopted, in which the encoder consists of multiple convolutional layers, the latent space is usually a fully connected layer, and its output is used to obtain the mean and variance of the latent feature vector; the decoder is usually the opposite of the encoder, it maps the points of the latent space back to the original data space to generate the first motion animation.

[0292] To illustrate, the input to the encoder is the sequence of first action frames and the number of first screen frames; the output of the encoder is a latent feature representation, or the output of the encoder is the mean and variance, which are used to obtain the posterior distribution from which the latent feature representation can be sampled.

[0293] (3) Training process

[0294] When training the VAE model, not only is the reconstruction error (i.e., the first prediction loss value mentioned above, representing the difference between the original action and the action reconstructed by the latent features) minimized, but the KL divergence of the latent space (i.e., the second prediction loss value mentioned above) is also minimized to ensure that the distribution of the latent space is close to the prior distribution (usually a Gaussian distribution).

[0295] (4) Length adjustment

[0296] To generate actions of varying lengths, a conditional variable, the target length (either the first duration or the number of second frames corresponding to the first duration), can be introduced during the decoding process. This target length is calculated by multiplying the duration t (in seconds) of the virtual human playing the speech by the rendering frame rate (e.g., 30fps). This means that if the speech length is 2 seconds, the target length will be 60 frames.

[0297] (5) Decoding and post-processing

[0298] During the decoding stage, given the target length and the latent variables obtained from the encoder, the decoder generates a 3D motion sequence of the corresponding length, i.e., the first motion animation; the generated motion requires further post-processing, such as smoothing filtering, to ensure the naturalness and coherence of the motion.

[0299] (6) Loss Function

[0300] As an illustration, when training a VAE model, the loss function may include reconstruction loss (such as mean squared error MSE) and regularization loss of the latent space (such as KL divergence).

[0301] Step 640: Smooth 3D motion frame interpolation.

[0302] In illustrative terms, in practical applications such as live streaming of intelligent non-player characters (NPCs) or 3D virtual objects, the actions of virtual objects must be continuous. After acquiring the next action data, the action needs to be added to the buffer of the currently rendered action data. Frame interpolation is required at the transition points to smooth the transition. Optionally, select certain data before and after the transition (e.g., 10 action frames before and after), and then use SLERP and / or LERP algorithms to interpolate the bone rotation angle and global displacement respectively, regenerating the action data for this transition interval to obtain smooth, non-abrupt 3D action data.

[0303] In some embodiments, players interact with virtual objects by triggering in-game dialogue functions, and the virtual objects' responses to the players constitute the text content to be analyzed. When generating the first action animation based on the text content, it may be generated by the virtual object performing a single action, such as generating a first action animation by the virtual object performing a chin-resting thought action; or it may be generated by the virtual object performing multiple actions consecutively, such as generating a first action animation by the virtual object performing a chin-resting thought action followed by an invitation action.

[0304] Optionally, the actions performed by the virtual object are determined based on the object type of the virtual object. For example, when the virtual object is implemented as a virtual character, the actions performed by the virtual object include at least one of several anthropomorphic actions such as standing, running, jumping, resting chin on hand in thought, waving, inviting, shaking head, smiling, and crying; when the virtual object is implemented as a virtual animal, the actions performed by the virtual object include at least one of several actions simulated by animals, such as running and wagging tail.

[0305] like Figure 7 The diagram shows an interface for displaying the first action animation based on text content. During dialogue interaction between the player and the virtual object 710, the player initiates a dialogue with the virtual object 710 by inputting text in the dialog box 720. In one dialogue, the virtual object 710 replies with the text "I have a special fondness for X place; this is the Jianghu in my heart." Based on this text content, the aforementioned action animation generation method is executed to generate the first action animation. For example, the first action animation might be: the virtual object performs a thoughtful gesture with its chin resting on its hand, then lowers its virtual arm, and subsequently performs a series of actions such as smiling and placing its virtual arm on its chest. Figure 7 A screenshot of a schematic interface is shown when the virtual object 710 performs a chin-resting thought action.

[0306] It is worth noting that the actions performed by the virtual object in the first motion animation above are at least one action generated based on the text content. The embodiments of this application do not limit the type of action.

[0307] In summary, when virtual objects narrate text content, using the first duration corresponding to the text content as the generation condition in the animation generation process, and using the first action sequence matching the text content as auxiliary content in the animation generation process, the animation duration of the first action animation generated based on the first action frame sequence matches the first duration. This makes the first action animation presented by the virtual object when narrating text content more relevant to the text content, and strengthens the contextual fit between the text content and the first action animation. This not only greatly enhances the expressiveness of virtual objects, but also improves the efficiency of human-computer interaction to a certain extent.

[0308] In the embodiments of this application, rich and varied smooth 3D actions of virtual objects that conform to the context can be generated in real time, which enhances the action generation capability of virtual objects and improves the overall performance capability of virtual objects.

[0309] Please refer to Figure 8 The diagram illustrates a structural block diagram of a motion animation generation apparatus according to an exemplary embodiment of this application, the apparatus comprising the following modules:

[0310] The acquisition module 810 is used to acquire text content, which is reference information for generating the first action animation of the virtual object. The text content corresponds to a first duration, which is the duration for the virtual object to narrate the text content.

[0311] The acquisition module 810 is further configured to acquire a first action frame sequence that matches the text content, the first action frame sequence being used to generate the action state of the virtual object narrating the text content;

[0312] The adjustment module 820 is used to adjust the sequence of the first action frame based on the first duration to generate a first action animation of the virtual object based on the text content. The animation duration of the first action animation matches the first duration. The first action animation is the animation performance presented when the virtual object narrates the text content.

[0313] In an optional embodiment, the adjustment module 820 is further configured to obtain the number of first screen frames corresponding to the first action frame sequence, wherein the number of first screen frames is the number of action screen frames that make up the first action frame sequence; map the first action frame sequence and the number of first screen frames to the latent space corresponding to the encoder, and extract latent feature representations, wherein the latent feature representations are key information representing the action changes of multiple action screen frames from the time dimension and the action change dimension, wherein the multiple action screen frames make up the first action frame sequence; decode the latent feature representations with the first duration as the action generation condition corresponding to the decoder, and generate the first action animation of the virtual object based on the text content, wherein the encoder and the decoder are pre-trained neural network layers involved in the process of generating the first action animation.

[0314] In an optional embodiment, the adjustment module 820 is further configured to obtain the number of second screen frames corresponding to the text content based on the rendering frame rate used when generating the first motion animation and the first duration; extract the quantity feature representation corresponding to the number of second screen frames; concatenate the quantity feature representation and the latent feature representation to obtain a concatenated feature representation; decode the concatenated feature representation and generate the first motion animation of the virtual object based on the text content.

[0315] In an optional embodiment, the adjustment module 820 is further configured to: obtain the number of second frame frames corresponding to the text content based on the rendering frame rate used when generating the first motion animation and the first duration; extract the quantity feature representation corresponding to the number of second frame frames; obtain the attention weights corresponding to the quantity feature representation and the multiple sub-feature representations in the latent feature representation, wherein the attention weights are used to represent the degree of attention paid by the multiple sub-feature representations to the quantity feature representation during the decoding process; perform a weighted summation on the multiple sub-feature representations based on the attention weights to obtain a weighted feature representation; decode the weighted feature representation to generate the first motion animation of the virtual object based on the text content.

[0316] In an optional embodiment, the acquisition module 810 is further configured to acquire multiple action data, the action data including action frame sequences and action text, the action frame sequences describing multiple action screen frames corresponding to the action data, and the action text representing the semantic information represented by the action frame sequences; match the text content with the action text corresponding to the multiple action data respectively, and determine the first action data matching the text content from the multiple action data, the first action data including the first action frame sequence and the first action text, and the first action text having a semantic matching relationship with the text content.

[0317] In an optional embodiment, the acquisition module 810 is further configured to extract the semantic feature representation corresponding to the text content, the semantic feature representation being used to characterize the semantic information of the text content; acquire the text feature representation of the action text corresponding to the plurality of action data respectively, the text feature representation being used to characterize the semantic information of the action text; match the semantic feature representation with the plurality of text feature representations, and determine the first action data that matches the text content from the plurality of action data.

[0318] In an optional embodiment, the acquisition module 810 is further configured to determine the similarity between the semantic feature representation and the plurality of text feature representations, obtain the feature similarity corresponding to the plurality of text feature representations respectively, the feature similarity being used to characterize the degree of semantic matching between the text content and the action content; and determine the first action data from the plurality of action data based on the plurality of feature similarities.

[0319] In an optional embodiment, the acquisition module 810 is further configured to determine the first action data corresponding to the first text feature representation based on the first text feature representation corresponding to the maximum feature similarity among the plurality of feature similarities; or, if there is a first feature similarity among the plurality of feature similarities that reaches a preset similarity threshold, the first action data is determined based on the first text feature representation corresponding to the first feature similarity, wherein the first text feature representation is the feature representation corresponding to the first action text in the first action data.

[0320] In an optional embodiment, the acquisition module 810 is further configured to generate action description text based on the text content under action restriction conditions, the action description text being used to trigger the first action animation of generating a virtual object, the action description text including restriction information and text content, the restriction information being text information used to describe the action restriction conditions; and to acquire the first action frame sequence matching the action description text, the first action frame sequence having a semantic matching relationship with the text content.

[0321] In an optional embodiment, such as Figure 9 As shown, the device further includes:

[0322] Training module 830 is used to acquire sample text, the sample text corresponding to sample duration; acquire sample action data based on the sample text, the sample action data including sample action frame sequence, the sample action frame sequence corresponding to the number of sample action frames, the number of sample action frames used to represent the number of action screen frames in the sample action frame sequence; encode the sample action frame sequence and the number of sample action frames through the encoder in the neural network model to obtain sample latent feature representation; use the sample duration as the action generation condition corresponding to the decoder, decode the sample latent feature representation through the decoder in the neural network model to obtain predicted action data; train the neural network model based on the difference between the sample action data and the predicted action data.

[0323] In an optional embodiment, the training module 830 is further configured to obtain a first prediction loss value between the sample reference data and the predicted action data; obtain a second prediction loss value between the prior distribution and the posterior distribution corresponding to the encoder, wherein the prior distribution is a distribution obtained by the encoder based on prior knowledge, and the posterior distribution is a distribution learned by the encoder based on the latent feature representation of the sample; and train the neural network model with the goal of reducing the first prediction loss value and the second prediction loss value.

[0324] In an optional embodiment, the first motion animation frame is used to present a first object motion of the virtual object;

[0325] The adjustment module 820 is further configured to, when the virtual object performs a second object action within a preset time period after performing a first object action, obtain a second action animation corresponding to the second object action; perform frame interpolation processing on the first action animation and the second action animation to generate an action animation in which the virtual object performs the first object action and the second object action consecutively.

[0326] In an optional embodiment, the adjustment module 820 is further configured to determine the continuation time of the second object action following the first object action performed by the virtual object, wherein the continuation time is the time when the first action animation finishes playing and the second action animation begins playing; obtain a first number of action animation frames before the continuation time from the first action animation, and obtain a second number of action animation frames after the continuation time from the second action animation; and perform frame interpolation processing on the first number of action animation frames and the second number of action animation frames to obtain the action animation.

[0327] In summary, when virtual objects narrate text content, using the first duration corresponding to the text content as the generation condition in the animation generation process, and using the first action sequence matching the text content as auxiliary content in the animation generation process, the animation duration of the first action animation generated based on the first action frame sequence matches the first duration. This makes the first action animation presented by the virtual object when narrating text content more relevant to the text content, and strengthens the contextual fit between the text content and the first action animation. This not only greatly enhances the expressiveness of virtual objects, but also improves the efficiency of human-computer interaction to a certain extent.

[0328] It should be noted that the motion animation generation device provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the motion animation generation device and the motion animation generation method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0329] Figure 10 This illustration shows a schematic diagram of the structure of a server provided in an exemplary embodiment of this application. Specifically, it includes the following structure.

[0330] Server 1000 includes a Central Processing Unit (CPU) 1001, a system memory 1004 including Random Access Memory (RAM) 1002 and Read Only Memory (ROM) 1003, and a system bus 1005 connecting the system memory 1004 and the CPU 1001. Server 1000 also includes a mass storage device 1006 for storing the operating system 1013, application programs 1014, and other program modules 1015.

[0331] Mass storage device 1006 is connected to central processing unit 1001 via a mass storage controller (not shown) connected to system bus 1005. Mass storage device 1006 and its associated computer-readable media provide non-volatile storage for server 1000. That is, mass storage device 1006 may include computer-readable media (not shown) such as hard disk or compact disc read-only memory (CD-ROM) drive.

[0332] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types. The system memory 1004 and mass storage device 1006 described above can be collectively referred to as memory.

[0333] According to various embodiments of this application, server 1000 can also be connected to a remote computer on a network, such as the Internet. That is, server 1000 can be connected to network 1012 via network interface unit 1011 connected to system bus 1005, or it can use network interface unit 1011 to connect to other types of networks or remote computer systems (not shown).

[0334] The aforementioned memory also includes one or more programs, which are stored in the memory and configured to be executed by the CPU.

[0335] Embodiments of this application also provide a computer device including a processor and a memory. The memory stores at least one instruction, at least one program, code set, or instruction set. The processor loads and executes the at least one instruction, at least one program, code set, or instruction set to implement the motion animation generation method provided in the above-described method embodiments. Optionally, the computer device may be a terminal or a server.

[0336] Embodiments of this application also provide a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the motion animation generation method provided in the above-described method embodiments.

[0337] Embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the motion animation generation methods described in the above embodiments.

[0338] Optionally, the computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSDs), or optical discs, etc. The random access memory may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM). The sequence numbers of the embodiments in this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0339] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0340] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for generating motion animation, characterized in that, The method includes: Obtain text content, which is reference information for generating the first motion animation of the virtual object. The text content corresponds to a first duration, which is the duration for the virtual object to narrate the text content. Obtain a first action frame sequence that matches the text content, the first action frame sequence being used to generate the action state of the virtual object narrating the text content; Based on the first duration, the sequence of the first action frames is adjusted to generate a first action animation of the virtual object based on the text content. The animation duration of the first action animation matches the first duration. The first action animation is the animated performance presented when the virtual object narrates the text content.

2. The method according to claim 1, characterized in that, The step of adjusting the sequence of the first action frame sequence based on the first duration to generate the first action animation of the virtual object based on the text content includes: Obtain the number of first screen frames corresponding to the first action frame sequence, where the number of first screen frames is the number of action screen frames that make up the first action frame sequence. The first action frame sequence and the number of first screen frames are mapped to the latent space corresponding to the encoder, and the latent feature representation is extracted. The latent feature representation is the key information that characterizes the action change from the time dimension and the action change dimension for multiple action screen frames. The multiple action screen frames constitute the first action frame sequence. The latent feature representation is decoded using the first duration as the action generation condition corresponding to the decoder, and the first action animation of the virtual object based on the text content is generated. The encoder and the decoder are pre-trained neural network layers involved in the process of generating the first action animation.

3. The method according to claim 2, characterized in that, The step of decoding the latent feature representation using the first duration as the action generation condition corresponding to the decoder, and generating the first action animation of the virtual object based on the text content, includes: Based on the rendering frame rate used when generating the first action animation and the first duration, the number of second screen frames corresponding to the text content is obtained; Extract the quantitative feature representation corresponding to the number of frames in the second frame; By concatenating the quantitative feature representation and the latent feature representation, a concatenated feature representation is obtained; Decode the splicing feature representation and generate the first motion animation of the virtual object based on the text content.

4. The method according to claim 2, characterized in that, The step of decoding the latent feature representation using the first duration as the action generation condition corresponding to the decoder, and generating the first action animation of the virtual object based on the text content, includes: Based on the rendering frame rate used when generating the first action animation and the first duration, the number of second screen frames corresponding to the text content is obtained; Extract the quantitative feature representation corresponding to the number of frames in the second frame; Obtain attention weights corresponding to multiple sub-feature representations in the quantitative feature representation and the latent feature representation, respectively. The attention weights are used to represent the degree of attention that the multiple sub-feature representations pay to the quantitative feature representation during the decoding process. The multiple sub-feature representations are weighted and summed based on the attention weights to obtain a weighted feature representation; Decode the weighted feature representation to generate the first motion animation of the virtual object based on the text content.

5. The method according to any one of claims 1 to 4, characterized in that, The step of obtaining the first action frame sequence that matches the text content includes: Acquire multiple action data, the action data including action frame sequences and action text, the action frame sequences being used to describe multiple action screen frames corresponding to the action data, and the action text being used to characterize the semantic information represented by the action frame sequences; The text content is matched with the action text corresponding to the plurality of action data respectively. The first action data that matches the text content is determined from the plurality of action data. The first action data includes the first action frame sequence and the first action text. The first action text and the text content have a semantic matching relationship.

6. The method according to claim 5, characterized in that, The step of matching the text content with the action text corresponding to the plurality of action data, and determining the first action data matching the text content from the plurality of action data, includes: Extract the semantic feature representation corresponding to the text content, and the semantic feature representation is used to characterize the semantic information of the text content; Obtain the text feature representations of the action texts corresponding to the multiple action data respectively, and the text feature representations are used to characterize the semantic information of the action texts; The semantic feature representation is matched with multiple text feature representations to determine the first action data that matches the text content from the multiple action data.

7. The method according to claim 6, characterized in that, The step of matching the semantic feature representation with multiple text feature representations to determine the first action data that matches the text content from the multiple action data includes: Determine the similarity between the semantic feature representation and the plurality of text feature representations to obtain the feature similarity corresponding to each of the plurality of text feature representations. The feature similarity is used to characterize the degree of semantic matching between the text content and the action content. The first action data is determined from the multiple action data based on multiple feature similarities.

8. The method according to claim 7, characterized in that, Determining the first action data from the multiple action data based on multiple feature similarities includes: Based on the first text feature representation corresponding to the largest feature similarity among the multiple feature similarities, the first action data corresponding to the first text feature representation is determined; or... If a first feature similarity among the plurality of feature similarities reaches a preset similarity threshold, the first action data is determined based on the first text feature representation corresponding to the first feature similarity, wherein the first text feature representation is the feature representation corresponding to the first action text in the first action data.

9. The method according to any one of claims 1 to 4, characterized in that, The step of obtaining the first action frame sequence that matches the text content includes: Based on the text content, an action description text is generated under action restriction conditions. The action description text is used to trigger the first action animation of generating a virtual object. The action description text includes restriction information and text content. The restriction information is text information used to describe the action restriction conditions. Obtain the first action frame sequence that matches the action description text, wherein the first action frame sequence and the text content have a semantic matching relationship.

10. The method according to any one of claims 1 to 4, characterized in that, The method is performed by a pre-trained neural network model; The training methods for the neural network model include: Obtain sample text, the sample text corresponding to the sample duration; Based on the sample text, sample action data is obtained. The sample action data includes a sample action frame sequence, and the sample action frame sequence corresponds to a number of sample action frames. The number of sample action frames is used to represent the number of action screen frames in the sample action frame sequence. The encoder in the neural network model encodes the sequence of sample action frames and the number of sample action frames to obtain the latent feature representation of the sample. Using the sample duration as the action generation condition corresponding to the decoder, the latent feature representation of the sample is decoded by the decoder in the neural network model to obtain predicted action data. The neural network model is trained based on the difference between the sample action data and the predicted action data.

11. The method according to claim 10, characterized in that, Training the neural network model based on the difference between the sample action data and the predicted action data includes: Obtain the first prediction loss value between the sample reference data and the predicted action data; Obtain the second prediction loss value between the prior distribution and the posterior distribution corresponding to the encoder, wherein the prior distribution is the distribution obtained by the encoder based on prior knowledge, and the posterior distribution is the distribution learned by the encoder based on the latent feature representation of the sample; The neural network model is trained with the goal of reducing the first prediction loss value and the second prediction loss value.

12. The method according to any one of claims 1 to 4, characterized in that, The first motion animation frame is used to present the first object motion of the virtual object; The method further includes: If a second object action is performed within a preset time period after the virtual object performs the first object action, the second action animation corresponding to the second object action is obtained; The first motion animation and the second motion animation are subjected to frame interpolation to generate motion animations in which the virtual object successively performs the actions of the first object and the second object.

13. The method according to claim 12, characterized in that, The step of performing frame interpolation on the first motion animation and the second motion animation to generate motion animation in which the virtual object successively performs the actions of the first object and the second object includes: The timing of the second object action following the first object action performed by the virtual object is determined, wherein the timing is the moment when the first action animation finishes playing and the second action animation begins playing; Obtain a first number of motion animation frames before the continuation moment from the first motion animation, and obtain a second number of motion animation frames after the continuation moment from the second motion animation; The first number of motion animation frames and the second number of motion animation frames are interpolated to obtain the motion animation.

14. A device for generating motion animation, characterized in that, The device includes: The acquisition module is used to acquire text content, which is reference information for generating the first motion animation of the virtual object. The text content corresponds to a first duration, which is the duration for the virtual object to narrate the text content. The acquisition module is also used to acquire a first action frame sequence that matches the text content, and the first action frame sequence is used to generate the action state of the virtual object describing the text content; The adjustment module is used to adjust the sequence of the first action frame based on the first duration to generate a first action animation of the virtual object based on the text content. The animation duration of the first action animation matches the first duration, and the first action animation is the animation performance presented by the virtual object when it narrates the text content.

15. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one program, which is loaded and executed by the processor to implement the motion animation generation method as described in any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that, The storage medium stores at least one program segment, which is loaded and executed by a processor to implement the motion animation generation method as described in any one of claims 1 to 13.

17. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the motion animation generation method as described in any one of claims 1 to 13.