Animation generation method and device, electronic equipment and computer readable storage medium
By encoding and fusing the multimodal input information of virtual objects, animations that conform to the target emotion are generated, which solves the problems of insufficient robustness and versatility of existing emotion analysis models and achieves highly accurate virtual object animation generation.
Patent Information
- Application Number
- CN202410696273.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-30
- Publication Date
- 2025-12-02
AI Technical Summary
Existing sentiment analysis models are not robust and lack versatility, making it difficult to achieve highly accurate virtual object animation generation.
By acquiring multiple modal input information of virtual objects, encoding each modality separately, fusing the encoded features, and performing sentiment prediction, virtual object animations that conform to the target sentiment are generated.
It improves the accuracy and robustness of sentiment prediction, generates vivid and realistic animation effects, and enhances user immersion and satisfaction.
Smart Images

Figure CN121053258A_ABST
Abstract
Description
Technical Field
[0001] This application relates to artificial intelligence technology, and more particularly to an animation generation method, apparatus, electronic device, and computer-readable storage medium. Background Technology
[0002] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0003] Sentiment analysis technology is one of the important research directions in the field of artificial intelligence. Virtual object technology based on sentiment analysis is widely used in various human-computer interaction scenarios. While saving a lot of manpower, it improves the human-computer interaction experience and gradually changes the way humans interact with machines, bringing more convenience and progress to human society. However, the sentiment analysis models in related technologies are not robust and lack versatility. Summary of the Invention
[0004] This application provides an animation generation method, apparatus, electronic device, and computer-readable storage medium, which can realize a highly versatile and robust method for generating animations of virtual objects.
[0005] The technical solution of this application embodiment is implemented as follows:
[0006] This application provides an animation generation method, the method comprising:
[0007] Obtain input information from multiple modalities of virtual objects;
[0008] The input information of each modality is encoded to obtain the encoding features of the input information of each modality;
[0009] By fusing the encoding features corresponding to the input information of the multiple modalities, a fused feature of the input information of the multiple modalities is obtained;
[0010] Based on the fusion features, sentiment prediction is performed on the input information of the multiple modalities to obtain the target sentiment expressed by the input information of the multiple modalities;
[0011] Based on the target emotion and at least some of the input information from the multiple modalities, the virtual object is animated to obtain a virtual object animation.
[0012] This application provides an animation generation apparatus, including:
[0013] The data management module is used to acquire input information from various modalities of virtual objects;
[0014] An encoding module is used to encode the input information of each modality to obtain the encoding features of the input information of each modality.
[0015] The fusion module is used to fuse the encoding features corresponding to the input information of the multiple modalities to obtain the fusion features of the input information of the multiple modalities.
[0016] The emotion generation module is used to predict the emotion of the input information of the multiple modalities based on the fusion features, so as to obtain the target emotion expressed by the input information of the multiple modalities.
[0017] An animation generation module is used to generate an animation for the virtual object based on the target emotion and at least some of the input information of the multiple modalities, thereby obtaining an animation of the virtual object.
[0018] This application provides an electronic device for animation generation, the electronic device comprising:
[0019] Memory is used to store executable instructions for a computer;
[0020] The processor, when executing computer-executable instructions stored in the memory, implements the animation generation method provided in the embodiments of this application.
[0021] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the animation generation method provided in this application when executed by a processor.
[0022] This application provides a computer program product, including a computer program or computer executable instructions, which, when executed by a processor, implements the animation generation method provided in this application.
[0023] The embodiments of this application have the following beneficial effects:
[0024] By fusing the encoded features of input information from multiple modalities, and then using the fused features for sentiment prediction, this approach ensures the rich emotional information contained in the multimodal input information while effectively processing and correctly classifying the complex and diverse emotional expressions within the multimodal input information. This improves the accuracy and robustness of sentiment prediction. Based on the target emotion and at least some multimodal input information, animations are generated for virtual objects. This allows for the identification of the target emotion based on artificial intelligence sentiment recognition, generating virtual object animations that match the target emotion. This makes the virtual object animations more vivid and realistic, significantly enhancing user immersion and satisfaction. Attached Figure Description
[0025] Figure 1 This is a schematic diagram of the architecture of the animation generation system provided in the embodiments of this application;
[0026] Figure 2 This is a schematic diagram of the structure of the animation generation device provided in the embodiments of this application;
[0027] Figure 3 This is a first flowchart illustrating the animation generation method provided in this application embodiment;
[0028] Figure 4 This is a second flowchart illustrating the animation generation method provided in this application embodiment;
[0029] Figure 5 This is a schematic diagram of the third process of the animation generation method provided in the embodiments of this application;
[0030] Figure 6 This is a schematic diagram of the fourth process of the animation generation method provided in the embodiments of this application;
[0031] Figure 7 This is a schematic diagram of a first product driven by virtual object emotion, provided in an embodiment of this application;
[0032] Figure 8 This is a schematic diagram of a second product driven by virtual object emotions, provided in an embodiment of this application.
[0033] Figure 9 This is a schematic diagram of the structure of the multimodal emotion recognition model provided in the embodiments of this application.
[0034] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0036] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0037] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0038] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0039] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.
[0040] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0041] 1) Modality: Every source or form of information can be called a modality, such as information in the form of speech, video, and text. Each form of information can be called a modality of information. Modality can be divided into unimodality and multimodality; unimodality represents information as a numerical vector that a computer can process or further abstracts it into a higher-level feature vector, while multimodality learns better feature representations by utilizing the complementarity between multiple modalities and eliminating redundancy between modalities.
[0042] 2) Zero-shot learning (ZSL): This is a machine learning technique that allows neural network models to classify or identify samples that have never been seen before. It is commonly used in fields such as cross-species identification, climate change prediction, and disease transmission pattern prediction.
[0043] 3) One-hot vectors are a commonly used encoding method, typically used to represent categorical variables. A one-hot vector is a sparse vector where only one element is 1 and the others are 0. In machine learning and deep learning, one-hot vectors are often used to represent categorical features, corresponding to class labels in classification tasks.
[0044] 4) Hidden layer vectors: Hidden layer vectors are key intermediate representations in deep learning models. They carry rich information learned by the model from the input data and are crucial to the model's functionality and performance.
[0045] 5) Cross-Attention: This is an attention mechanism widely used in machine learning and deep learning, particularly in sequence-to-sequence models, to understand the relationships between two sequences. For example, in machine translation tasks, when generating target language words, the model needs to pay attention to relevant information in the source language sentence in order to translate accurately.
[0046] 6) Convolutional Neural Networks (CNNs): A class of feedforward neural networks (FNNs) that incorporate convolutional computations and have a deep structure, CNNs are one of the representative algorithms of deep learning. CNNs possess representation learning capabilities, enabling them to perform shift-invariant classification of input images according to their hierarchical structure.
[0047] This application provides an animation generation method, apparatus, electronic device, and computer-readable storage medium, which can realize a highly versatile and robust method for generating animations of virtual objects.
[0048] The animation generation method provided in this application embodiment can be implemented by a terminal or a server alone; it can also be implemented by a terminal and a server in collaboration. For example, the terminal can undertake the animation generation method described below alone, or the terminal can send an animation generation request containing multiple modal input information to the server. The server can execute the animation generation method according to the received animation generation request containing multiple modal input information, determine the input information of multiple modalities and the corresponding target emotion, and generate virtual object animation based on the target emotion.
[0049] The electronic device for animation generation provided in this application can be various types of terminals or servers. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smart TV, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.
[0050] Taking servers as an example, such as server clusters deployed in the cloud, AI as a Service (AIaaS) is provided to users. The AIaaS platform breaks down several common AI services and provides them as independent or packaged services in the cloud. This service model is similar to an AI-themed marketplace. All users can access and use one or more artificial intelligence services provided by the AIaaS platform through application programming interfaces.
[0051] For example, one type of AI cloud service could be an animation generation service, where a cloud server encapsulates the animation generation program provided in this application embodiment. Users invoke the animation generation service in the cloud service via a terminal, causing the cloud-deployed server to call the encapsulated animation generation program. Based on received input information containing multiple modalities, the program determines the target emotion corresponding to the multiple modalities of the input information and generates a virtual object animation based on the target emotion.
[0052] See Figure 1 , Figure 1 This is a schematic diagram of an application scenario for the animation generation system 10 provided in this application embodiment. The terminal 200 is connected to the server 100 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.
[0053] Terminal 200 can be used to obtain virtual object animation generation requests for input information in multiple modalities. For example, if a user inputs input information in multiple modalities through terminal 200, terminal 200 will automatically obtain the input information in multiple modalities and automatically generate virtual object animation generation requests for the input information in multiple modalities.
[0054] In some embodiments, an animation generation plugin may be embedded in the client running on the terminal to generate virtual object animations locally on the client based on multimodal input information. For example, after the terminal 200 obtains a virtual object animation generation request for input information of multiple modalities, it calls the animation generation plugin to encode the input information of each modality separately to obtain the encoding features of the input information of each modality. The encoding features corresponding to the input information of multiple modalities are fused to obtain the fusion features of the input information of multiple modalities. Based on the fusion features, sentiment prediction is performed on the input information of multiple modalities to obtain the target sentiment expressed by the input information of multiple modalities, and virtual object animations are generated based on the target sentiment.
[0055] In some embodiments, after the terminal 200 obtains a virtual object animation generation request for input information of multiple modalities, it calls the animation generation interface of the server 100 through the network 300. The server 100 encodes the input information of each modality separately to obtain the encoding features of the input information of each modality. It fuses the encoding features corresponding to the input information of multiple modalities to obtain the fusion features of the input information of multiple modalities. Based on the fusion features, it performs sentiment prediction on the input information of multiple modalities to obtain the target sentiment expressed by the input information of multiple modalities, generates virtual object animation based on the target sentiment, and sends the virtual object animation to the terminal 200.
[0056] See Figure 2 , Figure 2 This is a schematic diagram of the structure of an electronic device 500 for animation generation provided in an embodiment of this application. The example given is that the electronic device 500 is a server. Figure 2 The illustrated electronic device 500 for animation generation includes at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. The various components of the electronic device 500 are coupled together via a bus system 540. It is understood that the bus system 540 is used to implement communication between these components. In addition to a data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus; however, for clarity, ... Figure 2 The general labeled all buses as Bus System 540.
[0057] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0058] Memory 550 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 550 described in this application embodiment is intended to include any suitable type of memory. Memory 550 may optionally include one or more storage devices physically located away from processor 510.
[0059] In some embodiments, memory 550 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0060] Operating system 551 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;
[0061] The network communication module 552 is used to reach other computing devices via one or more (wired or wireless) network interfaces 520, exemplary network interfaces 520 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.
[0062] In some embodiments, the animation generation apparatus provided in this application can be implemented in software, for example, through the animation generation service in the server described above. Of course, this is not the limitation; the animation generation apparatus provided in this application can be provided in various software embodiments, including applications, software, software modules, scripts, or code.
[0063] Figure 2An animation generation device 555 stored in a memory 550 is shown. It can be software in the form of programs and plugins, such as an animation generation plugin, and includes a series of modules, including a data management module 5551, an encoding module 5552, a fusion module 5553, an emotion generation module 5554, an animation generation module 5555, and a training module 5556. The data management module 5551, encoding module 5552, fusion module 5553, emotion generation module 5554, and animation generation module 5555 are used to implement the animation generation function provided in the embodiments of this application, and the training module 5556 is used to train the neural network model in the embodiments of this application.
[0064] In other embodiments, the apparatus provided in this application can be implemented in hardware. As an example, the apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the animation generation method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0065] As mentioned above, the electronic device implementing the animation generation method of this application embodiment can be a terminal, a server, or a combination of both. Therefore, the executing entity of each step will not be described again below. See [link to relevant documentation]. Figure 3 , Figure 3 This is a first flowchart illustrating the animation generation method provided in this application embodiment, which will be combined with... Figure 3 The steps shown are explained.
[0066] In step 101, input information for multiple modalities of the virtual object is obtained.
[0067] As an example, the input information of a virtual object in multiple modalities can be at least two of the following: text modal input information, audio modal input information, and image modal input information. For example, the input information in multiple modalities can be text and audio; the input information in multiple modalities can also be text and images; the input information in multiple modalities can also be images and audio; the input information in multiple modalities can also be text, audio, and images, etc.
[0068] In this application embodiment, the optional modalities of the input information are text modality, audio modality, and image modality. Even with the input information of some modalities missing, the neural network model of this application can still accurately and stably achieve the output of the target emotion, exhibiting high robustness. Furthermore, users can flexibly adjust the modal combination of the input information, making it easy to deploy and apply in different application scenarios, such as customer service robots, emotion detection systems, and dialogue interaction virtual humans. At the same time, the input information of this application embodiment supports the simultaneous input of information from multiple modalities, thereby utilizing the complexity of multimodal input information to make comprehensive emotion prediction, resulting in high accuracy of emotion analysis.
[0069] In step 102, the input information of each modality is encoded to obtain the encoded features of the input information of each modality.
[0070] As an example, based on the encoder in the neural network model of this application embodiment, the input information of each modality is encoded. The encoder includes a text encoder, an audio encoder, and an image encoder. Based on the modality of the input information, the encoding operation is performed through the encoder corresponding to the modality, thereby obtaining the encoded features of the input information of each modality.
[0071] For example, a text encoder encodes the input information of a text modality to obtain the text encoding features of the text modality input information (i.e., the encoding features of the text modality input information); an audio encoder encodes the input information of an audio modality to obtain the audio encoding features of the audio modality input information (i.e., the encoding features of the audio modality input information); and an image encoder encodes the input information of an image modality to obtain the image encoding features of the image modality input information (i.e., the encoding features of the text modality input information).
[0072] In some embodiments, the text encoding features of the text modality input information are obtained by encoding the text modality input information through a text encoder. This can be achieved by: embedding the text modality input information to obtain the embedding features of the text modality input information; encoding the embedding features through the first text encoding layer in a series of cascaded text encoding layers; outputting the encoding result of the first text encoding layer to the subsequent cascaded text encoding layers, and continuing to perform text encoding and output the encoding results through the subsequent cascaded text encoding layers until the last text encoding layer is reached, and using the encoding result output by the last text encoding layer as the text encoding feature.
[0073] As an example, the text encoder includes an input layer and multiple cascaded text encoding layers. First, the input layer decomposes the input information of the text modality into multiple word fragments, and transforms each word fragment into a word embedding vector containing semantic, type, and positional information. The word embedding vectors corresponding to all word fragments are used as the text embedding features of the input information of the text modality. Then, the text embedding features are encoded by multiple cascaded text encoding layers. For example, the text embedding features are input into the first text encoding layer to obtain the encoding result of the first text encoding layer. The encoding result of the first text encoding layer is input into the second text encoding layer to obtain the encoding result of the second text encoding layer, and so on. The output results of the previous cascaded text encoding layer are encoded by subsequent cascaded text encoding layers until the output is reached by the last text encoding layer. The encoding result of the last text encoding layer is used as the text encoding feature. Each text encoding layer includes a multi-head attention layer and a fully connected feedforward network layer. After the multi-head attention layer and the fully connected feedforward network layer, there is a residual connection layer and a normalization layer.
[0074] The text encoder in this embodiment of the application effectively improves the encoding efficiency of the text encoder by processing each word fragment of the input information of the text modality. At the same time, by converting each word fragment into a word embedding vector containing semantic information, type information and position information, it can better retain more subtle information in the input information of the text modality, such as semantic changes.
[0075] In some embodiments, when the input information of multiple modalities includes the input information of an audio modality, the audio modality input information is encoded by an audio encoder to obtain the audio encoding features of the audio modality input information. This can be achieved by: extracting features from the audio modality input information to obtain an audio feature sequence of the audio modality input information; encoding the audio feature sequence to obtain the audio embedding features of the audio modality input information; and performing context encoding on the audio embedding features to obtain the audio encoding features.
[0076] As an example, the audio encoder includes an audio feature extraction layer, an audio coding layer, and a context network layer. First, based on the audio feature extraction layer, features are extracted frame by frame from the input information of the audio modality. The audio waveform of each frame of the input information of the audio modality is transformed into a corresponding feature vector. By combining the feature vector representations of each frame of the input information of the audio modality, the feature vector sequence of the input information of the audio modality is determined. Then, the feature vector sequence of the input information of the audio modality is encoded by the audio coding layer to obtain the audio embedding features of the input information of the audio modality. Specifically, the audio coding layer consists of multiple cascaded convolutional layers. The feature vector sequence of the input information is input into the first convolutional layer to obtain the encoding result of the first convolutional layer. The encoding and encoding results are continued through subsequent cascaded convolutional layers until the output is reached by the last convolutional layer. The output of the last convolutional layer is used as the audio embedding feature of the input information of the audio modality. Finally, since the audio embedding features obtained by the audio coding layer do not contain context information, context information needs to be added to the audio embedding features through the context network layer to obtain the audio encoded features. The context network layer includes multiple self-attention layers.
[0077] In real-life interactive scenarios, emotions are often multifaceted and complex, meaning that a single sentence may contain multiple emotions, and these emotions are not entirely independent of each other. In this application embodiment, contextual network layers are used to add contextual information to audio embedding features, which improves the audio encoder's ability to model speech sequences and avoids the neural network model from being confused or misunderstanding the context of the input information, thereby affecting the results of emotion recognition.
[0078] In some embodiments, when the input information of multiple modalities includes the input information of the image modality, the image modality input information is encoded by an image encoder composed of multiple cascaded convolutional layers to obtain the image encoding features of the image modality input information. This can be achieved as follows: the image modality input information is image encoded by the first convolutional layer in the multiple cascaded convolutional layers; the encoding result of the first convolutional layer is output to the subsequent cascaded convolutional layers, and image encoding and encoding result output are continued through the subsequent cascaded convolutional layers until the last convolutional layer is output, and the encoding result output by the last convolutional layer is used as the image encoding feature.
[0079] As an example, for detailed processing, please refer to the processing of multiple cascaded convolutional layers in the audio coding layer above, which will not be repeated here.
[0080] In step 103, the encoding features corresponding to the input information of multiple modalities are fused to obtain the fused features of the input information of multiple modalities.
[0081] As an example, before fusing the encoding features corresponding to the input information of multiple modalities, the encoding features corresponding to the input information of multiple modalities can be linearly transformed by the projection layer in the neural network model of this application embodiment to ensure that the tensor dimensions of the encoding features corresponding to the input information of multiple modalities are unified. The text encoding features, audio encoding features and image encoding features mentioned below are all tensors of the same dimension after linear transformation.
[0082] This application embodiment performs a linear transformation on the encoded features (encoded features of multiple modalities) corresponding to the input information of multiple modalities, converting the encoded features of multiple modalities into the same dimension. This allows the neural network model to process the encoded feature data of all modalities in a unified manner, reducing the memory and computing resources required by the neural network model in subsequent calculations (fusion processing of the encoded features of each modality), thereby improving the overall performance of the neural network model.
[0083] In some embodiments, the fusion of encoded features corresponding to input information from multiple modalities is achieved through N cascaded cross-attention layers. Figure 3 The step 103 shown can be implemented through step 201: In step 201, attention processing is performed on the encoding features corresponding to the input information of multiple modalities through N cascaded cross attention layers to obtain the fusion features of the input information of multiple modalities, where N is a positive integer greater than 1.
[0084] As an example, the multimodal input information can be any two or three of the following: text modal input information, image modal input information, and audio modal input information.
[0085] In this embodiment, the optional modalities of the input information are text, audio, and image. Users can flexibly adjust the modal combination of the input information. For example, a user can choose to input text and audio modal information, or text and image modal information, or audio and image modal information, or text, audio, and image modal information. Therefore, the model in this embodiment is easy to implement and deploy in different application scenarios.
[0086] In some embodiments, when the input information of multiple modalities includes two of the input information of text modality, audio modality, and image modality, step 201 can be implemented in the following way: through the first cross-attention layer, attention processing is performed on the encoding features corresponding to the input information of the two modalities respectively to obtain the attention features of the first cross-attention layer; through the nth cross-attention layer, attention processing is performed on the attention features of the (n-1)th cross-attention layer and the encoding features of the input information of one modality to obtain the attention features of the nth cross-attention layer, and the attention features of the Nth cross-attention layer are used as the fusion features of the input information of multiple modalities, where n is a positive integer that increases sequentially, 1 < n ≤ N, and one modality is any one of the two modalities.
[0087] As an example, when the input information of multiple modalities consists of two modalities, the first encoding feature can be any one of text encoding features, audio encoding features, and image encoding features, and the second encoding feature can be any one of text encoding features, audio encoding features, and image encoding features other than the first encoding feature. Attention processing is performed on the first and second encoding features. Specifically, the first and second encoding features are input into the first cross-attention layer for attention processing to obtain the attention features of the first cross-attention layer. Similarly, the attention features of the (n-1)th cross-attention layer and the second encoding feature are input into the nth cross-attention layer to obtain the attention features of the nth cross-attention layer, until the attention features of the Nth cross-attention layer (the last cross-attention layer) are used as the fusion features of the input information of multiple modalities.
[0088] For example, the specific process of attention processing for text-encoded features and audio-encoded features through six cascaded cross-attention layers is as follows: The text-encoded features are treated as the query (Q), and the audio-encoded features are treated as the key (K) and value (V). Based on the cross-attention calculation formula, keeping K and V unchanged, six stacked operations are performed on Q.
[0089] The formula for calculating cross-attention is as shown in formula (1.1):
[0090]
[0091] Among them, Q text For text encoding features, K audio and V audio For audio coding features, It is obtained by transposing the audio coding features, d k Q represents the dimension of the audio coding features. iThe attention features output by each cross-attention layer will be used to perform 6 stacked operations on the Q-factor. 6 Features that fuse input information from multiple modalities.
[0092] The embodiments of this application effectively achieve deep fusion of the encoding features of two modal input information through multiple cascaded cross-attention layers, which is also beneficial for discovering the correlation between input information of different modalities.
[0093] In some embodiments, when the input information of multiple modalities includes input information of text modality, input information of audio modality, and input information of image modality, step 201 can be implemented in the following way: through the first M cross-attention layers of N cascaded cross-attention layers, attention processing is performed on the encoding features corresponding to the input information of the two modalities respectively to obtain the fusion features of the input information of the two modalities; through the last NM cross-attention layers of N cascaded cross-attention layers, attention processing is performed on the fusion features of the input information of the two modalities and the encoding features of the input information of the third modality to obtain the fusion features of the input information of multiple modalities, where M is a positive integer less than N, the two modalities are any two modalities among text modality, audio modality, and image modality, and the third modality is a modality other than the two modalities among text modality, audio modality, and image modality.
[0094] As an example, when the input information of multiple modalities consists of three modalities, attention processing is performed on the third, fourth, and fifth encoded features through N cascaded cross-attention layers to obtain the fusion feature of the input information of multiple modalities. Specifically, the third and fourth encoded features are cross-attention processed through the first to the Mth cross-attention layers to obtain the third fusion feature, and the third fusion feature and the fifth encoded feature are cross-attention processed through the (M+1)th to the Nth cross-attention layers to obtain the fusion feature of the input information of multiple modalities.
[0095] For example, attention processing is performed on text encoding features, audio encoding features, and image encoding features through 12 cascaded cross-attention layers. The processing can be divided into two stages. In the first stage, the text encoding features are used as the query (Q), and the representation of the audio encoding features are used as the key (K) and value (V). Based on the calculation formula of cross-attention (1.1), keeping K and V unchanged, Q is stacked 6 times to obtain the fused features of text encoding features and audio encoding features. In the second stage, the fused features of text encoding features and audio encoding features are used as Q, and the representation of the image encoding features are used as the key (K) and value (V). Keeping K and V unchanged, based on the calculation formula of cross-attention, the fused features of text encoding features and audio encoding features are stacked 6 times to obtain the fused features of input information of multiple modalities.
[0096] In some embodiments, "attention processing is performed on the encoding features corresponding to the input information of the two modalities through the first M cross-attention layers of N cascaded cross-attention layers to obtain the fusion features of the input information of the two modalities" can be achieved in the following way: Attention processing is performed on the encoding features corresponding to the input information of the two modalities through the first cross-attention layer to obtain the attention features of the first cross-attention layer; Attention processing is performed on the attention features of the (m-1)th cross-attention layer and the encoding features of the input information of one modality through the mth cross-attention layer to obtain the attention features of the mth cross-attention layer; The attention features of the Mth cross-attention layer are used as the fusion features of the input information of the two modalities, where m is a positive integer that increases sequentially, 1 < m ≤ M, and one modality is any one of the two modalities.
[0097] As an example, the third and fourth encoded features are cross-attention processed through the first to the Mth cross-attention layers to obtain the third fusion feature. Specifically, the third and fourth encoded features are input into the first cross-attention layer for attention processing to obtain the attention feature of the first cross-attention layer; and so on, the attention feature of the (m-1)th cross-attention layer and the fourth encoded feature are input into the mth cross-attention layer to obtain the attention feature of the mth cross-attention layer, until the attention feature of the Mth cross-attention layer is used as the third fusion feature.
[0098] For example, text-encoded features are used as queries (Q, Query), and audio-encoded features are used as keys (K, Key) and values (V, Value). Based on the calculation formula (1.1) for cross-attention, the fused features of text-encoded features and audio-encoded features are obtained.
[0099] In some embodiments, "attention processing is performed on the fusion features of the input information of two modalities and the encoding features of the input information of a third modality through the last NM cross-attention layers in N cascaded cross-attention layers to obtain the fusion features of the input information of multiple modalities," which can be achieved in the following way: Attention processing is performed on the fusion features of the input information of two modalities and the encoding features of the input information of a third modality through the (M+1)th cross-attention layer to obtain the attention features of the (M+1)th cross-attention layer; attention processing is performed on the attention features of the (n-1)th cross-attention layer and the encoding features of the input information of a third modality through the nth cross-attention layer to obtain the attention features of the nth cross-attention layer, and the attention features of the Nth cross-attention layer are used as the fusion features of the input information of multiple modalities, where n is a positive integer that increases sequentially, and M < n ≤ N.
[0100] As an example, the third fusion feature and the fifth encoding feature are cross-attention processed through the (M+1)th to the Nth cross-attention layers to obtain the fusion features of the input information of multiple modalities. Specifically, the third fusion feature and the fifth encoding feature are input into the (M+1)th cross-attention layer for attention processing to obtain the attention features of the (M+1)th cross-attention layer. Similarly, the attention features of the (n-1)th cross-attention layer and the fifth encoding feature are input into the nth cross-attention layer to obtain the attention features of the nth cross-attention layer, until the attention features of the Nth cross-attention layer are used as the fusion features of the input information of multiple modalities.
[0101] For example, based on the cross-attention calculation formula (1.1), the fusion feature Q of text coding features and audio coding features is obtained through 6 cascaded cross-attention layers. 6 (The fusion features of the input information from the two modalities), and then through another 6 cascaded cross-attention layers, the specific process of attention processing on the fusion features of the input information from the two modalities and the image encoding features is as follows: The fusion features of the input information from the two modalities Q combine As a query (Q, Query), the image encoding features are used as keys (K, Key) and values (V, Value). Based on the calculation formula (1.2) of cross attention, K and V are kept unchanged, and Q is stacked 6 times.
[0102] The formula for calculating cross-attention is as shown in formula (1.2):
[0103]
[0104] Among them, Q combine K is a fusion feature of text encoding features and audio encoding features. pic and Vpic Encoding features for images It is obtained by transposing the image encoding features, d k Q represents the dimension of the image encoding features. j The attention features output by each cross-attention layer will be used to perform 6 stacked operations on the Q-factor. 6 Features that fuse input information from multiple modalities.
[0105] This application embodiment fuses the encoding features of multimodal input information, thereby protecting the rich emotional information contained in the multimodal input information and achieving effective processing and correct classification of complex and diverse emotional expressions in the multimodal input information, thus improving the accuracy and robustness of the neural network model's emotion recognition.
[0106] In some embodiments, when the input information of multiple modalities includes input information of text modality, input information of audio modality, and input information of image modality, step 201 can be implemented in the following way: Attention processing is performed on the encoding features of the first modality input information and the encoding features of the second modality input information through the first M cross-attention layers of N cascaded cross-attention layers to obtain a first fusion feature; attention processing is performed on the encoding features of the first modality input information and the encoding features of the third modality input information through the last NM cross-attention layers of N cascaded cross-attention layers to obtain a second fusion feature, where M is a positive integer less than N, the first modality and the second modality are any two of the text modality, audio modality, and image modality, and the third modality is a modality other than any two of the text modality, audio modality, and image modality; the first fusion feature and the second fusion feature are fused to obtain the fusion feature of the input information of multiple modalities.
[0107] For example, the first six cross-attention layers in the 12 cascaded cross-attention layers are used to process the text encoding features and audio encoding features to obtain the first fusion feature. The last six cross-attention layers in the 12 cascaded cross-attention layers are used to process the text encoding features and image encoding features to obtain the second fusion feature. The first fusion feature and the second fusion feature are then concatenated or summed to obtain the fusion feature of the input information of multiple modalities.
[0108] When the data throughput of the neural network model in this application embodiment is large, the computational efficiency of the neural network model is improved by performing simple splicing or summation on the first fusion feature and the second fusion feature. And after verification, the output result of the neural network model still has a high accuracy.
[0109] In step 104, based on the fusion features, sentiment prediction is performed on the input information of multiple modalities to obtain the target sentiment expressed by the input information of multiple modalities.
[0110] As an example, different emotions of virtual objects correspond to an emotion label. The emotion label is a one-hot vector with an encoding length equal to the total number of emotions. Through the feedforward neural network layer in the neural network model of this application embodiment, the fused features are transformed into an emotion probability vector. The feedforward neural network layer includes a first linear transformation layer, an activation layer, a dropout layer, and a second linear transformation layer.
[0111] In some embodiments, see Figure 4 , Figure 4 This is a schematic diagram of the second process of the animation generation method provided in the embodiments of this application. Figure 3 Step 104 shown can be implemented through the following steps 1041 to 1044, which are explained in detail below.
[0112] In step 1041, the fused features are linearly transformed to obtain the linearly transformed fused features.
[0113] As an example, based on the first linear transformation layer, a linear transformation is performed on the fused features to obtain the linearly transformed fused features. The formula for the linear transformation function of the first linear transformation layer is shown in formula (1.3):
[0114] z1=ω1x+b1(1.3)
[0115] Where x is the fused feature, ω1 is the weight tensor used for linear transformation, b1 is the bias tensor, and the output z1 is the fused feature after linear transformation.
[0116] In step 1042, the fusion features after linear transformation are activated to obtain the activated fusion features.
[0117] As an example, activation of the fused features after linear transformation is achieved through activation functions, which can be step functions, sigmoid functions, hyperbolic tangent functions, rectified linear units (ReLU functions), etc., without limitation here.
[0118] For example, the tanh function can be used to activate the fused features after linear transformation. The tanh function is a symmetric activation function, and its activation function formula is shown in formula (1.4):
[0119] a1=tanh(z1)(1.4)
[0120] Where a1 is the activation fusion feature after processing by a nonlinear activation function.
[0121] In step 1043, the fusion features after activation are reduced in dimensionality to obtain the sentiment probability vectors of input information in multiple modalities.
[0122] As an example, before performing dimensionality reduction on the fused features after activation, some elements of the fused features after activation can be randomly set to 0 through a dropout layer to obtain the fused features after dropout. Then, the fused features after dropout are dimensionality reduced to obtain the sentiment probability vectors of input information for multiple modalities.
[0123] The formula for the discard layer is shown in formula (1.5):
[0124] a2 = dropout(a1, p) (1.5)
[0125] Where p is the probability of discarding an element (the probability of setting an element to 0), and a2 is the fused feature after the discarding process.
[0126] As an example, dimensionality reduction of the fused features after discarding is achieved through a second linear transformation layer, where the linear transformation function of the second linear transformation layer is shown in Equation (1.6):
[0127] y=ω2a2+b2(1.6)
[0128] Where ω2 is a dimension-reduced tensor with dimensions of fused feature dimension * sentiment tag dimension, b2 is a one-dimensional bias vector, and y is the sentiment probability vector of input information from multiple modalities.
[0129] In step 1044, the target emotion expressed by the input information of multiple modalities is determined based on the emotion probability vector.
[0130] As an example, the sentiment label is determined by the maximum probability in the sentiment probability vector, and then the target sentiment expressed by the input information of multiple modalities is determined based on the sentiment label.
[0131] For example, when the emotions of virtual objects include joy, sadness, anger, surprise, neutrality, disgust, fear, and doubt, the emotional labels are: joy [1,0,0,0,0,0,0,0], sadness [0,1,0,0,0,0,0,0], anger [0,0,1,0,0,0,0,0], surprise [0,0,0,1,0,0,0,0], neutrality [0,0,0,0,1,0,0,0], disgust [0,0,0,0,0,1,0,0], fear [0,0,0,0,0,0,1,0], and doubt [0,0,0,0,0,0,1,0]. When the output sentiment probability vector is [0.01,0.01,0.01,0.2,0.01,0.01,0.05,0.7], the probability of the target sentiment expressed by the input information of the current multiple modalities is 0.01 for joy, 0.01 for sadness, 0.01 for anger, 0.2 for surprise, 0.01 for neutrality, 0.01 for disgust, 0.05 for fear, and 0.7 for doubt. Therefore, the maximum probability in the output sentiment probability vector is 0.7, and its corresponding sentiment label is [0,0,0,0,0,0,0,1], and the target sentiment is doubt.
[0132] This application embodiment achieves dimensionality reduction of the fused features by performing a linear transformation on the fused features. While retaining the important information in the fused features, it effectively removes redundant information, reducing the high-dimensional fused features to a 1-dimensional sentiment probability vector. This facilitates intuitive observation and analysis of the sentiment probability vector, thereby obtaining the target sentiment.
[0133] In step 105, based on the target emotion and at least some of the input information from multiple modalities, an animation is generated for the virtual object to obtain the virtual object animation.
[0134] In some embodiments, see Figure 5 , Figure 5 This is a schematic diagram of the third process of the animation generation method provided in the embodiments of this application. Figure 3 Step 105 shown can be implemented through the following steps 1051 to 1052, which are explained in detail below.
[0135] In step 1051, motion data of the virtual object is determined based on the target emotion and at least some of the input information from multiple modalities;
[0136] As an example, the motion data of the virtual object includes facial expression motion data and lip movement data. The facial expression motion data of the virtual object is determined based on the target emotion, and the lip movement data of the virtual object is determined based on input information from at least some of the multiple modalities.
[0137] In some embodiments, step 1051 can be implemented by the following steps: when the input information of multiple modalities includes text modal input information, at least a portion of the input information of multiple modalities is determined as text modal input information; when the input information of multiple modalities does not include text modal input information, text modal input information is generated based on the input information of multiple modalities; and motion data of the virtual object is generated based on the target emotion and the text modal input information.
[0138] As an example, the meaning of the above-mentioned input information of at least some multiple modalities is that the current input information of multiple modalities needs to include at least text modal input information. It is known that the input information of multiple modalities can be at least two of the following: text modal input information, audio modal input information, and image modal input information. When the input information of multiple modalities does not include text modal input information, it means that the current input information of multiple modalities is audio modal input information and image modal input information. In this case, it is necessary to first construct the corresponding text modal input information based on the audio modal input information, and then generate facial expression motion data of the virtual object that matches the target emotion based on the text modal input information, generate lip movement data of the virtual object that matches the text modal input information, and then generate motion data of the virtual object based on the facial expression motion data and lip movement data of the virtual object.
[0139] For example, if the target expression is joy, the text modal input is "The weather is so nice today", the virtual object's facial expression motion data is joyful facial expression motion data, and the virtual object's lip movement data is lip movement data that matches "The weather is so nice today".
[0140] By using a neural network model to achieve emotion recognition based on multimodal input information, the target emotion is obtained, which then drives virtual objects to make feedback animations that match the target emotion and input information, making the virtual objects more vivid and realistic, and greatly improving the user's immersion and satisfaction.
[0141] In some embodiments, step 1051 can be achieved by the following steps: determining the response emotion and response text of the virtual object to respond to the input information based on the target emotion and input information of multiple modalities; and generating motion data of the virtual object based on the response emotion and response text.
[0142] As an example, based on the target emotion, the response emotion of the virtual object to respond to the input information is determined. Based on the input information of multiple modalities, the response text of the virtual object to respond to the input information is determined. Then, facial expression motion data of the virtual object that matches the response emotion is generated, and lip movement data of the virtual object that matches the response text is generated. Then, the facial expression motion data and lip movement data of the virtual object are combined to generate motion data of the virtual object.
[0143] For example, if the target expression is anger, the text modal input is "Why did you take my spot!", the determined response emotion is fear, the determined response text is "Please don't yell at me.", the virtual object's facial expression motion data is fear facial expression motion data, and the virtual object's lip movement data is lip movement data that matches "Please don't yell at me."
[0144] The embodiments of this application can also be applied to human-computer interaction scenarios that require dialogue functions, such as customer service robots. The virtual object can make appropriate response emoticons and response text based on multimodal input information, which enhances the interaction quality between the virtual object and the user, thereby greatly improving user satisfaction.
[0145] In step 1052, motion data is bound to the initial object model of the virtual object to obtain a virtual object animation that conforms to the target emotion and matches the input information of multiple modalities.
[0146] As an example, based on the facial bone structure of the initial object model, the facial expression motion data and lip movement data in the motion data are matched with the corresponding facial bones of the initial object model. Based on the mapping relationship of the matching results, the motion data is bound to the initial object model of the virtual object to obtain a virtual object animation that conforms to the target emotion and matches the input information of multiple modalities.
[0147] In some embodiments, see Figure 6 , Figure 6 This is a schematic diagram of the fourth process of the animation generation method provided in this application embodiment. The animation generation method in this application embodiment is implemented through a neural network model, and the training process of the neural network model is through... Figure 6 Steps 301 to 304 shown are implemented as follows, and will be explained in detail below.
[0148] In step 301, anchor point samples of input information for multiple modalities of the virtual object, positive samples of input information for multiple modalities, and negative samples of input information for multiple modalities are obtained.
[0149] As an example, a random strategy is used to select multiple modal input information anchor samples from the training sample set corresponding to each sentiment label of the neural network model, or the center point of the feature space of the training sample set corresponding to each sentiment label is selected as the multiple modal input information anchor samples. Based on the multiple modal input information anchor samples, multiple modal positive input samples and multiple modal negative input samples are constructed. Among them, the multiple modal positive input samples are samples with sentiment labels consistent with the multiple modal input information anchor samples, and the multiple modal negative input samples are samples with sentiment labels inconsistent with the multiple modal input information anchor samples.
[0150] As an example, in Figure 6 Before step 301, training samples for the deep neural network model can be constructed. These training samples include input information samples from multiple modalities and output sentiment labels. The process of constructing training samples is as follows: Select textual data rich in emotional expression from different application scenarios, audio data rich in emotion from different accents, and image data rich in explicit facial expressions or body language as input information samples for multiple modalities. Additionally, if audio data is insufficient, a TTS system can be used to generate emotionally rich audio data in batches, effectively enhancing the diversity and richness of the input information samples for multiple modalities. Through the construction of the above training samples, the number and diversity of training samples are significantly increased. This significantly improves the model's sentiment recognition performance for different input information in different environments during neural network model training, thereby enhancing the model's generalization ability in different application scenarios.
[0151] In step 302, the neural network model is used to perform sentiment prediction on the anchor samples of input information of multiple modalities, the positive samples of input information of multiple modalities, and the negative samples of input information of multiple modalities, respectively, to obtain the first sentiment probability vector corresponding to the anchor samples of input information, the second sentiment probability vector corresponding to the positive samples of input information, and the third sentiment probability vector corresponding to the negative samples of input information.
[0152] As an example, a neural network model is used to predict the sentiment of anchor samples of input information from multiple modalities to obtain a first sentiment probability vector. Sentiment prediction is then performed on positive samples of input information from multiple modalities to obtain a second sentiment probability vector. Sentiment prediction is then performed on negative samples of input information to obtain a third sentiment probability vector.
[0153] In step 303, a triplet loss function is constructed based on the first sentiment probability vector, the second sentiment probability vector, and the third sentiment probability vector.
[0154] In some embodiments, Figure 6Step 303 shown can be implemented through the following steps: calculate the similarity between the first sentiment probability vector and the second sentiment probability vector to obtain the first similarity between the first sentiment probability vector and the second sentiment probability vector; calculate the similarity between the first sentiment probability vector and the third sentiment probability vector to obtain the second similarity between the first sentiment probability vector and the third sentiment probability vector; and construct a triplet loss function based on the first similarity and the second similarity.
[0155] As an example, the similarity between two different sentiment vectors can be calculated by measuring the Euclidean distance, cosine similarity, or Manhattan distance between sentiment probability vector samples. The similarity between sentiment probability vector samples represents the similarity between the input information samples corresponding to the sentiment probability vector samples. For example, the greater the similarity between two sentiment probability vector samples, the more similar the two input information samples corresponding to the sentiment probability are; the smaller the similarity between two sentiment probability vector samples, the greater the difference between the two input information samples corresponding to the sentiment probability.
[0156] The triplet loss function is shown in formula (1.7):
[0157] L triplet =max(0,sim(z) a ,z n )-sim(z a ,z p +margin(1.7)
[0158] Among them, z a z is the first sentiment probability vector corresponding to the input information anchor sample. p Let z be the second sentiment probability vector corresponding to the positive samples of the input information. n The third sentiment probability vector is the negative sample of the input information. Margin is a hyperparameter representing the boundary value of similarity. It is used to ensure that the similarity between the second sentiment probability vector and the first sentiment probability vector is at least greater than the similarity between the third sentiment probability vector and the first sentiment probability vector by a fixed boundary value. Sim is the similarity between two sentiment probability vector samples, and its value ranges from [0,1]. The greater the similarity, the more similar the two sentiment probability vector samples are.
[0159] By minimizing the triplet loss, the model learns to group the vector representations of input information samples with the same emotion together and separate the vector representations of input information samples with different emotions in the feature space. This enables the model to more accurately identify subtle emotional changes, thereby improving the performance of the neural network model in complex emotion recognition tasks. Furthermore, based on the triplet loss function, the model can robustly handle intra-class variation and inter-class similarity issues that may exist in the high-dimensional feature space.
[0160] In step 304, the parameters of the neural network model are updated based on the triplet loss function, and the updated parameters of the neural network model are used as the parameters of the trained neural network model.
[0161] In some embodiments, Figure 6 Step 304 shown can be implemented through the following steps: fusing the encoded features corresponding to the input information anchor samples of multiple modalities through a neural network model to obtain fused feature samples of the input information anchor samples of multiple modalities, and constructing a cross-entropy loss function based on the fused feature samples; integrating the cross-entropy loss function and the triplet loss function to obtain the total loss function of the neural network model; and updating the parameters of the neural network model based on the total loss function.
[0162] As an example, the cross-entropy loss value is only related to the input information anchor samples of multiple modalities. Through the neural network model, the fused feature samples of the input information anchor samples of multiple modalities are obtained, and the cross-entropy loss function is constructed based on the fused feature samples. The formula for the cross-entropy loss function is as shown in formula (1.8):
[0163]
[0164] Where C is the total set of emotion categories, including all emotion categories c and p. c Let y be the probability that the anchor sample belongs to each sentiment category. c It is the correlation coefficient of the sentiment label of the anchor sample. If the sentiment label of the anchor sample is consistent with the sentiment label of the current sentiment category c, then y c If it is 1, then y c It is 0.
[0165] As an example, the total loss function can be expressed as the sum of the cross-entropy loss function and the triplet loss function according to a certain proportional coefficient, as shown in formula (1.9):
[0166] L combined =αL triplet +(1-α)L cross-entropy (1.9)
[0167] Where α and 1-α are the weight parameters for triplet loss and cross-entropy loss, respectively.
[0168] The embodiments of this application calculate the total loss function, minimize the cross-entropy loss and triplet loss, and perform backpropagation to adjust the model parameters. This enables the model to gradually learn to distinguish the features of different emotional states in complex emotion recognition tasks, achieving higher accuracy in multimodal emotion recognition tasks. At the same time, it provides more possibilities for the in-depth development of emotion analysis in academic research, and opens up new ideas for the future application of emotion analysis technology in industry.
[0169] This application can be applied to various human-computer interaction scenarios, such as in the application scenario of virtual object role-playing, or in the application scenario of simulating virtual characters (virtual objects include virtual characters) in games. Its core function is to predict and drive the emotional expression of virtual objects, making the expressions of virtual objects more realistic and interactive. Compared with traditional virtual objects, through this application, virtual objects can comprehensively analyze voice input and text input to predict the emotions corresponding to voice input and text input. Then, by adjusting the dimensions of the virtual object's expression, tone, and actions, the virtual object can express these emotional states more vividly, providing users with a richer and more realistic human-computer interaction experience. For example, in the application scenario of simulating virtual characters in a game, the virtual characters may need to react accordingly based on the development of the plot or the player's choices. Through the embodiments of this application, these virtual characters can more accurately express corresponding emotions, such as surprise, joy, or sadness. When the plot of the game develops to a critical moment, the virtual characters can show tense or excited emotions by changing the tone and speed of their voice, adjusting their facial expressions and body language. Similarly, when the plot of the game becomes peaceful or warm, the virtual characters can also express calmness and joy through a gentle tone, warm expression, and relaxed posture.
[0170] The following will describe exemplary applications of the embodiments of this application in application scenarios where virtual objects are used.
[0171] With the development of artificial intelligence technology, the demand for sentiment analysis technology in various real-time human-computer interaction scenarios is increasing. For example, in scenarios such as customer service robots, sentiment detection systems, and virtual object interactive dialogue, sentiment analysis technology, which is fast-responding, highly accurate, and robust, can save a lot of human labor costs. While improving the user's human-computer interaction experience, it gradually changes the way humans interact with machines, thereby bringing more convenience and progress to human society. However, the sentiment analysis models in related technologies can often only perform sentiment recognition on input information of a specific modality in a specific scenario, and are not accurate in sentiment recognition on input information of other scenarios or other modalities.
[0172] To address the aforementioned issues, this application proposes an animation generation method. This method incorporates a large amount of emotional text data from speakers with different styles and in different scenarios, and employs TTS (Text-to-Speech) technology, which injects emotional styles into synthesized speech, to enhance the data. This data is then used to train a multimodal emotion recognition model (equivalent to the neural network model mentioned above), improving the model's generalization ability across different scenarios, enhancing its performance on various datasets, reducing bias, and increasing stability. This addresses the shortcomings of emotion analysis in zero-shot tasks. The trained multimodal emotion recognition model receives multimodal information input, enabling it to effectively process and correctly classify complex and diverse emotional expressions. Furthermore, this application also designs a virtual object expression-driven system. When the virtual object interacts with the user, it provides emotion analysis functionality and promptly sends emotion-driven signals to drive the virtual object to provide appropriate emotional feedback, improving the accuracy and robustness of emotion analysis technology and laying a solid foundation for the future development of this field.
[0173] The virtual object expression driving system in the embodiments of this application will be introduced next.
[0174] In applications using virtual objects, this embodiment can drive the virtual object to display different expressions based on real-time input information. (See also...) Figure 7 , Figure 7This is a schematic diagram of the first product for virtual object emotion driving provided in this application embodiment. In order to realize the animation driving of virtual object 701, the user only needs to input text in the text input area 702 and audio in the audio input area 703 in the user interface to generate instructions to drive the multimodal emotion recognition model. When it is not possible to input audio through human dubbing, text-to-speech (TTS) technology can be used to convert the input text into emotional audio. Based on the received instructions, the background automatically performs sentence segmentation processing on the text (if there is speech, it will first recognize the speech), splitting the input text into multiple independent short sentences. Based on the text segmentation method, the input audio is also segmented to obtain multiple independent short audios corresponding to multiple independent short sentences. Then, the multimodal information composed of independent short sentences and short audios is input into the trained multimodal emotion recognition model for emotion recognition, and the emotion label of each short sentence is output. After receiving the expression labels corresponding to each independent short sentence of the input text, the central control program of the animation generation algorithm will use them as the access data for the expression animation of the virtual object. By loading the expression configuration information corresponding to the emotion label, the virtual object's face will present the corresponding emotion. This application embodiment allows users to freely edit the dialogue and plot of virtual objects, making the role-playing experience of virtual objects more realistic. Moreover, compared with the existing technology that drives the emotions of virtual objects through text recognition, this application embodiment can automatically identify and align multimodal input information, and identify fine-grained emotions in the input information based on sentence segmentation, so that the virtual object has coherent emotional fluctuations when speaking a whole sentence, improving the accuracy and diversity of emotional expression. Furthermore, by freely editing the dialogue and plot of virtual objects, the role-playing experience of virtual objects becomes more realistic.
[0175] The following explains the speech recognition function mentioned above: It identifies the text characters in the speech and assigns a range of corresponding audio timestamps to each character in the speech. Finally, for each sentence of speech information, a range sequence of audio timestamps corresponding to each character is obtained, which is convenient for subsequent input into the multimodal emotion recognition model.
[0176] The automatic sentence segmentation function mentioned above is explained below: First, it searches for specific punctuation marks in the text input, such as periods, question marks, and exclamation marks. Then, it searches for specific transition or conjunction words, such as "but" and "however." Finally, it segments the text into multiple fragments based on these punctuation marks and words. Each fragment usually corresponds to a complete sentence. If a sentence fragment is too short, it may be merged with the previous or next sentence to avoid generating too many fragmented sentences.
[0177] For example, such as Figure 7As shown, the text "The weather is so nice today, the clothes I hung out to dry this morning are dry by noon!" is input into the text input area 702 on the right side of the user interface, and the audio of "The weather is so nice today, the clothes I hung out to dry this morning are dry by noon!" is input into the audio input area 703. Based on symbols (such as periods, question marks, and exclamation marks) and conjunctions indicating transitions (such as but, however, etc.), the input text is split into multiple independent short sentences: "today," "weather," "so nice," and "the clothes I hung out to dry this morning are dry by noon." Based on the text splitting method, the input audio is also split into short audio segments: "today," "weather," "so nice," and "the clothes I hung out to dry this morning are dry by noon." Emotion recognition is performed on the multimodal input information composed of each independent short sentence and short audio segment, and the emotion label for each short sentence is output. Based on the output emotion label, the virtual character 701's face displays a happy emotion. See also Figure 8 , Figure 8 This is a schematic diagram of a second product with virtual object emotion-driven processing provided in this application embodiment. The text "How could you do this! You took my spot!" is input into the text input area 802 on the right side of the user interface, and the audio of "How could you do this! You took my spot!" is input into the audio input area 803. The input text is split into "How could you", "How could you do this", "You took", "My", "Spot", and "Place". At the same time, the input audio is split based on the text splitting method. Emotion recognition is performed on the multimodal input information composed of each independent short sentence and short audio, and the emotion label of each short sentence is output. Based on the output emotion label, the face of the virtual object 801 is made to show the emotion of anger.
[0178] The emotion-driven mechanism in this application makes virtual characters and idols more attractive and realistic in their interactions with users, significantly enhancing user immersion and satisfaction. Through accurate emotion prediction and expression, virtual objects can behave more vividly and realistically in various interactive scenarios, bringing users a lifelike interactive experience.
[0179] The following section will introduce the overall architecture of the multimodal emotion recognition model.
[0180] The multimodal emotion recognition model in this application embodiment can receive input data in multiple modalities, such as text, audio, and images. When the model receives bimodal input of speech and text, see [link to relevant documentation]. Figure 9 , Figure 9This is a schematic diagram of the structure of the multimodal emotion recognition model provided in this application embodiment. Text and audio are encoded by an encoder module corresponding to each modality to obtain text-encoded features and audio-encoded features. A projection layer makes the tensor dimensions of the text-encoded features and audio-encoded features consistent. Then, based on a cross-attention mechanism, the modality fusion of the text-encoded features and audio-encoded features is achieved to obtain fused features. The fused features are input into a feedforward neural network to obtain an emotion probability vector. The label corresponding to the highest probability in the probability vector is taken. For example, when the emotion probability vector is [0.1, 0.01, 0.84, 0.01, 0.01, 0.01, 0.01, 0.01], the highest probability in the emotion probability vector is 0.84. The label corresponding to the highest probability is converted into the corresponding expression label. Alternatively, multiple emotion labels can be selected as the final output label according to the probability distribution.
[0181] Meanwhile, multimodal emotion recognition models can also handle classification tasks with single-modal inputs, such as emotion prediction tasks with only audio or text inputs. The model will automatically mask the missing modalities, and the model has been verified to still have a certain degree of robustness and high emotion prediction accuracy.
[0182] Next, we will take the case where the input data of the multimodal emotion recognition model is bimodal data of speech and text as an example. Before training the multimodal emotion recognition model, it is necessary to build a training dataset. The training data of the multimodal emotion recognition model includes three parts: text samples, audio samples corresponding to the text samples, and output emotion labels.
[0183] For text and audio samples, this application embodiment selects a variety of text materials with rich emotional expression. In particular, this application embodiment adopts a TTS system that can inject emotional style into synthesized speech. By batch generating corresponding emotional audio data through input text, the diversity and richness of audio data are effectively enhanced, thereby making up for the lack of quantity of original audio plus text data.
[0184] For emotion labels, this application uses one-hot vector encoding as emotion labels, encoding eight emotion states: joy, sadness, anger, surprise, neutral, disgust, fear, and suspicion into one-hot vectors of length 8. For example, the one-hot vector representing the emotion label of joy is [1,0,0,0,0,0,0,0]. These one-hot vectors of emotion labels are used as target labels for multi-classification problems to supervise the training process of the multimodal emotion recognition model.
[0185] The following section uses the case where the input data for a multimodal emotion recognition model is bimodal data of speech and text as an example to introduce the training process of the multimodal emotion recognition model.
[0186] First, the multimodal input information is encoded using a multimodal model encoding module.
[0187] For different modalities of information input, the input information can be converted into corresponding encoded features by encoders of different modalities. When the input of the model is text and audio, the text is encoded by a text encoder and the audio is encoded by an audio encoder, converting the information input to the model into hidden layer vectors.
[0188] For the audio encoder, the embodiments of this application employ a self-supervised pre-trained model, which can capture rich features in the sound signal and effectively convert the audio spectrum signal into a tensor containing speech information and emotional features, as audio coding features. For a detailed description of the audio encoder, please refer to the above description of audio encoding; it will not be repeated here.
[0189] For the text encoder, this embodiment employs a pre-trained language model capable of understanding subtle semantic and sentiment differences in text data and converting the text into corresponding tensor representations as text encoding features. For a detailed description of the text encoder, please refer to the above description of text encoding; it will not be repeated here.
[0190] Then, the output of the multimodal model encoding module is fused through the multimodal fusion module.
[0191] The multimodal fusion module includes a projection layer, a feature fusion module, and a feedforward neural network module. The following will introduce the multimodal fusion module of this application embodiment from these three aspects.
[0192] Regarding the projection layer, the projection layer performs linear transformations on the text coding features and audio coding features respectively, resulting in text coding features and audio coding features with the same dimension after linear transformation.
[0193] Regarding the feature fusion module, in order to achieve effective fusion between input data from different modalities, this application embodiment designs a feature fusion module based on a cross-attention mechanism. The representation of the linearly transformed text encoding features is used as the query (Q, Query), and the representation of the linearly transformed audio encoding features is used as the key (K, Key) and value (V, Value). Based on the calculation formula of cross-attention, keeping K and V unchanged, Q is stacked 6 times. This operation is conducive to the deep fusion of the two modalities and discovers the correlation between text and audio. While ensuring the accuracy of emotion recognition, it realizes the modal fusion of multimodal inputs and finally outputs fused features. The cross-attention calculation formula is shown in Formula (1.1) above.
[0194] The feature fusion module enables each text feature to dynamically focus on audio features, thereby extracting the most relevant audio coding features for the current text content. This allows the model to better understand information within each modality and achieve cross-modal alignment, facilitating the interpretation of complex human emotions from both speech and text perspectives.
[0195] Regarding the feedforward neural network module, the feedforward neural network module processes the fused features through a feedforward neural network (FNN) to obtain the predicted sentiment probability vector. The feedforward neural network module includes a linear transformation function and a nonlinear activation function: First, the fused features are subjected to a first linear transformation through a first linear transformation layer to obtain the fused features after the first linear transformation, where the linear transformation function is shown in formula (1.3) above; Second, the fused features after the first linear transformation are subjected to a nonlinear transformation through an activation layer to obtain the activated fused features, where the activation function formula is shown in formula (1.4) above; Third, a dropout layer is used to randomly set some elements of the input activated fused features to 0 to obtain the fused features after the dropout process, where the dropout layer formula is shown in formula (1.5) above; Finally, the sentiment probability vector is obtained through a second linear transformation layer or a dimension reduction layer using a dimension reduction tensor, where the formula for the second linear transformation layer is shown in formula (1.6) above.
[0196] Finally, comparative learning was performed on the emotion recognition model.
[0197] In the training process of emotion recognition models, the contrastive learning module also plays a crucial role. The purpose of the contrastive learning module is to optimize the feature space so that the model can effectively identify and distinguish different emotional states. By using the contrastive learning strategy, the model is trained to recognize and amplify the subtle differences between emotional features.
[0198] In the training step, this embodiment of the application selects an anchor sample from the samples in the training set using a random strategy, or uses the center point of the existing feature space of each label sample as the anchor sample. Based on the anchor sample, corresponding positive and negative samples are constructed. A triplet loss function is constructed based on the anchor sample, positive sample, and negative sample. Then, the triplet loss is calculated based on the triplet loss function formula. The triplet loss function not only compares the similarity between the positive and negative samples and the anchor point, but also enforces a boundary value to ensure that the similarity between the positive sample and the anchor point is significantly higher than that between the negative sample and the anchor point. The goal of the triplet loss function is to minimize the distance (similarity) between the positive sample and the anchor sample, while maximizing the distance between the negative sample and the anchor sample. The triplet loss function formula is shown in formula (1.7) above. A higher similarity value between two samples indicates greater similarity, while a lower value indicates less similarity. This can be represented by a distance or similarity metric, such as Euclidean distance, cosine similarity, or Manhattan distance.
[0199] The embodiments of this application can calculate the cross-entropy loss function based on the fusion features of the anchor samples, wherein the calculation formula of the cross-entropy loss function is referred to formula (1.8) above.
[0200] This application embodiment constructs a total loss function based on the cross-entropy loss function and the triplet loss function. The total loss function can be expressed as the sum of the cross-entropy loss function and the triplet loss function according to a certain proportional coefficient, wherein the formula for the total loss function is given in formula (1.9) above.
[0201] This application embodiment calculates the total loss function, minimizes the cross-entropy loss and triplet loss for backpropagation, and adjusts the model parameters so that the model gradually learns to distinguish the features of different emotional states during training, thereby achieving higher accuracy in multimodal emotion recognition tasks. Specifically, by minimizing the triplet loss, the model learns to group the vector representations of input information samples with the same emotion together and separate the vector representations of input information samples with different emotions in the feature space, enabling it to more accurately identify subtle emotional changes and thus improve the model's performance in complex emotion recognition tasks. In addition, based on the triplet loss function, the model can robustly handle intra-class variation and inter-class similarity issues that may exist in the high-dimensional feature space.
[0202] This application utilizes an advanced multimodal emotion recognition model and a virtual object emotion-driven system to predict the emotional state of a virtual object based on pre-set input responses and voice information. It then adjusts the virtual object's behavior, such as facial expressions, body language, and voice tone, according to the prediction results to enhance the realism and appeal of its performance.
[0203] In summary, this application provides a new dynamic driving method for the emotional expression of virtual objects in industries such as entertainment interaction and games. It enriches the interactivity and realism of virtual objects, bringing users a higher level of entertainment experience. Specifically, in interactive entertainment application scenarios, it enables virtual objects to provide more accurate and vivid emotional expressions in various situations, not only enhancing user immersion but also increasing the quality of interaction between virtual objects and users. In game application scenarios, it endows virtual characters with more complex emotional dimensions, enabling characters to make more context-appropriate emotional responses based on changes in game plot and player interaction. This not only improves the playability of the game but also adds emotional layers to e-sports or role-playing games, further enhancing the gaming experience.
[0204] The following description continues to illustrate the exemplary structure of the animation generation device 555 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, a software module stored in an animation generation device 555 in a memory 550 may include:
[0205] The data management module 5551 is used to acquire input information of multiple modalities of the virtual object; the encoding module 5552 is used to encode the input information of each modality respectively to obtain the encoding features of the input information of each modality; the fusion module 5553 is used to fuse the encoding features corresponding to the input information of the multiple modalities respectively to obtain the fusion features of the input information of the multiple modalities; the emotion generation module 5554 is used to perform emotion prediction on the input information of the multiple modalities based on the fusion features to obtain the target emotion expressed by the input information of the multiple modalities; and the animation generation module 5555 is used to generate an animation of the virtual object based on the target emotion and at least some of the input information of the multiple modalities to obtain the virtual object animation.
[0206] In some embodiments, the encoding features corresponding to the input information of the multiple modalities are fused through N cascaded cross-attention layers, where N is a positive integer greater than 1. The fusion module 5553 is further configured to perform attention processing on the encoding features corresponding to the input information of the multiple modalities through N cascaded cross-attention layers to obtain the fused features of the input information of the multiple modalities.
[0207] In some embodiments, when the input information of the multiple modalities includes input information of two modalities among text modal input information, audio modal input information, and image modal input information, the fusion module 5553 is further configured to perform attention processing on the encoding features corresponding to the input information of the two modalities respectively through a first cross-attention layer to obtain the attention features of the first cross-attention layer; and to perform attention processing on the attention features of the (n-1)th cross-attention layer and the encoding features of the input information of one modality through a nth cross-attention layer to obtain the attention features of the nth cross-attention layer, and use the attention features of the Nth cross-attention layer as the fusion features of the input information of the multiple modalities, where n is a positive integer that increases sequentially, 1 < n ≤ N, and the one modality is any one of the two modalities.
[0208] In some embodiments, when the input information of the multiple modalities includes input information of text modality, input information of audio modality, and input information of image modality, the fusion module 5553 is further configured to perform attention processing on the encoding features corresponding to the input information of the two modalities respectively through the first M cross-attention layers of the N cascaded cross-attention layers to obtain the fusion features of the input information of the two modalities; and to perform attention processing on the fusion features of the input information of the two modalities and the encoding features of the input information of the third modality through the last NM cross-attention layers of the N cascaded cross-attention layers to obtain the fusion features of the input information of the multiple modalities, wherein M is a positive integer less than N, the two modalities are any two modalities among the text modality, the audio modality, and the image modality, and the third modality is a modality other than the two modalities among the text modality, the audio modality, and the image modality.
[0209] In some embodiments, the fusion module 5553 is further configured to perform attention processing on the encoding features corresponding to the input information of the two modalities respectively through the first cross-attention layer to obtain the attention features of the first cross-attention layer; perform attention processing on the attention features of the (m-1)th cross-attention layer and the encoding features of the input information of one modality through the mth cross-attention layer to obtain the attention features of the mth cross-attention layer; and use the attention features of the Mth cross-attention layer as the fusion features of the input information of the two modalities, where m is a positive integer that increases sequentially, 1 < m ≤ M, and the one modality is any one of the two modalities.
[0210] In some embodiments, the fusion module 5553 is further configured to: perform attention processing on the fusion features of the input information of the two modalities and the encoding features of the input information of the third modalities through the (M+1)th cross-attention layer to obtain the attention features of the (M+1)th cross-attention layer; perform attention processing on the attention features of the (n-1)th cross-attention layer and the encoding features of the input information of the third modalities through the nth cross-attention layer to obtain the attention features of the nth cross-attention layer; and use the attention features of the Nth cross-attention layer as the fusion features of the input information of the multiple modalities, where n is a positive integer that increases sequentially, and M < n ≤ N.
[0211] In some embodiments, when the input information of the multiple modalities includes input information of text modality, input information of audio modality, and input information of image modality, the fusion module 5553 is further configured to perform attention processing on the encoding features of the input information of the first modality and the encoding features of the input information of the second modality through the first M cross-attention layers of the N cascaded cross-attention layers to obtain a first fusion feature; and to perform attention processing on the encoding features of the input information of the first modality and the encoding features of the input information of the third modality through the last NM cross-attention layers of the N cascaded cross-attention layers to obtain a second fusion feature, wherein M is a positive integer less than N, the first modality and the second modality are any two of the text modality, the audio modality and the image modality, and the third modality is a modality other than any two of the text modality, the audio modality and the image modality; and to fuse the first fusion feature and the second fusion feature to obtain the fusion feature of the input information of the multiple modalities.
[0212] In some embodiments, the emotion generation module 5554 is further configured to perform a linear transformation on the fusion feature to obtain the linearly transformed fusion feature; perform activation processing on the linearly transformed fusion feature to obtain the activated fusion feature; perform dimensionality reduction processing on the activated fusion feature to obtain the emotion probability vector of the input information of the multiple modalities; and determine the target emotion expressed by the input information of the multiple modalities based on the emotion probability vector.
[0213] In some embodiments, the animation generation module 5555 is further configured to determine motion data of the virtual object based on the target emotion and at least some of the input information of the multiple modalities; bind the motion data to the initial object model of the virtual object to obtain a virtual object animation that conforms to the target emotion and matches the input information of the multiple modalities.
[0214] In some embodiments, the animation generation module 5555 is further configured to: determine at least a portion of the input information of the multiple modalities as the input information of the text modal when the input information of the multiple modalities includes the input information of the text modality; generate the input information of the text modality based on the input information of the multiple modalities when the input information of the multiple modalities does not include the input information of the text modality; and generate motion data of the virtual object based on the target emotion and the input information of the text modality.
[0215] In some embodiments, the animation generation module 5555 is further configured to determine, based on the target emotion and the input information of the multiple modalities, the response emotion and response text of the virtual object in response to the input information; and to generate motion data of the virtual object based on the response emotion and the response text.
[0216] In some embodiments, the animation generation method is implemented using a neural network model, and the apparatus further includes:
[0217] Training module 5556 is used to acquire multiple modal input information anchor point samples, multiple modal input information positive samples, and multiple modal input information negative samples of the virtual object; perform sentiment prediction on the multiple modal input information anchor point samples, multiple modal input information positive samples, and multiple modal input information negative samples through the neural network model to obtain a first sentiment probability vector corresponding to the input information anchor sample, a second sentiment probability vector corresponding to the input information positive sample, and a third sentiment probability vector corresponding to the input information negative sample; construct a triplet loss function based on the first sentiment probability vector, the second sentiment probability vector, and the third sentiment probability vector; update the parameters of the neural network model based on the triplet loss function, and use the updated parameters of the neural network model as the parameters of the trained neural network model.
[0218] In some embodiments, the training module 5556 is further configured to perform fusion processing on the encoded features corresponding to the input information anchor samples of the multiple modalities through the neural network model to obtain fused feature samples of the input information anchor samples of the multiple modalities, and construct a cross-entropy loss function based on the fused feature samples; integrate the cross-entropy loss function and the triplet loss function to obtain the total loss function of the neural network model; and update the parameters of the neural network model based on the total loss function.
[0219] In some embodiments, the training module 5556 is further configured to perform similarity calculation on the first emotion probability vector and the second emotion probability vector to obtain a first similarity between the first emotion probability vector and the second emotion probability vector; perform similarity calculation on the first emotion probability vector and the third emotion probability vector to obtain a second similarity between the first emotion probability vector and the third emotion probability vector; and construct the triplet loss function based on the first similarity and the second similarity.
[0220] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the animation generation method described above in this application.
[0221] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the animation generation method provided in this application. For example, ... Figures 3 to 6 The animation generation method is shown.
[0222] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0223] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0224] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).
[0225] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located in one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0226] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. An animation generation method, characterized in that, The method includes: Obtain input information from multiple modalities of virtual objects; The input information of each modality is encoded to obtain the encoding features of the input information of each modality; By fusing the encoding features corresponding to the input information of the multiple modalities, a fused feature of the input information of the multiple modalities is obtained; Based on the fusion features, sentiment prediction is performed on the input information of the multiple modalities to obtain the target sentiment expressed by the input information of the multiple modalities; Based on the target emotion and at least some of the input information from the multiple modalities, the virtual object is animated to obtain a virtual object animation.
2. The method according to claim 1, characterized in that, The encoding features corresponding to the input information of the multiple modalities are integrated through N cascaded cross-attention layers, where N is a positive integer greater than 1; The process of fusing the encoding features corresponding to the input information of the multiple modalities to obtain the fused features of the input information of the multiple modalities includes: By using N cascaded cross-attention layers, attention processing is performed on the encoded features corresponding to the input information of the multiple modalities to obtain the fused features of the input information of the multiple modalities.
3. The method according to claim 2, characterized in that, When the input information from the multiple modalities includes two of the following modalities: text modal input information, audio modal input information, and image modal input information, the process involves performing attention processing on the encoded features corresponding to the input information from the multiple modalities through N cascaded cross-attention layers to obtain the fused features of the input information from the multiple modalities, including: Through the first cross-attention layer, attention processing is performed on the encoded features corresponding to the input information of the two modalities respectively to obtain the attention features of the first cross-attention layer; Through the nth cross-attention layer, attention features of the (n-1)th cross-attention layer and the encoding features of the input information of one modality are processed to obtain the attention features of the nth cross-attention layer. The attention features of the Nth cross-attention layer are then used as the fusion features of the input information of the multiple modalities. Here, n is a positive integer that increases sequentially, 1 < n ≤ N, and the one modality is any one of the two modalities.
4. The method according to claim 2, characterized in that, When the input information from the multiple modalities includes text modal input information, audio modal input information, and image modal input information, the process involves performing attention processing on the encoded features corresponding to the input information from the multiple modalities through N cascaded cross-attention layers to obtain the fused features of the input information from the multiple modalities, including: Attention processing is performed on the encoded features corresponding to the input information of the two modalities through the first M cross-attention layers of the N cascaded cross-attention layers to obtain the fused features of the input information of the two modalities. Through the last NM cross-attention layers of the N cascaded cross-attention layers, attention processing is performed on the fusion features of the input information of the two modalities and the encoding features of the input information of the third modality to obtain the fusion features of the input information of the multiple modalities, where M is a positive integer less than N, the two modalities are any two of the text modality, the audio modality and the image modality, and the third modality is a modality other than the two of the text modality, the audio modality and the image modality.
5. The method according to claim 4, characterized in that, The process involves applying attention to the encoded features corresponding to the input information of the two modalities through the first M cross-attention layers of the N cascaded cross-attention layers to obtain the fused features of the input information of the two modalities, including: Through the first cross-attention layer, attention processing is performed on the encoded features corresponding to the input information of the two modalities respectively to obtain the attention features of the first cross-attention layer; Through the m-th cross-attention layer, attention features of the (m-1)-th cross-attention layer and the encoding features of the input information of one modality are processed to obtain the attention features of the m-th cross-attention layer; The attention features of the Mth cross-attention layer are used as the fusion features of the input information of the two modalities, where m is a positive integer that increases sequentially, 1 < m ≤ M, and the modality is any one of the two modalities.
6. The method according to claim 4, characterized in that, The process involves applying attention to the fusion features of the input information from the two modalities and the encoded features of the input information from the third modality through the last NM cross-attention layers of the N cascaded cross-attention layers, to obtain the fusion features of the input information from the multiple modalities, including: Through the (M+1)th cross-attention layer, attention processing is performed on the fusion features of the input information of the two modalities and the encoding features of the input information of the third modality to obtain the attention features of the (M+1)th cross-attention layer. Through the nth cross-attention layer, attention features of the (n-1)th cross-attention layer and the encoding features of the input information of the third modality are processed to obtain the attention features of the nth cross-attention layer. The attention features of the Nth cross-attention layer are then used as the fusion features of the input information of the multiple modalities, where n is a positive integer that increases sequentially, and M < n ≤ N.
7. The method according to claim 2, characterized in that, When the input information from the multiple modalities includes text modal input information, audio modal input information, and image modal input information, the process involves performing attention processing on the encoded features corresponding to the input information from the multiple modalities through N cascaded cross-attention layers to obtain the fused features of the input information from the multiple modalities, including: The first fusion feature is obtained by performing attention processing on the encoding features of the input information of the first modality and the encoding features of the input information of the second modality through the first M cross-attention layers of the N cascaded cross-attention layers. Through the last NM cross-attention layers of the N cascaded cross-attention layers, attention processing is performed on the encoded features of the input information of the first modality and the encoded features of the input information of the third modality to obtain the second fusion feature, where M is a positive integer less than N, the first modality and the second modality are any two of the text modality, the audio modality and the image modality, and the third modality is a modality other than any two of the text modality, the audio modality and the image modality; By fusing the first fusion feature and the second fusion feature, a fusion feature of the input information of the multiple modalities is obtained.
8. The method according to claim 1, characterized in that, The step of performing sentiment prediction on the input information from the multiple modalities based on the fusion features to obtain the target sentiment expressed by the input information from the multiple modalities includes: The fused features are linearly transformed to obtain the linearly transformed fused features. The fusion feature after linear transformation is activated to obtain the activated fusion feature. The fusion features after activation are subjected to dimensionality reduction to obtain the sentiment probability vector of the input information of the multiple modalities; Based on the emotion probability vector, the target emotion expressed by the input information of the multiple modalities is determined.
9. The method according to claim 1, characterized in that, The step of generating an animation for the virtual object based on the target emotion and at least some of the input information from the multiple modalities, to obtain the virtual object animation, includes: Based on the target emotion and at least some of the input information from the multiple modalities, determine the motion data of the virtual object; The motion data is bound to the initial object model of the virtual object to obtain a virtual object animation that conforms to the target emotion and matches the input information of the multiple modalities.
10. The method according to claim 9, characterized in that, Determining the motion data of the virtual object based on the target emotion and at least some of the input information from the multiple modalities includes: When the input information of the multiple modalities includes the input information of the text modality, at least a portion of the input information of the multiple modalities is determined to be the input information of the text modality; When the input information of the multiple modalities does not include the input information of the text modality, the input information of the text modality is generated based on the input information of the multiple modalities; Based on the target emotion and the input information of the text modality, motion data of the virtual object is generated.
11. The method according to claim 9, characterized in that, Determining the motion data of the virtual object based on the target emotion and at least some of the input information from the multiple modalities includes: Based on the target emotion and the input information of the multiple modalities, the response emotion and response text of the virtual object in response to the input information are determined; Based on the emotion and text of the response, motion data of the virtual object is generated.
12. The method according to any one of claims 1 to 11, characterized in that, The animation generation method is implemented through a neural network model; the training process of the neural network model includes: Obtain anchor point samples, positive samples, and negative samples of input information for multiple modalities of the virtual object; The neural network model is used to perform sentiment prediction on the anchor point samples, positive samples, and negative samples of the input information of the multiple modalities, respectively, to obtain the first sentiment probability vector corresponding to the anchor point sample, the second sentiment probability vector corresponding to the positive sample, and the third sentiment probability vector corresponding to the negative sample. Based on the first sentiment probability vector, the second sentiment probability vector, and the third sentiment probability vector, a triplet loss function is constructed. Based on the triplet loss function, the parameters of the neural network model are updated, and the updated parameters of the neural network model are used as the parameters of the trained neural network model.
13. The method according to claim 12, characterized in that, The step of updating the parameters of the neural network model based on the triplet loss function includes: The neural network model is used to fuse the encoded features corresponding to the input information anchor samples of the multiple modalities to obtain fused feature samples of the input information anchor samples of the multiple modalities, and a cross-entropy loss function is constructed based on the fused feature samples. By integrating the cross-entropy loss function and the triplet loss function, the total loss function of the neural network model is obtained; The parameters of the neural network model are updated based on the total loss function.
14. The method according to claim 12, characterized in that, The construction of a triplet loss function based on the first sentiment probability vector, the second sentiment probability vector, and the third sentiment probability vector samples includes: A similarity calculation is performed on the first emotion probability vector and the second emotion probability vector to obtain a first similarity between the first emotion probability vector and the second emotion probability vector; A similarity calculation is performed on the first emotion probability vector and the third emotion probability vector to obtain a second similarity between the first emotion probability vector and the third emotion probability vector. Based on the first similarity and the second similarity, the triplet loss function is constructed.
15. An animation generation device, characterized in that, The device includes: The data management module is used to acquire input information from various modalities of virtual objects; An encoding module is used to encode the input information of each modality to obtain the encoding features of the input information of each modality. The fusion module is used to fuse the encoding features corresponding to the input information of the multiple modalities to obtain the fusion features of the input information of the multiple modalities. The emotion generation module, based on the fusion features, performs emotion prediction on the input information of the multiple modalities to obtain the target emotion expressed by the input information of the multiple modalities; The animation generation module generates animations for the virtual object based on the target emotion and at least some of the input information of the multiple modalities, thereby obtaining virtual object animations.
16. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions for a computer; A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the animation generation method according to any one of claims 1 to 14.
17. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the animation generation method according to any one of claims 1 to 14 is implemented.
18. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the animation generation method according to any one of claims 1 to 14 is implemented.