Digital human generation method and device based on multi-modal large model
Through the multimodal large model combining audio and visual features, a delicate and natural dynamic digital human image is generated, which solves the problem of insufficient natural and smooth digital human movements in the existing technology, and achieves higher detailed expression and natural expression.
Patent Information
- Application Number
- CN202510569988.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-06
AI Technical Summary
The existing digital life generation methods rely on posture key points control and cannot fully capture the changes in the details of the human body, resulting in the generated digital human movements not being natural and smooth enough.
The digital human generation method based on the multimodal large model is adopted. By obtaining the multimodal data input by the user, intent recognition and emotion analysis are performed, combined with the visual characteristics of the audio sequence data, reference image data and mask characteristics of the target part, the diffusion model is used for denoising to generate a delicate and natural dynamic digital human image.
The dynamic digital human image that expresses delicate and natural in the target part is achieved, and the natural fluency and detailed expression ability of digital human movements are improved.
Smart Images

Figure CN120107427A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of digital human generation, and in particular to a method and device for generating a digital human based on a multimodal large model. Background Art
[0002] With the rapid development of artificial intelligence technology, digital humans have become an important bridge and application carrier connecting virtuality and reality. Digital humans are a highly complex and multidisciplinary technical system. With the help of technologies such as natural language processing, speech recognition, image recognition and sentiment analysis, digital humans can understand the user's deep intentions and emotions, thereby achieving accurate and efficient interaction. Digital human technology covers multiple core technical fields such as computer vision, image generation, and speech generation. Its goal is to generate virtual characters with realistic appearance, natural behavior, rich expressions and clear voices.
[0003] At present, most digital human generation schemes rely on posture key points to control the generated digital human. However, posture key points are usually a discrete representation of human body posture. For example, common posture estimation models may only mark the coordinates of several major joints of the human body, such as the head, shoulders, elbows, hands, hips, knees and feet. This discrete representation cannot fully capture the detailed change characteristics of specific parts of the human body, such as the facial expression change characteristics, the detailed change characteristics of the hands, etc. Therefore, the current key point-based representation can only present a limited, preset posture combination when generating digital humans, resulting in the digital human's movements not being natural and smooth, and ignoring the detailed expression of specific parts of the human body.
[0004] Therefore, how to generate delicate and natural dynamic digital human images has become an urgent problem to be solved. Summary of the invention
[0005] The present application provides a method and device for generating a digital human based on a multimodal large model, which breaks through the limitations of traditional generation based on posture key points or three-dimensional representation, and can generate a dynamic digital human image with delicate and natural performance in the target area.
[0006] In a first aspect, a method for generating a digital human based on a multimodal large model is provided, comprising: Get multimodal data input by the user; Based on the multimodal big model, perform intent recognition and sentiment analysis on multimodal data to determine the response text data; Determine corresponding audio sequence data according to the response text data; Obtaining a vector representation of visual features of reference image data, where the content of the reference image data includes a person and a background; According to the position of the target part of the person in the reference image data, determining the vector representation of the mask feature corresponding to the target part; Based on the diffusion model, at least one noise vector representation is denoised according to the audio sequence data, the vector representation of the visual features of the reference image data, and the vector representation of the mask features corresponding to the target part, to generate at least one frame of digital human image data, wherein the noise vector representation is obtained by adding noise to the vector representation of the visual features of the reference image data; A digital human animation with voice is generated based on at least one frame of digital human image data and audio sequence data.
[0007] In a feasible design, based on the diffusion model, according to the audio sequence data, the vector representation of the visual features of the reference image data, and the vector representation of the mask features corresponding to the target part, at least one noise vector representation is denoised to generate at least one frame of digital human image data, including: Extracting features from the audio sequence data using a first neural network model to obtain a vector representation of initial audio features; Processing the vector representation of the initial audio features based on a second neural network model composed of two layers of multi-layer perceptrons to obtain a vector representation of the audio features aligned with the visual features of the reference image data; Adding a vector representation of the mask feature corresponding to the target part and at least one noise vector representation to obtain a vector representation of at least one initial digital human image data; Based on a third neural network model, denoising the vector representation of at least one initial digital human image data according to the vector representation of the audio features aligned with the visual features of the reference image data and the vector representation of the visual features of the reference image data, to obtain a vector representation of the image data of at least one digital human, wherein the third neural network model is a diffusion model using an attention mechanism; The vector representation of at least one digital human image data is decoded to generate at least one frame of digital human image data.
[0008] In a feasible design, based on the third neural network model, according to the vector representation of the audio features aligned with the visual features of the reference image data and the vector representation of the visual features of the reference image data, a denoising process is performed on the vector representation of at least one initial digital human image data to obtain the vector representation of at least one digital human image data, including: The reference image data is processed by a fourth neural network model to obtain a vector representation of the semantic features of the reference image data, wherein the fourth neural network model is a self-supervised training model using an attention mechanism; Based on the third neural network model, an attention mechanism is used to fuse the vector representation of the audio features aligned with the visual features of the reference image data, the vector representation of the semantic features of the reference image data, and the vector representation of the visual features of the reference image data to obtain fused features; Based on the third neural network model, with the fused features as guidance information, iteratively performing denoising and noise addition processing on the vector representation of at least one initial digital human image data, to obtain the vector representation of at least one digital human image data.
[0009] In a feasible design, according to the position of the target part of the person in the reference image data, the vector representation of the mask feature corresponding to the target part is determined, including: Determine corresponding mask image data according to the position of the target part of the person in the reference image data; The features of the mask region in the mask image data are analyzed based on the region guidance model to obtain a vector representation of the mask features corresponding to the target part.
[0010] In a feasible design, obtaining a vector representation of visual features of reference image data includes: Encoding the reference image data to obtain a vector representation of the reference image data; The vector representation of the reference image data is input into the fifth neural network model using the attention mechanism to obtain the vector representation of the visual features of the reference image data.
[0011] In a feasible design, a vector representation of reference image data is input into the fifth neural network model using the attention mechanism to obtain a vector representation of visual features of the reference image data, including: The vector representation of the reference image data and the vector representation of the semantic features of the reference image data are input into the fifth neural network model using the attention mechanism, so that the fifth neural network model generates a vector representation of the visual features of the reference image data by fusing the semantic features of the reference image data.
[0012] In a feasible design, the multimodal big model includes an audio language model, a big language model, and a visual language model. Based on the multimodal big model, intent recognition and sentiment analysis are performed on the multimodal data to determine the response text data, including: If the multimodal data includes audio data, the audio data is processed by an audio language model to obtain text data converted from the audio data and speech emotion information of the audio data; intention recognition and emotion analysis are performed on data other than audio data in the multimodal data, text data converted from the audio data, and speech emotion information of the audio data to determine response text data; If the multimodal data only includes text data, the text data is subjected to intent recognition and sentiment analysis through a large language model to obtain response text data; If the multimodal data includes image data but does not include audio data, the multimodal data is subjected to intent recognition and sentiment analysis through a visual language model to obtain response text data.
[0013] In a feasible design, intention recognition and sentiment analysis are performed on data other than audio data in the multimodal data, text data converted from audio data, and speech sentiment information of audio data to determine response text data, including: If the data other than the audio data in the multimodal data only includes text data, the text data, the text data converted from the audio data, and the speech emotion information of the audio data are subjected to intention recognition and emotion analysis through the large language model to obtain response text data; If the data other than audio data in the multimodal data includes image data, the visual language model is used to perform intent recognition and sentiment analysis on the data other than audio data in the multimodal data, the text data converted from the audio data, and the speech sentiment information of the audio data to obtain response text data.
[0014] In a feasible design, the response text data is obtained through a large language model, including: After performing intent recognition and sentiment analysis on text data through a large language model, or performing intent recognition and sentiment analysis on text data, text data converted from audio data, and speech sentiment information of audio data through a large language model, if the large language model cannot generate response text data, a corresponding intelligent agent is determined according to the target text data, the intelligent agent is a third-party application platform, and the target text data includes text data, or text data converted from text data and audio data; Query the agent based on the target text data to obtain reference data output by the agent; The target text data and the reference data are processed by a large language model to obtain response text data; Alternatively, the response text data can be obtained through the visual language model, including: After performing intent recognition and sentiment analysis on the multimodal data through a visual language model, or after performing intent recognition and sentiment analysis on data other than audio data in the multimodal data, text data converted from audio data, and speech sentiment information of audio data through a visual language model, if the visual language model cannot generate response text data, determining the intent category or professional field according to the target multimodal data based on the visual language model, the target multimodal data includes the multimodal data, or data other than audio data in the multimodal data and text data converted from audio data; Determine the corresponding intelligent agent based on the intent category or professional field. The intelligent agent is a third-party application platform. Query the agent based on the target multimodal data to obtain reference data output by the agent; The target multimodal data and reference data are processed through a visual language model to obtain response text data.
[0015] In a second aspect, a digital human generation device based on a multimodal large model is provided, comprising: An input data acquisition module is used to acquire multimodal data input by a user; The text response module is used to perform intent recognition and sentiment analysis on multimodal data based on the multimodal large model to determine the response text data; An audio sequence data generating module, used to determine corresponding audio sequence data according to the response text data; A visual feature acquisition module, used to acquire a vector representation of visual features of reference image data, where the content of the reference image data includes a person and a background; A mask feature determination module, used to determine the vector representation of the mask feature corresponding to the target part according to the position of the target part of the person in the reference image data; A digital human image generation module, configured to denoise at least one noise vector representation based on the audio sequence data, the vector representation of the visual features of the reference image data, and the vector representation of the mask features corresponding to the target part based on the diffusion model, so as to generate at least one frame of digital human image data, wherein the noise vector representation is obtained by adding noise to the vector representation of the visual features of the reference image data; The digital human animation generation module is used to generate a digital human animation with voice according to at least one frame of digital human image data and audio sequence data.
[0016] The embodiment of the present application accurately extracts the response text data that conforms to the dialogue scene by performing intention recognition and emotion analysis on multimodal data, and converts it into audio sequence data with emotional characteristics, so as to establish an accurate emotional tone and audio-video synchronization benchmark for the subsequent digital human generation process, so that the subsequent diffusion model can generate a dynamic digital human image that integrates voice emotions based on the audio features and visual features of the audio sequence data. And it can make the posture of the digital human at the target part highly consistent with the audio features. In addition, since the vector representation of the mask feature determined according to the position of the target part of the character in the reference image data can include the features of the area where the target part is located, and the relationship features between the area where the target part is located and other areas. Therefore, the present application scheme can guide the diffusion model through the mask feature in the process of generating the image data of the digital human, and strengthen the ability to capture the natural details of the target part (such as face, hands, etc.), as well as the coordination with the overall posture of the character. So that the digital human is given a more delicate and natural expressiveness at the target part. In addition, if you want to improve the expressiveness of specific parts of the digital human, you can guide the diffusion model through the vector representation of the mask features corresponding to the specific parts. Therefore, compared with the current digital human movements generated by relying on posture key point control, which are not natural and smooth enough, the present application scheme can flexibly generate delicate and natural dynamic digital human images. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solution of the present application, the drawings required for use in the embodiments are briefly introduced below. Obviously, for ordinary technicians in this field, other related drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 is a schematic flow chart of a method for generating a digital human based on a multimodal large model provided by an exemplary embodiment of the present application; Figure 2 is a schematic flow chart of another method for generating a digital human based on a multimodal large model provided by an exemplary embodiment of the present application; Figure 3 is a schematic diagram of a network framework of a digital human generation model provided by an exemplary embodiment of the present application; Figure 4 It is a schematic diagram of a digital human generation device based on a multimodal large model provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0020] In order to flexibly generate delicate and natural dynamic digital human images, this application breaks away from the limitations of traditional reliance on posture key point control to generate digital humans. By introducing mask features of the target parts of the character (such as face, hands, etc.), the neural network model's ability to capture details in the target parts is enhanced, thereby enabling the neural network model to generate delicate and expressive dynamic digital human images.
[0021] Furthermore, the present application also optimizes the generation algorithm of matching the target area with the audio features by injecting audio features into the neural network model, thereby achieving a highly consistent fusion of the target area with the voice emotion, and the overall posture of the character with the voice emotion.
[0022] In addition, although digital humans have shown great potential in multiple scenarios, current digital human generation solutions have shortcomings in multimodal interaction scenarios, especially the limited in-depth understanding of the intentions, motivations and emotions contained in multimodal information. In order to improve the performance of digital humans in multimodal interaction scenarios, this application also uses a multimodal large model to integrate multimodal data such as text, audio, and images input by users, significantly enhancing the perception and understanding of emotional intentions in the digital human generation solution of this application. At the same time, combined with a multi-agent system, the intelligent performance of the digital human service algorithm is further optimized, bringing users a more professional and more demand-oriented service experience. Furthermore, combined with enhanced knowledge base technology and multimodal data retrieval mechanism, the high accuracy and real-time nature of the generated response content are ensured.
[0023] The digital human generation solution provided by this application is described in detail below. Figure 1 As shown, the digital human generation method based on a multimodal large model provided by the present application includes: S110, obtaining multimodal data input by a user.
[0024] The types of data in multimodal data include, but are not limited to, text, audio, pictures, and videos. Pictures and videos can be collectively referred to as images.
[0025] S120, performing intent recognition and sentiment analysis on the multimodal data based on the multimodal large model to determine response text data.
[0026] Exemplarily, the multimodal large model includes an audio language model, a large language model and a visual language model. It should be noted that the audio language model, the large language model and the visual language model can be regarded as various parts of the multimodal large model, and each part can independently process the input data.
[0027] In a feasible design, the following methods are used to implement intent recognition and sentiment analysis on multimodal data based on a multimodal large model to determine the response text data: If the multimodal data includes audio data, the audio data is processed by an audio language model to obtain text data converted from the audio data and speech emotion information of the audio data; intention recognition and emotion analysis are performed on data other than audio data in the multimodal data, text data converted from the audio data, and speech emotion information of the audio data to determine response text data; If the multimodal data only includes text data, the text data is subjected to intent recognition and sentiment analysis through a large language model to obtain response text data. The large language model is used to process the content and text data retrieved from the multimodal knowledge base to generate response text data. If the multimodal data includes image data but does not include audio data, the visual language model is used to perform intent recognition and sentiment analysis on the multimodal data to obtain response text data. The visual language model is used to process the content and multimodal data retrieved from the multimodal knowledge base to generate response text data.
[0028] It should be understood that if the multimodal data only includes audio data, it means that there is no data other than audio data in the multimodal data. The intention recognition and sentiment analysis can be performed directly on the text data converted from the audio data and the speech emotion information of the audio data through a large language model to determine the response text data.
[0029] It should be understood that multimodal data includes image data but does not include audio data, covering two situations: (1) multimodal data only includes image data; (2) multimodal data includes image data and text data.
[0030] Among them, the multimodal knowledge base is a dynamic knowledge base constructed by this application, which supports retrieval enhancement algorithms and dynamic updates. The multimodal knowledge base stores a large amount of text data, image data and audio data, providing a reliable basis for generating response data and ensuring that the generated content is accurate and contextually relevant.
[0031] This application does not limit the network framework of the large language model. It can realize retrieval in a multimodal knowledge base based on a retrieval enhancement algorithm, perform intent recognition and sentiment analysis on the retrieved content and text data, and generate reply text data. For example, the large language model can be the Tongyi Qianwen large model (Qwen).
[0032] This application does not limit the network framework of the visual language model. It can realize retrieval in a multimodal knowledge base based on a retrieval enhancement algorithm, perform intent recognition and sentiment analysis on the retrieved content and multimodal data, and generate reply text data. For example, the visual language model can be the Tongyi Qianwen Visual Language Model (Qwen-VL).
[0033] This application does not limit the network framework of the audio language model. It only needs to be able to convert audio data into text data and analyze the speech emotion information in the audio data. For example, the audio language model can be the Tongyi Qianwen Audio Language Model (Qwen-Audio).
[0034] That is to say, Figure 2 As shown, before processing the multimodal data through a large language model or a visual language model, it is first determined whether the multimodal data includes audio data.
[0035] Since audio data contains language and emotional information, if multimodal data includes audio data, it is necessary to first convert the audio data into text data through an audio language model and extract the speech and emotional information therein. Then, subsequent intent recognition and emotional analysis processing is performed through a large language model or a visual language model.
[0036] If the multimodal data does not include audio data, it can be directly processed through a large language model or a visual language model.
[0037] In the above example, when the multimodal data includes audio data, the audio data is converted into text data through the audio language model, so that the subsequent neural network model can integrate the text information in the audio data and the information of other data in the multimodal data except the audio data, and capture the user's intention more comprehensively and accurately. And by extracting the voice emotion information in the audio data and injecting it into the neural network model, the model's emotional analysis of the input multimodal data is more accurate, which is conducive to generating emotionally rich and contextual response text data, improving the ability to understand the intentions, motivations and emotions contained in the multimodal data, and providing more accurate information for the subsequent digital human animation generation. In the case where audio data is not included, by continuing to determine whether the included data form contains image data, in order to determine the type of model used to process the multimodal data, the adaptability of the digital human generation solution of the present application to multimodal scenarios can be enhanced.
[0038] Next, determine whether the data other than audio data in the multimodal data includes image data. If image data is included, the text data, image data and speech emotion information converted from the audio data are input into the visual language model for corresponding processing to obtain response text data. If image data is not included, the text data converted from the audio data, the data other than audio data in the multimodal data and speech emotion information are input into the large language model for corresponding processing to obtain response text data.
[0039] In actual application, if the multimodal data only includes text data, the text data is directly input into the large language model. The large language model searches in the multimodal knowledge base based on the retrieval enhancement algorithm, and then quickly generates response text data by performing intent recognition and sentiment analysis on the retrieved content and text data. If the multimodal data includes image data but not audio data (for example, only image data, or image data and text data), the multimodal data is input into the visual language model. The visual language model searches in the multimodal knowledge base based on the retrieval enhancement algorithm, and then generates response text data by performing intent recognition and sentiment analysis on the retrieved content and multimodal data.
[0040] In the above example, when the multimodal data includes image data, the visual language model is used to perform joint intent recognition and sentiment analysis on the input text data, image data, etc., to generate response text data, thereby providing accurate text information for replying to users for subsequent digital human animation generation, solving the problem that the current digital human generation solution cannot adapt to scenarios with multimodal input data.
[0041] Furthermore, since the large language model responds faster than the visual language model, compared with performing intent recognition and sentiment analysis on the input multimodal data based solely on the visual language model, the above example can improve the processing efficiency of input data by processing different types of input data based on different neural network models.
[0042] In a feasible design, intention recognition and sentiment analysis are performed on data other than audio data in the multimodal data, text data converted from audio data, and speech sentiment information of audio data to determine response text data, including: If the data other than the audio data in the multimodal data only includes text data, the text data, the text data converted from the audio data, and the speech emotion information of the audio data are subjected to intention recognition and emotion analysis through the large language model to obtain response text data; If the data other than audio data in the multimodal data includes image data, the visual language model is used to perform intent recognition and sentiment analysis on the data other than audio data in the multimodal data, the text data converted from the audio data, and the speech sentiment information of the audio data to obtain response text data.
[0043] If the multimodal data includes other types of data in addition to audio data, the above example further determines the types of other data to effectively select the corresponding model to process the multimodal data, thereby further improving the adaptability of the digital human generation solution to multimodal input data scenarios.
[0044] In one possible design, Figure 2 As shown, the response text data is obtained through the large language model in the following way: After performing intent recognition and sentiment analysis on text data through a large language model, or performing intent recognition and sentiment analysis on text data, text data converted from audio data, and speech sentiment information of audio data through a large language model, if the large language model cannot generate response text data, then determining a corresponding intelligent agent based on target text data, the intelligent agent may be a third-party application platform, and the target text data includes text data, or text data converted from text data and audio data; Query the agent based on the target text data to obtain reference data output by the agent; The target text data and reference data are processed through a large language model to obtain response text data.
[0045] It should be understood that, in this example, obtaining response text data through the large language model includes two situations: (1) performing intent recognition and sentiment analysis on the text data input by the user through the large language model to obtain response text data; (2) performing intent recognition and sentiment analysis on the text data input by the user, the text data converted from the audio data, and the speech sentiment information through the large language model to obtain response text data.
[0046] Exemplarily, querying the agent according to the target text data to obtain the reference data output by the agent is achieved in the following manner: Generate request information based on target text data through a large language model; Send request information to the agent; Receives reference data from the agent.
[0047] It should be noted that this application does not limit the interface or protocol used by the large language model to call the intelligent agent.
[0048] For example, the target text data is "How is the weather today?" Since the large language model lacks this real-time information, it is necessary to call the agent to query the weather data. The large language model generates the request information as "Get the weather conditions on XX / XX / XX." After the large language model obtains the weather data "Sunny, 22℃" (i.e., reference data) from the agent, it processes the target text data and the reference data to obtain the response text data as "Today's weather is sunny, suitable for going out for fun."
[0049] Exemplarily, the corresponding agent is determined according to the target text data by the following method: Establishing a first association relationship, where the first association relationship is used to associate the keyword with agent information, where the agent information includes the name of the agent, a functional description, an area of expertise, a task type that the agent is expertise in, or a type of data provided; Perform keyword recognition on target text data; The keywords in the target text data are matched with the keywords in the first association relationship to determine the corresponding intelligent agent.
[0050] For example, the target text data contains "weather", and after matching with the first association relationship, the agent is determined to be a "weather query program".
[0051] Exemplarily, the corresponding agent is determined according to the target text data by the following method: Establishing a second association relationship between the intent category and the agent information, the second association relationship is used to associate the intent category with the agent information, the agent information includes the name of the agent, function description, areas of expertise, task types that expertise, or data types provided; Obtain the intent category obtained after the large language model performs intent recognition on the target text data; The intent category obtained by the large language model is matched with the intent category in the second association relationship to determine the corresponding agent.
[0052] For example, the large language model performs intent recognition on the target text data and obtains the intent category of weather query. After matching with the second association relationship, the intelligent agent is determined to be a "weather query program."
[0053] Exemplarily, the corresponding agent is determined according to the target text data in the following manner: Establishing a third association relationship between the professional field and the agent information, the third association relationship is used to associate the name of the professional field with the agent information, the agent information includes the name of the agent, function description, areas of expertise, task types of expertise, or data types provided; Obtain the professional field name obtained after the large language model performs professional field recognition on the target text data; The professional field name obtained by the large language model is matched with the professional field name in the third association relationship to determine the corresponding intelligent agent.
[0054] For example, the large language model identifies the professional field of the target text data and obtains the name of the professional field as the weather field. After matching with the third association relationship, the intelligent agent is determined to be the "weather query program."
[0055] The above examples provide multiple ways to determine the corresponding intelligent agent based on text data or multimodal data, optimize the capabilities of multiple agents in task distribution and collaboration, and effectively improve the reliability of user services.
[0056] Exemplarily, after determining that an abnormality has occurred, the intelligent agent generates an alarm response and then sends the alarm response to business personnel, thereby improving the reliability and security of the system.
[0057] For example, a user asks the big language model why his mobile phone number cannot receive calls. The big language model then asks the intelligent agent provided by the operator to check the information and reasons of the mobile phone number. As a result, the intelligent agent finds that the problem is not caused by the mobile phone signal but by the problem of the local base station signal tower. The intelligent agent can then generate an alarm for business personnel to handle.
[0058] In one possible design, Figure 2 As shown, the response text data is obtained through the visual language model in the following way: After performing intent recognition and sentiment analysis on the multimodal data through a visual language model, or after performing intent recognition and sentiment analysis on data other than audio data in the multimodal data, text data converted from audio data, and speech sentiment information of audio data through a visual language model, if the visual language model cannot generate response text data, determining the intent category or professional field according to the target multimodal data based on the visual language model, the target multimodal data includes the multimodal data, or data other than audio data in the multimodal data and text data converted from audio data; Determine the corresponding intelligent agent based on the intent category or professional field. The intelligent agent can be a third-party application platform. Query the agent based on the target multimodal data to obtain reference data output by the agent; The target multimodal data and reference data are processed through a visual language model to obtain response text data.
[0059] It should be understood that, in this example, obtaining response text data through the visual language model includes two situations: (1) If the multimodal data does not include audio data, the visual language model is used to perform intent recognition and sentiment analysis on the multimodal data to obtain response text data; (2) If the multimodal data includes audio data, the visual language model is used to perform intent recognition and sentiment analysis on the data other than audio data in the multimodal data, the text data converted from the audio data, and the speech emotion information to obtain response text data.
[0060] Among them, the method of determining the corresponding intelligent agent according to the intention category or professional field can be found in the description of the example of determining the corresponding intelligent agent according to the target text data, which will not be repeated here.
[0061] Exemplarily, querying the agent according to the target multimodal data to obtain the reference data output by the agent is achieved in the following manner: Generate request information based on target multimodal data through visual language model; Send request information to the agent; Receives reference data from the agent.
[0062] It should be noted that this application does not limit the interface or protocol used by the visual language model to call the intelligent agent.
[0063] For example, after the visual language model recognizes the intent, it determines that the intent category is "query the weather". Since the visual language model lacks such real-time information, it is necessary to call the agent to query the weather data. The request information generated by the visual language model is "get the weather conditions on XX / XX / XX". After obtaining the weather data from the agent as "sunny, 22℃" (i.e. reference data), the response text data is obtained by processing the intent category and reference data as "today's weather is sunny, suitable for going out for fun".
[0064] In practical applications, large language models or visual language models may be subject to various limitations when processing certain specific multimodal data and may not be able to directly generate appropriate response text. These limitations may be due to the requirements of relevant policies and regulations, or due to the lack of certain necessary information in the model itself. For example, for user privacy data (such as information related to mobile phone numbers), real-time data in specific industries (such as real-time weather queries require professional meteorological data interfaces), etc., the model may not be able to directly obtain and process it. At this time, introducing an intelligent agent to provide reference data can help the model overcome these limitations, generate response text more effectively, and improve the ability to handle problems in multimodal scenarios.
[0065] In addition, since there can be multiple agents, the above example also supports multiple agents to process tasks at the same time, realizing multi-task parallel processing and improving the response efficiency of the model.
[0066] S130: Determine corresponding audio sequence data according to the response text data.
[0067] In practice, the response text data can be converted into audio sequence data with emotional characteristics through Text To Speech (TTS) technology.
[0068] S140, obtaining a vector representation of visual features of the reference image data.
[0069] The content of the reference image data includes characters and background.
[0070] In a feasible design, obtaining a vector representation of the visual features of the reference image data is achieved in the following manner: Encoding the reference image data to obtain a vector representation of the reference image data; The vector representation of the reference image data is input into the fifth neural network model using the attention mechanism to obtain the vector representation of the visual features of the reference image data.
[0071] Exemplarily, the reference image data is encoded based on an encoder, and the encoder is an encoder part of a Variational Autoencoder Encoder (VAE) model.
[0072] Exemplarily, the fifth neural network model is a pre-trained two-dimensional U-shaped Convolutional Network (U-Net) model.
[0073] The above example can convert the reference image data into a vector representation that is easier for the fifth neural network model to process and understand by encoding the reference image data. The vector representation contains the key features of the image, such as color, texture, shape, etc. Then, the fifth neural network model further processes the encoded vector representation through the attention mechanism, and can perform a deeper analysis and conversion of the key features therein, thereby extracting higher-level visual features, so that the subsequent neural network model can better capture the background details and identity information in the reference image based on these higher-level visual features, thereby ensuring that the generated digital human maintains consistency with the reference image in terms of character identity and background.
[0074] S150, determining a vector representation of a mask feature corresponding to the target part according to the position of the target part of the person in the reference image data.
[0075] The target part may include multiple parts, such as face and hand, etc. The vector representation of the mask feature is a feature vector generated by processing the mask image data through image segmentation technology.
[0076] In a feasible design, the vector representation of the mask feature corresponding to the target part is determined according to the position of the target part of the person in the reference image data in the following manner: Determine corresponding mask image data according to the position of the target part of the person in the reference image data; The features of the mask region in the mask image data are analyzed based on the regional guidance model to obtain a vector representation of the mask features corresponding to the target part. Specifically, the mask image data is input into the regional guidance model to obtain a vector representation of the mask image data; the regional guidance model can analyze the features of the mask region in the mask image data, and the vector representation of the output mask image data is the vector representation of the mask features corresponding to the target part.
[0077] The vector representation of the mask image data and each noise vector representation have the same dimension, so that the vector representation of the mask image data and each noise vector representation can be added subsequently.
[0078] The noise vector representation is obtained by adding noise to the vector representation of the visual features of the reference image data. Specifically, after the reference image data is encoded and converted into a latent representation through the VAE model, Gaussian noise is added to the latent representation to generate a noise vector representation.
[0079] It should be understood that, in the present application, the mask area refers to the area corresponding to the target part in the mask image data.
[0080] It should be understood that the vector representation of the mask image data includes not only the mask features of the area where the target part is located, but also the relationship features between the area where the target part is located and other areas, the mask features of other areas, etc. Therefore, the vector representation of the mask features corresponding to the target part does not only refer to the vector representation of the mask features of the area where the target part is located in the mask image data.
[0081] The mask image data can adopt a binary mask structure (the target area is 1 and the background is 0), which can accurately represent the contour information of the target part. The above example uses the mask image data to highlight the position of the target part in the reference image data, so that the vector representation of the mask image data output by the regional guidance model can include the mask features of the area where the target part is located, the relationship features between the area where the target part is located and other areas, etc., so that the neural network model used to generate the digital human can accurately locate the target part and pay more attention to the generation of details of the key area where the target part of the digital human is located, thereby improving the restoration of details of the target part of the digital human. In addition, the relationship features between the area where the target part is located and other areas in the vector representation of the mask image data enable the model to consider the coordination between the target part and other parts of the character when generating the digital human, and generate a digital human that is more in line with the proportions and natural postures of the human body, thereby improving the coordination of the overall expressiveness of the digital human.
[0082] Exemplarily, the region guidance model consists of four convolutional network layers, the convolution kernel size of each convolutional network layer is 4×4 (in pixels), the stride is set to 2×2 (in pixels), and the number of channels of the four convolutional network layers are 16, 32, 64, and 128 (in pixels), respectively.
[0083] The regional guidance model of the structure shown in the above example can gradually abstract and extract higher-level features through the gradually increasing number of channels, mine deep information, and improve the ability of the model to capture details in the target area of the generated digital human. And through the moderate convolution kernel size and stride setting, it helps to control the amount of data and reduce the computational complexity.
[0084] S160, based on the diffusion model, denoising is performed on at least one noise vector representation according to the audio sequence data, the vector representation of the visual features of the reference image data, and the vector representation of the mask features corresponding to the target part to generate at least one frame of digital human image data.
[0085] In a feasible design, the following method is used to generate at least one frame of digital human image data by denoising at least one noise vector representation based on the diffusion model according to the audio sequence data, the vector representation of the visual features of the reference image data, and the vector representation of the mask features corresponding to the target part: Extracting features from the audio sequence data using a first neural network model to obtain a vector representation of initial audio features; Processing the vector representation of the initial audio features based on a second neural network model composed of two layers of multilayer perceptrons (MLP) to obtain a vector representation of the audio features aligned with the visual features of the reference image data; Adding a vector representation of the mask feature corresponding to the target part and at least one noise vector representation to obtain a vector representation of at least one initial digital human image data; Based on a third neural network model, denoising the vector representation of at least one initial digital human image data according to the vector representation of the audio features aligned with the visual features of the reference image data and the vector representation of the visual features of the reference image data, to obtain a vector representation of the image data of at least one digital human, wherein the third neural network model is a diffusion model using an attention mechanism; The vector representation of at least one digital human image data is decoded to generate at least one frame of digital human image data.
[0086] Exemplarily, the first neural network model is a pre-trained Waveform-to-Vector (WAV2VEC) model.
[0087] Exemplarily, the third neural network model is a denoising Unet model.
[0088] Exemplarily, the second neural network model (also referred to as the Audio Project model in this application) also adopts a cross-attention mechanism, and uses the cross-attention mechanism to input the vector representation of the audio features aligned with the visual features of the reference image data into the third neural network model. Correspondingly, the third neural network model can use the audio attention mechanism to process the vector representation of the audio features aligned with the visual features of the reference image data, and realize the fusion of the audio features with other features. Since the denoising U-type network model cannot understand the vector representation of the initial audio features, the second neural network model is required to convert the vector representation of the initial audio features into the vector representation of the audio features that the denoising U-type network model can understand, that is, to achieve alignment with the visual features of the reference image data. Specifically, the "alignment" refers to the nonlinear transformation of the initial audio features by the second neural network model, so that its dimension and distribution match the input requirements of the denoising U-type network model, and at the same time, it is semantically associated with the visual features of the reference image (such as facial expressions, hand gestures). For example, the emotional intensity in the audio is mapped to a numerical range corresponding to the curvature of the mouth corners and the position of the eyebrows in the visual features through the second neural network model, thereby realizing the coordinated control of audio and video features.
[0089] Exemplarily, the fifth neural network model also adopts a spatial attention mechanism, and uses the spatial attention mechanism to input the vector representation of the visual features of the reference image data into the third neural network model. Accordingly, the third neural network model can use the spatial attention mechanism to process the vector representation of the visual features of the reference image data, and realize the fusion of the visual features with other features.
[0090] The above example uses the first neural network model to extract features from the rich semantic and acoustic information in the audio sequence data and converts it into a vector representation. Then, the powerful nonlinear mapping ability of the second neural network model composed of two layers of multi-layer perceptrons is used to learn the complex relationship between audio features and image features, so as to accurately align audio features and image features. In this way, the audio features and visual features jointly guide the digital human generation process of the third neural network model to ensure that the generated digital human's voice and dynamic expression movements are highly consistent. The present application scheme abandons the traditional multi-condition injection method, realizes a single-stage digital human generation technology based on audio and reference images, and can directly synthesize a stable digital human video based on audio features and visual features of reference images, thereby improving generation efficiency and flexibility. In addition, the above example also obtains a vector representation of at least one initial digital human image data by adding the vector representation of the mask feature corresponding to the target part to at least one noise vector representation. The target part is made more prominent and clear in the vector representation of each initial digital human image data, so as to help the third neural network model better understand the structure and details of the target part, thereby generating a more delicate and realistic digital human image.
[0091] In a feasible design, the method is implemented in the following manner: based on the third neural network model, according to the vector representation of the audio features aligned with the visual features of the reference image data and the vector representation of the visual features of the reference image data, the vector representation of at least one initial digital human image data is subjected to denoising processing to obtain the vector representation of the image data of at least one digital human: The reference image data is processed by a fourth neural network model to obtain a vector representation of the semantic features of the reference image data, wherein the fourth neural network model is a self-supervised training model using an attention mechanism; Based on the third neural network model, an attention mechanism is used to fuse the vector representation of the audio features aligned with the visual features of the reference image data, the vector representation of the semantic features of the reference image data, and the vector representation of the visual features of the reference image data to obtain fused features; Based on the third neural network model, with the fused features as guidance information, iteratively performing denoising and noise addition processing on the vector representation of at least one initial digital human image data, to obtain the vector representation of at least one digital human image data.
[0092] It should be noted that an iterative process includes first performing a noise addition process and then performing a noise removal process.
[0093] Exemplarily, the fourth neural network model is a Deformable Detection Transformer (DINO) model. The DINO model has the ability to capture rich detail features and can effectively separate and identify tiny detail differences in images. The DINO model can accurately extract the detailed postures of the target parts of the reference image (such as facial expressions and hand postures), thereby providing key guidance for the third neural network model to generate detailed and realistic digital humans in different scenarios.
[0094] Exemplarily, the fourth neural network model uses a cross-attention mechanism to input the vector representation of the semantic feature into the third neural network model. Accordingly, the third neural network model uses a cross-attention mechanism to process the vector representation of the semantic feature to achieve the fusion of the semantic feature and other features.
[0095] The semantic features of image data contain high-level, abstract information. The above example uses the fourth neural network model to extract semantic features of the reference image data, which can deeply understand the detailed information of the person in the reference image data and separate the feature differences. Inputting the vector representation of the semantic features of the reference image data into the third neural network model can guide the third neural network model to accurately generate the detailed posture of the reference image at the target part, such as detailed facial expressions, hand movements, etc.
[0096] Exemplarily, since the vector representation of an initial digital human image data corresponds to a frame of digital human image data, in order to ensure the temporal coherence of the frame sequence of the generated digital human animation, the third neural network model captures the complex temporal dependencies between the vector representations of consecutive initial digital human image data through the temporal attention mechanism, and fuses the temporal features of the vector representation of at least one initial digital human image data on the frame sequence timeline with other features, thereby facilitating the generation of smooth dynamic digital human videos.
[0097] In a feasible design, the vector representation of the reference image data is input into the fifth neural network model using the attention mechanism to obtain the vector representation of the visual features of the reference image data in the following manner: The vector representation of the reference image data and the vector representation of the semantic features of the reference image data are input into the fifth neural network model using the attention mechanism, so that the fifth neural network model generates a vector representation of the visual features of the reference image data by fusing the semantic features of the reference image data.
[0098] Exemplarily, the fourth neural network model uses a cross-attention mechanism to input the vector representation of the semantic features of the reference image data into the fifth neural network model. Accordingly, the fifth neural network model uses a cross-attention mechanism to process the vector representation of the semantic features to achieve the fusion of semantic features and visual features.
[0099] The above example inputs the vector representation of the semantic features of the reference image data into the fifth neural network model, which can help the fifth neural network model deeply understand the detailed information of the person in the reference image data and separate the feature differences, so as to capture the detailed visual features in the reference image data and obtain deeper visual features. In turn, the subsequent third neural network model can further accurately generate the detailed posture of the reference image at the target part based on the deeper visual features.
[0100] S170, generating a digital human animation with voice according to at least one frame of digital human image data and audio sequence data.
[0101] Exemplarily, by calling video editing software, at least one frame of digital human image data and audio sequence data are synthesized to obtain a digital human animation with voice.
[0102] Exemplarily, the input audio sequence data may be the response text data determined after performing intent recognition and sentiment analysis on the multimodal data based on the aforementioned multimodal large model, and obtained via TTS (Text To Speech) technology.
[0103] It should be noted that the input and output dimensions of each neural network model can be set according to actual needs, and this application does not limit this.
[0104] Based on the above embodiments, Figure 3 As shown, the present application provides a digital human generation model, which can generate a half-body digital human video by combining audio sequence data and reference image data. In many scenes, such as customer service scenes, half-body digital human animation has a higher affinity than full-body digital human or facial animation.
[0105] The overall framework of the digital human generation model is based on the VAE model, DINO model, two-dimensional U-Net model, denoising U-type network model, regional guidance model, WAV2VEC model and a second neural network model composed of two-layer multi-layer perceptron.
[0106] The following is an example of a case where the reference image data is a half-body image of a person and the target parts are the face and hands of the person to illustrate the process of data processing for the digital human generation model.
[0107] After receiving audio sequence data, reference image data, mask image data and at least one noise vector representation, the digital human generation model extracts features from the audio sequence data through the WAV2VEC model, and aligns the audio features with the visual features of the reference image data through the second neural network model. Then, the vector representation of the audio features aligned with the visual features of the reference image data is input into the denoising U-type network model using the cross attention mechanism in the second neural network model.
[0108] The digital human generation model is based on the encoder part of the VAE model (i.e. Figure 3 The VAE enc shown in FIG. 1 encodes the reference image data to obtain a vector representation of the reference image data. The vector representation of the reference image data is input into a two-dimensional U-Net model. The vector representation of the semantic features of the reference image data is extracted by the self-supervised DINO model. The DINO model uses a cross-attention mechanism to input the vector representation of the semantic features of the reference image data into a two-dimensional U-Net model and a denoising U-type network model respectively. The vector representation of the reference image data and the vector representation of the semantic features of the reference image data are processed by the two-dimensional U-Net model to obtain a vector representation of the visual features of the reference image data. The two-dimensional U-Net model uses a spatial attention mechanism to input the vector representation of the visual features of the reference image data into the denoising U-type network model. The mask image data is processed by a regional guidance model to obtain a vector representation of the mask features corresponding to the face and the hand. The vector representation of the mask feature is added to at least one noise vector representation to obtain a vector representation of at least one initial digital human image data. The vector representation of at least one initial digital human image data is input into the denoising U-type network model. The denoising U-type network model uses a spatial attention mechanism to extract features from the vector representation of the visual features of the reference image data and fuses them with other features; uses a cross-attention mechanism to extract features from the vector representation of the semantic features of the reference image data and fuses them with other features; uses an audio attention mechanism to extract features from the vector representation of the audio features aligned with the visual features of the reference image data and fuses them with other features; uses a temporal attention mechanism to generate temporal features of the vector representation of at least one initial digital human image data on the time axis of the frame sequence and fuses them with other features. The denoising U-type network model obtains fused features by fusing visual features, audio features, semantic features and temporal features. Using the fused features as guidance information, it iteratively performs denoising and denoising on the vector representation of at least one initial digital human image data fused with mask features to obtain a vector representation of at least one digital human image data. Using the decoder part of the VAE model (i.e. Figure 3The VAE dec shown in the figure decodes the vector representation of the image data of at least one digital human, and finally generates at least one frame of image data of the digital human with smooth voice and highly consistent dynamic expression and movement.
[0109] The embodiment of the present application accurately extracts the response text data that conforms to the dialogue scene by performing intention recognition and emotion analysis on multimodal data, and converts it into audio sequence data with emotional characteristics, so as to establish an accurate emotional tone and audio-video synchronization benchmark for the subsequent digital human generation process, so that the subsequent diffusion model can generate a dynamic digital human image that integrates voice emotions based on the audio features and visual features of the audio sequence data. And it can make the posture of the digital human at the target part highly consistent with the audio features. In addition, since the vector representation of the mask feature determined according to the position of the target part of the character in the reference image data can include the features of the area where the target part is located, and the relationship features between the area where the target part is located and other areas. Therefore, the present application scheme can guide the diffusion model through the mask feature in the process of generating the image data of the digital human, and strengthen the ability to capture the natural details of the target part (such as face, hands, etc.), as well as the coordination with the overall posture of the character. So that the digital human is given a more delicate and natural expressiveness at the target part. In addition, if you want to improve the expressiveness of specific parts of the digital human, you can guide the diffusion model through the vector representation of the mask features corresponding to the specific parts. Therefore, compared with the current digital human movements generated by relying on posture key point control, which are not natural and smooth enough, the present application scheme can flexibly generate delicate and natural dynamic digital human images.
[0110] like Figure 4 As shown, the present application also provides a digital human generation device, including: An input data acquisition module is used to acquire multimodal data input by a user; The text response module is used to perform intent recognition and sentiment analysis on multimodal data based on the multimodal large model to determine the response text data; An audio sequence data generating module, used to determine corresponding audio sequence data according to the response text data; A visual feature acquisition module, used to acquire a vector representation of visual features of reference image data, where the content of the reference image data includes a person and a background; A mask feature determination module, used to determine the vector representation of the mask feature corresponding to the target part according to the position of the target part of the person in the reference image data; A digital human image generation module, configured to denoise at least one noise vector representation based on the audio sequence data, the vector representation of the visual features of the reference image data, and the vector representation of the mask features corresponding to the target part based on the diffusion model, so as to generate at least one frame of digital human image data, wherein the noise vector representation is obtained by adding noise to the vector representation of the visual features of the reference image data; The digital human animation generation module is used to generate a digital human animation with voice according to at least one frame of digital human image data and audio sequence data.
[0111] In a feasible design, the digital human image generation module is implemented in the following manner: based on the diffusion model, according to the audio sequence data, the vector representation of the visual features of the reference image data, and the vector representation of the mask features corresponding to the target part, at least one noise vector representation is denoised to generate at least one frame of digital human image data, including: Extracting features from the audio sequence data using a first neural network model to obtain a vector representation of initial audio features; Processing the vector representation of the initial audio features based on a second neural network model composed of two layers of multi-layer perceptrons to obtain a vector representation of the audio features aligned with the visual features of the reference image data; Adding a vector representation of the mask feature corresponding to the target part and at least one noise vector representation to obtain a vector representation of at least one initial digital human image data; Based on a third neural network model, denoising the vector representation of at least one initial digital human image data according to the vector representation of the audio features aligned with the visual features of the reference image data and the vector representation of the visual features of the reference image data, to obtain a vector representation of the image data of at least one digital human, wherein the third neural network model is a diffusion model using an attention mechanism; The vector representation of at least one digital human image data is decoded to generate at least one frame of digital human image data.
[0112] In a feasible design, the digital human image generation module is implemented in the following manner: based on the third neural network model, according to the vector representation of the audio features aligned with the visual features of the reference image data and the vector representation of the visual features of the reference image data, denoising is performed on the vector representation of at least one initial digital human image data to obtain the vector representation of at least one digital human image data, including: The reference image data is processed by a fourth neural network model to obtain a vector representation of the semantic features of the reference image data, wherein the fourth neural network model is a self-supervised training model using an attention mechanism; Based on the third neural network model, an attention mechanism is used to fuse the vector representation of the audio features aligned with the visual features of the reference image data, the vector representation of the semantic features of the reference image data, and the vector representation of the visual features of the reference image data to obtain fused features; Based on the third neural network model, with the fused features as guidance information, iteratively performing denoising and noise addition processing on the vector representation of at least one initial digital human image data, to obtain the vector representation of at least one digital human image data.
[0113] In a feasible design, the mask feature determination module is implemented in the following manner, and the vector representation of the mask feature corresponding to the target part is determined according to the position of the target part of the person in the reference image data: Determine corresponding mask image data according to the position of the target part of the person in the reference image data; The features of the mask region in the mask image data are analyzed based on the region guidance model to obtain a vector representation of the mask features corresponding to the target part.
[0114] In a feasible design, the visual feature acquisition module is implemented in the following way to obtain the vector representation of the visual features of the reference image data: Encoding the reference image data to obtain a vector representation of the reference image data; The vector representation of the reference image data is input into the fifth neural network model using the attention mechanism to obtain the vector representation of the visual features of the reference image data.
[0115] In a feasible design, the visual feature acquisition module is implemented by inputting a vector representation of the reference image data into the fifth neural network model using the attention mechanism to obtain a vector representation of the visual features of the reference image data, including: The vector representation of the reference image data and the vector representation of the semantic features of the reference image data are input into the fifth neural network model using the attention mechanism, so that the fifth neural network model generates a vector representation of the visual features of the reference image data by fusing the semantic features of the reference image data.
[0116] In a feasible design, the multimodal big model includes an audio language model, a big language model, and a visual language model. The text response module is implemented in the following way: based on the multimodal big model, the multimodal data is subjected to intent recognition and sentiment analysis to determine the response text data: If the multimodal data includes audio data, the audio data is processed by an audio language model to obtain text data converted from the audio data and speech emotion information of the audio data; intention recognition and emotion analysis are performed on data other than audio data in the multimodal data, text data converted from the audio data, and speech emotion information of the audio data to determine response text data; If the multimodal data only includes text data, the text data is subjected to intent recognition and sentiment analysis through a large language model to obtain response text data; If the multimodal data includes image data but does not include audio data, the multimodal data is subjected to intent recognition and sentiment analysis through a visual language model to obtain response text data.
[0117] In a feasible design, the text response module performs intention recognition and sentiment analysis on data other than audio data in the multimodal data, text data converted from audio data, and speech sentiment information of audio data to determine the response text data in the following manner: If the data other than the audio data in the multimodal data only includes text data, the text data, the text data converted from the audio data, and the speech emotion information of the audio data are subjected to intention recognition and emotion analysis through the large language model to obtain response text data; If the data other than audio data in the multimodal data includes image data, the visual language model is used to perform intent recognition and sentiment analysis on the data other than audio data in the multimodal data, the text data converted from the audio data, and the speech sentiment information of the audio data to obtain response text data.
[0118] In a feasible design, the text response module obtains response text data through a large language model in the following way: After performing intent recognition and sentiment analysis on text data through a large language model, or performing intent recognition and sentiment analysis on text data, text data converted from audio data, and speech sentiment information of audio data through a large language model, if the large language model cannot generate response text data, a corresponding intelligent agent is determined according to the target text data, the intelligent agent is a third-party application platform, and the target text data includes text data, or text data converted from text data and audio data; Query the agent based on the target text data to obtain reference data output by the agent; The target text data and the reference data are processed by a large language model to obtain response text data; Alternatively, the text response module obtains the response text data through the visual language model in the following way: After performing intent recognition and sentiment analysis on the multimodal data through a visual language model, or after performing intent recognition and sentiment analysis on data other than audio data in the multimodal data, text data converted from audio data, and speech sentiment information of audio data through a visual language model, if the visual language model cannot generate response text data, determining the intent category or professional field according to the target multimodal data based on the visual language model, the target multimodal data includes the multimodal data, or data other than audio data in the multimodal data and text data converted from audio data; Determine the corresponding intelligent agent based on the intent category or professional field. The intelligent agent is a third-party application platform. Query the agent based on the target multimodal data to obtain reference data output by the agent; The target multimodal data and reference data are processed through a visual language model to obtain response text data.
[0119] The descriptions of other implementations of the above-mentioned device and the embodiments of the method for generating an effect digital human are not repeated here.
[0120] The basic principles of the present application are described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, effects, etc. mentioned in the present application are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. are required by each embodiment of the present application. In addition, the specific details disclosed above are only for the purpose of illustration and ease of understanding, not for limitation, and the above details do not limit the present application to being implemented by adopting the above specific details.
[0121] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.
[0122] The block diagrams of the devices, apparatuses, equipment, and systems involved in this application are only illustrative examples, and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagram. As will be appreciated by those skilled in the art, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open words, referring to "including but not limited to", and can be used interchangeably with them. The words "or" and "and" used here refer to the words "and / or" and can be used interchangeably with them, unless the context clearly indicates otherwise. The words "such as" used here refer to the phrase "such as but not limited to", and can be used interchangeably with them.
[0123] It should also be noted that in the apparatus, device and method of the present application, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present application.
[0124] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
[0125] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.
Claims
1. A method for generating digital human based on a multimodal large model, characterized in that: include: Get multimodal data input by the user; Performing intent recognition and sentiment analysis on the multimodal data based on the multimodal big model to determine response text data; Determine corresponding audio sequence data according to the response text data; Obtaining a vector representation of visual features of reference image data, wherein the content of the reference image data includes a person and a background; Determining, according to the position of the target part of the person in the reference image data, a vector representation of the mask feature corresponding to the target part; Based on the diffusion model, denoising is performed on at least one noise vector representation according to the audio sequence data, the vector representation of the visual features of the reference image data, and the vector representation of the mask features corresponding to the target part, so as to generate at least one frame of digital human image data, wherein the noise vector representation is obtained by adding noise to the vector representation of the visual features of the reference image data; A digital human animation with voice is generated according to the at least one frame of digital human image data and the audio sequence data.
2. The method according to claim 1, characterized in that The method of performing denoising on at least one noise vector representation based on the audio sequence data, the vector representation of the visual features of the reference image data, and the vector representation of the mask features corresponding to the target part based on the diffusion model to generate at least one frame of digital human image data includes: Extracting features from the audio sequence data using a first neural network model to obtain a vector representation of initial audio features; Processing the vector representation of the initial audio features based on a second neural network model composed of two layers of multi-layer perceptrons to obtain a vector representation of the audio features aligned with the visual features of the reference image data; Adding the vector representation of the mask feature corresponding to the target part and at least one noise vector representation to obtain a vector representation of at least one initial digital human image data; Based on a third neural network model, denoising the vector representation of the at least one initial digital human image data according to the vector representation of the audio features aligned with the visual features of the reference image data and the vector representation of the visual features of the reference image data, to obtain a vector representation of the image data of at least one digital human, wherein the third neural network model is a diffusion model using an attention mechanism; The vector representation of the image data of the at least one digital human is decoded to generate at least one frame of image data of the digital human.
3. The method according to claim 2, characterized in that The method of performing denoising on the vector representation of at least one initial digital human image data based on the third neural network model and the vector representation of the audio features aligned with the visual features of the reference image data and the vector representation of the visual features of the reference image data to obtain the vector representation of at least one digital human image data comprises: Processing the reference image data through a fourth neural network model to obtain a vector representation of the semantic features of the reference image data, wherein the fourth neural network model is a model using an attention mechanism for self-supervised training; Based on the third neural network model, using an attention mechanism to perform feature fusion on the vector representation of the audio features aligned with the visual features of the reference image data, the vector representation of the semantic features of the reference image data, and the vector representation of the visual features of the reference image data to obtain fused features; Based on the third neural network model, and taking the fused features as guidance information, iteratively performing denoising and noise addition processing on the vector representation of the at least one initial digital human image data, to obtain a vector representation of the at least one digital human image data.
4. The method according to any one of claims 1 to 3, characterized in that The step of determining the vector representation of the mask feature corresponding to the target part according to the position of the target part of the person in the reference image data comprises: Determining corresponding mask image data according to the position of the target part of the person in the reference image data; The features of the mask region in the mask image data are analyzed based on the region guidance model to obtain a vector representation of the mask features corresponding to the target part.
5. The method according to claim 3, characterized in that: The step of obtaining a vector representation of visual features of the reference image data comprises: Encoding the reference image data to obtain a vector representation of the reference image data; The vector representation of the reference image data is input into a fifth neural network model that adopts an attention mechanism to obtain a vector representation of the visual features of the reference image data.
6. The method according to claim 5, characterized in that The step of inputting the vector representation of the reference image data into the fifth neural network model using the attention mechanism to obtain the vector representation of the visual features of the reference image data comprises: The vector representation of the reference image data and the vector representation of the semantic features of the reference image data are input into a fifth neural network model that adopts an attention mechanism, so that the fifth neural network model generates a vector representation of the visual features of the reference image data by fusing the semantic features of the reference image data.
7. The method according to any one of claims 1 to 3, characterized in that The multimodal large model includes an audio language model, a large language model and a visual language model. The multimodal data is subjected to intention recognition and sentiment analysis based on the multimodal large model to determine the response text data, including: If the multimodal data includes audio data, the audio data is processed by the audio language model to obtain text data converted from the audio data and speech emotion information of the audio data; intention recognition and emotion analysis are performed on data other than the audio data in the multimodal data, the text data converted from the audio data, and the speech emotion information of the audio data to determine response text data; If the multimodal data only includes text data, performing intent recognition and sentiment analysis on the text data through the large language model to obtain response text data; If the multimodal data includes image data but does not include audio data, the visual language model is used to perform intent recognition and sentiment analysis on the multimodal data to obtain response text data.
8. The method according to claim 7, characterized in that The performing intention recognition and sentiment analysis on the data other than the audio data in the multimodal data, the text data converted from the audio data, and the speech sentiment information of the audio data to determine the response text data includes: If the data other than the audio data in the multimodal data only includes text data, performing intention recognition and sentiment analysis on the text data, the text data converted from the audio data, and the speech sentiment information of the audio data through the large language model to obtain response text data; If the data other than the audio data in the multimodal data includes image data, the visual language model is used to perform intent recognition and sentiment analysis on the data other than the audio data in the multimodal data, the text data converted from the audio data, and the voice sentiment information of the audio data to obtain response text data.
9. The method according to claim 8, characterized in that The response text data is obtained through the large language model, including: After performing intent recognition and sentiment analysis on the text data through a large language model, or after performing intent recognition and sentiment analysis on the text data, the text data converted from the audio data, and the voice sentiment information of the audio data through the large language model, if the large language model cannot generate response text data, determining a corresponding intelligent agent according to the target text data, the intelligent agent is a third-party application platform, and the target text data includes the text data, or the text data and the text data converted from the audio data; Querying the agent according to the target text data to obtain reference data output by the agent; Processing the target text data and the reference data by using the large language model to obtain response text data; Alternatively, the response text data can be obtained through the visual language model, including: After performing intent recognition and sentiment analysis on the multimodal data through a visual language model, or after performing intent recognition and sentiment analysis on data other than the audio data in the multimodal data, text data converted from the audio data, and voice sentiment information of the audio data through the visual language model, if the visual language model cannot generate response text data, determining the intent category or professional field according to target multimodal data based on the visual language model, the target multimodal data including the multimodal data, or data other than the audio data in the multimodal data and text data converted from the audio data; Determine a corresponding intelligent agent according to the intent category or professional field, wherein the intelligent agent is a third-party application platform; querying the agent according to the target multimodal data to obtain reference data output by the agent; The target multimodal data and the reference data are processed by the visual language model to obtain response text data.
10. A digital human generation device based on a multimodal large model, characterized in that: include: An input data acquisition module is used to acquire multimodal data input by a user; A text response module, used to perform intent recognition and sentiment analysis on the multimodal data based on the multimodal large model to determine response text data; An audio sequence data generating module, used to determine corresponding audio sequence data according to the response text data; A visual feature acquisition module, used to acquire a vector representation of visual features of reference image data, wherein the content of the reference image data includes a person and a background; A mask feature determination module, used to determine a vector representation of a mask feature corresponding to a target part of the person according to a position of the target part in the reference image data; a digital human image generation module, configured to denoise at least one noise vector representation based on the audio sequence data, the vector representation of the visual features of the reference image data, and the vector representation of the mask features corresponding to the target part, so as to generate at least one frame of digital human image data, wherein the noise vector representation is obtained by adding noise to the vector representation of the visual features of the reference image data; The digital human animation generation module is used to generate a digital human animation with voice according to the at least one frame of digital human image data and the audio sequence data.
Citation Information
Patent Citations
Virtual human generation method and device based on audio control, equipment and medium
CN117152285A
Image generation method and device, electronic equipment, storage medium and program product
CN117437317A
Speaking face video generation method and device based on multi-modal information control
CN117456587A
Image redrawing method and device, electronic equipment and readable storage medium
CN119379561A
Video generation method and device based on multi-modal information fusion, equipment and medium
CN119906872A
Cited By
Digital human video generation method, training method and device of digital human generation model, and computer equipment
CN121330444A
Normal map generation method and digital human video generation method
CN121544760A
Normal map generation method and digital human video generation method
CN121544760B