Digital human generation method and device based on multi-modal large model
By combining multimodal large models and diffusion models, and utilizing the mask features and audio features of the target parts, a delicate and natural dynamic digital human image was generated. This solved the problem of unnatural digital human movements under posture key point control, and improved the performance and service quality in multimodal interaction scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2026-04-17
AI Technical Summary
Existing digital human generation solutions rely on pose key point control, which cannot fully capture the detailed changes in human body features, resulting in digital human movements that are not natural and smooth, and neglecting the detailed expression of specific parts.
A multimodal large model-based approach is adopted to acquire multimodal data for intent recognition and sentiment analysis. By combining diffusion models and neural networks, a delicate and natural dynamic digital human image is generated. Noise reduction is performed using mask features of target parts and audio features to generate highly consistent speech emotions and postures.
It has achieved the ability to capture details of the target parts of the digital human and coordinate with the overall posture, generating a more delicate and natural dynamic digital human image, improving the ability to perceive and understand emotional intentions in multimodal interaction scenarios, and providing a more professional and demand-oriented service experience.
Smart Images

Figure CN120107427B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital human generation technology, and in particular to a method and apparatus for digital human generation based on a multimodal large model. Background Technology
[0002] With the rapid development of artificial intelligence technology, digital humans have become an important bridge and application carrier connecting the virtual and the real world. Digital humans are a highly complex and multidisciplinary technological system. Utilizing technologies such as natural language processing, speech recognition, image recognition, and sentiment analysis, digital humans can understand users' deep-seated intentions and emotions, thereby achieving precise and efficient interaction. Digital human technology encompasses multiple core technology areas, including computer vision, image generation, and speech generation. Its goal is to generate virtual characters with realistic appearances, natural behaviors, rich expressions, and clear voices.
[0003] Currently, most digital human generation solutions rely on pose keypoints to control the generated digit. However, pose keypoints are typically a discretized representation of human posture. For example, common pose estimation models may only mark the coordinates of a few major joints, such as the head, shoulders, elbows, hands, hips, knees, and feet. This discretized representation cannot fully capture the detailed changes in specific parts of the human body, such as facial expressions and hand movements. Therefore, current keypoint-based representations can only present a limited set of preset pose combinations when generating digits, resulting in less natural and fluid movements and neglecting the detailed expression of specific human body parts.
[0004] Therefore, how to generate detailed and natural dynamic digital human images has become an urgent problem to be solved. Summary of the Invention
[0005] This application provides a method and apparatus for generating digital humans based on a multimodal large model, which breaks through the limitations of traditional generation methods based on pose key points or 3D representation, and can generate dynamic digital human images with delicate and natural expressions in the target parts.
[0006] Firstly, a digital human generation method based on a multimodal large model is provided, including:
[0007] Acquire multimodal data input by the user;
[0008] Based on a multimodal large model, intent recognition and sentiment analysis are performed on multimodal data to determine response text data;
[0009] Determine the corresponding audio sequence data based on the response text data;
[0010] Obtain vector representations of the visual features of the reference image data, which includes people and background;
[0011] Based on the position of the target part of the person in the reference image data, determine the vector representation of the mask feature corresponding to the target part;
[0012] Based on the diffusion model, the vector representations of the visual features of the audio sequence data, the reference image data, and the vector representations of the mask features corresponding to the target parts are used to denoise at least one noise vector representation to generate at least one frame of digital human image data. The noise vector representation is obtained by adding noise to the vector representations of the visual features of the reference image data.
[0013] Generate a digital human animation with voice based on at least one frame of digital human image data and audio sequence data.
[0014] In one feasible design, based on a diffusion model, denoising is performed on at least one noise vector representation using vector representations of visual features from audio sequence data, reference image data, and mask features corresponding to the target area, to generate at least one frame of digital human image data, including:
[0015] The first neural network model is used to extract features from the audio sequence data to obtain a vector representation of the initial audio features;
[0016] The vector representation of the initial audio features is processed based on a second neural network model consisting of two multilayer perceptrons to obtain a vector representation of the audio features aligned with the visual features of the reference image data.
[0017] Add the vector representation of the mask feature corresponding to the target part to at least one noise vector representation to obtain a vector representation of at least one initial digital human image data;
[0018] Based on the third neural network model, the vector representation of at least one initial digital human image data is denoised according to the vector representation of audio features aligned with the visual features of the reference image data and the vector representation of the visual features of the reference image data, so as to obtain the vector representation of at least one digital human image data. The third neural network model is a diffusion model that adopts an attention mechanism.
[0019] Decode the vector representation of image data of at least one digital human to generate at least one frame of image data of the digital human.
[0020] In one feasible design, based on a third neural network model, and according to the vector representations of audio features aligned with the visual features of the reference image data and the vector representations of the visual features of the reference image data, denoising is performed on the vector representation of at least one initial digital human image data to obtain the vector representation of at least one digital human image data, including:
[0021] The reference image data is processed by the fourth neural network model to obtain a vector representation of the semantic features of the reference image data. The fourth neural network model is a self-supervised training model that uses an attention mechanism.
[0022] Based on the third neural network model, an attention mechanism is used to fuse the vector representations of audio features aligned with the visual features of the reference image data, the vector representations of semantic features of the reference image data, and the vector representations of visual features of the reference image data to obtain the fused features.
[0023] Based on the third neural network model, and guided by the fused features, the vector representation of at least one initial digital human image data is iteratively subjected to noise addition and denoising processes to obtain the vector representation of at least one digital human image data.
[0024] In one feasible design, based on the position of the target part of the person in the reference image data, the vector representation of the mask feature corresponding to the target part is determined, including:
[0025] Determine the corresponding mask image data based on the position of the target part of the person in the reference image data;
[0026] Based on the region guidance model, the characteristics of the masked region in the masked image data are analyzed to obtain the vector representation of the masked features corresponding to the target part.
[0027] In a feasible design, obtaining a vector representation of the visual features of the reference image data includes:
[0028] The reference image data is encoded to obtain a vector representation of the reference image data;
[0029] By inputting a vector representation of the reference image data into a fifth neural network model employing an attention mechanism, a vector representation of the visual features of the reference image data is obtained.
[0030] In one feasible design, a vector representation of the reference image data is input into a fifth neural network model employing an attention mechanism to obtain a vector representation of the visual features of the reference image data, including:
[0031] The fifth neural network model employing the attention mechanism is input with a vector representation of the reference image data and a vector representation of the semantic features of the reference image data, so that the fifth neural network model can generate a vector representation of the visual features of the reference image data by fusing the semantic features of the reference image data.
[0032] In a feasible design, the multimodal large model includes an audio language model, a large language model, and a visual language model. Intent recognition and sentiment analysis are performed on the multimodal data based on the multimodal large model to determine the response text data, including:
[0033] If the multimodal data includes audio data, the audio data is processed through an audio language model to obtain text data converted from the audio data and speech emotion information of the audio data; intent recognition and emotion analysis are performed on the data other than audio data in the multimodal data, the text data converted from the audio data, and the speech emotion information of the audio data to determine the response text data;
[0034] If the multimodal data only includes text data, the response text data can be obtained by performing intent recognition and sentiment analysis on the text data through a large language model.
[0035] If the multimodal data includes image data but not audio data, the visual language model is used to perform intent recognition and sentiment analysis on the multimodal data to obtain response text data.
[0036] In one feasible design, intent recognition and sentiment analysis are performed on multimodal data other than audio data, text data converted from audio data, and speech sentiment information from audio data to determine response text data, including:
[0037] If the multimodal data, excluding audio data, only includes text data, then the large language model is used to perform intent recognition and sentiment analysis on the text data converted from audio data and the speech sentiment information of the audio data to obtain the response text data.
[0038] If the multimodal data includes image data in addition to audio data, the visual language model is used to perform intent recognition and sentiment analysis on the data other than audio data in the multimodal data, the text data converted from audio data, and the speech emotion information of audio data to obtain response text data.
[0039] In a feasible design, the response text data is obtained through a large language model, including:
[0040] After performing intent recognition and sentiment analysis on text data using a large language model, or after performing intent recognition and sentiment analysis on the speech sentiment information of text data and audio data converted from text data and audio data using a large language model, if the large language model cannot generate response text data, a corresponding intelligent agent is determined based on the target text data. The intelligent agent is a third-party application platform. The target text data includes text data, or text data converted from text data and audio data.
[0041] The agent is queried based on the target text data to obtain the reference data output by the agent;
[0042] The target text data and reference data are processed using a large language model to obtain the response text data;
[0043] Alternatively, response text data can be obtained through a visual language model, including:
[0044] After performing intent recognition and sentiment analysis on multimodal data using a visual language model, or after performing intent recognition and sentiment analysis on data other than audio data, text data converted from audio data, and speech sentiment information of audio data using a visual language model, if the visual language model cannot generate response text data, the intent category or professional field is determined based on the target multimodal data according to the visual language model. The target multimodal data includes multimodal data, or data other than audio data and text data converted from audio data in multimodal data.
[0045] The appropriate intelligent agent is determined based on the intent category or professional field, and the intelligent agent is a third-party application platform;
[0046] The agent is queried based on the target multimodal data to obtain the reference data output by the agent;
[0047] The target multimodal data and reference data are processed by a visual language model to obtain response text data.
[0048] Secondly, a digital human generation device based on a multimodal large model is provided, including:
[0049] The input data acquisition module is used to acquire multimodal data input by the user;
[0050] The text response module is used to perform intent recognition and sentiment analysis on multimodal data based on a multimodal large model to determine the response text data;
[0051] The audio sequence data generation module is used to determine the corresponding audio sequence data based on the response text data;
[0052] The visual feature acquisition module is used to acquire vector representations of the visual features of the reference image data, which includes people and background.
[0053] The mask feature determination module is used to determine the vector representation of the mask features corresponding to the target part based on the position of the target part of the person in the reference image data.
[0054] The digital human image generation module is used to perform denoising processing on at least one noise vector representation based on the vector representation of the visual features of the audio sequence data, the vector representation of the visual features of the reference image data, and the vector representation of the mask features corresponding to the target part, according to the diffusion model, to generate at least one frame of digital human image data. The noise vector representation is obtained by adding noise to the vector representation of the visual features of the reference image data.
[0055] The digital human animation generation module is used to generate digital human animation with voice based on at least one frame of digital human image data and audio sequence data.
[0056] This application's embodiments, through intent recognition and sentiment analysis of multimodal data, accurately extract response text data consistent with the dialogue scenario and convert it into audio sequence data with emotional features. This establishes an accurate emotional tone and audio-visual synchronization benchmark for the subsequent digital human generation process, enabling the subsequent diffusion model to generate a dynamic digital human image incorporating vocal emotions based on the audio and visual features of the audio sequence data. Furthermore, it ensures a high degree of consistency between the digital human's posture at the target location and the audio features. Moreover, since the vector representation of the mask features determined based on the location of the target location in the reference image data can include the features of the region where the target location is located, as well as the relationship features between the target location and other regions, this application's solution, through mask features, can guide the diffusion model in enhancing the capture of natural details of target locations (such as the face and hands) and their coordination with the overall posture of the person during the generation of digital human image data. This results in a more delicate and natural expressiveness for the digital human at the target location. Furthermore, to enhance the expressiveness of specific parts of a digital human, the diffusion model can be guided by the vector representation of the mask features corresponding to those parts. Therefore, compared to the current digital human motion generated by relying on pose keypoint control, which is not natural and smooth enough, the solution proposed in this application can flexibly generate delicate and natural dynamic digital human images. Attached Figure Description
[0057] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, those skilled in the art can obtain other related drawings based on these drawings without creative effort.
[0058] Figure 1This is a schematic flowchart illustrating an exemplary embodiment of a digital human generation method based on a multimodal large model provided in this application;
[0059] Figure 2 This is a schematic flowchart illustrating yet another example of a digital human generation method based on a multimodal large model provided in an exemplary embodiment of this application;
[0060] Figure 3 This is a schematic diagram of the network framework of an exemplary digital human generation model provided in an exemplary embodiment of this application;
[0061] Figure 4 This is a schematic diagram of an exemplary embodiment of the present application of a digital human generation device based on a multimodal large model. Detailed Implementation
[0062] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0063] In order to flexibly generate delicate and natural dynamic digital human images, this application breaks away from the limitations of traditional digital human generation that relies on pose key point control. By introducing mask features of target parts of the person (such as face, hands, etc.), the neural network model is enhanced in its ability to capture details of target parts, thereby enabling the neural network model to generate delicate and expressive dynamic digital human images.
[0064] Furthermore, this application optimizes the generation algorithm for matching target body parts with audio features by injecting audio features into the neural network model, thereby achieving a high degree of consistency between target body parts and voice emotion, and between the overall posture of the person and voice emotion.
[0065] Furthermore, while digital humans have demonstrated strong potential in various scenarios, current digital human generation solutions have shortcomings in multimodal interaction scenarios, particularly in their limited depth of understanding of the intentions, motivations, and emotions contained in multimodal information. To improve the performance of digital humans in multimodal interaction scenarios, this application utilizes a multimodal large model to significantly enhance the perception and understanding of emotional intentions in its digital human generation solution by fusing user-input text, audio, and image data. Simultaneously, by combining a multi-agent system, the intelligent performance of the digital human service algorithm is further optimized, providing users with a more professional and tailored service experience. Furthermore, by combining enhanced knowledge base technology and a multimodal data retrieval mechanism, the high accuracy and real-time nature of the generated response content are ensured.
[0066] The digital human generation scheme provided in this application will be described in detail below, such as... Figure 1 As shown, the digital human generation method based on a multimodal large model provided in this application includes:
[0067] S110, acquire multimodal data input by the user.
[0068] Multimodal data includes, but is not limited to, text, audio, images, and video. Images and videos can be collectively referred to as images.
[0069] S120 performs intent recognition and sentiment analysis on multimodal data based on a multimodal large model to determine response text data.
[0070] For example, a multimodal large model includes an audio language model, a large language model, and a visual language model. It should be noted that the audio language model, the large language model, and the visual language model can be regarded as different parts of the multimodal large model, each of which can independently process the input data.
[0071] In a feasible design, intent recognition and sentiment analysis based on a multimodal large model are performed on multimodal data to determine response text data, thereby determining the response text data:
[0072] If the multimodal data includes audio data, the audio data is processed through an audio language model to obtain text data converted from the audio data and speech emotion information of the audio data; intent recognition and emotion analysis are performed on the data other than audio data in the multimodal data, the text data converted from the audio data, and the speech emotion information of the audio data to determine the response text data;
[0073] If the multimodal data only includes text data, the large language model is used to perform intent recognition and sentiment analysis on the text data to obtain response text data. The large language model is used to process the content and text data retrieved from the multimodal knowledge base to generate response text data.
[0074] If the multimodal data includes image data but not audio data, the visual language model is used to perform intent recognition and sentiment analysis on the multimodal data to obtain response text data. The visual language model is used to process the content and multimodal data retrieved from the multimodal knowledge base to generate response text data.
[0075] It should be understood that if multimodal data only includes audio data, it means that there is no data other than audio data in the multimodal data. In this case, the intention recognition and sentiment analysis can be performed directly on the text data converted from the audio data and the speech emotion information of the audio data through the large language model to determine the response text data.
[0076] It should be understood that multimodal data includes image data but excludes audio data, covering two cases: (1) multimodal data includes only image data; (2) multimodal data includes both image data and text data.
[0077] The multimodal knowledge base, a dynamic knowledge base constructed in this application, supports retrieval enhancement algorithms and dynamic updates. This multimodal knowledge base stores a large amount of text, image, and audio data, providing a reliable foundation for generating response data and ensuring that the generated content is accurate and context-relevant.
[0078] This application does not limit the network framework of the large language model. It can realize retrieval in a multimodal knowledge base based on retrieval enhancement algorithm, perform intent recognition and sentiment analysis on the retrieved content and text data, and generate response text data. For example, the large language model can be the Tongyi Qianwen large model (Qwen).
[0079] This application does not limit the network framework of the visual language model. It can realize retrieval in a multimodal knowledge base based on retrieval enhancement algorithm, perform intent recognition and sentiment analysis on the retrieved content and multimodal data, and generate response text data. For example, the visual language model can be the Qwen-VL visual language model.
[0080] This application does not limit the network framework of the audio language model, as long as it can convert audio data into text data and analyze the speech emotion information in the audio data. For example, the audio language model can be the Qwen-Audio audio language model.
[0081] In other words, such as Figure 2 As shown, before processing multimodal data through a large language model or a visual language model, it is necessary to determine whether the multimodal data includes audio data.
[0082] Since audio data contains linguistic and emotional information, if multimodal data includes audio data, it is necessary to first convert the audio data into text data using an audio language model and extract the linguistic and emotional information. Then, subsequent intent recognition and sentiment analysis processing can be performed using a large language model or a visual language model.
[0083] If the multimodal data does not include audio data, it can be processed directly using a large language model or a visual language model.
[0084] In the above example, when multimodal data includes audio data, an audio language model converts the audio data into text data. This allows the subsequent neural network model to integrate textual information from the audio data with information from other data besides audio in the multimodal data, enabling a more comprehensive and accurate capture of the user's intent. Furthermore, by extracting emotional information from the audio data and injecting it into the neural network model, the model's sentiment analysis of the input multimodal data becomes more accurate. This facilitates the generation of emotionally rich and context-appropriate response text data, enhancing the understanding of the intent, motivation, and emotion contained in the multimodal data and providing more accurate information for subsequent digital human animation generation. In the case where audio data is not included, further determination of whether the included data format contains image data helps identify the type of model used to process the multimodal data, enhancing the adaptability of this application's digital human generation solution to multimodal scenarios.
[0085] Next, it is determined whether the multimodal data, excluding audio data, includes image data. If image data is included, the text data converted from the audio data, the image data, and the speech emotion information are input into the visual language model for appropriate processing to obtain the response text data. If image data is not included, the text data converted from the audio data, the multimodal data excluding audio data, and the speech emotion information are input into the large language model for appropriate processing to obtain the response text data.
[0086] In practical applications, if the multimodal data only includes text data, it is directly input into the large language model. The large language model uses retrieval enhancement algorithms to search the multimodal knowledge base, and then performs intent recognition and sentiment analysis on the retrieved content and text data to quickly generate response text data. If the multimodal data includes image data but not audio data (e.g., only image data, or both image and text data), it is input into the visual language model. The visual language model uses retrieval enhancement algorithms to search the multimodal knowledge base, and then performs intent recognition and sentiment analysis on the retrieved content and multimodal data to generate response text data.
[0087] In the above example, when multimodal data includes image data, a visual language model is used to perform joint intent recognition and sentiment analysis on the input text data, image data, etc., to generate response text data. This provides accurate text information for responding to users in subsequent digital human animation generation, solving the problem that current digital human generation solutions cannot adapt to scenarios with multimodal input data.
[0088] Furthermore, since large language models respond faster than visual language models, the above example can improve the processing efficiency of input data by processing different types of input data based on different neural network models, compared to performing intent recognition and sentiment analysis on multimodal input data solely based on visual language models.
[0089] In one feasible design, intent recognition and sentiment analysis are performed on multimodal data other than audio data, text data converted from audio data, and speech sentiment information from audio data to determine response text data, including:
[0090] If the multimodal data, excluding audio data, only includes text data, then the large language model is used to perform intent recognition and sentiment analysis on the text data converted from audio data and the speech sentiment information of the audio data to obtain the response text data.
[0091] If the multimodal data includes image data in addition to audio data, the visual language model is used to perform intent recognition and sentiment analysis on the data other than audio data in the multimodal data, the text data converted from audio data, and the speech emotion information of audio data to obtain response text data.
[0092] If multimodal data includes not only audio data but also other types of data, the above example can effectively select the appropriate model to process the multimodal data by further identifying the types of other data, thereby further improving the adaptability of the digital human generation solution to multimodal input data scenarios.
[0093] In a feasible design, such as Figure 2 As shown, the response text data is obtained from the large language model in the following way:
[0094] After performing intent recognition and sentiment analysis on text data using a large language model, or after performing intent recognition and sentiment analysis on the speech sentiment information of text data and audio data converted from text data and audio data using a large language model, if the large language model cannot generate response text data, then the corresponding intelligent agent is determined based on the target text data. The intelligent agent can be a third-party application platform. The target text data includes text data, or text data converted from text data and audio data.
[0095] The agent is queried based on the target text data to obtain the reference data output by the agent;
[0096] The target text data and reference data are processed using a large language model to obtain the response text data.
[0097] It should be understood that in this example, obtaining response text data through the large language model includes two cases: (1) obtaining response text data by performing intent recognition and sentiment analysis on the user-input text data through the large language model; (2) obtaining response text data by performing intent recognition and sentiment analysis on the user-input text data, the text data converted from audio data, and the speech sentiment information through the large language model.
[0098] For example, the following method can be used to query the agent based on the target text data and obtain the reference data output by the agent:
[0099] Request information is generated based on the target text data using a large language model;
[0100] Send a request message to the intelligent agent;
[0101] Receive reference data from the intelligent agent.
[0102] It should be noted that this application does not impose any restrictions on the interface or protocol used by the large language model to call the intelligent agent.
[0103] For example, if the target text is "What's the weather like today?", since the large language model lacks this real-time information, it needs to call the agent to query the weather data. The large language model generates a request message as "Get the weather conditions for XX year XX month XX day". After obtaining the weather data "Sunny, 22℃" (i.e., reference data) from the agent, the large language model processes the target text and the reference data to obtain the response text "The weather is sunny today, suitable for going out."
[0104] For example, the appropriate agent is determined based on the target text data in the following way:
[0105] Establish the first association relationship, which is used to associate keywords with agent information. Agent information includes the agent's name, functional description, domain it is good at, type of task it is good at or type of data it provides.
[0106] Keyword identification is performed on the target text data;
[0107] Match the keywords in the target text data with the keywords in the first association relationship to determine the corresponding intelligent agent.
[0108] For example, if the target text data contains "weather", after matching it with the first association, the agent is determined to be "weather query program".
[0109] For example, the appropriate agent is determined based on the target text data in the following way:
[0110] Establish a second association between intent category and agent information. The second association is used to associate intent category with agent information. Agent information includes the agent's name, functional description, domain it is good at, task type it is good at, or data type it provides.
[0111] Obtain the intent category obtained after the large language model performs intent recognition on the target text data;
[0112] The intent categories obtained from the large language model are matched with the intent categories in the second association relationship to determine the corresponding intelligent agent.
[0113] For example, the large language model identifies the intent category of the target text data as "weather query" and, after matching it with the second association relationship, determines the agent as "weather query program".
[0114] For example, the appropriate agent is determined based on the target text data in the following way:
[0115] Establish a third association between the professional domain and the agent information. The third association is used to associate the name of the professional domain with the agent information. The agent information includes the agent's name, functional description, domain that it is good at handling, type of task that it is good at handling, or type of data it provides.
[0116] Obtain the professional field name after the large language model performs professional field recognition on the target text data;
[0117] The professional domain names obtained from the large language model are matched with the professional domain names in the third association relationship to determine the corresponding intelligent agents.
[0118] For example, the large language model identifies the professional field of the target text data as the weather field. After matching it with the third association relationship, the intelligent agent is determined to be "weather query program".
[0119] The examples above provide multiple ways to determine the appropriate agents based on text data or multimodal data, optimizing the ability of multi-agents in task distribution and collaboration, and effectively improving the reliability of user services.
[0120] For example, after determining that an anomaly has occurred, the intelligent agent generates an alarm response and then sends the alarm response to the business personnel, thereby improving the reliability and security of the system.
[0121] For example, a user asks a large language model if their mobile phone number cannot receive calls. The large language model then uses an agent provided by the operator to check the information of the mobile phone number and the reason. The agent finds that the problem is not with the mobile phone signal but with the local base station signal tower. The agent can then generate an alarm to prompt the business personnel to handle the issue.
[0122] In a feasible design, such as Figure 2 As shown, the response text data is obtained through a visual language model in the following way:
[0123] After performing intent recognition and sentiment analysis on multimodal data using a visual language model, or after performing intent recognition and sentiment analysis on data other than audio data, text data converted from audio data, and speech sentiment information of audio data using a visual language model, if the visual language model cannot generate response text data, the intent category or professional field is determined based on the target multimodal data according to the visual language model. The target multimodal data includes multimodal data, or data other than audio data and text data converted from audio data in multimodal data.
[0124] The appropriate intelligent agent is determined based on the intent category or professional field, and the intelligent agent can be a third-party application platform;
[0125] The agent is queried based on the target multimodal data to obtain the reference data output by the agent;
[0126] The target multimodal data and reference data are processed by a visual language model to obtain response text data.
[0127] It should be understood that in this example, obtaining response text data through the visual language model includes two cases: (1) if the multimodal data does not include audio data, the visual language model is used to perform intent recognition and sentiment analysis on the multimodal data to obtain response text data; (2) if the multimodal data includes audio data, the visual language model is used to perform intent recognition and sentiment analysis on the data other than audio data, the text data converted from audio data, and the speech emotion information in the multimodal data to obtain response text data.
[0128] The method for determining the appropriate agent based on the intent category or professional field can be found in the example of determining the appropriate agent based on the target text data, and will not be repeated here.
[0129] For example, the reference data output by the agent can be obtained by querying the agent based on the target multimodal data in the following way:
[0130] Request information is generated based on target multimodal data using a visual language model;
[0131] Send a request message to the intelligent agent;
[0132] Receive reference data from the intelligent agent.
[0133] It should be noted that this application does not limit the interface or protocol used by the visual language model to call the intelligent agent.
[0134] For example, after the visual language model identifies the intent, it determines the intent category to be "query the weather." Since the visual language model lacks this real-time information, it needs to call the agent to query weather data. The request information generated by the visual language model is "Get the weather conditions for XX year XX month XX day." After obtaining the weather data from the agent as "sunny, 22℃" (i.e., reference data), the response text data is obtained by processing the intent category and the reference data: "Today's weather is sunny, suitable for going out."
[0135] In practical applications, large language models or visual language models may face various limitations when processing certain specific multimodal data, preventing them from directly generating appropriate response text. These limitations may stem from relevant policy and regulatory requirements, or from the model's own lack of certain necessary information. For example, models may be unable to directly acquire and process user privacy data (such as mobile phone number information) or real-time data specific to certain industries (such as real-time weather queries requiring specialized meteorological data interfaces). In such cases, introducing an intelligent agent to provide reference data can help the model overcome these limitations, generate response text more effectively, and improve its ability to handle problems in multimodal scenarios.
[0136] Furthermore, since there can be multiple agents, the above example also supports multiple agents processing tasks simultaneously, enabling multi-task parallel processing and improving the model's response efficiency.
[0137] S130, determine the corresponding audio sequence data based on the response text data.
[0138] In practice, text-to-speech (TTS) technology can be used to convert response text data into audio sequence data with emotional features.
[0139] S140, Obtain a vector representation of the visual features of the reference image data.
[0140] The reference image data includes both people and background.
[0141] In a feasible design, the vector representation of the visual features of the reference image data is achieved through the following method:
[0142] The reference image data is encoded to obtain a vector representation of the reference image data;
[0143] By inputting a vector representation of the reference image data into a fifth neural network model employing an attention mechanism, a vector representation of the visual features of the reference image data is obtained.
[0144] For example, reference image data is encoded based on an encoder, which is the encoder part of a Variational Autoencoder Encoder (VAE) model.
[0145] For example, the fifth neural network model is a pre-trained two-dimensional U-shaped convolutional network (U-Net) model.
[0146] The example above, by encoding the reference image data, transforms it into a vector representation that is easier for the fifth neural network model to process and understand. This vector representation contains key features of the image, such as color, texture, and shape. Then, based on the fifth neural network model, the encoded vector representation is further processed through an attention mechanism. This allows for deeper analysis and transformation of the key features, extracting higher-level visual features. These higher-level visual features enable the subsequent neural network model to better capture background details and identity information in the reference image, ensuring that the generated digital human maintains consistency with the reference image in terms of person identity and background.
[0147] S150, based on the position of the target part of the person in the reference image data, determine the vector representation of the mask feature corresponding to the target part.
[0148] The target area can include multiple areas, such as the face and hands. The vector representation of the mask features is a feature vector generated by processing the masked image data using image segmentation techniques.
[0149] In one feasible design, the vector representation of the mask features corresponding to the target body part is determined based on the position of the target body part in the reference image data, in the following way:
[0150] Determine the corresponding mask image data based on the position of the target part of the person in the reference image data;
[0151] Based on the region guidance model, the features of the masked regions in the masked image data are analyzed to obtain the vector representation of the mask features corresponding to the target location. Specifically, the masked image data is input into the region guidance model to obtain a vector representation of the masked image data; the region guidance model can analyze the features of the masked regions in the masked image data, and the output vector representation of the masked image data is the vector representation of the mask features corresponding to the target location.
[0152] The vector representation of the mask image data and the representation of each noise vector have the same dimension, so that the vector representation of the mask image data and the representation of each noise vector can be added together in the future.
[0153] The noise vector representation is obtained by adding noise to the vector representation of the visual features of the reference image data. Specifically, after the reference image data is encoded into a latent representation using a VAE model, Gaussian noise is added to this latent representation to generate the noise vector representation.
[0154] It should be understood that, in this application, the mask region refers to the region corresponding to the target part in the mask image data.
[0155] It should be understood that the vector representation of masked image data includes not only the mask features of the region where the target area is located, but also the relationship features between the target area and other regions, as well as the mask features of other regions. Therefore, the vector representation of the mask features corresponding to the target area does not simply refer to the vector representation of the mask features of the region where the target area is located in the masked image data.
[0156] Masked image data can employ a binary mask structure (target area as 1, background as 0), which can accurately represent the contour information of the target area. The example above uses masked image data to highlight the position of the target area in the reference image data. This allows the vector representation of the masked image data output by the region-guided model to include the mask features of the target area, the relationship between the target area and other areas, etc. This enables the neural network model used to generate the digital human to accurately locate the target area and focus more on the detailed generation of the key areas where the target area is located, thereby improving the detail reproduction of the target area in the digital human. Furthermore, the relationship between the target area and other areas in the vector representation of the masked image data allows the model to consider the coordination between the target area and other parts of the human body when generating the digital human, resulting in a digital human that better conforms to human proportions and natural posture, thus improving the overall expressiveness and coordination of the digital human.
[0157] For example, the region guidance model consists of four convolutional network layers, each with a kernel size of 4×4 (in pixels) and a stride of 2×2 (in pixels). The number of channels in the four convolutional network layers are 16, 32, 64, and 128 (in pixels), respectively.
[0158] The region-guided model shown in the example above can progressively abstract and extract higher-level features by gradually increasing the number of channels, mining deeper information, and improving the ability of the digital human generation model to capture details in target areas. Furthermore, by using appropriate convolutional kernel size and stride settings, it helps control the amount of data and reduce computational complexity.
[0159] S160, based on the diffusion model, according to the vector representation of the visual features of the audio sequence data, the vector representation of the visual features of the reference image data, and the vector representation of the mask features corresponding to the target part, denoise at least one noise vector representation to generate at least one frame of digital human image data.
[0160] In one feasible design, this is achieved by denoising at least one noise vector representation based on a diffusion model, using vector representations of visual features from audio sequence data, reference image data, and mask features corresponding to the target area, to generate at least one frame of digital human image data:
[0161] The first neural network model is used to extract features from the audio sequence data to obtain a vector representation of the initial audio features;
[0162] The vector representation of the initial audio features is processed by a second neural network model consisting of two multilayer perceptrons (MLPs) to obtain a vector representation of the audio features aligned with the visual features of the reference image data.
[0163] Add the vector representation of the mask feature corresponding to the target part to at least one noise vector representation to obtain a vector representation of at least one initial digital human image data;
[0164] Based on the third neural network model, the vector representation of at least one initial digital human image data is denoised according to the vector representation of audio features aligned with the visual features of the reference image data and the vector representation of the visual features of the reference image data, so as to obtain the vector representation of at least one digital human image data. The third neural network model is a diffusion model that adopts an attention mechanism.
[0165] Decode the vector representation of image data of at least one digital human to generate at least one frame of image data of the digital human.
[0166] For example, the first neural network model is a pre-trained Waveform-to-Vector (WAV2VEC) model.
[0167] For example, the third neural network model is a denoising Unet model.
[0168] For example, the second neural network model (also referred to as the Audio Project model in this application) also employs a cross-attention mechanism, and uses this mechanism to input the vector representation of audio features aligned with the visual features of the reference image data into the third neural network model. Correspondingly, the third neural network model can utilize the audio attention mechanism to process the vector representation of audio features aligned with the visual features of the reference image data, achieving the fusion of audio features with other features. Since the denoising U-shaped network model cannot understand the vector representation of the initial audio features, the second neural network model needs to transform the vector representation of the initial audio features into a vector representation of audio features that the denoising U-shaped network model can understand, thus achieving alignment with the visual features of the reference image data. Specifically, "alignment" refers to the nonlinear transformation of the initial audio features by the second neural network model, making its dimension and distribution match the input requirements of the denoising U-shaped network model, while semantically associating it with the visual features of the reference image (such as facial expressions and hand gestures). For example, the emotional intensity in the audio is mapped by the second neural network model to a numerical range corresponding to the curvature of the corners of the mouth and the position of the eyebrows in the visual features, thereby achieving coordinated control of audio and video features.
[0169] For example, the fifth neural network model also employs a spatial attention mechanism, which is used to input the vector representation of the visual features of the reference image data into the third neural network model. Correspondingly, the third neural network model can utilize the spatial attention mechanism to process the vector representation of the visual features of the reference image data, enabling the fusion of visual features with other features, etc.
[0170] The above example uses a first neural network model to extract features from the rich semantic and acoustic information in the audio sequence data and convert them into vector representations. Then, leveraging the powerful nonlinear mapping capabilities of a second neural network model composed of two multilayer perceptrons, it learns the complex relationship between audio and image features, achieving precise alignment between them. This allows audio and visual features to jointly guide the digital human generation process of the third neural network model, ensuring a high degree of consistency between the generated digital human's speech and dynamic facial expressions. This application's solution abandons the traditional multi-condition injection method, realizing a single-stage digital human generation technology based on audio and reference images. It can directly synthesize stable digital human videos based on audio features and the visual features of the reference image, improving generation efficiency and flexibility. Furthermore, the above example also adds the vector representation of the mask features corresponding to the target area to at least one noise vector representation to obtain a vector representation of at least one initial digital human image data. This makes the target area more prominent and clear in the vector representation of each initial digital human image data, helping the third neural network model better understand the structure and details of the target area, thereby generating a more detailed and realistic digital human image.
[0171] In one feasible design, this is achieved by denoising the vector representation of at least one initial digital human image data based on a third neural network model, using the vector representation of audio features aligned with the visual features of the reference image data and the vector representation of the visual features of the reference image data, to obtain the vector representation of the image data of at least one digital human:
[0172] The reference image data is processed by the fourth neural network model to obtain a vector representation of the semantic features of the reference image data. The fourth neural network model is a self-supervised training model that uses an attention mechanism.
[0173] Based on the third neural network model, an attention mechanism is used to fuse the vector representations of audio features aligned with the visual features of the reference image data, the vector representations of semantic features of the reference image data, and the vector representations of visual features of the reference image data to obtain the fused features.
[0174] Based on the third neural network model, and guided by the fused features, the vector representation of at least one initial digital human image data is iteratively subjected to noise addition and denoising processes to obtain the vector representation of at least one digital human image data.
[0175] It should be noted that an iterative process includes first adding noise and then removing noise.
[0176] For example, the fourth neural network model is a Deformable Detection Transformer (DINO) model. The DINO model has the ability to capture rich detailed features, effectively separating and identifying subtle differences in details within an image. Through the DINO model, detailed poses (such as facial expressions and hand gestures) of target parts in a reference image can be accurately extracted, thus providing crucial guidance for the third neural network model to generate detailed and realistic digital humans in different scenarios.
[0177] For example, the fourth neural network model uses a cross-attention mechanism to input the vector representation of semantic features into the third neural network model. Correspondingly, the third neural network model uses the cross-attention mechanism to process the vector representation of semantic features to achieve the fusion of semantic features with other features.
[0178] The semantic features of image data contain high-level, abstract information. The example above uses a fourth neural network model to extract semantic features from the reference image data, enabling a deep understanding of the detailed information of the person in the reference image data and the separation of feature differences. Inputting the vector representation of the semantic features of the reference image data into a third neural network model can guide the third neural network model to accurately generate the fine pose of the reference image at the target location, such as fine facial expressions and hand movements.
[0179] For example, since the vector representation of an initial digital human image data corresponds to a frame of digital human image data, in order to ensure the temporal continuity of the frame sequence of the generated digital human animation, the third neural network model captures the complex temporal dependencies between the vector representations of consecutive initial digital human image data through a temporal attention mechanism, and fuses the temporal features of the vector representation of at least one initial digital human image data on the frame sequence time axis with other features, thereby facilitating the generation of smooth dynamic digital human videos.
[0180] In one feasible design, the vector representation of the reference image data input to the fifth neural network model employing an attention mechanism is used to obtain the vector representation of the visual features of the reference image data:
[0181] The fifth neural network model employing the attention mechanism is input with a vector representation of the reference image data and a vector representation of the semantic features of the reference image data, so that the fifth neural network model can generate a vector representation of the visual features of the reference image data by fusing the semantic features of the reference image data.
[0182] For example, the fourth neural network model uses a cross-attention mechanism to input the vector representation of the semantic features of the reference image data into the fifth neural network model. Accordingly, the fifth neural network model uses the cross-attention mechanism to process the vector representation of the semantic features to achieve the fusion of semantic features and visual features, etc.
[0183] The example above inputs the vector representation of the semantic features of the reference image data into the fifth neural network model. This helps the fifth neural network model to deeply understand the detailed information of the person in the reference image data and separate feature differences, thereby capturing the detailed visual features in the reference image data and obtaining deeper visual features. This, in turn, enables the subsequent third neural network model to more accurately generate the fine pose of the reference image at the target location based on the deeper visual features.
[0184] S170, Based on at least one frame of digital human image data and audio sequence data, generate a digital human animation with voice.
[0185] For example, by calling video editing software, image data and audio sequence data of at least one frame of a digital human are synthesized to obtain a digital human animation with voice.
[0186] For example, the input audio sequence data can be the response text data determined by the aforementioned multimodal big model after performing intent recognition and sentiment analysis on multimodal data, and obtained through TTS (Text To Speech) technology.
[0187] It should be noted that the input and output dimensions of each neural network model can be set according to actual needs, and this application does not impose any restrictions on this.
[0188] Based on the above embodiments, such as Figure 3 As shown, this application provides a digital human generation model that can generate half-body digital human videos by combining audio sequence data and reference image data. In many scenarios, such as customer service scenarios, half-body digital human animation has higher affinity than full-body digital humans or facial animation.
[0189] The overall framework of this digital human generation model is based on a VAE model, a DINO model, a two-dimensional U-Net model, a denoised U-shaped network model, a region guidance model, a WAV2VEC model, and a second neural network model composed of two multi-layer perceptrons.
[0190] The following section uses a half-body human image as the reference image data and the human's face and hands as the target parts as an example to illustrate the data processing process for the digital human generation model.
[0191] This digital human generation model receives audio sequence data, reference image data, masked image data, and at least one noise vector representation. It then extracts features from the audio sequence data using a WAV2VEC model and aligns the audio features with the visual features of the reference image data using a second neural network model. Finally, the cross-attention mechanism in the second neural network model inputs the vector representation of the audio features aligned with the visual features of the reference image data into a denoising U-shaped network model.
[0192] Digital human generation models utilize the encoder portion of the VAE model (i.e., Figure 3 The VAE (enc) shown encodes the reference image data to obtain a vector representation of the reference image data. This vector representation is then input into a 2D U-Net model. A self-supervised DINO model extracts the vector representation of the semantic features of the reference image data. The DINO model uses a cross-attention mechanism to input these vector representations into both the 2D U-Net model and the denoising U-shaped network model. The 2D U-Net model processes both the vector representations of the reference image data and their semantic features to obtain a vector representation of the visual features of the reference image data. The 2D U-Net model uses a spatial attention mechanism to input this vector representation of the visual features into the denoising U-shaped network model. A region-guided model processes the masked image data to obtain vector representations of the mask features corresponding to the face and hands. The vector representation of the mask features is added to at least one noise vector representation to obtain a vector representation of at least one initial digital human image data. This vector representation of at least one initial digital human image data is then input into the denoising U-shaped network model. The denoising U-shaped network model utilizes a spatial attention mechanism to extract features from the vector representation of visual features of the reference image data and fuses them with other features; it utilizes a cross-attention mechanism to extract features from the vector representation of semantic features of the reference image data and fuses them with other features; it utilizes an audio attention mechanism to extract features from the vector representation of audio features aligned with the visual features of the reference image data and fuses them with other features; and it utilizes a temporal attention mechanism to generate the temporal features of the vector representation of at least one initial digital human image data on the frame sequence time axis and fuses them with other features. The denoising U-shaped network model obtains fused features by fusing visual, audio, semantic, and temporal features. Using these fused features as guidance, it iteratively performs noise addition and denoising processing on the vector representation of at least one initial digital human image data, which incorporates mask features, to obtain the vector representation of at least one digital human image data. The decoder part of the VAE model (i.e., Figure 3The VAE (dec) shown decodes the vector representation of image data of at least one digital human, and can ultimately generate at least one frame of image data of the digital human that is fluent and highly consistent with speech and dynamic facial expressions.
[0193] This application's embodiments, through intent recognition and sentiment analysis of multimodal data, accurately extract response text data consistent with the dialogue scenario and convert it into audio sequence data with emotional features. This establishes an accurate emotional tone and audio-visual synchronization benchmark for the subsequent digital human generation process, enabling the subsequent diffusion model to generate a dynamic digital human image incorporating vocal emotions based on the audio and visual features of the audio sequence data. Furthermore, it ensures a high degree of consistency between the digital human's posture at the target location and the audio features. Moreover, since the vector representation of the mask features determined based on the location of the target location in the reference image data can include the features of the region where the target location is located, as well as the relationship features between the target location and other regions, this application's solution, through mask features, can guide the diffusion model in enhancing the capture of natural details of target locations (such as the face and hands) and their coordination with the overall posture of the person during the generation of digital human image data. This results in a more delicate and natural expressiveness for the digital human at the target location. Furthermore, to enhance the expressiveness of specific parts of a digital human, the diffusion model can be guided by the vector representation of the mask features corresponding to those parts. Therefore, compared to the current digital human motion generated by relying on pose keypoint control, which is not natural and smooth enough, the solution proposed in this application can flexibly generate delicate and natural dynamic digital human images.
[0194] like Figure 4 As shown, this application also provides a digital human generation apparatus, comprising:
[0195] The input data acquisition module is used to acquire multimodal data input by the user;
[0196] The text response module is used to perform intent recognition and sentiment analysis on multimodal data based on a multimodal large model to determine the response text data;
[0197] The audio sequence data generation module is used to determine the corresponding audio sequence data based on the response text data;
[0198] The visual feature acquisition module is used to acquire vector representations of the visual features of the reference image data, which includes people and background.
[0199] The mask feature determination module is used to determine the vector representation of the mask features corresponding to the target part based on the position of the target part of the person in the reference image data.
[0200] The digital human image generation module is used to perform denoising processing on at least one noise vector representation based on the vector representation of the visual features of the audio sequence data, the vector representation of the visual features of the reference image data, and the vector representation of the mask features corresponding to the target part, according to the diffusion model, to generate at least one frame of digital human image data. The noise vector representation is obtained by adding noise to the vector representation of the visual features of the reference image data.
[0201] The digital human animation generation module is used to generate digital human animation with voice based on at least one frame of digital human image data and audio sequence data.
[0202] In one feasible design, the digital human image generation module is implemented by performing denoising on at least one noise vector representation based on a diffusion model, using vector representations of visual features from audio sequence data, reference image data, and the mask features corresponding to the target area, to generate at least one frame of digital human image data, including:
[0203] The first neural network model is used to extract features from the audio sequence data to obtain a vector representation of the initial audio features;
[0204] The vector representation of the initial audio features is processed based on a second neural network model consisting of two multilayer perceptrons to obtain a vector representation of the audio features aligned with the visual features of the reference image data.
[0205] Add the vector representation of the mask feature corresponding to the target part to at least one noise vector representation to obtain a vector representation of at least one initial digital human image data;
[0206] Based on the third neural network model, the vector representation of at least one initial digital human image data is denoised according to the vector representation of audio features aligned with the visual features of the reference image data and the vector representation of the visual features of the reference image data, so as to obtain the vector representation of at least one digital human image data. The third neural network model is a diffusion model that adopts an attention mechanism.
[0207] Decode the vector representation of image data of at least one digital human to generate at least one frame of image data of the digital human.
[0208] In one feasible design, the digital human image generation module is implemented by, based on a third neural network model, denoising the vector representation of at least one initial digital human image data according to the vector representation of audio features aligned with the visual features of the reference image data and the vector representation of the visual features of the reference image data, to obtain the vector representation of at least one digital human image data, including:
[0209] The reference image data is processed by the fourth neural network model to obtain a vector representation of the semantic features of the reference image data. The fourth neural network model is a self-supervised training model that uses an attention mechanism.
[0210] Based on the third neural network model, an attention mechanism is used to fuse the vector representations of audio features aligned with the visual features of the reference image data, the vector representations of semantic features of the reference image data, and the vector representations of visual features of the reference image data to obtain the fused features.
[0211] Based on the third neural network model, and guided by the fused features, the vector representation of at least one initial digital human image data is iteratively subjected to noise addition and denoising processes to obtain the vector representation of at least one digital human image data.
[0212] In a feasible design, the mask feature determination module is implemented by determining the vector representation of the mask features corresponding to the target part of the person based on the position of the target part in the reference image data:
[0213] Determine the corresponding mask image data based on the position of the target part of the person in the reference image data;
[0214] Based on the region guidance model, the characteristics of the masked region in the masked image data are analyzed to obtain the vector representation of the masked features corresponding to the target part.
[0215] In a feasible design, the visual feature acquisition module obtains a vector representation of the visual features of the reference image data in the following way:
[0216] The reference image data is encoded to obtain a vector representation of the reference image data;
[0217] By inputting a vector representation of the reference image data into a fifth neural network model employing an attention mechanism, a vector representation of the visual features of the reference image data is obtained.
[0218] In one feasible design, the visual feature acquisition module is implemented by inputting a vector representation of the reference image data into a fifth neural network model employing an attention mechanism, thereby obtaining a vector representation of the visual features of the reference image data, including:
[0219] The fifth neural network model employing the attention mechanism is input with a vector representation of the reference image data and a vector representation of the semantic features of the reference image data, so that the fifth neural network model can generate a vector representation of the visual features of the reference image data by fusing the semantic features of the reference image data.
[0220] In a feasible design, the multimodal large model includes an audio language model, a large language model, and a visual language model. The text response module is implemented by performing intent recognition and sentiment analysis on the multimodal data based on the multimodal large model to determine the response text data:
[0221] If the multimodal data includes audio data, the audio data is processed through an audio language model to obtain text data converted from the audio data and speech emotion information of the audio data; intent recognition and emotion analysis are performed on the data other than audio data in the multimodal data, the text data converted from the audio data, and the speech emotion information of the audio data to determine the response text data;
[0222] If the multimodal data only includes text data, the response text data can be obtained by performing intent recognition and sentiment analysis on the text data through a large language model.
[0223] If the multimodal data includes image data but not audio data, the visual language model is used to perform intent recognition and sentiment analysis on the multimodal data to obtain response text data.
[0224] In a feasible design, the text response module performs intent recognition and sentiment analysis on data other than audio data in the multimodal data, text data converted from audio data, and speech sentiment information of audio data to determine the response text data:
[0225] If the multimodal data, excluding audio data, only includes text data, then the large language model is used to perform intent recognition and sentiment analysis on the text data converted from audio data and the speech sentiment information of the audio data to obtain the response text data.
[0226] If the multimodal data includes image data in addition to audio data, the visual language model is used to perform intent recognition and sentiment analysis on the data other than audio data in the multimodal data, the text data converted from audio data, and the speech emotion information of audio data to obtain response text data.
[0227] In a feasible design, the text response module obtains the response text data from a large language model in the following way:
[0228] After performing intent recognition and sentiment analysis on text data using a large language model, or after performing intent recognition and sentiment analysis on the speech sentiment information of text data and audio data converted from text data and audio data using a large language model, if the large language model cannot generate response text data, a corresponding intelligent agent is determined based on the target text data. The intelligent agent is a third-party application platform. The target text data includes text data, or text data converted from text data and audio data.
[0229] The agent is queried based on the target text data to obtain the reference data output by the agent;
[0230] The target text data and reference data are processed using a large language model to obtain the response text data;
[0231] Alternatively, the text response module can obtain response text data from a visual language model in the following way:
[0232] After performing intent recognition and sentiment analysis on multimodal data using a visual language model, or after performing intent recognition and sentiment analysis on data other than audio data, text data converted from audio data, and speech sentiment information of audio data using a visual language model, if the visual language model cannot generate response text data, the intent category or professional field is determined based on the target multimodal data according to the visual language model. The target multimodal data includes multimodal data, or data other than audio data and text data converted from audio data in multimodal data.
[0233] The appropriate intelligent agent is determined based on the intent category or professional field, and the intelligent agent is a third-party application platform;
[0234] The agent is queried based on the target multimodal data to obtain the reference data output by the agent;
[0235] The target multimodal data and reference data are processed by a visual language model to obtain response text data.
[0236] Other embodiments and effects of the above-described device are described in the embodiments of the digital human generation method, and will not be repeated here.
[0237] The basic principles of this application have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this application are merely examples and not limitations, and should not be considered as essential features of each embodiment of this application. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the application to the necessity of employing the aforementioned specific details for implementation.
[0238] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0239] The block diagrams of devices, apparatuses, devices, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0240] It should also be noted that in the apparatus, equipment, and methods of this application, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions of this application.
[0241] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0242] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A digital human generation method based on a multi-modal large model, characterized in that, include: Acquire multimodal data input by the user; Based on a multimodal large model, intent recognition and sentiment analysis are performed on the multimodal data to determine the response text data; Determine the corresponding audio sequence data based on the response text data; Obtain vector representations of the visual features of reference image data, the content of which includes people and background; Based on the position of the target part of the person in the reference image data, determine the corresponding mask image data; Based on the region guidance model, the features of the mask region in the mask image data are analyzed to obtain the vector representation of the mask features corresponding to the target part. The mask region is the region corresponding to the target part in the mask image data. Based on a diffusion model, denoising at least one noise vector representation is performed on the audio sequence data, the vector representation of the visual features of the reference image data, and the vector representation of the mask features corresponding to the target part, to generate at least one frame of digital human image data. The noise vector representation is obtained by adding noise to the vector representation of the visual features of the reference image data. The process of denoising at least one noise vector representation based on the diffusion model to generate at least one frame of digital human image data includes: The vector representation of the mask feature corresponding to the target part is added to at least one noise vector representation to obtain a vector representation of at least one initial digital human image data. Based on the vector representation of the visual features of the audio sequence data and the reference image data according to the diffusion model, the vector representation of the at least one initial digital human image data is denoised to obtain the vector representation of the image data of at least one digital human. The diffusion model is a denoising U-shaped network model. Decode the vector representation of the image data of the at least one digital human to generate at least one frame of digital human image data; A digital human animation with voice is generated based on the image data of at least one frame of the digital human and the audio sequence data.
2. The method of claim 1, wherein, The step of denoising the vector representation of the at least one initial digital human image data based on the vector representation of the visual features of the audio sequence data and the reference image data using the diffusion model to obtain the vector representation of the image data of at least one digital human includes: The first neural network model is used to extract features from the audio sequence data to obtain a vector representation of the initial audio features; The vector representation of the initial audio features is processed based on a second neural network model consisting of two multilayer perceptrons to obtain a vector representation of the audio features aligned with the visual features of the reference image data. Based on the third neural network model, the vector representation of the audio features aligned with the visual features of the reference image data and the vector representation of the visual features of the reference image data are used to denoise the vector representation of the at least one initial digital human image data to obtain the vector representation of the image data of at least one digital human. The third neural network model is a diffusion model that uses an attention mechanism.
3. The method of claim 2, wherein, The method based on the third neural network model, using the vector representation of the audio features aligned with the visual features of the reference image data and the vector representation of the visual features of the reference image data, performs denoising processing on the vector representation of the at least one initial digital human image data to obtain the vector representation of the image data of at least one digital human, including: The reference image data is processed by a fourth neural network model to obtain a vector representation of the semantic features of the reference image data. The fourth neural network model is a self-supervised training model that uses an attention mechanism. Based on the third neural network model, an attention mechanism is used to fuse the vector representations of the audio features aligned with the visual features of the reference image data, the vector representations of the semantic features of the reference image data, and the vector representations of the visual features of the reference image data to obtain the fused features. Based on the third neural network model, and using the fused features as guiding information, the vector representation of the at least one initial digital human image data is iteratively subjected to noise addition and denoising processing to obtain the vector representation of the image data of at least one digital human.
4. The method of claim 3, wherein, The vector representation of the visual features of the acquired reference image data includes: The reference image data is encoded to obtain a vector representation of the reference image data; The vector representation of the reference image data is input into a fifth neural network model employing an attention mechanism to obtain a vector representation of the visual features of the reference image data.
5. The method of claim 4, wherein, The step of inputting the vector representation of the reference image data into the fifth neural network model employing an attention mechanism to obtain the vector representation of the visual features of the reference image data includes: The vector representation of the reference image data and the vector representation of the semantic features of the reference image data are input into the fifth neural network model employing the attention mechanism, so that the fifth neural network model generates a vector representation of the visual features of the reference image data by fusing the semantic features of the reference image data.
6. The method according to any one of claims 1-3, characterized in that, The multimodal large model includes an audio language model, a large language model, and a visual language model. The process of performing intent recognition and sentiment analysis on the multimodal data based on the multimodal large model to determine the response text data includes: If the multimodal data includes audio data, the audio data is processed by the audio language model to obtain text data converted from the audio data and speech emotion information of the audio data; intent recognition and emotion analysis are performed on the data other than the audio data in the multimodal data, the text data converted from the audio data, and the speech emotion information of the audio data to determine the response text data; If the multimodal data only includes text data, the large language model is used to perform intent recognition and sentiment analysis on the text data to obtain response text data; If the multimodal data includes image data but does not include audio data, the visual language model is used to perform intent recognition and sentiment analysis on the multimodal data to obtain response text data.
7. The method of claim 6, wherein, The step of performing intent recognition and sentiment analysis on the multimodal data excluding the audio data, the text data converted from the audio data, and the speech emotion information of the audio data to determine the response text data includes: If the multimodal data, excluding the audio data, only includes text data, the large language model is used to perform intent recognition and sentiment analysis on the text data, the text data converted from the audio data, and the speech emotion information of the audio data to obtain response text data. If the multimodal data includes image data other than the audio data, the visual language model is used to perform intent recognition and sentiment analysis on the multimodal data other than the audio data, the text data converted from the audio data, and the speech emotion information of the audio data to obtain response text data.
8. The method of claim 7, wherein, The response text data is obtained through a large language model, including: After performing intent recognition and sentiment analysis on the text data using a large language model, or after performing intent recognition and sentiment analysis on the text data, the text data converted from the audio data, and the speech sentiment information of the audio data using the large language model, if the large language model cannot generate response text data, a corresponding intelligent agent is determined based on the target text data. The intelligent agent is a third-party application platform. The target text data includes the text data, or the text data converted from the text data and the audio data. The agent is queried based on the target text data to obtain the reference data output by the agent; The target text data and the reference data are processed using the large language model to obtain response text data; Alternatively, response text data can be obtained through a visual language model, including: After performing intent recognition and sentiment analysis on the multimodal data using a visual language model, or after performing intent recognition and sentiment analysis on the data other than the audio data, the text data converted from the audio data, and the speech sentiment information of the audio data using the visual language model, if the visual language model cannot generate response text data, the intent category or professional field is determined based on the target multimodal data according to the visual language model. The target multimodal data includes the multimodal data, or the data other than the audio data and the text data converted from the audio data in the multimodal data. The corresponding intelligent agent is determined based on the intent category or professional field, and the intelligent agent is a third-party application platform; The agent is queried based on the target multimodal data to obtain the reference data output by the agent; The target multimodal data and the reference data are processed by the visual language model to obtain response text data.
9. A digital human generation device based on a multi-modal large model, characterized by, include: The input data acquisition module is used to acquire multimodal data input by the user; The text response module is used to perform intent recognition and sentiment analysis on the multimodal data based on a multimodal big model to determine the response text data; An audio sequence data generation module is used to determine the corresponding audio sequence data based on the response text data; The visual feature acquisition module is used to acquire vector representations of the visual features of the reference image data, the content of which includes people and background; The mask feature determination module is used to determine the corresponding mask image data based on the position of the target part of the person in the reference image data; and to analyze the features of the mask region in the mask image data based on the region guidance model to obtain the vector representation of the mask feature corresponding to the target part, wherein the mask region is the region corresponding to the target part in the mask image data. A digital human image generation module is used to perform denoising processing on at least one noise vector representation based on a diffusion model, according to the vector representation of the visual features of the audio sequence data, the vector representation of the visual features of the reference image data, and the vector representation of the mask features corresponding to the target part, to generate at least one frame of digital human image data. The noise vector representation is obtained by adding noise to the vector representation of the visual features of the reference image data. The step of performing denoising processing on at least one noise vector representation based on the diffusion model, according to the vector representation of the visual features of the reference image data, and the vector representation of the mask features corresponding to the target part, to generate at least one frame of digital human image data includes: The vector representation of the mask feature corresponding to the target part is added to at least one noise vector representation to obtain a vector representation of at least one initial digital human image data. Based on the vector representation of the visual features of the audio sequence data and the reference image data according to the diffusion model, the vector representation of the at least one initial digital human image data is denoised to obtain the vector representation of the image data of at least one digital human. The diffusion model is a denoising U-shaped network model. Decode the vector representation of the image data of the at least one digital human to generate at least one frame of digital human image data; A digital human animation generation module is used to generate a digital human animation with voice based on the image data of the at least one frame of the digital human and the audio sequence data.
Citation Information
Patent Citations
Virtual human generation method and device based on audio control, equipment and medium
CN117152285A
Image redrawing method and device, electronic equipment and readable storage medium
CN119379561A
Video generation method and device based on multi-modal information fusion, equipment and medium
CN119906872A