Virtual human design and application platform and method based on artificial intelligence, equipment and medium
By performing anonymization processing and local analysis of multimodal input data on user equipment, combining cross-modal learning and knowledge graphs, dynamically adjusting the personality parameters of virtual people, the privacy protection and personalized modeling of virtual digital people system is solved, the naturalness and emotional resonance of interaction are improved, and the long-term interaction willingness of users is enhanced.
Patent Information
- Application Number
- CN202510221284.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-07-18
AI Technical Summary
The existing virtual digital human system has technical bottlenecks in privacy protection, emotional understanding and personalized modeling, and it is difficult to take into account data security, interaction accuracy and personalized adaptability, resulting in high risk of user privacy leakage, low interaction fluency and a decrease in long-term interaction willingness.
By anonymizing the multimodal input data on the user equipment, using cross-modal comparison learning and emotional computing models to perform data analysis locally, combining knowledge graphs and long-term memory networks for context perception, dynamically adjusting the personality parameters of virtual people, and driving the virtual digital people model for real-time rendering through the local rendering engine.
It realizes that while ensuring user privacy and security, it improves the naturalness and emotional resonance of virtual digital human interaction, reduces communication load, enhances the real-time and personalized adaptability of interaction, and improves the long-term retention rate of users.
Smart Images

Figure CN120339470A_ABST
Abstract
Description
Background Art
[0002] With the wide application of virtual digital human technology in fields such as social entertainment, online education, and telemedicine, related technologies have been continuously developed to improve the interaction experience and intelligence level.
[0003] Currently, related virtual human systems usually rely on cloud-based centralized processing architectures to process text, voice, image, or video data input by users. Although this method has strong computing power, it also brings significant privacy risks. During the data collection and processing process, users' biometric data (such as voice timbre, facial expressions, and behavior patterns) need to be directly uploaded to the cloud, increasing the risk of data leakage and abuse. In addition, due to the transmission of raw data over the network, its high data volume characteristics lead to a significant increase in communication latency, affecting the interaction fluency between virtual humans and users, and it is difficult to meet application scenarios with high real-time requirements.
[0004] In terms of emotion understanding, related technologies mainly perform emotion recognition based on unimodal data, such as only analyzing the emotional tendency of text or making emotion judgments based on facial expression classification. However, unimodal emotion analysis is easily affected by the quality of input data, such as voice background noise or text semantic ambiguity, resulting in a relatively high misjudgment rate of the system. According to relevant benchmark test data, the misjudgment rate of traditional unimodal emotion recognition methods can reach up to 32% in different scenarios. In addition, some studies have tried to improve intention recognition using attention mechanisms, but they have ignored the impact of emotional factors on the user interaction experience, making the system's feedback lack pertinence and affecting the user's emotional resonance and immersion.
[0005] In terms of personalized modeling, related virtual human systems generally adopt static personality parameter settings, that is, fixed personality parameters are preset during system initialization and remain unchanged during the interaction process. This method cannot dynamically adjust the personality characteristics of virtual humans according to the user's conversation context, resulting in a rigid interaction method and lack of adaptability. Related research shows that in the absence of personalized adjustment, users' long-term interaction willingness with virtual digital humans drops significantly. Although some studies have proposed to achieve personalized modeling through federated learning, they still have not effectively solved the problem of multimodal semantic alignment, resulting in the loss of key information during the information fusion process of different modal data, thus affecting the behavior generation quality of virtual digital humans.
[0006] In summary, related virtual digital human systems still have technical bottlenecks in privacy protection, emotion understanding, and personalized modeling, and it is difficult to balance data security, interaction accuracy, and personalized adaptability. Therefore, there is an urgent need to build a virtual digital human system architecture with privacy protection capabilities, multimodal emotion fusion analysis, and dynamic personality evolution characteristics to improve the naturalness of interaction, emotional resonance, and long-term user retention rate.
[0007] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0008] The purpose of the embodiments of the present disclosure is to provide a method for designing and applying virtual humans based on artificial intelligence, a platform for designing and applying virtual humans based on artificial intelligence, an electronic device, and a computer-readable storage medium, so as to improve the naturalness of virtual digital human interaction, emotional resonance, and long-term user retention rate, and protect the security of user privacy data.
[0009] According to the first aspect of the embodiments of the present disclosure, a method for designing and applying virtual humans based on artificial intelligence is provided, including:
[0010] Collect multimodal input data through an input interface on a user device, where the multimodal input data includes at least one of text data, voice data, image data, and video data;
[0011] Perform local anonymization operations on the multimodal input data on the user device to generate anonymized multimodal features;
[0012] Input the anonymized multimodal features into a multimodal encoder deployed in the cloud. The multimodal encoder maps the anonymized multimodal features to a unified semantic space through cross-modal contrast learning to obtain multimodal feature vectors;
[0013] Input the multimodal feature vectors into a pre-trained emotion computing model. The emotion computing model includes a text emotion classification sub-model, a voice emotion analysis sub-model, and a facial expression recognition sub-model, and outputs a user emotion intensity quantization value;
[0014] Input the multimodal feature vectors into a pre-trained context awareness model. The context awareness model extracts temporal features in multi-round guided conversations based on long short-term memory networks and combines knowledge graphs to generate context-corrected user intention vectors;
[0015] Adjust a preset personality parameter matrix according to the user emotion intensity quantization value and the user intention vector to generate an updated personality parameter matrix;
[0016] Input the updated personality parameter matrix into a pre-trained multimodal generation model. The multimodal generation model includes a voice synthesis sub-model, a facial action generation sub-model, and a limb action generation sub-model, and outputs voice waveform data, facial muscle movement parameters, and skeletal joint coordinate data;
[0017] On the user device, the voice waveform data, the facial muscle movement parameters, and the skeletal joint coordinate data are loaded through a local rendering engine to drive real-time rendering of the virtual digital human three-dimensional model.
[0018] According to a second aspect of the embodiments of the present disclosure, there is provided an artificial intelligence-based virtual human design and application platform, including:
[0019] A multimodal input interface module for collecting multimodal input data through an input interface on a user device, where the multimodal input data includes at least one of text data, voice data, image data, and video data;
[0020] A local anonymization processing module for performing local anonymization operations on the multimodal input data on the user device to generate anonymized multimodal features;
[0021] A cloud multimodal encoding module for inputting the anonymized multimodal features into a multimodal encoder deployed in the cloud. The multimodal encoder maps the output of the anonymized multimodal features into a unified semantic space through cross-modal contrast learning to obtain multimodal feature vectors;
[0022] An emotion computing module for inputting the multimodal feature vectors into a pre-trained emotion computing model. The emotion computing model includes a text emotion classification sub-model, a voice emotion analysis sub-model, and a facial expression recognition sub-model, and outputs a quantified value of the user's emotion intensity;
[0023] A context awareness module for inputting the multimodal feature vectors into a pre-trained context awareness model. The context awareness model extracts temporal features in multi-turn guided conversations based on a long short-term memory network and generates a user intention vector corrected by context in combination with a knowledge graph;
[0024] A personality parameter update module for adjusting a preset personality parameter matrix according to the quantified value of the user's emotion intensity and the user intention vector to generate an updated personality parameter matrix;
[0025] A multimodal generation module for inputting the updated personality parameter matrix into a pre-trained multimodal generation model. The multimodal generation model includes a voice synthesis sub-model, a facial action generation sub-model, and a limb action generation sub-model, and outputs voice waveform data, facial muscle movement parameters, and skeletal joint coordinate data;
[0026] A local rendering module for loading the voice waveform data, the facial muscle movement parameters, and the skeletal joint coordinate data through a local rendering engine on the user device to drive real-time rendering of the virtual digital human three-dimensional model.
[0027] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including: a processor; and a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method for designing and applying a virtual human based on artificial intelligence described in any one of the above is implemented.
[0028] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the method for designing and applying a virtual human based on artificial intelligence described in any one of the above is implemented.
[0029] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0030] In the method for designing and applying a virtual human based on artificial intelligence in the exemplary embodiments of the present disclosure, by performing anonymization processing on multi-modal input data on a user device, the text, voice, image, and video data of the user have been de-identified before being transmitted to the cloud, thus avoiding the privacy and security risks brought by the cloud centralized processing method; at the same time, the anonymized multi-modal features ensure their usability for subsequent semantic analysis through specific data transformation strategies, enabling data processing to avoid the leakage of user identity information without affecting the accurate understanding of the user's emotional state and interaction intention by the virtual human system. In addition, local anonymization reduces the direct transmission of raw data, which not only reduces the risk of data leakage but also significantly reduces the communication load and improves the real-time performance of data transmission.
[0031] The encoding of multi-modal data adopts cross-modal contrast learning, enabling text, voice, and visual information to be represented and associated in a unified semantic space, and solving the problem of modal misalignment in traditional multi-modal fusion methods. In this way, the virtual human system can more comprehensively analyze the user's emotional state and avoid emotional recognition biases caused by the distortion or loss of single-modal information. Since emotion computing involves multiple independent sub-models, cross-validation is performed by combining different modal features, enabling the system to be more robust in identifying the user's emotional intensity, thereby reducing the emotion misjudgment rate and improving the interaction accuracy between the virtual human and the user.
[0032] In terms of interactive intent inference, a long short-term memory network is adopted to extract the temporal features of multi-turn guided conversations, and a knowledge graph is combined for context correction, enabling the user's intent expression to be dynamically optimized based on the continuity of the conversation. This method can effectively avoid the problem of misjudging intents caused by insufficient context understanding and accurately identify the true needs of users in complex conversation environments. By introducing a knowledge graph for intent correction, the system can not only identify the intent expressed by the user currently, but also reason based on historical interaction patterns to ensure that the responses of the virtual human conform to the long-term interaction logic of the user, improving the coherence and naturalness of the user experience.
[0033] The adjustment of personality parameters is based on the quantified value of the user's emotional intensity and the user intent vector after context correction, making the personality parameters of the virtual human no longer fixed, but able to be dynamically adapted according to the actual interaction situation. Compared with the traditional static personality setting method, this dynamic adjustment method enables the behavior pattern of the virtual human to better fit the personalized needs of users and improves the adaptability of the virtual human in different interaction scenarios. The system can gradually adjust the behavior style of the virtual human during long-term interactions, making its performance more in line with user expectations, thus effectively enhancing the user's interaction experience and long-term usage willingness.
[0034] In terms of virtual human behavior generation, speech, facial expressions, and body movements are all driven based on a multi-modal generation model, enabling the expression mode of the virtual human to maintain coordination at multiple levels such as speech and vision. Through the collaborative effect of different sub-models, the facial expressions and speech rhythms of the virtual human can form a natural match, avoiding problems such as lagging facial expressions or mismatches between speech and lip movements that exist in traditional animation-driven methods. In addition, the generation of skeletal joints is based on a dynamic calculation method, making the movements of the virtual human smoother and avoiding the problem of rigid movements caused by the preset animation method.
[0035] In terms of real-time rendering, a local rendering engine is adopted to perform the dynamic loading and driving of the virtual human's three-dimensional model, making the rendering process no longer completely dependent on cloud computing, thereby reducing the impact of network fluctuations on the interaction fluency of the virtual human. Combining the environmental light perception and background segmentation strategies, the system can adjust the display effect of the virtual human according to the actual environment where the user is located, enabling the virtual human to blend more naturally with the real background and enhancing the interaction immersion of the virtual human in the real world. At the same time, the local rendering method reduces the cloud computing burden, enabling the system to run in a wider range of hardware environments and improving the applicability and scalability of the virtual human system.
[0036] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings
[0037] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments in line with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0038] Figure 1 A schematic diagram of the system architecture of an exemplary application environment of an AI-based virtual human design and application method and platform to which the embodiments of the present disclosure can be applied is shown.
[0039] Figure 2 A schematic flowchart of an AI-based virtual human design and application method according to some embodiments of the present disclosure is schematically shown.
[0040] Figure 3 A schematic flowchart of the modal encoder obtaining multi-modal feature vectors through cross-modal contrast learning according to some embodiments of the present disclosure is schematically shown.
[0041] Figure 4 A schematic flowchart of generating a user intention vector with context correction according to some embodiments of the present disclosure is schematically shown.
[0042] Figure 5 A schematic flowchart of generating an updated personality parameter matrix according to some embodiments of the present disclosure is schematically shown.
[0043] Figure 6 A schematic flowchart of rendering a virtual digital human with an augmented reality effect according to some embodiments of the present disclosure is schematically shown.
[0044] Figure 7 A schematic diagram of the structure of an AI-based virtual human design and application platform according to some embodiments of the present disclosure is schematically shown.
[0045] Figure 8 A schematic diagram of the structure of the computer system of an electronic device according to some embodiments of the present disclosure is schematically shown.
[0046] Figure 9 A schematic diagram of a computer-readable storage medium according to some embodiments of the present disclosure is schematically shown.
[0047] In the drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Embodiments
[0048] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.
[0049] In addition, the accompanying drawings are only schematic diagrams and are not necessarily drawn to scale. The block diagrams shown in the accompanying drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0050] Figure 1 A schematic diagram of a system architecture of an exemplary application environment to which the embodiments of the present disclosure for an AI-based virtual human design and application method and platform can be applied is shown.
[0051] As Figure 1 shown, the system architecture 100 may include one or more of the terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The terminal devices 101, 102, 103 may be various electronic devices with artificial intelligence (AI) computing capabilities, including but not limited to desktop computers, portable computers, smartphones, and tablet computers, etc. It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0052] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. For example, the server 105 may be a server cluster composed of multiple servers, etc.
[0053] In this exemplary embodiment, first, a method for designing and applying an AI-based virtual human is provided. This AI-based virtual human design and application method can be applied to a terminal device or a server device. This embodiment is not limited thereto, and hereinafter, the case where the terminal device executes this method will be taken as an example for illustration. Figure 2 A flowchart of a method for designing and applying an AI-based virtual human according to some embodiments of the present disclosure is schematically shown. Refer to Figure 2 As shown, this AI-based virtual human design and application method may include the following steps:
[0054] Step S210, collecting multimodal input data through an input interface on a user device, where the multimodal input data includes at least one of text data, voice data, image data, and video data;
[0055] Step S220, performing a local anonymization operation on the multimodal input data on the user device to generate anonymized multimodal features;
[0056] Step S230, inputting the anonymized multimodal features into a multimodal encoder deployed in the cloud. The multimodal encoder maps the output of the anonymized multimodal features into a unified semantic space through cross-modal contrast learning to obtain multimodal feature vectors;
[0057] Step S240, inputting the multimodal feature vectors into a pre-trained emotion computing model. The emotion computing model includes a text emotion classification sub-model, a voice emotion analysis sub-model, and a facial expression recognition sub-model, and outputs a quantified value of the user's emotion intensity;
[0058] Step S250, inputting the multimodal feature vectors into a pre-trained context-aware model. The context-aware model extracts temporal features in multi-turn guiding conversations based on a long short-term memory network and combines a knowledge graph to generate a context-corrected user intention vector;
[0059] Step S260, adjusting a preset personality parameter matrix according to the quantified value of the user's emotion intensity and the user intention vector to generate an updated personality parameter matrix;
[0060] Step S270, inputting the updated personality parameter matrix into a pre-trained multimodal generation model. The multimodal generation model includes a voice synthesis sub-model, a facial motion generation sub-model, and a limb motion generation sub-model, and outputs voice waveform data, facial muscle movement parameters, and skeletal joint coordinate data;
[0061] Step S280: On the user device, load the voice waveform data, the facial muscle movement parameters, and the skeletal joint coordinate data through the local rendering engine, and drive the real-time rendering of the virtual digital human three-dimensional model.
[0062] According to the virtual human design and application method based on artificial intelligence in this exemplary embodiment, by performing local anonymization processing on multi-modal input data on the user device, identifiable information is removed before the data is transmitted to the cloud, thereby reducing the risk of privacy leakage, reducing the data transmission volume, and improving the real-time performance of the interaction; cross-modal contrast learning is used to uniformly encode multi-modal data, aligning text, voice, and visual information in the semantic space, thereby improving the accuracy of emotion recognition and reducing emotion misjudgment caused by single-modal analysis; the long short-term memory network combined with the knowledge graph is used for context correction, enabling the system to extract user intentions based on multi-round conversations, thereby improving the virtual human's understanding ability of complex interactions and avoiding incorrect responses caused by context absence; the personality parameters are dynamically adjusted based on the user's emotion intensity quantization value and the intention vector of context correction, making the virtual human's behavior pattern adjustable with the interaction, thereby enhancing the personalized adaptation ability and improving the long-term interaction experience of the user; a multi-modal generation model is used to drive the voice, facial expressions, and body movements, making the voice rhythm, facial expressions, and body movements coordinated, thereby improving the naturalness of the virtual human's expression and avoiding the stiffness problem caused by traditional animation methods; the dynamic loading of the virtual human three-dimensional model is executed through the local rendering engine, combined with environmental light perception and background segmentation, to integrate the virtual human with the real background, thereby enhancing the immersion of the virtual human and improving the applicability of the system.
[0063] Next, the virtual human design and application method based on artificial intelligence in this exemplary embodiment will be further described.
[0064] In step S210, multi-modal input data is collected through the input interface on the user device, and the multi-modal input data includes at least one of text data, voice data, image data, and video data.
[0065] In an exemplary embodiment of the present disclosure, the multimodal input data refers to information input from different sensing channels to enhance the system's ability to understand the user's state. For example, the text data can be a character sequence input by the user, including but not limited to keyboard input, screen touch input, or text after natural language transcription; the voice data can be an audio signal input by the user through a microphone, containing features such as voice waveforms, pitch, and timbre; the image data can be collected by the camera of the user device, mainly used to extract facial expression and posture information; the video data can contain continuous image frames and audio streams, capable of providing more complete temporal information for behavior analysis. During the data collection process, hardware devices such as high-sampling-rate microphones and high-definition cameras can be used to ensure the integrity and accuracy of the input data. In addition, to reduce noise interference, audio and video processing technologies such as signal denoising and automatic gain control (AGC) can be combined during collection to optimize the input quality. The time synchronization of different modal data can be achieved through timestamp alignment methods to ensure the consistency of multimodal data in subsequent processing stages.
[0066] In step S220, on the user device, perform a local anonymization operation on the multimodal input data to generate anonymized multimodal features.
[0067] In an exemplary embodiment of the present disclosure, local anonymization refers to de-identifying the original data on the user device to avoid direct exposure of sensitive information to the cloud server and reduce the risk of privacy leakage.
[0068] For the anonymization of text data, a lightweight text feature extraction model can be used. This model can remove information directly pointing to the user's identity, such as names, addresses, etc., while retaining the context semantics; specifically, a pre-trained named entity recognition (NER) model can be used to replace sensitive entities with generalized labels. For example, "Zhang San" is replaced with "someone".
[0069] The anonymization of voice data involves de-identifying the timbre features. A lightweight voice conversion model can be used to adjust the timbre parameters so that it cannot be directly traced back to the user's identity. For example, methods such as spectral envelope transformation or pseudo-timbre synthesis can be adopted to retain the voice content but change the personalized features.
[0070] Anonymization of image data and video data can be achieved through facial feature blurring or key point mapping. For example, techniques such as Gaussian blur and Edge-Preserving Denoising (EPD) can be used to mask facial features, or the coordinates of facial key points can be randomly perturbed based on a feature transformation model to ensure that the anonymized data can still be used for emotion and intention recognition. To ensure the usability of the anonymized data, an adversarial training method can be introduced during the processing to enable the anonymization model to retain the necessary feature expression ability while protecting privacy. In addition, in some scenarios, a reversible perturbation method can be combined to make the anonymization process have a certain degree of reversibility, so as to restore the original information in a local secure environment.
[0071] In step S230, the anonymized multimodal features are input into a multimodal encoder deployed in the cloud. The multimodal encoder maps the output of the anonymized multimodal features into a unified semantic space through cross-modal contrastive learning to obtain multimodal feature vectors.
[0072] In an exemplary embodiment of the present disclosure, the core function of the multimodal encoder is to align data of different modalities and map them into the same semantic expression space for subsequent intelligent analysis. Cross-modal contrastive learning is an unsupervised or semi-supervised learning method that can construct a similarity metric between data of different modalities, so that inputs with the same semantics have similar feature expressions in different modalities. Specifically, a multimodal contrastive learning model based on the Transformer architecture can be adopted, and the encoder parameters can be optimized using a multimodal data contrast loss during training. For example, text data can be encoded using Bidirectional Encoder Representations from Transformers (BERT) or its lightweight variant, speech data can be converted into vector representations through a pre-trained speech feature extractor (such as Wav2Vec), and image data can be feature-extracted through a Residual Network (ResNet) or a Vision Transformer (ViT).
[0073] The alignment of different modality data can be achieved by the method of sharing a latent space, that is, converting the features of each modality into the same dimension through a projection network, and enhancing the alignment ability between modalities through positive and negative sample contrast training. In addition, to further optimize the robustness of the encoder, a self-attention mechanism can be introduced to enable the model to focus on the key feature regions in different modality data, thereby improving the quality and generalization ability of encoding. In this way, the multi-modal encoder can effectively fuse text, speech, and visual information, enabling the multi-modal feature vector to express the user's true intention and emotional state.
[0074] In step S240, the multi-modal feature vector is input into a pre-trained emotion calculation model, which includes a text emotion classification sub-model, a speech emotion analysis sub-model, and a facial expression recognition sub-model, and a user emotion intensity quantization value is output.
[0075] In an exemplary embodiment of the present disclosure, the role of the emotion computing model is to identify the user's current emotional state and output the emotional intensity in a quantitative manner. The text emotion classification sub-model usually adopts natural language processing (NLP) methods based on deep learning. Specifically, models such as bidirectional long short-term memory network (BiLSTM), BERT, or robustly optimized BERT (RoBERTa) can be used to classify text data, identify emotion polarities (such as positive, neutral, negative), and more fine-grained emotion categories (such as anger, joy, sadness, etc.). The speech emotion analysis sub-model is based on speech feature extraction and classification methods. Usually, acoustic feature extraction methods such as Mel-frequency cepstral coefficients (MFCCs) and chroma features are first used for data preprocessing, and then a convolutional neural network (CNN), long short-term memory network (LSTM), or pre-trained speech model is used to classify the audio signal for emotion. The facial expression recognition sub-model is mainly based on computer vision technology. Facial key points are extracted through CNN, and the expression category is recognized in combination with the attention mechanism. Multimodal emotion fusion usually adopts weighted fusion or multimodal attention mechanism to integrate text, speech, and facial expression information and generate a quantified value of the user's emotional intensity. This quantified value can represent the user's emotional state in the form of a vector or a scalar. For example, an interval scale (such as 0-1) or a multi-dimensional vector (such as a combination of emotional dimensions such as anger, sadness, excitement, etc.) is used for expression. In this way, the system can accurately capture the user's emotional changes and provide data support for the behavior adjustment of the virtual human.
[0076] In step S250, the multimodal feature vector is input into a pre-trained context-aware model. The context-aware model extracts temporal features in multiple rounds of guided conversations based on the long short-term memory network and generates a context-corrected user intention vector in combination with the knowledge graph.
[0077] In an exemplary embodiment of the present disclosure, the main function of the context-aware model is to understand the long-term interaction intention of the user and perform dynamic intention correction during the multi-turn conversation process. As a variant of the Recurrent Neural Network (RNN), the Long Short-Term Memory network (LSTM) can effectively capture the long-term dependencies in sequential data and is suitable for modeling the temporal information in multi-turn conversations. In a specific implementation, the LSTM unit gradually updates the historical multi-modal feature vectors and filters out key information through a gating mechanism (including an input gate, a forget gate, and an output gate) to retain the content most relevant to the current user intention in the conversation history. At the same time, a bidirectional LSTM or a variant model based on the Attention Mechanism can be introduced to further improve the understanding ability of temporal features.
[0078] The introduction of the knowledge graph can improve the accuracy of user intention recognition. The knowledge graph is a technology for storing and representing structured knowledge in the form of Entity-Relation-Attribute, which can provide semantic associations for user inputs. For example, in the context where the user provides fuzzy or omitted information, the knowledge graph can be used to infer and complete the missing information to enhance the accuracy of intention recognition. In a specific implementation, a method based on the Graph Neural Network (GNN) can be adopted to match the user input with the entity nodes in the knowledge graph and deduce the true intention of the user based on path search algorithms (such as random walk, attention-weighted neighbor aggregation). To improve the computational efficiency, the system can use a pre-trained Knowledge Embedding model. For example, through TransE, RotatE, or ComplEx, the relationship information in the knowledge graph can be converted into a computable vector representation to achieve efficient intention reasoning. By combining with the context-aware model, the system can not only recognize the user intention in the current turn but also adjust according to the historical interaction pattern to ensure that the response of the virtual human is consistent with the long-term behavior logic of the user, thereby enhancing the naturalness and continuity of the interaction.
[0079] In step S260, according to the user emotion intensity quantization value and the user intention vector, the preset personality parameter matrix is adjusted to generate an updated personality parameter matrix.
[0080] In an exemplary embodiment of the present disclosure, the personality parameter matrix can be used to control the personalized behavior of the virtual digital human, enabling it to exhibit personality characteristics that adapt to the user's expectations in different interaction scenarios. The adjustment of the personality parameters is based on the user's emotional and intention states, and the expression and behavior patterns of the virtual human are updated dynamically. In a specific implementation, the quantization value of the user's emotional intensity can be normalized to maintain consistency between different emotional scales (such as the activation-valence model, five-dimensional emotion model). Furthermore, methods such as Min-Max normalization or Z-score standardization can be used to map the emotional intensity value to a standard range.
[0081] It can be based on the weighted fusion of the user intention vector, which helps to take into account the long-term and short-term preferences of the user when adjusting the personality parameters. The weighted fusion can adopt methods such as Bayesian update or Weighted Moving Average (WMA) to smooth the newly input emotional and intention information, preventing short-term abnormal emotions from causing drastic fluctuations in the behavior of the virtual human. In addition, to ensure that the update of the personality parameter matrix conforms to the long-term interaction logic, an optimization strategy based on Reinforcement Learning (RL) can also be adopted, using methods such as Policy Gradient or Deep Q-Learning (DQN) to adjust the personality parameters according to the user's long-term feedback. Finally, based on the obtained adjusted values of the personality parameters, the preset personality parameter matrix is updated element-wise with weights to ensure that the personality of the virtual human can gradually adapt to the user's needs during the long-term interaction process, thereby enhancing the user experience.
[0082] In step S270, the updated personality parameter matrix is input into a pre-trained multi-modal generation model, which includes a speech synthesis sub-model, a facial action generation sub-model, and a limb movement generation sub-model, and speech waveform data, facial muscle movement parameters, and skeletal joint coordinate data are output.
[0083] In an exemplary embodiment of the present disclosure, the multi-modal generation model can generate the virtual human's speech, expressions, and actions that conform to its personalized settings according to the user's interaction needs, enabling the virtual human to exhibit Figure 1Consistent behavioral characteristics. In a specific implementation, the speech synthesis sub-model can generate multi-modal latent feature vectors based on a Conditional Variational Autoencoder (CVAE) and use an Autoregressive Neural Network to generate speech waveform data. Speech synthesis can adopt a neural network-based Text-to-Speech (TTS) model, such as Tacotron 2 or FastSpeech 2, and combine it with a Neural Vocoder such as WaveGlow or HiFi-GAN to improve the speech quality and naturalness. Of course, the methods adopted above are only illustrative examples, and this embodiment does not make special limitations on this.
[0084] The facial action generation sub-model can be used to calculate the virtual human facial muscle movement parameters to drive the virtual human expression changes. This sub-model can perform expression synthesis based on a Graph Convolutional Network (GCN) or a Conditional Generative Adversarial Network (CGAN). The specific implementation method can adopt an expression control method based on Facial Action Units (FAUs), and adjust the movement trajectory of facial muscles by encoding user emotion features. In addition, to ensure the synchronization of facial expressions and speech prosody, a time series modeling module, such as a method based on Variational RNN, can be added during the training process to learn the temporal pattern of expression changes.
[0085] The limb action generation sub-model can be used to calculate the movement coordinates of the virtual human skeletal joints so that its actions can meet the situational requirements. The limb action generation sub-model can adopt a motion calculation method based on Inverse Kinematics (IK) to ensure the coherence of limb actions; specifically, an LSTM or Transformer model can be trained in combination with motion capture data to predict the skeletal movement trajectory, and physical constraints can be added to prevent unreasonable deformation during the action generation process. Through the above methods, the virtual human can complete interactions with more natural speech, expressions, and limb actions, enhancing the user's immersive experience.
[0086] In step S280, on the user device, the speech waveform data, the facial muscle movement parameters, and the skeletal joint coordinate data are loaded through a local rendering engine to drive the real-time rendering of the virtual digital human three-dimensional model.
[0087] In an exemplary embodiment of the present disclosure, the role of the local rendering engine is to reduce the dependence on cloud computing and improve the real-time performance of virtual human generation. The local rendering engine can accelerate calculations based on a Graphics Processing Unit (GPU) and adopt Deferred Rendering or Forward Rendering techniques to optimize rendering efficiency. Facial mesh updates can use animation methods based on Blendshape or Bone-Driven to achieve natural and smooth expression changes. Limb movements can be based on Rigging technology, and the character animation can be driven by calculating the joint rotation angles frame by frame.
[0088] It can be understood that environmental light perception and background segmentation techniques can also be combined to optimize the integration of the virtual human and the real-world scene. Environmental light perception can use light mapping or Image-Based Lighting (IBL) methods to automatically adjust the lighting effects of the virtual human's material. Background segmentation can be based on deep learning models (such as DeepLabV3+ or MODNet) for foreground-background separation, enabling the virtual human to adaptively adjust its appearance in different scenes, thereby enhancing the interaction experience between the virtual human and the real environment.
[0089] The content in steps S210 to S260 will be described in detail below.
[0090] In an exemplary embodiment of the present disclosure, the local anonymization operation on the multi-modal input data in step S220 to generate anonymized multi-modal features can be achieved through the following steps, which can specifically include:
[0091] A lightweight text feature extraction model can be used to perform context semantic recognition on the text data to obtain a text feature vector; a lightweight speech recognition model can be used to extract timbre features from the speech data to obtain a speech timbre feature vector; a lightweight image recognition model can be used to extract facial key point coordinates from the image data and the video data to obtain an image feature vector; Laplace noise can be added to each of the text feature vectors, the speech timbre feature vectors, and the image feature vectors to generate anonymized multi-modal features.
[0092] Among them, the lightweight text feature extraction model can be used for the anonymization of text data. This model can identify and mask sensitive information in the text, such as privacy-related content like names, addresses, phone numbers, etc. Optionally, the Named Entity Recognition (NER) method can be adopted, and a sequence annotation model (such as a model based on Conditional Random Field (CRF) or a deep learning model) can be used to detect and mark privacy entities in the text. After marking, anonymization can be carried out in ways such as fixed replacement, random generation, or generalization processing. For example, a specific name can be replaced with "someone", or the data camouflage method can be used to randomly generate meaningless but structurally consistent text fragments to maintain the readability and context coherence of the text. Of course, semantic embedding methods can also be combined, such as calculating the semantic similarity of replacement candidates based on word vectors (Word2Vec) or sentence vectors (Sentence-BERT), to ensure that the text after anonymization still conforms to the context. Some implementation solutions can also adopt the Local Differential Privacy (LDP) method to add random noise to text features to increase the anti-inference ability of anonymized data and improve the privacy protection effect.
[0093] When processing voice data, the main goal is to remove or obfuscate timbre features that may reveal the user's identity while preserving key information to ensure that subsequent voice processing tasks (such as speech recognition or sentiment analysis) are not affected. Lightweight speech recognition models can be used to extract timbre features of speech, including Mel-Frequency Cepstral Coefficients (MFCCs), Linear Predictive Cepstral Coefficients (LPCCs), and Chroma features, etc., and anonymize the timbre based on these features. Specifically, based on the method of Spectral Envelope Warping, the spectral features of speech can be stretched or shifted to change the recognizability of the timbre. In addition, voice conversion methods can also be adopted, such as timbre conversion models based on Cycle-Consistent Generative Adversarial Network (CycleGAN) or Variational Autoencoder (VAE), to convert the input speech into a neutral timbre without specific identity features. Another method is that, based on statistical modeling, such as using Gaussian Mixture Model (GMM) or Mean Normalization method, the statistical distribution of the speech signal can be standardized so that the voice data of different users present similar acoustic features. Of course, perturbation techniques can also be introduced, such as methods based on Laplacian Noise or Gaussian Noise, to add random noise in the time domain or frequency domain of the speech signal to reduce the recoverability of personalized information and further improve the reliability of privacy protection.
[0094] When anonymizing image data and video data, the method of extracting facial key point coordinates is mainly adopted to identify and mask facial features related to user identity. A lightweight image recognition model can be used to detect the face region in an image or video frame and extract the key point coordinates. For example, the key point coordinates can include the eyes, nose, mouth, and contour points. Common facial key point detection methods can include models based on convolutional neural networks (CNNs), such as MTCNN (Multi-task Cascaded Convolutional Networks), FaceNet, or lightweight detection models based on self-supervised learning. The extracted key point coordinates can be used for further anonymization processing. For example, feature perturbation technology can be adopted to add random offsets to the key point coordinates to disrupt the original facial feature structure. In addition, geometric transformation methods, such as based on affine transformation or morphological transformation, can be used to distort the face structure so that it can still be used for facial expression analysis but cannot be directly associated with a specific user. Another implementation method is based on deep generative models, such as StyleGAN or U-Net, to generate virtual faces similar to the original image structure but without specific identity information to replace the original facial image. For video data, the anonymization process usually needs to be performed frame by frame, while considering temporal consistency. Optical flow can also be used to track the movement trajectory of key points to ensure the coherence of the anonymized video between frames. This embodiment is not limited thereto.
[0095] After anonymizing text, speech, image, and video data, the Laplace noise addition method can be further employed to enhance the anonymization effect and improve the anti-deanonymization ability. Laplace noise refers to the differential privacy mechanism commonly used for privacy protection, whose distribution has a high probability density concentration, capable of effectively introducing small perturbations in the data space, making it difficult to reverse-infer the anonymized data. Specifically, during implementation, the intensity of Laplace noise can be adjusted based on the Privacy Budget to strike a balance between privacy protection and data availability. For example, in text data, noise can be added in the word vector space to make the anonymized text semantically close to the original text but without containing the original sensitive information; in speech data, random perturbations can be introduced in the spectrogram or acoustic features to further reduce recognizability; in image data, noise can be added at the pixel level or in the feature space to make facial features more blurred without affecting subsequent expression analysis; in video data, random noise can be introduced in the keyframe or motion trajectory to reduce identity recognizability while ensuring that the video content can still be used for subsequent computing tasks. Through the above methods, the anonymization of multi-modal data is achieved, enabling the data to remain valid in the subsequent analysis stage while enhancing the user privacy protection ability.
[0096] In an exemplary embodiment of the present disclosure, step S230 can be implemented through the steps Figure 3 in which the multi-modal encoder maps the output of the anonymized multi-modal features into a unified semantic space through cross-modal contrast learning to obtain the content of the multi-modal feature vector. Refer to Figure 3 as shown, which specifically may include:
[0097] Step S310, calculating the feature alignment loss between the text feature vector, the voice timbre feature vector, and the image feature vector through cosine similarity;
[0098] Step S320, optimizing the weight parameters of the multi-modal encoder through the feature alignment loss in combination with gradient descent, and based on the multi-modal encoder with updated weight parameters, mapping the output of the anonymized multi-modal features into a unified semantic space to obtain the multi-modal feature vector.
[0099] Among them, the text feature vector refers to a high-dimensional vector extracted by a natural language processing (NLP) model, for example, a deep neural network based on a bidirectional encoder representation, or a lightweight text embedding method, such as a word vector model or a sentence vector model. This example embodiment does not specifically limit the method for extracting text feature vectors. These methods can capture the contextual relationship of the text and generate a stable semantic expression. The extraction of speech timbre feature vectors depends on the frequency domain and time domain features of the audio signal. Audio feature representation methods such as Mel spectrum and linear prediction cepstral coefficients can be used, and combined with pre-trained models based on self-supervised learning, such as Wav2Vec 2.0 or HuBERT, to improve the expressiveness of speech features. The extraction of image feature vectors can be mainly based on computer vision models. For example, a convolutional neural network or a visual converter can be used to extract features from facial areas to generate low-dimensional and compact image feature expressions.
[0100] When calculating the feature alignment loss, cosine similarity can be used as the alignment metric between feature vectors of different modalities. In the specific implementation process, contrastive learning or triplet loss methods can be used to optimize the feature alignment of different modal data. The core idea of contrastive learning is to bring semantically similar cross-modal feature vectors closer together and separate semantically irrelevant feature vectors to maximize information consistency. For example, the contrastive loss function can be used to optimize the parameters of the feature encoder by constructing positive sample pairs (i.e., information from the same semantics) and negative sample pairs (i.e., information from different semantics). Another feasible method is to use the Maximum Mean Discrepancy (MMD) method to reduce the statistical deviation between different modal distributions so that multimodal features can be represented in the same space.
[0101] Based on the calculation of the feature alignment loss, the weight parameters of the multimodal encoder can be optimized through gradient descent to enhance the cross-modal representation ability of the model. Gradient descent is an optimization algorithm that calculates the gradient of the loss function with respect to the model parameters and updates the parameters in the opposite direction of the gradient to minimize the value of the loss function. In specific implementation, an adaptive learning rate optimization algorithm can be adopted. For example, the Adam (Adaptive Moment Estimation) optimization algorithm or the RMSProp (Root Mean Square Propagation) optimization algorithm can be used to ensure that the model can adaptively adjust the learning rate at different training stages, thereby accelerating the convergence speed and reducing the oscillation of parameter updates. During the optimization process, the Batch Normalization technique can also be combined to reduce the problems of gradient vanishing or gradient explosion and improve the training stability. Of course, a regularization strategy such as L2 regularization (weight decay) or Dropout can also be introduced during training to reduce the risk of overfitting and further improve the generalization ability of the model. The optimization method for the training process in this exemplary embodiment is not specifically limited.
[0102] After completing the optimization of the weight parameters, based on the multimodal encoder with updated weight parameters, the anonymized multimodal features can be mapped into a unified semantic space to obtain high-quality multimodal feature vectors. The construction of the unified semantic space can be based on the principle of Shared Representation Learning, that is, features of different modalities are jointly mapped through a common representation layer, so that inputs with the same semantics can have similar feature expressions. In specific implementation, a cross-modal embedding method can be adopted. For example, a model based on Multimodal Transformer or CycleGAN can be used to learn the information conversion relationship between different modalities. The Self-Supervised Learning method can also be used to enhance the model's ability to align cross-modal features by designing mask prediction or pseudo-task learning.
[0103] After being mapped to a unified semantic space, the multi-modal feature vectors can be further used for subsequent tasks such as sentiment computing, intention recognition, etc. Since the multi-modal feature vectors maintain consistency among different modalities, the system can obtain stable feature expressions under different input types, reducing the misjudgment problems caused by inconsistent information among modalities in traditional multi-modal systems. In addition, it can effectively improve the accuracy of sentiment computing and intention recognition, enabling the virtual digital human to make more reasonable responses when facing complex interaction scenarios. Through the above optimization strategies, it is ensured that the multi-modal encoder can fully integrate text, speech, and image features, thereby enhancing the intelligent level of the virtual digital human system.
[0104] In an exemplary embodiment of the present disclosure, step S250 in which the context-aware model extracts temporal features from multi-round guided conversations based on a long short-term memory network and generates a user intention vector with context correction in combination with a knowledge graph can be implemented through the steps in Figure 4 As shown in Figure 4 and specifically may include:
[0105] Step S410, inputting the multi-modal feature vector into the long short-term memory network of the context-aware model for temporal feature extraction to determine the temporal features in multi-round guided conversations;
[0106] Step S420, inputting the temporal features into a pre-set knowledge graph and generating an initial user intention vector based on the structured information in the knowledge graph;
[0107] Step S430, determining a user intention vector with context correction through the temporal features and the initial user intention vector.
[0108] Among them, the long short-term memory network is an improved recurrent neural network that can capture long-term dependencies through a gating mechanism and is suitable for processing the dynamic changes of user intentions during the conversation process. In a specific implementation, the input of the long short-term memory network consists of feature vectors encoded by multi-modalities, which can include text, speech, and image features. Inside the model, an input gate, a forget gate, and an output gate can be used to control the storage and update of the information flow, enabling the system to effectively retain key information in multi-round conversations and filter out irrelevant features. Optionally, when processing long sequence inputs, an attention mechanism can also be combined to enable the model to assign different weights to the inputs at different time steps, thereby enhancing the attention to key interaction information. For example, a long-range dependence modeling method based on self-attention can be adopted to improve the model's perception ability of remote context information. Another alternative solution is to use a method based on the Gated Recurrent Unit (GRU). The GRU has a lower computational complexity and is suitable for terminal devices with limited computing resources and can replace the LSTM in some scenarios to reduce the computational overhead.
[0109] After extracting the temporal features, the temporal features can be input into a pre-set knowledge graph, and an initial user intention vector can be generated based on the structured information in the knowledge graph. A knowledge graph refers to a structured data representation for storing entities and their relationships, where its nodes can represent concepts or entities, and its edges can represent semantic associations between entities. Through knowledge graph reasoning, additional background knowledge can be provided to the system, enhancing the virtual human's ability to understand user intentions. In specific implementation, a knowledge graph oriented to user interaction can be constructed first, including common user requirements, situational relationships, interaction rules, etc. For example, in a customer service application, the knowledge graph can include user question categories and their corresponding solutions, and in an educational application, the knowledge graph can cover course knowledge points and their sequence relationships.
[0110] After inputting the temporal features into the knowledge graph, a graph neural network can be used for reasoning to generate an initial user intention vector. A graph neural network can aggregate information from different nodes through a recursive propagation mechanism, thereby achieving semantic enhancement of user intentions. In specific implementation, the Graph Attention Network (GAT) method can be adopted to aggregate relevant node information with different weights to ensure the accuracy of intention recognition. In addition, embedding-based methods such as TransE and RotatE can be combined to map entities and relationships in the knowledge graph to a low-dimensional vector space and perform intention reasoning in the vector space to improve computational efficiency. Of course, a Rule-Based Reasoning method can also be used to screen the knowledge fragments most relevant to the current conversation through predefined rules to ensure the interpretability of knowledge reasoning.
[0111] After generating the initial user intention vector, the context-corrected user intention vector is determined through the fusion of the temporal features and the initial user intention vector. The core goal of intention correction is to combine the user's current input and historical interaction information to adjust the expression of the user intention to make it more in line with the real needs. In specific implementation, a weighted fusion method can be adopted to dynamically adjust the intention correction weight by calculating the similarity between the historical intention vector and the current initial intention vector. For example, a weighted method based on Bayesian update can be used to calculate the influence degree of the historical interaction pattern on the current intention. In addition, a dynamic adjustment strategy based on Reinforcement Learning (RL) can be adopted to optimize the intention correction parameters through a reward mechanism, enabling the system to continuously optimize the accuracy of intention recognition during long-term interactions.
[0112] The context-aware model can generate a context-corrected user intention vector through a long short-term memory network, which can be used for subsequent tasks such as personality parameter adjustment, behavior generation, and virtual digital human rendering, ensuring that the interaction behavior of the virtual digital human can accurately reflect the real needs of users and improving the coherence and intelligence level of the user experience.
[0113] In an exemplary embodiment of the present disclosure, step S260 can be implemented by the steps in Figure 5 to adjust the preset personality parameter matrix according to the user emotion intensity quantization value and the user intention vector, and generate an updated personality parameter matrix. As shown in Figure 5 it specifically may include:
[0114] Step S510, perform weight normalization processing on the user emotion intensity quantization value to generate an emotion normalization vector;
[0115] Step S520, perform weighted fusion based on the context-corrected user intention vector to generate an intention enhancement vector;
[0116] Step S530, perform a linear transformation on the emotion normalization vector and the intention enhancement vector to calculate the personality parameter adjustment value;
[0117] Step S540, based on the personality parameter adjustment value, perform element-wise weighted update on the preset personality parameter matrix to obtain an updated personality parameter matrix.
[0118] Among them, the emotion intensity quantization value refers to representing the current emotional state of the user in the form of a scalar or multi-dimensional vector, covering emotion categories (such as anger, joy, sadness, etc.) and the corresponding emotion intensity. In specific implementation, the main goal of the normalization process is to ensure that emotional data from different sources are consistent in the numerical scale for effective fusion in subsequent calculations. The normalization method can adopt Min-Max Normalization, that is, the emotion intensity value can be mapped to the [0, 1] interval, so that the emotion values of different users have the same range. In addition, Z-score Normalization can also be used, adjusting the emotion value by calculating the mean and standard deviation to make the data conform to the standard normal distribution, thereby enhancing the stability of the model. Of course, a dynamic normalization method can also be combined to adaptively adjust the normalization parameters according to the user's historical emotion distribution, so that the long-term emotion change trend can be effectively modeled.
[0119] After completing the normalization of emotional intensity, weighted fusion can be performed based on the user intention vector corrected by context to generate an intention enhancement vector. The user intention vector can be used to represent the user's behavioral goal in the current interaction, and its generation method usually combines multi-modal features, context information, and knowledge graph reasoning results. During the fusion process, the correlation between the intention vector and the emotion normalization vector can be calculated first to determine the weighting coefficients of the two. For example, the dot product attention method can be used to dynamically allocate fusion weights by calculating the dot product similarity of the two vectors, or a weighting method based on Bayesian Inference can be used, with historical interaction data as prior information to dynamically adjust the fusion parameters, making the update of the intention vector more in line with the user's personalized behavior pattern. Of course, based on the Autoregressive Neural Network, time dependence can be introduced during the fusion process, so that the intention enhancement vector can reflect the user's long-term interaction trend. The weighted fusion method in this exemplary embodiment is not particularly limited.
[0120] After generating the intention enhancement vector, the emotion normalization vector and the intention enhancement vector can be linearly transformed to calculate the personality parameter adjustment value. The core goal of the linear transformation is to map feature vectors from different sources to the same parameter space for subsequent personality parameter adjustment. In specific implementation, a Fully Connected Neural Network (FCN) can be used for non-linear transformation to learn the optimal mapping relationship. In addition, matrix transformation methods such as Principal Component Analysis (PCA) can be combined to extract the main components of features in a dimensionality reduction manner, thereby improving the calculation efficiency. Of course, a mapping method based on a Multilayer Perceptron (MLP) can also be used, by introducing non-linear activation functions (such as ReLU, Leaky ReLU), to enhance the flexibility of personality parameter adjustment.
[0121] After calculating the personality parameter adjustment value, the preset personality parameter matrix can be updated element by element based on the adjustment value to generate an updated personality parameter matrix. The personality parameter matrix can be used to control the personality characteristics of the virtual human, including dimensions such as speech style, emotional expression, speech prosody, and facial expression intensity. During the update process, the Exponential Weighted Moving Average (EWMA) method can be used to make the influence of the newly input adjustment value on the personality parameters decrease over time to ensure the stability of personality adjustment; the Reinforcement Learning (RL) method can also be combined to optimize the update weights through Policy Gradient, making the changes in personality parameters more in line with long-term user interaction habits. Optionally, a personality modeling method based on the Graph Neural Network (GNN) can be used to represent the personality parameter matrix as a graph structure and learn the mutual influence between different personality parameters through the Node Embedding method, thereby improving the accuracy of personality modeling.
[0122] Finally, by generating the updated personality parameter matrix through the user emotion intensity quantization value and the user intention vector, it can effectively ensure that the personality parameters of the virtual digital human can be dynamically adjusted according to the user's emotional state and interaction intention, enabling the virtual digital human to exhibit personalized behaviors that meet the user's expectations in different scenarios.
[0123] In an exemplary embodiment of the present disclosure, the updated personality parameter matrix can be input into a pre-trained multi-modal generation model to output speech waveform data, facial muscle movement parameters, and skeletal joint coordinate data through the following steps, which can specifically include:
[0124] Based on the updated personality parameter matrix, a multi-modal latent feature vector can be generated using a conditional variational autoencoder; the multi-modal latent feature vector can be input into the speech synthesis sub-model of the multi-modal generation model to generate speech waveform data through the autoregressive neural network in the speech synthesis sub-model; the multi-modal latent feature vector can be input into the facial action generation sub-model of the multi-modal generation model to determine the facial muscle movement parameters through the graph convolutional neural network in the facial action generation sub-model; the multi-modal latent feature vector can be input into the limb action generation sub-model of the multi-modal generation model to perform inverse kinematics calculations through the limb action generation sub-model to determine the skeletal joint coordinate data.
[0125] Among them, the conditional variational autoencoder is a generative model based on probabilistic inference, which can efficiently encode and generate complex modal features by learning the hidden structure of the data distribution. In specific implementation, an encoder network can be constructed first to accept the updated personality parameter matrix, the user's current emotional state, and the intention vector as inputs, and model the probability distribution of the latent variables through a Gaussian distribution. The encoder network can adopt a multilayer perceptron (MLP) or a convolutional neural network structure to extract the high-dimensional representation of the personality features. Subsequently, the decoder network can decode through the sampled latent variables to generate a multimodal latent feature vector. To enhance the stability of generation, the KL (Kullback-Leibler) divergence constraint can be introduced to make the latent variable distribution consistent with the standard normal distribution, thereby improving the generalization ability of the model.
[0126] After generating the multimodal latent feature vector, this vector can be input into the speech synthesis sub-model of the multimodal generation model to generate speech waveform data. The core goal of speech synthesis is to generate a voice signal that meets the user's expectations based on personalized parameters, ensuring that the speech prosody and timbre features are consistent with the virtual human personality parameters set by the user. In specific implementation, the text can be first converted into a phoneme sequence by a text-to-speech system, and the generation of the phoneme sequence can be based on a sequence-to-sequence model, such as Tacotron 2 or FastSpeech 2. Subsequently, an autoregressive neural network or a non-autoregressive neural network can be used to predict speech features, and combined with a neural vocoder, such as WaveGlow, HiFi-GAN, to convert the predicted acoustic features into high-quality speech waveforms. To enhance the personalized performance of speech synthesis, the timbre conversion technology can also be combined, and through the timbre conversion method based on the Cycle-Consistent Generative Adversarial Network (CycleGAN), it is ensured that the generated speech conforms to the personalized settings of the virtual human. Optionally, to improve the emotional expression ability of the synthesized speech, the Emotion Embedding method can also be adopted to enhance the richness of the speech changes in different emotional states, making the speech performance of the virtual human more emotionally resonant.
[0127] While generating speech data, the multi-modal latent feature vector can also be input into the facial motion generation sub-model to determine the facial muscle movement parameters. The goal of facial motion generation is to drive the expression changes of the virtual human based on the user's emotional state, speech prosody, and personality parameters, so that it is consistent with the speech performance. In specific implementation, the basic expression features can be extracted first through a facial key point detection model, and feature modeling can be performed based on a graph convolutional neural network to learn the facial muscle movement pattern. Subsequently, a method based on a conditional generative adversarial network can be used to generate expression animation data that conforms to the input features. Another optional method is the expression-driven technology based on sparse representation, which represents the expression changes as a weighted combination of different Facial Action Units (FAUs) through sparse coding, thereby improving the controllability of expression generation. In addition, to enhance the real-time performance of facial expressions, a temporal modeling method can be combined, such as the method based on Variational Recurrent Neural Network (VRNN), to learn the temporal dependence relationship of expression changes, thereby ensuring the coherence of facial animation. Finally, the generated facial muscle movement parameters can be used for 3D expression animation driving of the virtual digital human to ensure that the expression changes conform to the natural human expression pattern.
[0128] After generating the facial motion parameters, the multi-modal latent feature vector can be input into the limb motion generation sub-model to determine the bone joint coordinate data. The generation of limb motion aims to generate virtual human limb motions that meet the user's expectations based on the user's interaction intention, emotional state, and personality settings. In specific implementation, the bone motion trajectory can be calculated first based on the Inverse Kinematics (IK) method. The inverse kinematics method solves the rotation angles of the entire joint chain through the constraints of the target joint positions to ensure that the generated limb motions conform to physical rules. To improve the naturalness of motion generation, a data-driven method based on Motion Capture (MoCap) can be combined to learn the real human motion pattern and perform motion prediction through a deep generative model. For example, an action generation model based on a recurrent neural network can be used to learn the temporal relationship of bone motion. In addition, an action style conversion method based on a generative adversarial network can also be used to make the generated limb motions adapt to different personality settings, such as adding personalized gait features, gesture habits, etc. To improve the stability of motion data, a pose filtering method can also be combined, such as a noise reduction method based on Kalman filtering or first-order Gaussian smoothing, to reduce the jitter phenomenon in the generated motions. Finally, the generated bone joint coordinate data can be used for 3D animation driving of the virtual digital human to ensure the natural and smooth limb motions.
[0129] During the multi-modal generation process, the above technical means are used to ensure the consistency of voice, facial expressions, and body movements, so that the expression of the virtual digital human conforms to the personalized settings and can be dynamically adjusted according to the user's emotional state and interaction intention, thereby enhancing the interaction experience and emotional resonance ability of the virtual digital human.
[0130] In an exemplary embodiment of the present disclosure, the following steps can be used to implement step S280. On the user device, load the voice waveform data, facial muscle movement parameters, and skeletal joint coordinate data through the local rendering engine, and drive the three-dimensional model of the virtual digital human for real-time rendering. Specifically, it can include:
[0131] Based on the voice waveform data, the local audio synthesis module of the user device can be used for voice playback; based on the facial muscle movement parameters, the expression driving engine of the user device can be used to determine the facial mesh deformation amount, and the facial mesh model corresponding to the three-dimensional model of the virtual digital human can be updated through the facial mesh deformation amount; based on the skeletal joint coordinate data, the skeletal binding module of the user device can be used to adjust the skeletal structure of the three-dimensional model of the virtual digital human; the facial mesh model and the adjusted skeletal structure can be fused through the graphics rendering pipeline of the user device to render the complete three-dimensional model of the virtual digital human in real time.
[0132] Among them, during the real-time rendering process of the three-dimensional model of the virtual digital human, based on the voice waveform data, the local audio synthesis module of the user device can be used for voice playback. The voice waveform data can be generated by the voice synthesis sub-model of the multi-modal generation model, including time-series waveform information, voice prosody parameters, and timbre characteristics. In order to ensure high-quality and low-latency voice playback, the audio data needs to be decoded and format-converted. Audio decoding can use the Fast Fourier Transform (FFT) or Short-Time Fourier Transform (STFT) to perform time-frequency domain conversion on the voice signal, and combine audio filtering techniques (such as mean filtering, band-pass filtering, etc.) to optimize the sound quality. Format conversion can be based on encoding format conversion algorithms (such as Linear Predictive Coding (LPC) or Pulse Code Modulation (PCM)) to ensure the device compatibility of the voice data.
[0133] While the audio is playing, based on the facial muscle movement parameters, the expression-driven engine of the user device can determine the amount of facial mesh deformation and update the facial mesh model corresponding to the virtual digital human three-dimensional model through the amount of facial mesh deformation. The facial muscle movement parameters can be generated by the facial action generation sub-model of the multi-modal generation model, mainly including expression unit parameters and key point coordinate information. The core task of the expression-driven engine is to map these parameters to the facial mesh structure of the virtual human to achieve real-time driving of facial animation. In specific implementation, based on the vertex blending deformation method, the facial expression parameters can be mapped to a predefined set of expression weights, thereby driving the deformation of the facial mesh. Another implementation method is that, based on the bone binding method, by adjusting the rotation angle of the control bones, the mesh deformation can be indirectly affected, thereby realizing the animation driving of facial expressions.
[0134] After the facial mesh deformation is completed, based on the bone joint coordinate data, the bone binding module of the user device can adjust the bone structure of the virtual digital human three-dimensional model. The bone joint coordinate data can be generated by the limb action generation sub-model of the multi-modal generation model, including the spatial positions and rotation angles of different joints. In specific implementation, the bone binding module first needs to perform bone mapping, that is, map the input joint data to the skeleton model of the virtual digital human. The mapping method can adopt direct coordinate mapping, that is, directly match the input data with the bone hierarchy of the 3D model according to the bone name or index. Of course, based on joint interpolation, by calculating the joint pose interpolation of adjacent frames, the animation transition can be smoothed to reduce the unnatural phenomenon caused by sudden action changes. Optionally, during the bone adjustment process, motion constraint optimization can also be combined to ensure that the generated joint rotations conform to the rules of human kinematics, such as restricting excessive torsion or avoiding unreasonable joint bending, thereby enhancing the physical authenticity of the action.
[0135] After the bone adjustment is completed, the facial mesh model and the adjusted bone structure can be fused through the graphics rendering pipeline of the user device and rendered in real time to ensure the animation fluency and visual realism of the virtual digital human. The core task of the graphics rendering pipeline is to convert the mesh data of the 3D model into the final image output through steps such as geometric transformation, lighting calculation, and pixel shading. In the specific implementation, vertex shading can be performed first, that is, based on the 3D vertex information of the virtual human face and body, the transformation matrix of the model in the world coordinate system is calculated to ensure the correct display of the model in the virtual environment. Subsequently, lighting calculation can be performed, and the Physically Based Rendering (PBR) method can be adopted to simulate the reflection and scattering effects of real lighting on the surface of the virtual digital human. In addition, to improve the computational efficiency of rendering, the deferred rendering method can be adopted to defer the lighting calculation to the fragment shading stage to reduce redundant calculations and increase the frame rate of real-time rendering. Finally, the image data generated by the rendering pipeline is output through the display device, enabling the 3D animation of the virtual digital human to be presented to the user in a high-quality and low-latency manner, thereby enhancing the immersion and realism of the virtual human interaction.
[0136] In an exemplary embodiment of the present disclosure, a virtual digital human with an augmented reality effect can be rendered through the steps in Figure 6 , as shown in reference to Figure 6 and specifically may include:
[0137] Step S610, collecting the current ambient light information through the sensor of the user device;
[0138] Step S620, obtaining the background image through the camera of the user device, and performing foreground segmentation on the background image to generate a virtual digital human layer and a real background layer;
[0139] Step S630, based on the current ambient light information and the virtual digital human layer, using a deep learning-driven augmented reality engine to adjust the appearance parameters of the virtual digital human to obtain an adjusted virtual digital human layer;
[0140] Step S640, synthesizing the adjusted virtual digital human layer and the real background layer to render an augmented reality virtual digital human.
[0141] Among them, during the rendering process of the virtual digital human in augmented reality, the current environmental light information can be collected through the sensors of the user device to ensure that the visual effect of the virtual digital human matches the real environment. The collection of environmental light information mainly relies on the built-in light sensors, cameras or RGB-D depth sensors of the device to measure parameters such as the light intensity, direction and color temperature in the scene. When specifically implemented, the dominant color of the scene can be calculated based on the Auto White Balance (AWB) algorithm, and the global light characteristics can be obtained through the Environmental Light Estimation method.
[0142] After obtaining the environmental light information, the background image can be obtained through the camera of the user device, and the foreground segmentation is performed on the background image to generate a virtual digital human layer and a real background layer. The goal of foreground segmentation is to separate the virtual digital human from the real environment to ensure a high visual consistency in the subsequent rendering and fusion process. When specifically implemented, a foreground separation method based on Alpha matting can be adopted, and the details can be retained by calculating the transparency channel of the image, such as the separation of complex regions such as hair and shadows, to enhance the natural transition effect of the foreground edge. Of course, a method based on background modeling can also be adopted, where the static background is detected through frame difference and the moving object is extracted using a foreground detection algorithm to reduce the computational overhead and improve the real-time performance. The method adopted for foreground segmentation in this exemplary embodiment is not specifically limited.
[0143] After completing the foreground segmentation, based on the current environmental light information and the virtual digital human layer, the appearance parameters of the virtual digital human can be adjusted using a deep learning-driven augmented reality engine to ensure the degree of fusion between the virtual and the real scene. The adjustment of appearance parameters can include aspects such as light matching, shadow synthesis, and color correction. When specifically implemented, a light matching method based on image enhancement can be adopted, and the light direction, intensity and color temperature of the virtual digital human are adjusted to be consistent with the environmental light. Of course, a method based on real-time light estimation can also be adopted, where the environmental light is predicted by calculating the spherical harmonic light coefficients of the scene, and the light response of the virtual human is calculated based on the light transport matrix to achieve more accurate light adaptation.
[0144] After adjusting the appearance parameters of the virtual digital human, the adjusted virtual digital human layer can be synthesized with the real background layer to render an augmented reality virtual digital human. The core goal of layer synthesis is to achieve seamless integration of the virtual human and the real environment through foreground and background fusion technology. In specific implementation, a depth perception-based method can be adopted. By calculating the depth relationship between the virtual human and the background, the stacking order and transparency of the layers are adjusted to enhance the sense of space of the virtual human. Or the Screen-Space Reflections (SSR) technology can be combined. By simulating the reflection of ambient light, the virtual human can truly reflect the lighting changes of surrounding objects. This exemplary embodiment does not make special limitations on the method for simulating the augmented reality effect.
[0145] After the final rendering is completed, the image data of the augmented reality virtual digital human can be displayed through the user device, enabling the user to intuitively observe the interaction effect between the virtual human and the real world in the real environment. Through the above method, the system can achieve high-quality fusion of the virtual human and the real scene under different lighting conditions and complex background environments, thereby enhancing the immersion and interaction naturalness of the augmented reality experience.
[0146] It should be noted that although the steps of the methods in this disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0147] In addition, in this exemplary embodiment, a virtual human design and application system based on artificial intelligence is also provided. Referring to Figure 7 As shown, the virtual human design and application system 700 based on artificial intelligence includes: a multimodal input interface module 710, a local anonymization processing module 720, a cloud multimodal encoding module 730, an emotion computing module 740, a context awareness module 750, a personality parameter update module 760, a multimodal generation module 770, and a local rendering module 780. Among them:
[0148] The multimodal input interface module 710 is used to collect multimodal input data through the input interface on the user device, and the multimodal input data includes at least one of text data, voice data, image data, and video data;
[0149] The local anonymization processing module 720 is used to perform local anonymization operations on the multimodal input data on the user device to generate anonymized multimodal features;
[0150] A cloud multi-modal encoding module 730, configured to input the anonymized multi-modal features into a multi-modal encoder deployed in the cloud. The multi-modal encoder maps the output of the anonymized multi-modal features into a unified semantic space through cross-modal contrastive learning, obtaining multi-modal feature vectors;
[0151] An emotion computing module 740, configured to input the multi-modal feature vectors into a pre-trained emotion computing model. The emotion computing model includes a text emotion classification sub-model, a speech emotion analysis sub-model, and a facial expression recognition sub-model, and outputs a quantified value of the user's emotion intensity;
[0152] A context awareness module 750, configured to input the multi-modal feature vectors into a pre-trained context awareness model. The context awareness model extracts temporal features in multi-round guiding conversations based on a long short-term memory network, and generates a user intention vector with context correction in combination with a knowledge graph;
[0153] A personality parameter update module 760, configured to adjust a preset personality parameter matrix according to the quantified value of the user's emotion intensity and the user intention vector, generating an updated personality parameter matrix;
[0154] A multi-modal generation module 770, configured to input the updated personality parameter matrix into a pre-trained multi-modal generation model. The multi-modal generation model includes a speech synthesis sub-model, a facial action generation sub-model, and a limb action generation sub-model, and outputs speech waveform data, facial muscle movement parameters, and skeletal joint coordinate data;
[0155] A local rendering module 780, configured to load the speech waveform data, the facial muscle movement parameters, and the skeletal joint coordinate data on the user device through a local rendering engine, driving a virtual digital human three-dimensional model for real-time rendering.
[0156] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the local anonymization processing module 720 is configured to: perform context semantic recognition on the text data using a lightweight text feature extraction model, obtaining text feature vectors; perform timbre feature extraction on the speech data using a lightweight speech recognition model, obtaining speech timbre feature vectors; perform facial key point coordinate extraction on the image data and the video data using a lightweight image recognition model, obtaining image feature vectors; add Laplace noise to each of the text feature vectors, the speech timbre feature vectors, and the image feature vectors, generating anonymized multi-modal features.
[0157] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the cloud multi-modal encoding module 730 is configured to: calculate the feature alignment loss between the text feature vector, the voice timbre feature vector, and the image feature vector through cosine similarity; optimize the weight parameters of the multi-modal encoder through the feature alignment loss and in combination with gradient descent, and based on the multi-modal encoder with updated weight parameters, map the anonymized multi-modal feature output into a unified semantic space to obtain a multi-modal feature vector.
[0158] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the context awareness module 750 is configured to: input the multi-modal feature vector into the long short-term memory network of the context awareness model for temporal feature extraction to determine the temporal features in the multi-round guided conversation; input the temporal features into a pre-set knowledge graph, and generate an initial user intention vector based on the structured information in the knowledge graph; determine the context-corrected user intention vector through the temporal features and the initial user intention vector.
[0159] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the personality parameter update module 760 is configured to: perform weight normalization processing on the user emotion intensity quantization value to generate an emotion normalization vector; perform weighted fusion based on the context-corrected user intention vector to generate an intention enhancement vector; perform a linear transformation on the emotion normalization vector and the intention enhancement vector to calculate a personality parameter adjustment value; based on the personality parameter adjustment value, perform element-by-element weighted update on the pre-set personality parameter matrix to obtain an updated personality parameter matrix.
[0160] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the multi-modal generation module 770 is configured to: generate a multi-modal latent feature vector by using a conditional variational autoencoder based on the updated personality parameter matrix; input the multi-modal latent feature vector into the speech synthesis sub-model of the multi-modal generation model to generate speech waveform data through the autoregressive neural network in the speech synthesis sub-model; input the multi-modal latent feature vector into the facial action generation sub-model of the multi-modal generation model to determine facial muscle movement parameters through the graph convolutional neural network in the facial action generation sub-model; input the multi-modal latent feature vector into the limb action generation sub-model of the multi-modal generation model to perform inverse kinematics calculation through the limb action generation sub-model to determine bone joint coordinate data.
[0161] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the local rendering module 780 is configured to: based on the voice waveform data, use the local audio synthesis module of the user device to play voice; based on the facial muscle movement parameters, use the expression driving engine of the user device to determine the facial mesh deformation amount, and update the facial mesh model corresponding to the virtual digital human three-dimensional model through the facial mesh deformation amount; based on the bone joint coordinate data, use the bone binding module of the user device to adjust the bone structure of the virtual digital human three-dimensional model; and fuse the facial mesh model and the adjusted bone structure through the graphics rendering pipeline of the user device to render the complete virtual digital human three-dimensional model in real time.
[0162] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the artificial intelligence-based virtual human design and application system 700 further includes an augmented reality rendering module, and the augmented reality rendering module is configured to: collect the current ambient light information through the sensor of the user device; obtain a background image through the camera of the user device, and perform foreground segmentation on the background image to generate a virtual digital human layer and a real background layer; based on the current ambient light information and the virtual digital human layer, use the deep learning-driven augmented reality engine to adjust the appearance parameters of the virtual digital human to obtain an adjusted virtual digital human layer; and synthesize the adjusted virtual digital human layer and the real background layer to render an augmented reality virtual digital human.
[0163] The specific details of each module of the artificial intelligence-based virtual human design and application system described above have been described in detail in the corresponding artificial intelligence-based virtual human design and application method, and thus will not be elaborated herein.
[0164] It should be noted that although several modules or units of the artificial intelligence-based virtual human design and application system are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.
[0165] In addition, in an exemplary embodiment of the present disclosure, there is also provided an electronic device capable of implementing the above artificial intelligence-based virtual human design and application method.
[0166] Those skilled in the art to which the present disclosure pertains will appreciate that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Accordingly, various aspects of the present disclosure may be embodied in the following forms, namely: a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which may be collectively referred to herein as a "circuit", a "module", or a "system".
[0167] Reference will now be made to Figure 8 describe the electronic device 800 according to such an embodiment of the present disclosure. Figure 8 The illustrated electronic device 800 is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present disclosure.
[0168] As Figure 8 shown, the electronic device 800 is presented in the form of a general-purpose computing device. The components of the electronic device 800 may include, but are not limited to: at least one of the above-mentioned processing units 810, at least one of the above-mentioned storage units 820, a bus 830 connecting different system components (including the storage unit 820 and the processing unit 810), and a display unit 840.
[0169] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 810, so that the processing unit 810 executes the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification. For example, the processing unit 810 may execute as Figure 2Step S210 shown in the figure collects multimodal input data through an input interface on the user device, where the multimodal input data includes at least one of text data, voice data, image data, and video data; Step S220 performs a local anonymization operation on the multimodal input data on the user device to generate anonymized multimodal features; Step S230 inputs the anonymized multimodal features into a multimodal encoder deployed in the cloud. The multimodal encoder maps the output of the anonymized multimodal features into a unified semantic space through cross-modal contrast learning to obtain multimodal feature vectors; Step S240 inputs the multimodal feature vectors into a pre-trained emotion computing model, where the emotion computing model includes a text emotion classification sub-model, a voice emotion analysis sub-model, and a facial expression recognition sub-model, and outputs a user emotion intensity quantization value; Step S250 inputs the multimodal feature vectors into a pre-trained context-aware model. The context-aware model extracts temporal features in multi-turn guided conversations based on a long short-term memory network and generates a context-corrected user intention vector in combination with a knowledge graph; Step S260 adjusts a preset personality parameter matrix according to the user emotion intensity quantization value and the user intention vector to generate an updated personality parameter matrix; Step S270 inputs the updated personality parameter matrix into a pre-trained multimodal generation model, where the multimodal generation model includes a voice synthesis sub-model, a facial motion generation sub-model, and a limb motion generation sub-model, and outputs voice waveform data, facial muscle movement parameters, and skeletal joint coordinate data; Step S280 loads the voice waveform data, the facial muscle movement parameters, and the skeletal joint coordinate data through a local rendering engine on the user device to drive real-time rendering of the virtual digital human three-dimensional model.
[0170] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 821 and / or a cache storage unit 822, and may further include a read-only storage unit (ROM) 823.
[0171] The storage unit 820 may further include a program / utilities 824 having a set (at least one) of program modules 825. Such program modules 825 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.
[0172] The bus 830 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the multiple bus structures.
[0173] The electronic device 800 can also communicate with one or more external devices 870 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 800, and / or communicate with any device that enables the electronic device 800 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 850. Moreover, the electronic device 800 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 860. As shown in the figure, the network adapter 860 communicates with other modules of the electronic device 800 through the bus 830. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0174] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by the way of software combined with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0175] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above method of the present specification is stored. In some possible embodiments, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of the present specification.
[0176] Reference Figure 9 As shown, a program product 900 for implementing the above method for designing and applying an AI-based virtual human according to an embodiment of the present disclosure is described. It can adopt a portable compact disc read-only memory (CD-ROM) and include program code, and can run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.
[0177] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0178] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0179] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0180] The program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0181] In addition, the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, and are not for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0182] From the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0183] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed herein. The specification and embodiments are only to be regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0184] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for the design and application of virtual humans based on artificial intelligence, characterized in that, Including: Collecting multi-modal input data through an input interface on a user device, where the multi-modal input data includes at least one of text data, voice data, image data, and video data; Performing local anonymization operations on the multi-modal input data on the user device to generate anonymized multi-modal features; Inputting the anonymized multi-modal features into a multi-modal encoder deployed in the cloud, and the multi-modal encoder maps the output of the anonymized multi-modal features into a unified semantic space through cross-modal contrast learning to obtain multi-modal feature vectors; Inputting the multi-modal feature vectors into a pre-trained emotion computing model, where the emotion computing model includes a text emotion classification sub-model, a voice emotion analysis sub-model, and a facial expression recognition sub-model, and outputting a user emotion intensity quantization value; Inputting the multi-modal feature vectors into a pre-trained context awareness model, and the context awareness model extracts temporal features in multi-round guided conversations based on a long short-term memory network and generates a context-corrected user intention vector in combination with a knowledge graph; Adjusting a preset personality parameter matrix according to the user emotion intensity quantization value and the user intention vector to generate an updated personality parameter matrix; Inputting the updated personality parameter matrix into a pre-trained multi-modal generation model, where the multi-modal generation model includes a speech synthesis sub-model, a facial motion generation sub-model, and a limb motion generation sub-model, and outputting speech waveform data, facial muscle movement parameters, and skeletal joint coordinate data; On the user device, loading the speech waveform data, the facial muscle movement parameters, and the skeletal joint coordinate data through a local rendering engine to drive real-time rendering of a virtual digital human three-dimensional model.
2. The virtual human design and application method according to claim 1, characterized in that The performing local anonymization operations on the multi-modal input data to generate anonymized multi-modal features includes: Using a lightweight text feature extraction model to perform context semantic recognition on the text data to obtain text feature vectors; Using a lightweight speech recognition model to extract timbre features from the voice data to obtain voice timbre feature vectors; Using a lightweight image recognition model to extract facial key point coordinates from the image data and the video data to obtain image feature vectors; Adding Laplace noise to each of the text feature vectors, the voice timbre feature vectors, and the image feature vectors to generate anonymized multi-modal features.
3. The virtual human design and application method according to claim 1, characterized in that, The multi-modal encoder maps the output of the anonymized multi-modal features into a unified semantic space through cross-modal contrast learning to obtain multi-modal feature vectors, including: Calculating the feature alignment loss between the text feature vectors, the voice timbre feature vectors, and the image feature vectors through cosine similarity; Optimizing the weight parameters of the multi-modal encoder through the feature alignment loss and in combination with gradient descent, and based on the multi-modal encoder with updated weight parameters, mapping the output of the anonymized multi-modal features into a unified semantic space to obtain multi-modal feature vectors.
4. The virtual human design and application method according to claim 1, characterized in that, The context awareness model extracts temporal features in multi-round guided conversations based on a long short-term memory network and generates a context-corrected user intention vector in combination with a knowledge graph, including: Input the multi-modal feature vector into the long short-term memory network of the context-aware model for temporal feature extraction to determine the temporal features in the multi-round guided dialogue; Input the temporal features into a pre-set knowledge graph, and generate an initial user intention vector based on the structured information in the knowledge graph; Determine the context-corrected user intention vector through the temporal features and the initial user intention vector.
5. The virtual human design and application method according to claim 1, characterized in that The adjusting the pre-set personality parameter matrix according to the user emotion intensity quantization value and the user intention vector to generate an updated personality parameter matrix includes: Perform weight normalization processing on the user emotion intensity quantization value to generate an emotion normalization vector; Perform weighted fusion based on the context-corrected user intention vector to generate an intention enhancement vector; Perform a linear transformation on the emotion normalization vector and the intention enhancement vector to calculate the personality parameter adjustment value; Based on the personality parameter adjustment value, perform element-wise weighted update on the pre-set personality parameter matrix to obtain the updated personality parameter matrix.
6. The virtual human design and application method according to claim 1, characterized in that The inputting the updated personality parameter matrix into a pre-trained multi-modal generation model to output speech waveform data, facial muscle movement parameters, and skeletal joint coordinate data includes: Based on the updated personality parameter matrix, use a conditional variational autoencoder to generate a multi-modal latent feature vector; Input the multi-modal latent feature vector into the speech synthesis sub-model of the multi-modal generation model to generate speech waveform data through the autoregressive neural network in the speech synthesis sub-model; Input the multi-modal latent feature vector into the facial action generation sub-model of the multi-modal generation model to determine facial muscle movement parameters through the graph convolutional neural network in the facial action generation sub-model; Input the multi-modal latent feature vector into the limb action generation sub-model of the multi-modal generation model to perform inverse kinematics calculation through the limb action generation sub-model to determine skeletal joint coordinate data.
7. The virtual human design and application method according to claim 1, characterized in that The loading the speech waveform data, the facial muscle movement parameters, and the skeletal joint coordinate data on the user device through a local rendering engine to drive real-time rendering of the virtual digital human three-dimensional model includes: Based on the speech waveform data, use the local audio synthesis module of the user device to play the speech; Based on the facial muscle movement parameters, use the expression driving engine of the user device to determine the facial mesh deformation amount, and update the facial mesh model corresponding to the virtual digital human three-dimensional model through the facial mesh deformation amount; Based on the skeletal joint coordinate data, use the bone binding module of the user device to adjust the bone structure of the virtual digital human three-dimensional model; Fuse the facial mesh model and the adjusted bone structure through the graphics rendering pipeline of the user device to render the complete virtual digital human three-dimensional model in real time.
8. The virtual human design and application method according to claim 1, characterized in that The method further includes: Collect the current environmental light information through the sensor of the user device; Obtain a background image through the camera of the user device, and perform foreground segmentation on the background image to generate a virtual digital human layer and a real background layer; Based on the current environmental light information and the virtual digital human layer, use a deep learning-driven augmented reality engine to adjust the appearance parameters of the virtual digital human, and obtain an adjusted virtual digital human layer; Composite the adjusted virtual digital human layer with the real background layer, and render to obtain an augmented reality virtual digital human.
9. A virtual human design and application platform based on artificial intelligence, characterized in that, Including: A multimodal input interface module for collecting multimodal input data through an input interface on a user device, where the multimodal input data includes at least one of text data, voice data, image data, and video data; A local anonymization processing module for performing local anonymization operations on the multimodal input data on the user device to generate anonymized multimodal features; A cloud multimodal encoding module for inputting the anonymized multimodal features into a multimodal encoder deployed in the cloud. The multimodal encoder maps the output of the anonymized multimodal features into a unified semantic space through cross-modal contrast learning to obtain multimodal feature vectors; An emotion calculation module for inputting the multimodal feature vectors into a pre-trained emotion calculation model. The emotion calculation model includes a text emotion classification sub-model, a voice emotion analysis sub-model, and a facial expression recognition sub-model, and outputs a user emotion intensity quantization value; A context awareness module for inputting the multimodal feature vectors into a pre-trained context awareness model. The context awareness model extracts temporal features in multi-round guided conversations based on a long short-term memory network and combines a knowledge graph to generate a context-corrected user intention vector; A personality parameter update module for adjusting a preset personality parameter matrix according to the user emotion intensity quantization value and the user intention vector to generate an updated personality parameter matrix; A multimodal generation module for inputting the updated personality parameter matrix into a pre-trained multimodal generation model. The multimodal generation model includes a voice synthesis sub-model, a facial motion generation sub-model, and a limb motion generation sub-model, and outputs voice waveform data, facial muscle movement parameters, and skeletal joint coordinate data; A local rendering module for loading the voice waveform data, the facial muscle movement parameters, and the skeletal joint coordinate data through a local rendering engine on the user device to drive real-time rendering of a virtual digital human three-dimensional model.
10. An electronic device, characterized in that, Including: A processor; And A memory, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, the virtual human design and application method described in any one of claims 1 to 8 is implemented.
Citation Information
Cited By
AI digital human expression and facial feature migration method and system
CN120997351A
AI digital human interactive response method based on large language model
CN121144484A
Tablet dynamic authority management method and system based on biological characteristics and behavior patterns
CN121145187A
Digital human face binding generation method based on expression capture
CN121214518A
Interaction method and device based on intelligent agent, intelligent agent and electronic equipment
CN121434452A