Digital human intelligent interaction and posture expression synthesis method based on multi-modal synchronization

Through the time synchronization and structured annotation of multimodal data, combined with timing Transformer and generative adversarial network, the problems of insufficient emotional consistency, nature and synchronization in the traditional virtual character generation method are solved, and high-quality digital human intelligent interaction and posture expression synthesis are achieved.

CN120068923APending Publication Date: 2025-05-30BEIJING YINGTAI LICHEN INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510560688.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The traditional virtual character expressions and posture generation methods lack emotional consistency, nature and synchronization, resulting in the user interaction experience not being realistic enough and unable to meet the needs of highly immersive virtual scenes.

Method used

Voice signals are collected through microphone arrays, expression image data captured by cameras and attitude sensors are acquired by body posture data, and time stamp alignment technology is used to form a multimodal data set under a unified timestamp reference, and a standardized multimodal feature sequence is generated through adaptive noise filtering and standardized processing. Dynamic modeling and multimodal semantic fusion are used to use timing Transformer, and finally multimodal synchronization features are input to the emotion-driven generation model, and natural expressions and pose sequences are generated based on the generative adversarial network.

Benefits of technology

It significantly improves the real-time and accuracy of digital human interaction, enhances the naturalness of posture and expression synthesis, improves the accuracy and emotional consistency of multimodal data processing, and improves user experience and interaction quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068923A_ABST
    Figure CN120068923A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a digital human intelligent interaction and posture expression synthesis method based on multi-modal synchronization, which comprises the following steps: acquiring voice, expression and posture data through a multi-modal acquisition device, and generating a multi-modal feature sequence after noise filtering and standardization processing; a self-attention mechanism and a time sequence Transform are adopted to carry out time alignment and semantic fusion on the features, and multi-modal synchronization features are generated; generating parameters by utilizing an emotion-driven generative model and generative adversarial network optimization, generating a natural expression and posture sequence, and realizing real-time rendering and output through edge computing equipment; and continuously optimizing the multi-modal model and generating parameters based on user interaction data. The method improves the real-time performance of interaction and the authenticity of emotion expression, has high expansibility and adaptive optimization capability, and can be widely applied to the fields of virtual assistants, immersive experience, distance education and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a digital human intelligent interaction and gesture expression synthesis method based on multi-modal synchronization. Background Art

[0002] With the rapid development of artificial intelligence and virtual reality technologies, digital human technology has gradually become an important research direction in the field of human-computer interaction. Digital humans not only need to have natural and fluent voice and behavior interactions, but also need to be able to simulate real facial expressions and body postures, so as to enhance the user experience in the virtual environment. In the application scenarios of digital humans, how to achieve synchronous processing of multi-modal data such as voice, expression, and posture, and how to generate natural perception expressions and postures based on these data have become one of the current research difficulties.

[0003] Traditional virtual character expression and posture synthesis methods usually use static data for generation, ignoring the changes in emotions and the synchronization between multi-modal data. Therefore, in digital human interactions, the performance of virtual characters often lacks naturalness and emotional consistency, resulting in an insufficiently realistic user interaction experience and being unable to meet the high requirements for the performance of digital characters in highly immersive virtual scenarios.

[0004] In addition, existing emotion-driven virtual character generation methods have insufficient synchronous processing of multiple modal data (such as voice, expression, posture, etc.), and are unable to accurately capture and integrate emotional information between different modalities. Generating expression and posture sequences with consistent emotions through the synchronization and integration of multi-modal data, and then realizing natural interactions of virtual characters, is still a major challenge in digital character technology.

[0005] Therefore, the present invention provides a digital human intelligent interaction and gesture expression synthesis method based on multi-modal synchronization, aiming to solve the problems of emotional consistency, naturalness, and synchronization in traditional virtual character expression and posture generation. Summary of the Invention

[0006] The present invention provides a digital human intelligent interaction and gesture expression synthesis method based on multi-modal synchronization, which is used to help solve the problems mentioned in the above background art.

[0007] The present invention provides the following technical solution: A digital human intelligent interaction and gesture expression synthesis method based on multi-modal synchronization, comprising: Collecting user voice signals through a microphone array ; Capturing user's expression image data through a camera ; Obtaining user's body posture data through a posture sensor ; Wherein, Represents the time step, , and are all data in the form of time series; Respectively, , and are aligned based on a unified timestamp benchmark The specific timestamp alignment process is as follows: Respectively obtain , and corresponding timestamps; If the timestamps corresponding to the speech signal and the video image data are inconsistent, timestamp alignment is performed through linear interpolation: Calculate the weighted average between adjacent data points and through linear interpolation to fill in the missing time step data. The specific linear interpolation calculation formula is as follows: ; where, and represent adjacent time points, and respectively represent and corresponding values, represents the time point for interpolation, represents the interpolation result; If the timestamp corresponding to the pose data is inconsistent with the timestamp corresponding to the speech signal or the video image data , timestamp alignment operation is performed through the time difference correction algorithm, which is as follows: Set the original timestamp of the pose data to ; Set the reference timestamp of the video data to ; Set the time difference correction algorithm, which is as follows: ; where, represents the timestamp of the pose data after correction, represents the systematic offset, , and represent the corresponding and under a common reference point; After the timestamp alignment operation, , and , to form an initial multi-modal data set under a unified time stamp benchmark, denoted as ; Annotate to construct a multi-modal structured data set containing time stamps and modal tags, denoted as ; The above modal tags specifically include , and , used to distinguish different types of data.

[0008] Optionally, it also includes detecting and filtering abnormal signals in the multi-modal structured data set, eliminating background interference through an adaptive noise filtering algorithm, and generating multi-modal data after noise filtering, specifically: Mark the abnormal signals in the multi-modal structured data set; Apply the adaptive noise filtering algorithm for filtering, and its transfer function is: ; where is the cut-off frequency, is the filtering order, represents the signal frequency.

[0009] Recombine the data after noise filtering to generate a filtered multi-modal data set .

[0010] Optionally, it also includes normalizing the multi-modal data set after noise filtering based on a standardization parameter library to generate a standardized multi-modal feature sequence, specifically: Read the normalization parameters of each modality from the standardization parameter library , and calculate the normalization formula: ; where, and are the mean and standard deviation respectively, represents the multi-modal data set after normalization processing.

[0011] Concatenate the standardized data in and in time series to generate a standardized multi-modal feature matrix : ; where, Represents the nth data in the time series; For Slice by row to generate a standardized feature sequence with a time step of T ; Sort all the standardized feature sequences in ascending order of time steps; Determine the time step : ; where represents the minimum sampling frequency corresponding to different types of data.

[0012] Optionally, it also includes dynamic modeling of the feature sequence based on the temporal Transformer, specifically: Set the feature sequence , represents the kth feature vector; Embed the standardized feature sequence into the feature sequence and input it into the encoding layer of the temporal Transformer; For each time step t, the feature representation is , where represents the number of modalities, represents the feature vector of the mth modality at time step ; Use the self-attention mechanism of the temporal Transformer to calculate the attention weight matrix between time step and t : ; where is a mathematical function that transforms each element in the input vector into a value between 0 and 1, and the sum of all values is 1. The formula is: ; where is the exponential operation on the jth input value , is the number of elements, is the sum of the exponents of all elements; The above represents the query vector, specifically as follows: ; where is obtained by linearly transforming the feature vector with ; is the query weight matrix, which is learned through optimization during the model training process. is the time step of the input feature vector; The above represents the key vector, is the dimension of the key vector, represents the transpose of the key vector, specifically as follows: ; Among them, is obtained by linearly transforming the feature vector with , is the key weight matrix, similar to , is the weight matrix learned through the training process and is used for feature transformation; According to the attention weight and the feature vector , calculate and output : ; Among them, is the final fused feature representation at time step , representing the unified feature vector generated after the attention weighting mechanism and is used for subsequent feature decoding and classification judgment, is the value vector; Combine the corresponding to each time step to generate the feature sequence after dynamic modeling, represents the num th feature vector.

[0013] Optionally, it also includes multi-modal semantic fusion of the features after dynamic modeling, specifically: Align the feature dimensions of different modalities through linear transformation of the dynamic feature sequence : ; Among them is the parameter matrix automatically learned during the model training process, is the bias of modality , represents the feature vector of the th modality in time step , and the aligned feature sequence is represented as ; Use the weighting mechanism to perform semantic fusion on the modality features and calculate the fused feature : ; Among them, the weight represents the modality at the time step The importance is calculated by the following formula: ; Among them, is the semantic vector of the modality , represents the transpose of; Obtain the fusion features of all time steps to get the feature sequence after semantic fusion ; For the feature sequence after semantic fusion , each time step The fusion feature of is introduced into the synchronous mapping layer to map it to the synchronous feature space: ; Among them, represents the synchronous feature formed at the time step t, The specific expression of is: ; Among them, is the synchronous mapping function, which is implemented through a fully connected layer and a normalization operation; Combine the synchronous features of each time step to generate a time-synchronous feature sequence .

[0014] Optionally, it further includes inputting the multi-modal synchronous features into an emotion-driven generation model, specifically: The multi-modal synchronous feature sequence Specifically: ; Input the elements in the multi-modal synchronous feature into the initial layer of the emotion-driven generation model in chronological order; According to the multi-modal synchronous feature and the current emotion label (for example, the emotion label input by the user or the detected emotion state) are combined to form a comprehensive input denoted as, ; Input into the emotion-driven generation model to generate a preliminary expression and gesture sequence , that is: ; The emotion-driven generation model adopts a generative adversarial network architecture, including a generator and a discriminator ; wherein, represents the generator function; According to the expression and gesture sequence and the real expression and gesture sequence , the generative adversarial network is iteratively optimized; Output the natural expression and gesture sequence according to the optimization result .

[0015] Optionally, it further includes real-time rendering and output through an edge computing device based on the expression and gesture sequence, and capturing user interaction data as feedback. Specifically: Transmit and to the edge computing tool for real-time rendering to generate the dynamic performance of the digital human, denoted as ; Real-time collect the interaction data between the user and the digital human through a multimodal sensor; Perform time synchronization processing on the interaction data: ; wherein, is the user interaction data after time synchronization processing, is the synchronization operation; Transmit to the feedback processing module to generate feedback information, and adjust the generator according to the feedback information; When the adjustment of the generator by the feedback information terminates, optimize the expression and gesture sequence of the digital human through and the generator.

[0016] Optionally, it further includes inputting the user interaction data to a continuous learning module to iteratively optimize the emotion-driven generation model. Specifically: Input to the continuous learning module to optimize the emotion-driven generation model; Synchronize the time of and , and denote the processed as ; Train the emotion-driven generation model according to and update the parameters of the generator ; Use the updated parameters of the generator to optimize .

[0017] The present invention has the following beneficial effects: 1. For the digital human intelligent interaction and gesture expression synthesis method based on multimodal synchronization, voice signals are collected through a microphone array, facial expression image data is captured by a camera, and body posture data is obtained by a posture sensor, and these data are recorded in the form of a time series. Through timestamp alignment technology, the voice signals and video image data use linear interpolation to fill in missing data, and the posture data is synchronized through a time difference correction algorithm, and finally a multimodal dataset under a unified timestamp reference is formed. The dataset is labeled to construct a structured dataset containing timestamps and modality tags for subsequent multimodal feature fusion and processing. In this way, through the time synchronization and structured annotation of multimodal data, the real-time performance and accuracy of digital human interaction are significantly improved, the naturalness of gesture and expression synthesis is enhanced, and at the same time, an efficient basis for multimodal data fusion and processing is provided, which is applicable to intelligent interaction applications in complex scenarios.

[0018] 2. For the digital human intelligent interaction and gesture expression synthesis method based on multimodal synchronization, by detecting and filtering abnormal signals in the multimodal structured dataset, an adaptive noise filtering algorithm is used to effectively eliminate background interference, thereby generating noise-filtered data. By marking abnormal signal intensity, spectral features, time series, equipment failures, and data inconsistencies, the accuracy and reliability of the data are ensured. Then, the data is normalized based on a standardized parameter library to generate a standardized multimodal feature sequence, and these sequences are concatenated to form a standardized feature matrix. By precisely controlling the relationship between the time step and the data acquisition frequency, the unity and timeliness of different modality data in the time series are ensured. This method helps to improve the accuracy of multimodal data processing, optimize the data fusion effect, and provide a high-quality data basis for subsequent intelligent modeling and applications, especially in scenarios such as emotion-driven generation and digital human interaction, where the performance is more natural and realistic.

[0019] 3. The digital human intelligent interaction and gesture expression synthesis method based on multimodal synchronization. In the present invention, a temporal Transformer is used to dynamically model the multimodal feature sequence, and the relationship between each time step is calculated through the self-attention mechanism, so as to capture long-term dependencies and the correlation of cross-modal features. At each time step, the feature vector is weighted through the mechanism of query, key, and value to generate a fused feature representation. This process enables the model to automatically adjust and optimize the attention to different modal data by learning the query weight matrix and the key weight matrix, enhancing the connection between features. This method effectively improves the expressive ability of multimodal data, can better process complex time series features, and provides a more accurate input for subsequent generation and classification tasks. In applications such as digital human interaction and emotion-driven generation, this technology can improve the accuracy and effect of data fusion and enhance the system's response ability to complex dynamic information.

[0020] 4. The digital human intelligent interaction and gesture expression synthesis method based on multimodal synchronization effectively integrates the features after dynamic modeling through multimodal semantic fusion technology. First, linear transformation is used to align the feature dimensions of different modalities (such as speech, video, gesture, etc.) and unify them into a common feature space. Through the parameter matrix automatically learned during the training process, as well as the weights and biases of the modalities, the modal features at each time step are further weighted and adjusted. The weighting mechanism calculates the fused features using the semantic vectors of the modalities according to the importance of the modalities to ensure the synergistic effect and semantic consistency of different modal data within the time step. Then, the fused features of all time steps are combined to form a feature sequence after semantic fusion, forming a rich and compact multimodal representation. On this basis, for the fused feature sequence, a synchronous mapping layer is used to map it to a unified synchronous feature space to ensure the synchronization and consistency of different modal data in time. The synchronous mapping process is achieved through a fully connected layer and a normalization operation. The fully connected layer sums the weighted features and passes them to the next layer, and the normalization operation standardizes the data, enhancing the stability and training efficiency of the model. Finally, through the generation of the time-synchronized feature sequence, a precise and highly consistent input feature is provided for subsequent generation and classification tasks. This method can effectively improve the expressive ability of multimodal data, enhance the model's processing ability for cross-modal data, provide a more accurate and natural feature representation for applications such as digital human interaction and emotion-driven generation, and significantly improve the performance and reliability of the multimodal system.

[0021] 5. The digital human intelligent interaction and gesture expression synthesis method based on multimodal synchronization realizes the generation of natural expressions and gestures based on the generative adversarial network (GAN) by inputting multimodal synchronization features into the emotion-driven generation model. This method innovatively drives based on the multimodal time synchronization feature sequence, rather than the traditional single modality (such as speech), and introduces the user's personalized emotional state vector to support dynamic style control. The emotion-driven generation model generates an expression and gesture sequence highly consistent with the emotion label through the adversarial training of the generator and discriminator. During the training process, the adversarial loss, discriminator loss, emotion consistency loss, and reconstruction loss are combined to ensure that the generated expressions and gestures are both natural and consistent with the emotion label. Through continuous iterative optimization, the generator gradually improves the quality of the generated results, and the output expression and gesture sequence can accurately reflect the user's emotional state and multimodal features. This method not only improves the interaction experience of digital humans or virtual characters, making them more emotionally resonant, but also provides strong support for personalized customization and style control, and is widely applied in fields such as virtual reality, emotion computing, and digital entertainment, significantly improving the naturalness and emotion consistency of the generated content.

[0022] 6. The digital human intelligent interaction and gesture expression synthesis method based on multi-modal synchronization performs real-time rendering and output through edge computing devices, enabling the digital human to present natural and expressive dynamic expressions and gesture sequences according to the emotions and interactions of users. The system collects the interaction data between the user and the digital human in real time through multi-modal sensors, and performs time synchronization processing on these data to ensure that the interaction data at each moment is consistent with the performance of the digital human. The rendering process involves pixel-level adjustment of each frame of expression and gesture to ensure that the performance conforms to the emotion labels, and uses transformation matrices to map facial muscle control parameters, joint angles, etc. to three-dimensional space or two-dimensional image space, while combining emotion labels for detailed adjustment. The emotional intensity of expressions and gestures is adjusted through linear interpolation, and the physical-based rendering (PBR) method is used to simulate lighting and material effects to make the rendered images more realistic. Finally, the generated expression and gesture sequences are displayed under the viewing perspective through projection transformation to generate displayable images or animation frames. At the same time, the system captures the emotional expressions, speech content, and gesture changes of the user as feedback, and optimizes the adjustment coefficients in the generator and rendering process in real time to ensure that the expressions and gestures of the digital human are consistent with the emotional state of the user. The optimization objective is adjusted through feedback signals, where the feedback signals are based on user interaction data to help the system optimize the parameters of the generation model and enhance the emotional consistency with the user. The discriminator further promotes the optimization of the generator and rendering process by evaluating the authenticity of the rendered expressions and gestures. Finally, the parameters of the generator and rendering process are continuously adjusted to output natural expression and gesture sequences that are highly consistent with the emotions and interactions of the user. This method can not only significantly improve the interaction experience between the digital human and the user, making the virtual character's performance more emotionally resonant, but also provide stronger accuracy and flexibility for real-time emotion-driven generation and character performance in virtual interaction scenarios, and is widely applied in fields such as virtual reality, augmented reality, digital entertainment, and emotion computing.

[0023] 7. The digital human intelligent interaction and gesture expression synthesis method based on multi-modal synchronization iteratively optimizes the emotion-driven generation model through a continuous learning module to improve the performance consistency and accuracy of virtual characters in user interaction. After rendering and feedback, the system inputs user interaction data (such as voice, facial expressions, and gesture modality features) into the continuous learning module and performs time synchronization processing. Specifically, the interaction data and the expression and gesture sequences are aligned by timestamps, and data with different sampling frequencies are synchronized to the same time step using linear interpolation or resampling methods to ensure the precise alignment of all modality data. After synchronization, the interaction data is fused with the expression and gesture features into multi-modal input features for further training and optimization of the emotion-driven generation model. By updating the generator parameters, the system can continuously optimize the performance of the generated virtual characters based on user interaction data, thereby achieving more natural and realistic interactions. This method effectively enhances the emotional consistency and dynamic performance between virtual characters and users, making the generated expression and gesture sequences more in line with the emotional state and interaction needs of users, and is widely applicable to fields such as digital humans, virtual reality, and augmented reality, significantly improving the user experience and interaction quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a schematic flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0026] Example 1, refer to Figure 1 , a digital human intelligent interaction and gesture expression synthesis method based on multi-modal synchronization, characterized by including: The specific process is as follows: S1. Based on multi-modal acquisition devices, obtain voice, expression, and gesture multi-modal data, and perform noise filtering and normalization processing on the data to generate multi-modal feature sequences.

[0027] S2. Perform time alignment processing on the multi-modal feature sequences, and generate semantically fused multi-modal synchronization features based on the self-attention mechanism and temporal Transformer.

[0028] S3. Input the multi-modal synchronization features into the emotion-driven generation model, optimize the generation parameters based on the generative adversarial network and emotion labels, and generate natural expression and gesture sequences.

[0029] S4. Based on the expression and gesture sequences, perform real-time rendering and output through an edge computing device, and capture user interaction data as feedback.

[0030] S5. Input the user interaction data into a continuous learning module to iteratively optimize the emotion-driven generation model and update the generation parameters and feature processing weights.

[0031] Collect the user voice signal through a microphone array ; Capture the user's facial expression image data through a camera ; Obtain the user's body posture data through a posture sensor ; where represents the time step, , and are all data in the form of time series; Respectively align , and according to a unified time stamp benchmark . The specific time stamp alignment process is as follows: Respectively obtain , and corresponding time stamps; If the time stamps corresponding to the voice signal and the video image data are inconsistent, perform time stamp alignment through linear interpolation: Calculate the weighted average between adjacent data points and through linear interpolation to fill in the missing time step data. The specific linear interpolation calculation formula is as follows:

[0032] where and represent adjacent known time points, and respectively represent and corresponding values, represents the time point to be interpolated, represents the interpolation result; If the time stamp corresponding to the posture data is inconsistent with the time stamp corresponding to the voice signal or the video image data , perform time stamp alignment operation through a time difference correction algorithm, as follows: Set the original timestamp of the pose data to ; Set the reference timestamp of the video data to ; Set the time difference correction algorithm as follows: ; Wherein, represents the timestamp of the corrected pose data, represents the systematic offset, , and represent the corresponding and under a common reference point; Align the , and after the timestamp alignment operation to form an initial multi-modal dataset under a unified timestamp benchmark, denoted as ; Annotate to construct a multi-modal structured dataset containing timestamps and modality tags, denoted as ; The above modality tags specifically include , and , which are used to distinguish different types of data. For example, voice data will be annotated as "audio", video data will be annotated as "video", and pose data will be annotated as "pose". These tags facilitate subsequent multi-modal feature fusion and processing.

[0033] Collect voice signals through a microphone array, capture facial expression image data with a camera, and obtain body pose data with a pose sensor, and record these data in the form of a time series. Through timestamp alignment technology, linear interpolation is used to fill in the missing data for voice signals and video image data, and pose data is synchronized through the time difference correction algorithm, finally forming a multi-modal dataset under a unified timestamp benchmark. Annotate the dataset to construct a structured dataset containing timestamps and modality tags for subsequent multi-modal feature fusion and processing. This can significantly improve the real-time performance and accuracy of digital human interaction, enhance the naturalness of pose and expression synthesis, and at the same time provide an efficient basis for the fusion and processing of multi-modal data, and is applicable to intelligent interaction applications in complex scenarios.

[0034] It also includes detecting and filtering abnormal signals in the multi-modal structured dataset, eliminating background interference through an adaptive noise filtering algorithm, and generating noise-filtered multi-modal data, specifically as follows: Mark abnormal signals in the multimodal structured dataset; The criteria for abnormal signals include the following aspects: Abnormal signal intensity: If the amplitude value of the signal exceeds the predetermined threshold range, it may indicate that the signal is strongly interfered or there is an incorrect acquisition. For example, if the amplitude of the voice signal is too large or too small, extremely bright values appear in the video signal, or unreasonable angles appear in the pose signal, they will be regarded as abnormal.

[0035] Abnormal spectral characteristics: High-frequency or low-frequency components that do not conform to the normal signal pattern appear in the signal spectrum, which may indicate signal interference or acquisition errors. For example, spectral components that do not conform to language characteristics appear in the voice signal, unnatural frequency fluctuations appear in the video signal, and unconventional periodic changes appear in the pose signal.

[0036] Abnormal time series: If there are mutations or losses in the signal in the time series, it is usually regarded as an abnormal signal. For example, sudden silence appears in the voice signal, a certain frame of the video signal is missing, or a sudden spike or sharp change occurs at a certain moment in the pose data.

[0037] Abnormal signals caused by acquisition device failures: Abnormal signals caused by device failures (such as sensor errors, microphone or camera failures). Such abnormalities are usually inconsistent with the normal operating state of the device and can be identified through device health status monitoring.

[0038] Data inconsistency: If the relationships between multimodal data are inconsistent, it may also be regarded as abnormal. For example, the expression does not match the pose in the video signal, or the voice is not coordinated with the pose.

[0039] Apply an adaptive noise filtering algorithm for filtering, and its transfer function is: ; where is the cut-off frequency, is the filter order, represents the signal frequency.

[0040] Recombine the data after noise filtering to generate a filtered multimodal dataset .

[0041] It also includes normalizing the multimodal dataset after noise filtering based on a standardized parameter library to generate a standardized multimodal feature sequence, specifically: Read the normalization parameters of each modality from the standardized parameter library , and calculate the normalization formula: ; Among them, and are the mean and standard deviation respectively, representing the multimodal dataset after normalization processing.

[0042] The standardized parameter library is used to store the statistical parameters of each modality data, including the mean and the standard deviation . This parameter library can be obtained by calculating the pre-collected historical training data. The normalization parameters are organized by modality to ensure that multimodal features are processed and fused on a unified numerical scale.

[0043] The standardized data in and are concatenated according to the time series to generate a standardized multimodal feature matrix : ; Among them, represents the nth data in the time series; The is sliced by row to generate a standardized feature sequence with a time step of T ; All the standardized feature sequences are sorted in the order of time steps from early to late; Determine the time step : ; Among them, represents the minimum acquisition frequency corresponding to different types of data.

[0044] The time step should be closely related to the acquisition frequency of the data. Assume that the data acquisition frequencies of voice signals, video signals, and pose signals are , and (unit: Hz). Usually, should be selected according to these frequencies to ensure that an appropriate number of sample points are included in each time step.

[0045] For example, if the frame rate of the video signal is 30 frames per second ( Hz), and the sampling rate of the voice signal is 16 kHz ( Hz), then can be set according to the video frame rate because the frame rate of the video is lower, and the corresponding time step It can be used to ensure that the number of voice and gesture data points included in each time step is neither too large nor too small.

[0046] By detecting and filtering abnormal signals in the multimodal structured dataset, the adaptive noise filtering algorithm is used to effectively eliminate background interference, thereby generating noise-filtered data. By marking abnormal signal intensity, spectral characteristics, time series, equipment failures, and data inconsistencies, the accuracy and reliability of the data are ensured. Then, based on the standardized parameter library, the data is normalized to generate a standardized multimodal feature sequence, and these sequences are concatenated to form a standardized feature matrix. By precisely controlling the relationship between the time step length and the data acquisition frequency, the unity and timeliness of different modality data in the time series are ensured. This method helps to improve the accuracy of multimodal data processing, optimize the data fusion effect, and provide a high-quality data foundation for subsequent intelligent modeling and applications, especially in scenarios such as emotion-driven generation and digital human interaction, where the performance is more natural and realistic.

[0047] It also includes dynamic modeling of the feature sequence based on the temporal Transformer, specifically: Set the feature sequence , representing the k-th feature vector; Embed the standardized feature sequence into the feature sequence and input it into the encoding layer of the temporal Transformer; For each time step t, the feature representation is , where represents the number of modalities, represents the feature vector of the m-th modality at time step ; Using the self-attention mechanism of the temporal Transformer, calculate the attention weight matrix between time steps and t : ; t represents another time step. In the self-attention mechanism of time series data, it is another time point related to the current time step . In the Transformer model, the calculation of the self-attention mechanism not only depends on the information of the current time step but also considers the information of other time steps. Therefore, t as a time step identifier indicates the relationship between the current time step and other time steps t when calculating the attention weights.

[0048] Among them, is a mathematical function used to convert each element in the input vector into a value between 0 and 1, and the sum of all values is 1. The formula is: ; Among them, is the exponential operation of the j-th input value , is the number of elements, is the sum of the exponents of all elements; The above represents the query vector, specifically as follows: ; Among them, is obtained by linearly transforming the feature vector with . is the query weight matrix, which is learned through optimization during the training of the model, is the time step of the input feature vector; The above represents the key vector, is the dimension of the key vector, represents the transpose of the key vector, specifically as follows: ; Among them, is obtained by linearly transforming the feature vector with . is the key weight matrix, similar to , is the weight matrix learned through the training process for feature transformation; According to the attention weight and the feature vector , calculate and output : ; Among them, is the final fused feature representation at time step , representing the unified feature vector generated after the attention weighting mechanism, which is used for subsequent feature decoding and classification judgment, is the value vector; Combine the corresponding to each time step to generate the feature sequence after dynamic modeling, represents the num th feature vector; The present invention utilizes a sequential Transformer to dynamically model multi-modal feature sequences, calculates the relationships between each time step through the self-attention mechanism, thereby capturing long-term dependencies and the associations of cross-modal features. At each time step, the feature vectors are weighted through the query, key, and value mechanism to generate a fused feature representation. This process enables the model to automatically adjust and optimize the attention to different modal data by learning the query weight matrix and the key weight matrix, enhancing the connection between features. This method effectively improves the expressive ability of multi-modal data, can better handle complex time series features, and provides a more accurate input for subsequent generation and classification tasks. In applications such as digital human interaction and emotion-driven generation, this technology can improve the accuracy and effect of data fusion and enhance the system's response ability to complex dynamic information.

[0049] It also includes multi-modal semantic fusion of the dynamically modeled features, specifically: Align the feature dimensions of different modalities through linear transformation of the dynamic feature sequence: ; ; where is the parameter matrix automatically learned during model training, is the bias of modality , represents the feature vector of the rd modality (e.g., speech, video, or gesture) at time step , and the aligned feature sequence is denoted as ; Semantically fuse the modal features using a weighting mechanism to calculate the fused feature : ; where the weight represents the importance of modality at time step , and is calculated by the following formula: ; where, is the semantic vector of modality , represents the transpose of ; Obtain the fused features of all time steps to get the semantically fused feature sequence ; For the semantically fused feature sequence , for each time step , the fused feature Introduce a synchronous mapping layer and map it to the synchronous feature space: ; Among them, The specific expression of is: ; Among them, represents the synchronous feature formed at time step t, is the synchronous mapping function, which is implemented through a fully connected layer and a normalization operation; A fully connected layer is a layer structure in a neural network, where each neuron is connected to all neurons in the previous layer. This layer is usually used to perform a weighted sum of the input features and pass them to the next layer, usually followed by an activation function (such as ReLU, Sigmoid, etc.). In the synchronous mapping function, the fully connected layer maps the input features to a new feature space. The normalization operation usually refers to processing the input data so that it has a unified scale and range. Common normalization operations include Batch Normalization, which normalizes the input for each batch so that its mean is 0 and variance is 1. Normalization helps to accelerate training and improve the stability of the model.

[0050] Combine the synchronous features of each time step to generate a time-synchronous feature sequence .

[0051] Through the multi-modal semantic fusion technology, the features after dynamic modeling are effectively integrated. First, linear transformation is used to align the feature dimensions of different modalities (such as speech, video, posture, etc.) and unify them into a common feature space. Through the parameter matrix automatically learned during the training process, as well as the weights and biases of the modalities, the modal features at each time step are further weighted and adjusted. The weighting mechanism calculates the fusion features using the semantic vectors of the modalities according to the importance of the modalities to ensure the collaborative effect and semantic consistency of different modal data within the time step. Then, the fusion features of all time steps are combined to form a feature sequence after semantic fusion, forming a rich and compact multi-modal representation. On this basis, for the fused feature sequence, a synchronous mapping layer is used to map it to a unified synchronous feature space to ensure the synchronization and consistency of different modal data in time. The synchronous mapping process is achieved through a fully connected layer and a normalization operation. The fully connected layer performs weighted summation on the features and passes them to the next layer, and the normalization operation normalizes the data, enhancing the stability and training efficiency of the model. Finally, through the generation of the time-synchronized feature sequence, accurate and highly consistent input features are provided for subsequent generation and classification tasks. This method can effectively improve the expressive ability of multi-modal data, enhance the model's processing ability for cross-modal data, provide more accurate and natural feature representations for applications such as digital human interaction and emotion-driven generation, and significantly improve the performance and reliability of multi-modal systems.

[0052] It also includes inputting the multi-modal synchronous features into an emotion-driven generation model, specifically: The multi-modal synchronous feature sequence Specifically: ; Input the elements of the multi-modal synchronous features into the initial layer of the emotion-driven generation model in the order of the time series; According to the multi-modal synchronous features and the current emotion label (for example, through the emotion label input by the user or the detected emotion state) combined, form a comprehensive input denoted as, ; Input into the emotion-driven generation model to generate a preliminary expression and posture sequence , that is: ; The emotion-driven generation model adopts a generative adversarial network architecture, including a generator and a discriminator ; Among them, represents the generator function; The emotion-driven generation model is different from traditional emotion generation networks or virtual human animation models. Its innovation points include: Driven by a multi-modal time-synchronized feature sequence instead of a single modality (such as speech); Introduce the user's personalized emotion state vector to support dynamic style control; The model structure can be generalized to a conditional generative adversarial network (Conditional GAN) or a Transformer-based generator structure, with stronger expression and control capabilities.

[0053] The training process and loss function design are as follows: The emotion-driven generation model adopts the architecture of a generative adversarial network (GAN), including a generator and a discriminator to improve the naturalness and emotion consistency of the generated content.

[0054] Generator output: ; The generator takes the fused features as input and generates a sequence of natural expressions and postures .

[0055] Discriminator output discriminant probability P: ; Used to determine whether the input is real data, where can be a real sample or .

[0056] Loss function design: Comprehensively use the following three types of losses: Adversarial loss (generator target) : ; Among them, represents the expected value of the generator input x, usually a symbol used in probability statistics; Discriminator loss : ; Among them, D(y) is the probability that the discriminator judges the real data y, represents the expected value of the real data , represents the expected value of the real data x ; Emotion consistency loss (optional, measuring the matching degree between the generated sequence and the emotion label) : ; Among them, represents the generated emotion prediction, CE represents the cross-entropy loss, and the cross-entropy loss is used to measure the difference between the generated data and the actual label; Reconstruction loss (for pose / expression accuracy control) : ; Among them, represents the real expression and pose sequence; Total optimization objective :

[0057] Among them is the weighting coefficient, which can be set by experimental parameter tuning.

[0058] According to the expression and pose sequence and the real expression and pose sequence , the generative adversarial network is iteratively optimized; The generator and discriminator play a game through continuous iteration. The iteration process is carried out alternately. In each iteration, the discriminator and generator are updated once, gradually improving the quality of the generated results. By regularly evaluating the quality of the generated samples, for example, through manual inspection or calculating some evaluation metrics such as FID score (Frechet Inception Distance) or IS score (Inception Score), if the quality of the generated samples steadily improves and reaches the expected standard, the iteration can be stopped.

[0059] Output the natural expression and pose sequence according to the optimization result .

[0060] Finally, through the optimization of GAN, the generator can output a natural expression and pose sequence that conforms to the emotion label and multi-modal features. Specifically, the process is as follows: After the optimization is completed, the generator generates a natural expression and pose sequence with emotion consistency according to the finally learned parameters , including the naturalness and emotion consistency optimized through generative adversarial training.

[0061] The generated expression and pose sequence as the final output can drive the digital human or virtual character to perform natural expressions and poses.

[0062] By inputting multi-modal synchronization features into an emotion-driven generation model, natural expression and gesture generation based on a generative adversarial network (GAN) is achieved. This method innovatively drives based on multi-modal time synchronization feature sequences instead of traditional single modalities (such as speech), and introduces a user's personalized emotional state vector to support dynamic style control. The emotion-driven generation model generates expression and gesture sequences highly consistent with emotion labels through adversarial training of a generator and a discriminator. During the training process, by combining adversarial loss, discriminator loss, emotion consistency loss, and reconstruction loss, it is ensured that the generated expressions and gestures are both natural and conform to emotion labels. Through continuous iterative optimization, the generator gradually improves the quality of the generated results, and the output expression and gesture sequences can accurately reflect the user's emotional state and multi-modal features. This method not only improves the interaction experience of digital humans or virtual characters, making them more emotionally resonant, but also provides strong support for personalized customization and style control, and is widely applied in fields such as virtual reality, emotion computing, and digital entertainment, significantly enhancing the naturalness and emotion consistency of the generated content.

[0063] It also includes, based on the expression and gesture sequences, performing real-time rendering and output through an edge computing device, and capturing user interaction data as feedback. Specifically: Transfer and to an edge computing tool for real-time rendering to generate the dynamic performance of the digital human, denoted as ; Real-time collect user interaction data with the digital human through multi-modal sensors; Perform time synchronization processing on the interaction data: ; Among them, is the user interaction data after time synchronization processing, is the synchronization operation; Transfer to a feedback processing module to generate feedback information, and adjust the generator according to the feedback information; When the adjustment of the generator by the feedback information terminates, optimize the expression and gesture sequences of the digital human through and the generator.

[0064] The rendering process involves pixel-level adjustment of each frame of expression and gesture to ensure that the performance is consistent with the emotion label and conforms to natural dynamic performance; For each moment , perform the following operations through an edge computing device:

[0065] Among them, is the rendered expression and pose sequence, is the final expression and pose sequence generated in the previous process, is the emotion label at each moment, represents the rendering operation; The generated expression and pose features (such as facial muscle control parameters, joint angles) are mapped to a three-dimensional space or a two-dimensional image space through a transformation matrix, and the emotion label is used to adjust the details of the expression and pose, and dynamic adjustment is performed through a parametric model or an interpolation method. For example, the emotion label may control the tension of the facial expression or adjust certain angles of the body pose. According to the emotion label at the current time step, the expression parameters and pose control parameters are dynamically adjusted.

[0066] Expression control (taking facial muscle AUs as an example) adjusts the emotion intensity in a linear interpolation manner:

[0067] where, represents the expression control parameter at time step , represents the neutral state parameter, represents the ideal parameter in the target emotion state, , and represents the interpolation weight mapped from the emotion label intensity; According to the three-dimensional model of, calculate the lighting and material effects to make the rendered image look realistic. Through the physically based rendering (PBR) method, simulate the optical phenomena of surface reflection and refraction.

[0068] Micro-surface reflection function (Cook-Torrance model) is:

[0069] where, represents the lighting direction; represents the viewing direction; represents the normal direction; , represents the half-angle vector; represents the normal distribution function; represents the Fresnel reflection coefficient; represents the geometric shadowing function.

[0070] The final pixel color :

[0071] where is the incident light intensity.

[0072] According to the camera's perspective and the scene settings, adjust the display angle of the final image, perform synthesis, and generate a displayable image or animation frame.

[0073] Project the rendered model onto the two-dimensional image coordinate system under the viewing perspective to generate an image frame. The projection transformation used is as follows:

[0074] where, represents the three-dimensional point coordinates; represents the model transformation matrix; represents the camera viewing matrix; represents the perspective projection or orthographic projection matrix; represents the pixel position in the image.

[0075] Finally, at each time step the rendering result is expressed as:

[0076] where represents the above complete rendering function.

[0077] The feedback signal mainly comes from the user's emotional expression, speech content, and gesture changes. The goal is to optimize the consistency between the performance of the digital human and the user's emotions.

[0078] The system uses the feedback signal by comparing the interaction data with the user to adjust the parameters in the generation model

[0079] and the adjustment coefficient during the rendering process. The optimization goal is as follows: is the optimization loss, is the feedback loss based on the user interaction data, is the adjustment coefficient, is the probability output by the discriminator.

[0080] is used to adjust and the relative contributions of. For example, if the optimization loss is crucial for the model performance, while the feedback loss is only used to enhance the interaction quality, then will be set to a smaller value, and vice versa. The value of is usually determined by experimental tuning. A common approach is to perform grid search, random search, or use cross-validation and other methods to test different values to select the value that best suits the current task and data.

[0081] represents the discriminator's evaluation of the time step The generated and rendered sequence of expressions and postures The authenticity score, which is usually the probability output by a binary classification model.

[0082] Process: Input generated sample: The generated sequence of expressions and postures is used as the input to the discriminator.

[0083] Feature extraction: The discriminator extracts features from the input and may convert the input image or sequence data into a feature representation through a convolutional neural network (CNN) or other feature extraction methods.

[0084] Discrimination process: Based on the input features, the discriminator processes them through fully connected layers or other neural network layers and finally outputs a scalar value, which represents the probability that the input sample is "real". Usually, the sigmoid activation function is used to output the probability:

[0085] where is the sigmoid function, and its mathematical expression is as follows:

[0086] is the discriminator network, is the input generated sample, and the output value ranges from 0 to 1. A value close to 1 indicates that the sample is more real, and a value close to 0 indicates that the sample is not real.

[0087] According to the optimization results, the parameters in the generator model and the rendering process are adjusted to generate a sequence of expressions and postures consistent with user interaction. The optimized sequence of expressions and postures is represented by the following formula:

[0088] where is the optimized generated sequence of expressions and postures, : The finally generated sequence of expressions and postures, : Time step The emotional label at time step is the optimized generator parameter.

[0089] Finally, the optimized sequence of expressions and postures is output to the edge computing device to drive the digital human or virtual character to present a natural performance consistent with the user's emotions and interactions.

[0090] Real - time rendering and output through edge - computing devices enable the digital human to present natural and expressive dynamic expressions and pose sequences according to the user's emotions and interactions. Through multi - modal sensors, the interaction data between the user and the digital human is collected in real - time. The system performs time - synchronization processing on this data to ensure that the interaction data at each moment is consistent with the digital human's performance. The rendering process involves pixel - level adjustment of each frame of expressions and poses to ensure that the performance matches the emotion labels, and uses transformation matrices to map facial muscle control parameters, joint angles, etc. to three - dimensional space or two - dimensional image space, while making detailed adjustments in combination with emotion labels. The emotional intensity of expressions and poses is adjusted through linear interpolation, and physically - based rendering (PBR) methods are used to simulate lighting and material effects to make the rendered images more realistic. Finally, the generated expression and pose sequences are displayed under the viewing perspective through projection transformation to generate displayable images or animation frames. At the same time, the system captures the user's emotional expressions, speech content, and pose changes as feedback, and real - time optimizes the adjustment coefficients in the generator and rendering process to ensure that the digital human's expressions and poses are consistent with the user's emotional state. The optimization target is adjusted through feedback signals, where the feedback signals are based on user interaction data to help the system optimize the parameters of the generation model and enhance emotional consistency with the user. The discriminator further promotes the optimization of the generator and rendering process by evaluating the authenticity of the rendered expressions and poses. Finally, the parameters of the generator and rendering process are continuously adjusted to output natural expression and pose sequences that are highly consistent with the user's emotions and interactions. This method can not only significantly improve the interaction experience between the digital human and the user, making the virtual character's performance more emotionally resonant, but also provide stronger accuracy and flexibility for real - time emotion - driven generation and character performance in virtual interaction scenarios, and is widely applied in fields such as virtual reality, augmented reality, digital entertainment, and emotion computing.

[0091] It also includes inputting the user interaction data to a continuous learning module to iteratively optimize the emotion - driven generation model, specifically: Input to the continuous learning module to optimize the emotion - driven generation model; Input and perform time - synchronization processing, and denote the processed as ; After completing the rendering and feedback steps, the user interaction data is input to the continuous learning module. At this time, the interaction data is organized and processed into a form suitable for model update, including speech, facial expressions, and pose modal features.

[0092] The user interaction data and the current expression and pose sequences Perform time synchronization processing. The steps of the synchronization processing include: First, extract user interaction data and the timestamps of the expression and gesture sequences and align them to a unified time reference . Then, use interpolation methods (such as linear interpolation) or resampling methods to synchronize data with different sampling frequencies so that they correspond at the same time step . Specifically, linear interpolation fills in the missing time step data by calculating the weighted average between adjacent data points. For known data points and , if the time point to be interpolated is , then the calculation formula for linear interpolation is:

[0093] Here, and are the known time points, and are the corresponding values, is the time point to be interpolated, is the interpolation result. Through this method, the lower-frequency data (such as user interaction data) can be interpolated to the time steps corresponding to the higher-frequency data (such as the rendered expression and gesture sequences) to ensure that the two are aligned at the same time step .

[0094] After synchronization is completed, and are concatenated or fused to generate multi-modal input features , and used as the input of the model for further training and optimization to improve the accuracy and consistency of the generated virtual character performance and user interaction, represents the synchronized user interaction data.

[0095] According to train the emotion-driven generation model and update the parameters of the generator , specifically as follows: ; wherein, are the parameters before the generator update, represents the parameters after the generator update; is the learning rate, which controls the step size of weight update; is the generation error loss function with respect to gradient; Generate an error loss function, which represents the error of the generation model on user interaction data:

[0096] Where: is the sequence of expressions and postures output by the generation model; is the target output (e.g., real user interaction data); is the number of time steps (total time steps of the training data); Use the parameters of the updated generator Optimize ; The optimized generated expression and posture sequence is represented by the following formula: ; Where, represents the optimized , represents the optimization operation; Optimized will be fed back to the continuous learning module for the input of the next iteration. This process forms a closed loop to ensure that the digital human or virtual character can continuously optimize its performance according to the user's interaction data.

[0097] Iteratively optimize the emotion-driven generation model through the continuous learning module to improve the performance consistency and accuracy of the virtual character in user interaction. After rendering and feedback, the system inputs the user interaction data (such as voice, facial expressions, and posture modal features) into the continuous learning module and performs time synchronization processing. Specifically, the interaction data and the expression and posture sequences are aligned by timestamps, and data with different sampling frequencies are synchronized to the same time step using linear interpolation or resampling methods to ensure the precise alignment of all modal data. After synchronization, the interaction data is fused with the expression and posture features into multi-modal input features for further training and optimization of the emotion-driven generation model. By updating the generator parameters, the system can continuously optimize the performance of the generated virtual character based on the user interaction data, thereby achieving more natural and realistic interactions. This method effectively enhances the emotional consistency and dynamic performance between the virtual character and the user, making the generated expression and posture sequences more in line with the user's emotional state and interaction needs, and is widely applicable to fields such as digital humans, virtual reality, and augmented reality, significantly improving the user experience and interaction quality.

[0098] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.

[0099] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A method for intelligent interaction and gesture expression synthesis of digital human based on multi-modal synchronization, characterized in that: include: Collect user voice signals through microphone array ; Capture user's facial expression image data through camera ; Acquire the user's body posture data through the posture sensor ; in, represents the time step, , and All data are in the form of time series; Respectively , and Based on a unified timestamp benchmark Perform alignment. The specific timestamp alignment process is as follows: Get separately , and The corresponding timestamp; If the voice signal and video image data If the corresponding timestamps are inconsistent, the timestamps are aligned through linear interpolation: Calculate adjacent data points by linear interpolation and The weighted average between fills the missing time step data. The linear interpolation calculation formula is as follows: in, and Represents adjacent time points, and Respectively and The corresponding value, Indicates the time point where interpolation is required. Represents the interpolation result; If the posture data Corresponding timestamps and speech signals or video image data If the corresponding timestamps are inconsistent, the timestamp alignment operation is performed through the time difference correction algorithm, as follows: Set the original timestamp of the pose data to ; Set the reference timestamp of the video data to ; Set the time difference correction algorithm as follows: ; in, Indicates the timestamp of the corrected posture data, represents the systematic offset, , and Indicates the corresponding and ; After the timestamp alignment operation , and , forming an initial multimodal dataset under a unified timestamp benchmark, denoted as ; right Annotate and construct a multimodal structured dataset containing timestamps and modality tags. ; The above modal tags specifically include , and , used to distinguish different types of data.

2. According to the method of intelligent interaction and gesture expression synthesis of digital human based on multi-modal synchronization according to claim 1, it is characterized in that: It also includes detecting and filtering abnormal signals in the multimodal structured data set, eliminating background interference through an adaptive noise filtering algorithm, and generating noise-filtered multimodal data, specifically: Marking anomalous signals in multimodal structured datasets; The adaptive noise filtering algorithm is used for filtering, and its transfer function for: ; in is the cut-off frequency, is the filter order, Indicates the signal frequency; Recombine the noise-filtered data to generate a filtered multimodal dataset .

3. According to the method of intelligent interaction and gesture expression synthesis of digital human based on multi-modal synchronization in claim 1, it is characterized in that: It also includes normalizing the noise-filtered multimodal data set based on a standardized parameter library to generate a standardized multimodal feature sequence, specifically: Read the normalized parameters of each mode from the standardized parameter library , and calculate the normalized formula: ; in, and are the mean and standard deviation, Represents a normalized multimodal dataset; Will Standardized data in and By time series Splicing to generate a standardized multimodal feature matrix : ; in, Represents the nth data in the time series; right Slice by row to generate a standardized feature sequence with a time step of T ; All standardized feature sequences Sort the time steps from earliest to latest; Determine the time step : ; in, Indicates the minimum collection frequency corresponding to different types of data.

4. The method for intelligent interaction and gesture expression synthesis of digital human based on multi-modal synchronization according to claim 1, characterized in that: It also includes dynamic modeling of feature sequences based on temporal Transformer, specifically: Set feature sequence , represents the kth eigenvector; Normalize the feature sequence Embedding feature sequence And input to the encoding layer of the temporal Transformer; For each time step t, the feature is expressed as ,in represents the number of modes, represents the mth mode at time step The eigenvector of Using the self-attention mechanism of the temporal Transformer, calculate the time step and t The attention weight matrix : ; in, Is a mathematical function that converts each element in the input vector into a value between 0 and 1, and the sum of all values ​​is 1. The formula is: ; in, is the jth input value The exponential operation of is the number of elements, is the exponential sum of all elements; Above Represents the query vector, as follows: ; in, By taking the feature vector and The linear transformation is obtained. To query the weight matrix, the model is learned through optimization during training. is the time step The input feature vector of Above represents the key vector, is the dimension of the key vector, represents the transposition of the key vector, as follows: ; in, is to transform the feature vector and The linear transformation is obtained. is the key weight matrix, and similar, It is the weight matrix learned through the training process and is used for feature transformation; According to the attention weight and the eigenvector , calculate and output : ; in, is the time step The final fusion feature representation represents the unified feature vector generated after the attention weighting mechanism, which is used for subsequent feature decoding and classification judgment. is a value vector; Each time step Corresponding Combine to generate a feature sequence after dynamic modeling , Indicates num feature vectors.

5. The method for intelligent interaction and gesture expression synthesis of digital human based on multi-modal synchronization according to claim 1, characterized in that: It also includes multimodal semantic fusion of dynamic modeling features, specifically: Through dynamic feature sequence Perform a linear transformation to align the feature dimensions of different modes: ; in is the parameter matrix automatically learned during model training. For modal The bias of Represents the time step In The feature vectors of the modalities, the aligned feature sequence is expressed as ; Use the weighting mechanism to semantically fuse the modal features and calculate the fused features : ; The weight Representing modality At time step The importance of is calculated by the following formula: ; in, For modal The semantic vector of express The transpose of Get the fusion features of all time steps and get the feature sequence after semantic fusion ; For the feature sequence after semantic fusion , each time step Fusion features Introduce the synchronous mapping layer and map it to the synchronous feature space: ; in, represents the synchronization feature formed at time step t, The specific expression is: ; in, It is a synchronous mapping function, which is implemented through a fully connected layer and a normalization operation; Combine the synchronization features of each time step to generate a time synchronization feature sequence .

6. The method for intelligent interaction and gesture expression synthesis of digital human based on multi-modal synchronization according to claim 1, characterized in that: It also includes inputting multimodal synchronization features into the emotion-driven generation model, specifically: The multimodal synchronization feature sequence Specifically: ; Synchronize multimodal features The elements in the graph are input into the initial layer of the emotion-driven generative model in a time series order; According to the multimodal synchronization characteristics and the current sentiment label Combined, the comprehensive input is recorded as, ; Will Connect to the emotion-driven generation model to generate preliminary expression and gesture sequences ,Right now: ; The emotion-driven generative model adopts a generative adversarial network architecture, including a generator With the discriminator ; in, represents a generator function; According to the sequence of facial expressions and the real expression pose sequence , iteratively optimize the generative adversarial network; Output natural expression and posture sequence based on the optimization results .

7. The method for intelligent interaction and gesture expression synthesis of digital human based on multi-modal synchronization according to claim 1, characterized in that: It also includes real-time rendering and output through edge computing devices based on the expression and gesture sequence, and capturing user interaction data as feedback, specifically: Will and Transmitted to the edge computing tool for real-time rendering, generating the dynamic performance of the digital human, recorded as ; Collect the interaction data between users and digital humans in real time through multimodal sensors; Perform time synchronization on interaction data: ; in, It is the user interaction data after time synchronization processing. For synchronous operation; Will Transmit to the feedback processing module, generate feedback information, and adjust the generator according to the feedback information; When the feedback information ends the adjustment of the generator, and generators to optimize the expression and posture sequences of digital humans.

8. The method for intelligent interaction and gesture expression synthesis of digital human based on multi-modal synchronization according to claim 1, characterized in that: It also includes inputting user interaction data into the continuous learning module to iteratively optimize the emotion-driven generation model, specifically: Will Input into the continuous learning module to optimize the emotion-driven generation model; Will and Perform time synchronization and process the processed Recorded as ; according to Train the emotion-driven generative model and update the parameters of the generator ; Use the updated generator's parameters optimization .

Citation Information

Patent Citations

  • Emotion analysis processing method, system and equipment for humanoid robot and medium

    CN119498844A

  • Interactive digital human generation method and system based on artificial intelligence

    CN119600159A

  • Method and system for collecting decision data by multi-modal large model driven intelligent agent

    CN119808006A

  • Dynamic and static detection data labeling method and system based on time key association

    CN119830229A

Cited By

  • Multi-modal driven virtual digital human face animation generation method and system

    CN120298559A

  • Digital human interaction method and system in native call process

    CN120567835A

  • Deep learning digital human live broadcast method and system

    CN120915973A

  • Digital human interaction control method and device based on multiple modes

    CN122389919A