AIGC-based Multimodal Digital Human Generation Method, System and Storage Medium
By enhancing and feature extraction of multimodal data, combining graph structure neural network and adaptive optimization, the multimodal inconsistency and adaptability problems in digital human generation are solved, personalized expression and stable performance are achieved, and the performance quality and user experience of digital humans are improved.
Patent Information
- Application Number
- CN202510280089.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The existing digital life generation technology has visual, voice and action in multimodal data processing, which is difficult to capture personality characteristics, lack of adaptability, and unstable performance in a long time series, which affects the authenticity and performance quality of digital people.
By enhancing and extracting multimodal character data, using graph structure neural network modeling and dynamic weight adjustment, combining target scene information for intelligent matching and behavioral command generation, and introducing behavioral bias recognition and adaptive optimization mechanisms to achieve coordination and personalized expression of multimodal digital human performance.
It improves the authenticity and uniqueness of digital people, enhances the performance ability in complex environments, ensures modal coordination and behavioral coherence, improves user interaction experience and immersion, and continuously changes the performance mode through adaptive optimization.
Smart Images

Figure CN119782828B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a multimodal digital human generation method, system and storage medium based on AIGC. Background Art
[0002] With the rapid development of artificial intelligence and computer graphics technology, digital humans have been widely used in entertainment, education, customer service and other fields. Traditional digital human generation methods mainly rely on pre-set animations and scripts, which makes it difficult to achieve natural and smooth multimodal interactions. In recent years, deep learning-based methods have made significant progress in the field of digital human generation, and can achieve the coordinated generation of vision, speech and action to a certain extent.
[0003] However, existing digital human generation technologies still have some significant limitations. First, the collaborative processing and fusion of multimodal data remains a huge challenge, especially in maintaining temporal consistency and semantic coordination between vision, speech, and action. For example, the generated digital human may have problems such as lip shape not matching speech, expression not coordinating with intonation, or action not matching context. Second, existing methods are difficult to effectively capture and express subtle personality characteristics of characters, such as specific facial expressions, unique voice rhythms, or personalized body language, resulting in the lack of uniqueness and realism of generated digital humans. In addition, for complex and changing scenes, the adaptability of existing technologies is insufficient, and it is difficult to dynamically adjust the behavior and performance of digital humans according to different social situations, cultural backgrounds, or emotional states, which limits the expressiveness and credibility of digital humans in real-world applications. Finally, when processing long-term multimodal data, existing technologies often have problems such as incoherent behavior and unstable emotional expression, which affects the overall performance quality of digital humans. Summary of the invention
[0004] The present application provides a method, system and storage medium for generating a multimodal digital human based on AIGC, which are used to improve the efficiency and accuracy of generating a multimodal digital human based on AIGC.
[0005] In a first aspect, the present application provides a multimodal digital human generation method based on AIGC. The multimodal digital human generation method based on AIGC includes: performing data augmentation processing on pre-acquired multimodal human data to obtain an augmented multimodal data set, and performing feature extraction on the augmented multimodal data set to obtain multimodal feature data; performing modeling and dynamic weight adjustment processing based on a graph-structured neural network on the multimodal feature data to obtain dynamic personality knowledge data; performing intelligent matching processing on preset target scenario information and the dynamic personality knowledge data to obtain contextualized behavior instruction data; generating a hierarchical action sequence, performing adaptive resolution optimization, and performing speech synthesis processing on a preset initial digital human according to the contextualized behavior instruction data to obtain multimodal digital human performance data; performing behavior deviation recognition on the multimodal digital human performance data to obtain behavior deviation data and performance feedback data, and performing adaptive optimization processing on the multimodal digital human performance data and the performance feedback data through the behavior deviation data to obtain a target digital human.
[0006] In combination with the first aspect, in a first implementation manner of the first aspect of the present application, the performing data augmentation processing on pre-acquired multimodal human data to obtain an augmented multimodal data set, and performing feature extraction on the augmented multimodal data set to obtain multimodal feature data includes: performing timestamp alignment processing on multimodal human data including video data, audio data, and motion capture data to obtain time-synchronized multimodal raw data, and performing high-dynamic range image synthesis processing on the synchronized video data in the time-synchronized multimodal raw data to obtain enhanced video data; performing spectral enhancement processing on the synchronized audio data in the time-synchronized multimodal raw data to obtain enhanced audio data, and performing joint trajectory smoothing processing on the synchronized motion capture data in the time-synchronized multimodal raw data to obtain enhanced motion data; performing facial feature point extraction processing on the enhanced video data to obtain facial key point data; performing Mel-frequency cepstral coefficient extraction processing on the enhanced audio data to obtain audio feature vectors; performing bone angle and speed calculation processing on the enhanced motion data to obtain a motion feature matrix; performing temporal alignment processing on the facial key point data, the audio feature vectors, and the motion feature matrix to obtain aligned feature data, and performing feature importance evaluation based on a random forest on the aligned feature data, and selecting the top K features sorted based on importance scores to obtain dimensionality-reduced feature data, where K is a positive integer greater than 1; performing feature scaling processing on the dimensionality-reduced feature data to obtain the multimodal feature data.
[0007] Combined with the first aspect, in the second implementation manner of the first aspect of the present application, the modeling and dynamic weight adjustment processing of the multi-modal feature data based on the graph-structured neural network to obtain dynamic personality knowledge data includes: performing personality feature node construction processing on the multi-modal feature data to obtain an initial personality feature graph, and performing inter-modal association weight initialization processing on the initial personality feature graph to obtain a weighted personality feature graph; performing cross-modal feature aggregation processing on the weighted personality feature graph to obtain aggregated personality data, and performing non-linear expression mapping processing on the aggregated personality data to obtain an expression feature vector; performing emotional residual connection processing on the expression feature vector to obtain enhanced emotional representation data, and performing multi-head personality attention processing on the enhanced emotional representation data to obtain attention-weighted personality features; performing scene adaptation pooling processing on the attention-weighted personality features to obtain hierarchical personality representation data, and performing global personality extraction processing on the hierarchical personality representation data to obtain a comprehensive personality vector; performing dynamic behavior weight update processing on the comprehensive personality vector to obtain an optimized personality structure; and performing iterative personality propagation processing on the optimized personality structure to obtain the dynamic personality knowledge data.
[0008] Combined with the first aspect, in the third implementation manner of the first aspect of the present application, the intelligent matching processing of the preset target scene information and the dynamic personality knowledge data to obtain situation-based behavior instruction data includes: performing multi-modal scene feature extraction processing on the target scene information to obtain a scene feature vector; performing cross-attention processing on the scene feature vector and the dynamic personality knowledge data to obtain personality-scene association data; performing situation-based behavior pattern generation processing on the personality-scene association data to obtain a candidate behavior sequence; performing filtering processing based on a preset behavior rule on the candidate behavior sequence to obtain a target behavior sequence that conforms to the preset behavior rule; performing emotional intensity parameterization processing on the target behavior sequence to obtain an emotion-parameterized behavior instruction; and performing timestamp alignment and interpolation smoothing processing on the emotion-parameterized behavior instruction to obtain the situation-based behavior instruction data.
[0009] Combined with the first aspect, in the fourth implementation manner of the first aspect of the present application, generating a hierarchical action sequence, optimizing the adaptive resolution, and performing speech synthesis processing on the preset initial digital human according to the context-aware behavior instruction data to obtain multi-modal digital human performance data, including: performing temporal decomposition processing on the context-aware behavior instruction data to obtain a single-frame behavior instruction sequence, and performing key frame extraction processing on the single-frame behavior instruction sequence to obtain key behavior node data; performing interpolation expansion processing on the key behavior node data to obtain complete action sequence data; performing skeleton mapping processing on the complete action sequence data to obtain initial skeleton animation data, and performing animation correction based on physical constraints on the initial skeleton animation data to obtain target skeleton animation data; performing skinning weight calculation processing on the target skeleton animation data to obtain deformed mesh data, and performing texture mapping processing on the deformed mesh data to obtain initial rendering data; performing adaptive resolution adjustment processing on the initial rendering data to obtain optimized rendering data; performing phoneme segmentation processing on the speech part in the context-aware behavior instruction data to obtain phoneme sequence data, and performing acoustic parameter generation processing on the phoneme sequence data and the optimized rendering data to obtain the multi-modal digital human performance data.
[0010] Combined with the first aspect, in the fifth implementation manner of the first aspect of the present application, the method for identifying behavioral deviations from the multi-modal digital human performance data to obtain behavioral deviation data and performance feedback data, and adaptively optimizing the multi-modal digital human performance data and the performance feedback data through the behavioral deviation data to obtain a target digital human includes: performing modal separation processing on the multi-modal digital human performance data to obtain visual component data, audio component data, and motion component data; performing time window segmentation processing on the visual component data, the audio component data, and the motion component data to obtain a multi-modal behavior segment set, where the multi-modal behavior segment set includes: visual behavior segments, audio behavior segments, and motion behavior segments; performing facial expression feature extraction processing on the visual behavior segments to obtain an expression feature sequence; performing phoneme segmentation processing on the audio behavior segments to obtain phoneme time series data; performing joint angle calculation processing on the motion behavior segments to obtain a bone motion sequence; performing time series alignment processing on the expression feature sequence, the phoneme time series data, and the bone motion sequence to obtain an aligned multi-modal sequence, and performing cross-modal consistency analysis on the aligned multi-modal sequence to obtain a modal coordination degree index; performing threshold segmentation processing on the modal coordination degree index to obtain a preliminary behavioral deviation mark; performing clustering analysis processing on the preliminary behavioral deviation mark to obtain the behavioral deviation data, and performing user experience mapping processing on the behavioral deviation data to obtain the performance feedback data; and adaptively optimizing the multi-modal digital human performance data and the performance feedback data through the behavioral deviation data to obtain a target digital human.
[0011] Combined with the first aspect, in the sixth implementation manner of the first aspect of the present application, the adaptive optimization processing of the multi-modal digital human performance data and the performance feedback data by the behavior deviation data to obtain the target digital human includes: performing importance scoring processing on the behavior deviation data to obtain a deviation priority list; performing semantic analysis processing on the performance feedback data to obtain a user preference feature vector, and performing weighted fusion processing on the deviation priority list and the user preference feature vector to obtain an adjustment target matrix; performing expression correction processing on the visual data in the multi-modal digital human performance data based on the adjustment target matrix to obtain a corrected expression sequence; performing pitch adjustment processing on the audio data in the multi-modal digital human performance data based on the adjustment target matrix to obtain adjusted speech data; performing posture fine-tuning processing on the action data in the multi-modal digital human performance data based on the adjustment target matrix to obtain an adjusted action sequence; performing temporal alignment processing on the corrected expression sequence, the adjusted speech data, and the adjusted action sequence to obtain preliminarily adjusted multi-modal data, and performing inter-modal correlation analysis processing on the preliminarily adjusted multi-modal data to obtain a correlation matrix; performing threshold screening processing on the correlation matrix to obtain modal coordination data, and performing multi-modal fusion processing on the modal coordination data to obtain the target digital human.
[0012] In a second aspect, the present application provides a multi-modal digital human generation system based on AIGC. The multi-modal digital human generation system based on AIGC includes:
[0013] An enhancement module, configured to perform data enhancement processing on pre-acquired multi-modal human data to obtain an enhanced multi-modal data set, and perform feature extraction on the enhanced multi-modal data set to obtain multi-modal feature data;
[0014] A processing module, configured to perform modeling and dynamic weight adjustment processing on the multi-modal feature data based on a graph structure neural network to obtain dynamic personality knowledge data;
[0015] A matching module, configured to perform intelligent matching processing on preset target scene information and the dynamic personality knowledge data to obtain contextualized behavior instruction data;
[0016] A generation module, configured to perform hierarchical action sequence generation, adaptive resolution optimization, and speech synthesis processing on a preset initial digital human according to the contextualized behavior instruction data to obtain multi-modal digital human performance data;
[0017] An identification module is used to identify behavioral deviations in the multi-modal digital human performance data, obtain behavioral deviation data and performance feedback data, and perform adaptive optimization processing on the multi-modal digital human performance data and the performance feedback data through the behavioral deviation data to obtain a target digital human.
[0018] In the third aspect of the present application, a computer-readable storage medium is provided. Instructions are stored in the computer-readable storage medium, and when it runs on a computer, it causes the computer to execute the above-mentioned AIGC-based multi-modal digital human generation method.
[0019] In the technical solution provided by the present application, through data augmentation processing and feature extraction of pre-acquired multi-modal human data, the quality and diversity of the original data are significantly improved, laying a solid foundation for subsequent personalized modeling. This not only enhances the realism of the generated digital human but also greatly improves the generalization ability of the model. Secondly, by adopting a modeling and dynamic weight adjustment processing method based on a graph-structured neural network, the accurate capture and dynamic expression of complex personality characteristics are realized, enabling the generated digital human to exhibit rich and diverse personalized characteristics, greatly enhancing the uniqueness and realism of the digital human. Furthermore, through intelligent matching processing of target scene information and dynamic personality knowledge data, this method can dynamically generate adaptive behavior instructions according to different scene requirements, greatly improving the performance ability and naturalness of the digital human in a complex and changing environment. In addition, when generating multi-modal digital human performance data, this method adopts advanced technologies such as hierarchical action sequence generation, adaptive resolution optimization, and speech synthesis, ensuring a high degree of coordination between the visual, action, and speech modalities and effectively solving the common modality inconsistency problem in traditional methods. It is worth noting that this method also introduces a behavioral deviation identification and adaptive optimization mechanism, which can detect and correct abnormal behaviors in the digital human performance in real time and make dynamic adjustments according to user feedback. This not only improves the performance quality of the digital human but also continuously optimizes and improves the behavior pattern of the digital human. At the same time, the modular design of this method makes the entire generation process highly flexible and scalable, and can be customized and optimized according to the needs of different application scenarios. In addition, through the deep fusion and collaborative processing of multi-modal data, this method can generate a more natural, fluent, and expressive digital human, greatly enhancing the user's interaction experience and immersion. It is worth mentioning that the adaptive optimization mechanism of this method can not only improve the performance quality of the digital human but also continuously learn and evolve according to long-term interaction data, making the behavior and performance of the digital human more and more close to real humans. In addition, this method performs well in processing long-time sequence multi-modal data, can maintain the coherence of behavior and the stability of emotional expression, solves the common long-term performance instability problem in the prior art, and improves the efficiency and accuracy of AIGC-based multi-modal digital human generation. Brief Description of the Drawings
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0021] Figure 1 It is a schematic diagram of an embodiment of the multi-modal digital human generation method based on AIGC in the embodiments of the present application;
[0022] Figure 2 It is a schematic diagram of an embodiment of the multi-modal digital human generation system based on AIGC in the embodiments of the present application. Detailed Embodiments
[0023] The embodiments of the present application provide a multi-modal digital human generation method, system and storage medium based on AIGC. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "include" or "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0024] For ease of understanding, the following describes the specific process of the embodiments of the present application. Please refer to Figure 1 An embodiment of the multi-modal digital human generation method based on AIGC in the embodiments of the present application includes:
[0025] Step S101: Perform data enhancement processing on the pre-acquired multi-modal character data to obtain an enhanced multi-modal data set, and perform feature extraction on the enhanced multi-modal data set to obtain multi-modal feature data;
[0026] It can be understood that the execution subject of the present application can be a multi-modal digital human generation system based on AIGC, or a terminal or a server. Specifically, it is not limited here. The embodiments of the present application will be described by taking the server as the execution subject as an example.
[0027] Specifically, the multi-modal human data includes video data, audio data, and motion capture data, which are simultaneously collected by high-definition cameras, high-fidelity microphones, and motion capture devices. To ensure data synchronization, timestamp alignment processing is performed on these data, and the linear interpolation method is used to align the data of different modalities to a unified timeline. Data augmentation processing is performed on the time-synchronized multi-modal raw data. For video data, high dynamic range image synthesis technology is adopted. By merging images with different exposure times, the dynamic range of the image is extended, and the detail expressiveness of the image is improved. Specifically, the multi-exposure fusion algorithm is used to fuse the image sequences with short, medium, and long exposures to obtain enhanced video data with a wider dynamic range. For audio data, spectral enhancement processing technology is adopted. By adjusting the spectral characteristics of the audio signal, the clarity and recognizability of the audio are improved. In this process, a Wiener filter is used to filter the original audio signal according to the estimated noise spectrum and signal spectrum to obtain enhanced audio data. For motion capture data, joint trajectory smoothing processing is adopted. The Savitzky-Golay filter is used to smooth the original motion data to reduce the influence of noise and obtain more smooth and natural enhanced motion data.
[0028] After data augmentation is completed, feature extraction is performed on the enhanced multi-modal data set. Facial feature point extraction is performed on the enhanced video data. Using the OpenFace algorithm, 68 facial key points are extracted to obtain facial key point data. Mel-frequency cepstral coefficients (MFCCs) are extracted from the enhanced audio data. Through short-time Fourier transform, Mel filter bank, and discrete cosine transform, a 13-dimensional audio feature vector is obtained. The bone angles and speeds are calculated for the enhanced motion data. Quaternions are used to represent joint rotations, and the relative angles and angular velocities between adjacent joints are calculated to obtain a motion feature matrix. Subsequently, temporal alignment processing is performed on the facial key point data, audio feature vector, and motion feature matrix to ensure the consistency of different modality features in the time dimension. A feature importance evaluation method based on random forest is adopted to select the top K features with the highest importance scores for feature dimensionality reduction. Specifically, by constructing multiple decision trees, the contribution degree of each feature to the model prediction accuracy is calculated, and the top K features with the highest contribution degree are selected. Feature scaling processing is performed on the dimension-reduced feature data. The Min-Max normalization method is used to map the feature values to the [0, 1] interval to obtain the final multi-modal feature data.
[0029] For example, assume that the originally captured video data has a resolution of 1080p at 30fps, the audio data is mono audio with a sampling rate of 44.1kHz, and the motion capture data is 42 key-point data at 60fps. Through timestamp alignment processing, all data is unified to a sampling rate of 60fps. High-dynamic range image synthesis is performed on the video data, expanding the original 8-bit color depth to 16-bit color depth, effectively improving the detail performance of the image. Spectral enhancement processing is performed on the audio data, and after using the Wiener filter, the signal-to-noise ratio of the speech is increased by approximately 5dB. Savitzky-Golay filtering is performed on the motion data, using a 5th-order polynomial and a window size of 11 frames, effectively smoothing the motion trajectory. In the feature extraction stage, 68 facial key points are extracted from the video, 13-dimensional MFCC features are extracted from the audio, and 126-dimensional angle and velocity features (42 key points, 3 dimensions for each point) are calculated from the motion data. After random forest feature selection, the top 100 most important features are selected from the original 207-dimensional features (68 + 13 + 126). Finally, through Min-Max normalization, all feature values are mapped to the [0, 1] interval to obtain the final 100-dimensional multi-modal feature data.
[0030] Step S102: Perform modeling and dynamic weight adjustment processing on the multi-modal feature data based on a graph-structured neural network to obtain dynamic personality knowledge data;
[0031] Specifically, in the multi-modal digital human generation method based on AIGC, performing modeling and dynamic weight adjustment processing on the multi-modal feature data is a key step, aiming to capture and express the personality characteristics of the digital human. First, perform personality feature node construction processing on the multi-modal feature data, regarding each feature as a node in the graph to construct an initial personality feature graph. Specifically, for the 100-dimensional multi-modal feature data, each dimension is regarded as a node, forming a graph structure containing 100 nodes. Then, perform inter-modal correlation weight initialization processing on the initial personality feature graph, setting the weights of the edges according to the correlation between the features to obtain a weighted personality feature graph. The initial weights are determined by calculating the Pearson correlation coefficient between the features. The greater the absolute value of the correlation coefficient, the greater the corresponding edge weight.
[0032] Subsequently, perform cross-modal feature aggregation processing on the weighted personality feature graph, using a graph convolutional network (GCN) to aggregate the node features. The aggregation operation of GCN can be expressed as:
[0033] ;
[0034] where is the feature of node at the +1 layer, is the node the neighbor set is the normalization constant is the weight matrix of the th layer, and is the activation function. Through multiple layers of GCN, the aggregated personality data is obtained. The aggregated personality data is processed by non-linear expression mapping, and a multi-layer perceptron (MLP) is used to map the aggregated features to the expression space to obtain the expression feature vector. Each layer of the MLP can be expressed as:
[0035] ; where is the weight matrix is the bias vector is the activation function. The expression feature vector is processed by emotional residual connection, and the original feature is added to the mapped feature to obtain the enhanced emotional representation data. The residual connection can be expressed as:
[0036] F(x) = H(x) + x;
[0037] where H(x) is the mapping function and x is the input feature. The enhanced emotional representation data is processed by multi-head personality attention, and the multi-head self-attention mechanism is used to capture the long-range dependence relationship between features to obtain the attention-weighted personality features. The calculation of multi-head attention can be expressed as:
[0038] ;
[0039] where
[0040] where Q, K, and V represent the query, key, and value matrices respectively. In this solution, they all originate from the same input sequence, that is, the self-attention mechanism.
[0041] is the number of attention heads represents the th attention head and is calculated as follows:
[0042] ;
[0043] where: Concat is the concatenation function used to concatenate the outputs of all attention heads; is the output linear transformation matrix used to map the concatenated vector to the final output space; is the query weight matrix of the th head; is the key weight matrix of the th head; is the The value weight matrix of the head; Attention is the attention calculation function. This multi-head attention mechanism allows the model to simultaneously focus on information in different subspaces, enhancing the ability to capture complex personality features.
[0044] Then, perform scene adaptation pooling on the attention-weighted personality features, and use graph pooling operations to reduce the complexity of the graph, obtaining hierarchical personality representation data. Perform global personality extraction on the hierarchical personality representation data, and use graph readout operations to convert graph-level features into vector representations, obtaining a comprehensive personality vector. The graph readout uses global average pooling:
[0045] ;
[0046] where N is the number of nodes, is the node feature.
[0047] Finally, perform dynamic behavior weight update on the comprehensive personality vector, dynamically adjust the importance of features according to the current task and context, and obtain an optimized personality structure. The weight update uses a gated update mechanism:
[0048] ;
[0049] where is the current feature, is the context information, is the gate value, represents elementwise multiplication. Perform iterative personality propagation on the optimized personality structure, and obtain the final dynamic personality knowledge data by propagating personality features multiple times.
[0050] For example, assume that the initial multi-modal feature data is a 100-dimensional vector, including 30-dimensional visual features, 20-dimensional audio features, and 50-dimensional action features. First, construct an initial personality feature graph containing 100 nodes. By calculating the Pearson correlation coefficient between features, a weighted personality feature graph is obtained. For example, the correlation coefficient between the visual feature "eyebrow height" and the audio feature "pitch" is 0.7, and the weight of the corresponding edge is set to 0.7. Use 3 layers of GCN for feature aggregation, and the output dimensions of each layer are 64, 32, and 16 respectively. In the first layer of GCN, for node , assume it has 5 neighbor nodes and the initial feature dimension is 100, then its update process is:
[0051] ;
[0052] where It is a 64×100 matrix. The non-linear expression mapping uses a 2-layer MLP, with a 16-dimensional input and an 8-dimensional expression feature vector output. After the emotional residual connection, the feature dimension remains 16. The multi-head attention uses 4 attention heads, each with a dimension of 4. The scene adaptability pooling reduces the number of nodes from 100 to 50. The global personality extraction obtains a 16-dimensional comprehensive personality vector. The dynamic behavior weight update assigns a weight between 0 and 1 to each dimension in the 16-dimensional vector according to the current scene information. Finally, through 5 iterative propagations, stable 16-dimensional dynamic personality knowledge data is obtained. The whole process not only captures the complex relationships between multi-modal features but also realizes the dynamic adjustment of personality features, providing rich personalized information for subsequent digital life generation.
[0053] Step S103: Perform intelligent matching processing on the preset target scene information and dynamic personality knowledge data to obtain contextualized behavior instruction data;
[0054] Specifically, performing intelligent matching processing on the preset target scene information and dynamic personality knowledge data to obtain contextualized behavior instruction data is a key step in the AIGC-based multi-modal digital life generation method. Perform multi-modal scene feature extraction processing on the target scene information to obtain a scene feature vector. This process uses a pre-trained multi-modal encoder, such as the CLIP model, to encode the image and text description of the scene into a unified feature space. Assume the dimension of the scene feature vector is 128, denoted as .
[0055] Next, perform cross-attention processing on the scene feature vector and dynamic personality knowledge data to obtain personality-scene association data. The calculation formula of the cross-attention mechanism is as follows:
[0056] ;
[0057] where , , , is the scene feature vector, is the dynamic personality knowledge data, is the value matrix mentioned above, , , are learnable weight matrices, is the dimension of the key vector. The output A of this step is the personality-scene association data, with the same dimension as the dynamic personality knowledge data.
[0058] Then, perform contextualized behavior pattern generation processing on the personality-scenario association data to obtain candidate behavior sequences. This step uses a recurrent neural network (RNN) or a Transformer decoder to generate the behavior sequences. The generation process can be expressed as:
[0059] ;
[0060] where is the hidden state of the RNN, is the representation of the personality-scenario association data at time step t, is the state update function of the RNN, g is the output function, is the generated candidate behavior. Perform predefined behavior rule filtering processing on the candidate behavior sequences to obtain behavior data that conforms to the rules. This step uses a predefined rule set ;
[0061] Filter the generated behaviors. Each rule is a boolean function that takes a behavior y as input and outputs whether it conforms to the rule. The filtering process can be expressed as:
[0062]
[0063] Perform emotional intensity parameterization processing on the behavior data that conforms to the rules to obtain emotionally parameterized behavior instructions. This step uses an emotion analysis model E to score each behavior y to obtain an emotional intensity vector e:
[0064] ;
[0065] where is the number of emotional dimensions (e.g., pleasantness, activation, etc.).
[0066] Finally, perform timestamp alignment and interpolation smoothing processing on the emotionally parameterized behavior instructions to obtain contextualized behavior instruction data. Timestamp alignment ensures that the behavior instructions are consistent with the target scenario timeline, and interpolation smoothing uses linear interpolation methods to fill the gaps between behavior instructions to ensure the continuity of behaviors.
[0067] For example, assume there is a social scenario where the goal is to generate the behavior of a digital human at a party. The scenario feature vector S is encoded multimodally to obtain a 128-dimensional vector. The dynamic personality knowledge data P is a 16-dimensional vector representing the personality characteristics of the digital human. Through cross-attention processing, 16-dimensional personality-scenario association data A is obtained. The Transformer decoder is used to generate candidate behavior sequences such as "walking towards the bar", "talking to others", "dancing", etc. The preset behavior rules include "not committing illegal acts", "being polite", etc., and behaviors that do not conform to the rules are filtered out. The remaining behaviors are parameterized with emotional intensity. For example, the emotional vector of "dancing" may be [0.8, 0.7], indicating a relatively high degree of pleasure and activation. Finally, these behavior instructions are aligned to the time axis of the scenario, such as "walking towards the bar at 19:30:00", "talking to others at 19:35:00", "dancing at 20:00:00", and linear interpolation is performed between these discrete time points to obtain continuous contextualized behavior instruction data. This process ensures that the generated digital human behavior not only meets the scenario requirements, but also reflects the personality characteristics, while maintaining the continuity and naturalness of the behavior.
[0068] Step S104: Generate a hierarchical action sequence, perform adaptive resolution optimization, and speech synthesis processing on the preset initial digital human according to the contextualized behavior instruction data to obtain multimodal digital human performance data;
[0069] Step S105: Identify behavior deviations from the multimodal digital human performance data to obtain behavior deviation data and performance feedback data, and perform adaptive optimization processing on the multimodal digital human performance data and performance feedback data through the behavior deviation data to obtain the target digital human.
[0070] Specifically, generating a hierarchical action sequence, optimizing the adaptive resolution, and performing speech synthesis processing on the pre-set initial digital human according to the contextual behavior instruction data are key steps. First, perform temporal decomposition processing on the contextual behavior instruction data to obtain a single-frame behavior instruction sequence. This process discretizes the continuous behavior instructions into specific instructions at a series of time points. Then, perform key frame extraction processing on the single-frame behavior instruction sequence to obtain key behavior node data. This step identifies the critical moments of behavior changes and reduces redundant information. Next, perform interpolation expansion processing on the key behavior node data to obtain a complete action sequence data. The actions between key frames are filled through the interpolation algorithm to ensure the continuity and smoothness of the actions. Perform bone mapping processing on the complete action sequence data to obtain initial bone animation data. This step converts the abstract action instructions into specific bone motion data. Perform animation correction based on physical constraints on the initial bone animation data to obtain target bone animation data. This process introduces physical simulation to make the actions more realistic and natural. Next, perform skinning weight calculation processing on the target bone animation data to obtain deformed mesh data. This step applies the bone animation to the 3D model to achieve the deformation of the model. Then, perform texture mapping processing on the deformed mesh data to obtain initial rendering data, adding surface details and materials to the 3D model.
[0071] Perform adaptive resolution adjustment processing on the initial rendering data to obtain optimized rendering data. This step dynamically adjusts the rendering quality according to different display devices and network conditions. At the same time, perform phoneme segmentation processing on the speech part of the contextual behavior instruction data to obtain phoneme sequence data, preparing for subsequent speech synthesis. Finally, perform acoustic parameter generation processing on the phoneme sequence data and the optimized rendering data to obtain multi-modal digital human performance data. This step generates speech data synchronized with the actions to complete the multi-modal performance of the digital human. After obtaining the multi-modal digital human performance data, perform behavior deviation identification to obtain behavior deviation data and performance feedback data. This process first performs modal separation processing on the multi-modal digital human performance data to obtain visual component data, audio component data, and action component data. Then, perform time window slicing processing on these component data to obtain a multi-modal behavior fragment set. Perform facial expression feature extraction processing on the visual behavior fragments to obtain an expression feature sequence; perform phoneme segmentation processing on the audio behavior fragments to obtain phoneme time series data; perform joint angle calculation processing on the action behavior fragments to obtain a bone motion sequence.
[0072] Perform temporal alignment processing on the expression feature sequence, phoneme temporal data, and skeletal motion sequence to obtain an aligned multi-modal sequence, and perform cross-modal consistency analysis on the aligned multi-modal sequence to obtain a modal coordination degree index. Perform threshold segmentation processing on the modal coordination degree index to obtain a preliminary behavior deviation label, and then perform clustering analysis processing on the preliminary behavior deviation label to obtain behavior deviation data. Finally, perform user experience mapping processing on the behavior deviation data to obtain performance feedback data. Perform adaptive optimization processing on the multi-modal digital human performance data and performance feedback data through the behavior deviation data to obtain a target digital human. This process includes performing importance scoring processing on the behavior deviation data to obtain a deviation priority list; performing semantic analysis processing on the performance feedback data to obtain a user preference feature vector; performing weighted fusion processing on the deviation priority list and the user preference feature vector to obtain an adjustment target matrix. Then, perform correction processing on the visual data, audio data, and action data in the multi-modal digital human performance data according to the adjustment target matrix to obtain a corrected expression sequence, adjusted speech data, and adjusted action sequence. Finally, perform temporal alignment processing on these corrected data to obtain preliminary adjusted multi-modal data, and perform inter-modal correlation analysis and threshold screening processing to finally obtain modal coordination data, and obtain the target digital human through multi-modal fusion processing.
[0073] For example, assume that the performance of a digital human in a social scenario needs to be generated and optimized. The initial contextualized behavior instruction data includes actions such as "smile", "wave", "speak", etc., with a duration of 10 seconds. Through temporal decomposition and key frame extraction, 5 key behavior nodes are obtained, such as "smile start" at 0 seconds, "wave start" at 2 seconds, "speak start" at 4 seconds, etc. After interpolation expansion, a complete action sequence of 300 frames (30fps) is generated. Skeletal mapping converts these actions into motion data of 20 key skeletal points. Physical constraint correction ensures that the actions conform to the physical laws of the real world, such as the inertial effect when the arm swings. Skinning weight calculation applies the skeletal animation to a 3D model containing 10,000 vertices. Texture mapping adds 4K resolution skin texture. Adaptive resolution optimization adjusts the rendering resolution to 1080p according to the target device. Speech synthesis generates 10 seconds of audio data that matches the actions.
[0074] In the behavioral deviation recognition stage, 10 seconds of data are divided into 20 time windows of 0.5 seconds each. The movement trajectories of 68 facial key points, the temporal data of 200 phonemes, and the angular changes of 20 skeletal points are extracted. Cross-modal consistency analysis reveals that at the 3rd second, the facial expression does not match the speech emotion, and the modal coordination index is lower than the threshold of 0.7. Cluster analysis classifies this deviation as the "emotion inconsistency" type. User experience mapping converts this deviation into a feedback score of -0.5 (range -1 to 1). In the adaptive optimization stage, based on the deviation data and user feedback, a 16x16 adjustment target matrix is generated. The expression near the 3rd second is corrected by increasing the upward angle of the corners of the mouth from 15 degrees to 20 degrees to match the pleasant emotion of the speech. At the same time, the speech pitch at the corresponding moment is adjusted by increasing the fundamental frequency from 220 Hz to 240 Hz. These corrections, through temporal alignment and multi-modal fusion, finally generate a target digital human performance with more coordinated behavior and meeting user expectations.
[0075] In the embodiments of the present application, through data augmentation processing and feature extraction on pre-acquired multimodal human data, the quality and diversity of the original data are significantly improved, laying a solid foundation for subsequent personalized modeling. This not only enhances the realism of the generated digital human but also greatly improves the generalization ability of the model. Secondly, by adopting a modeling and dynamic weight adjustment processing method based on a graph-structured neural network, the accurate capture and dynamic expression of complex personality characteristics are realized, enabling the generated digital human to exhibit rich and diverse personalized characteristics and greatly enhancing the uniqueness and realism of the digital human. Furthermore, through intelligent matching processing of target scene information and dynamic personality knowledge data, this method can dynamically generate adaptive behavior instructions according to different scene requirements, greatly improving the performance ability and naturalness of the digital human in complex and changing environments. In addition, when generating multimodal digital human performance data, this method adopts advanced technologies such as hierarchical action sequence generation, adaptive resolution optimization, and speech synthesis to ensure a high degree of coordination among the three modalities of vision, action, and speech, effectively solving the common modality inconsistency problem in traditional methods. It is worth noting that this method also introduces a behavior deviation identification and adaptive optimization mechanism, which can detect and correct abnormal behaviors in the digital human performance in real time and make dynamic adjustments according to user feedback. This not only improves the performance quality of the digital human but also continuously optimizes and improves the behavior pattern of the digital human. At the same time, the modular design of this method makes the entire generation process highly flexible and scalable, and can be customized and optimized according to the requirements of different application scenarios. In addition, through the deep fusion and collaborative processing of multimodal data, this method can generate a more natural, fluent, and expressive digital human, greatly enhancing the user's interaction experience and immersion. It is worth mentioning that the adaptive optimization mechanism of this method can not only improve the performance quality of the digital human but also continuously learn and evolve according to long-term interaction data, making the behavior and performance of the digital human more and more close to real humans. In addition, this method performs well in processing long-time sequence multimodal data, can maintain the coherence of behavior and the stability of emotional expression, solves the common long-term performance instability problem in the prior art, and improves the efficiency and accuracy of multimodal digital human generation based on AIGC.
[0076] In a specific embodiment, the process of executing step S101 may specifically include the following steps:
[0077] (1) Perform timestamp alignment processing on multimodal human data including video data, audio data, and motion capture data to obtain time-synchronized multimodal raw data, and perform high-dynamic range image synthesis processing on the synchronized video data in the time-synchronized multimodal raw data to obtain enhanced video data;
[0078] (2) Perform spectral enhancement processing on the synchronized audio data in the time-synchronized multimodal raw data to obtain enhanced audio data, and perform joint trajectory smoothing processing on the synchronized motion capture data in the time-synchronized multimodal raw data to obtain enhanced motion data;
[0079] (3) Perform facial feature point extraction processing on the enhanced video data to obtain facial key point data;
[0080] (4) Perform Mel-frequency cepstral coefficient extraction processing on the enhanced audio data to obtain audio feature vectors;
[0081] (5) Perform bone angle and speed calculation processing on the enhanced motion data to obtain a motion feature matrix;
[0082] (6) Perform temporal alignment processing on the facial key point data, audio feature vectors, and motion feature matrix to obtain aligned feature data, and perform feature importance evaluation based on random forest on the aligned feature data, and select the top K features sorted based on the importance scores to obtain dimensionality-reduced feature data, where K is a positive integer greater than 1;
[0083] (7) Perform feature scaling processing on the dimensionality-reduced feature data to obtain multimodal feature data.
[0084] Specifically, processing the multimodal human data containing video data, audio data, and motion capture data is a key initial step. First, perform timestamp alignment processing to synchronize data of different modalities to a unified timeline. This process uses linear interpolation method. According to the timestamps of each data point, map them to a common time series to obtain time-synchronized multimodal raw data.
[0085] For synchronized video data, adopt high dynamic range (HDR) image synthesis processing technology. This technology expands the dynamic range of the image and improves the detail expressiveness by merging images with different exposure times. The specific implementation uses the multi-exposure fusion algorithm, which can be expressed as:
[0086] ;
[0087] where, is the final HDR image, is the th image with different exposures, is the corresponding weight, is the number of images. The weight is usually calculated based on the brightness, contrast, and saturation of the image.
[0088] For synchronous audio data, spectral enhancement processing is performed. This process uses a Wiener filter to filter the original audio signal according to the estimated noise spectrum and signal spectrum. The frequency-domain expression of the Wiener filter is:
[0089] ;
[0090] where, is the frequency response of the filter, is the power spectrum of the signal, is the power spectrum of the noise. For synchronous motion capture data, a Savitzky-Golay filter is used for joint trajectory smoothing. This filter smooths the data through local polynomial fitting while preserving the higher-order moments of the signal. The filtered data points can be expressed as:
[0091] ;
[0092] where, is the smoothed data point, is the original data point, are the filter coefficients, and m is the half-width of the window. Next, facial feature point extraction processing is performed on the enhanced video data. This step uses an improved OpenFace algorithm to extract 68 standard facial key points to form facial key point data. Mel Frequency Cepstral Coefficient (MFCC) extraction processing is performed on the enhanced audio data. The calculation process of MFCC includes performing a short-time Fourier transform on the audio signal, mapping the frequency to the Mel scale, taking the logarithm, and performing a discrete cosine transform, finally obtaining a 13-dimensional audio feature vector.
[0093] Skeletal angle and speed calculation processing is performed on the enhanced motion data. Quaternions are used to represent joint rotations, and the relative angles and angular velocities between adjacent joints are calculated to form an action feature matrix. Then, temporal alignment processing is performed on the facial key point data, audio feature vector, and action feature matrix to ensure the consistency of different modality features in the time dimension, obtaining aligned feature data. Feature importance evaluation based on random forest is performed on the aligned feature data. The random forest algorithm constructs multiple decision trees and calculates the contribution degree of each feature to the model prediction accuracy. The feature importance score can be expressed as:
[0094] ;
[0095] where, is the importance score of feature , M is the number of decision trees, is the reduction in impurity of feature in decision tree . The top K features sorted by importance scores are selected to obtain the dimensionality-reduced feature data.
[0096] Finally, perform feature scaling on the dimensionality-reduced feature data. Use the Min-Max normalization method to map the feature values to the interval [0, 1] to obtain the final multi-modal feature data. The Min-Max normalization formula is:
[0097] ;
[0098] where is the normalized feature value, X is the original feature value, and are the minimum and maximum values of the feature respectively.
[0099] For example, assume that a 10-second multi-modal human data is collected, including 1080p video at 30fps, audio at a sampling rate of 44.1kHz, and motion capture data at 60fps. First, perform timestamp alignment to unify all data to a sampling rate of 60fps. Perform HDR processing on the video data to expand the original 8-bit color depth to 16-bit color depth, significantly improving the details in the dark and bright parts. Perform spectral enhancement on the audio data. After using the Wiener filter, the signal-to-noise ratio of the speech is increased from 15dB to 20dB. Smooth the motion data using the Savitzky-Golay filter with a window size of 11 frames, effectively reducing jitter. In the feature extraction stage, extract 68 facial key points from the video to obtain facial key point data of 68x2x600 (x coordinate and y coordinate, 600 frames). Extract 13-dimensional MFCC features from the audio to obtain an audio feature vector of 13x600. Calculate the angles and angular velocities of 20 key joints for the motion data to obtain an 80x600 motion feature matrix (4 dimensions for each joint: 3 angle components and 1 angular velocity). Evaluate the feature importance of these 161-dimensional features (68x2 + 13 + 80) using random forest with 1000 decision trees. Assume that K = 100 is set, and the top 100 most important features are selected. Finally, through Min-Max normalization, map all feature values to the interval [0, 1] to obtain the final 100x600 multi-modal feature data. This series of processes not only improves the data quality but also significantly reduces the data dimension, laying a foundation for subsequent personalized modeling.
[0100] In a specific embodiment, the process of executing step S102 may specifically include the following steps:
[0101] (1) Perform personalized feature node construction processing on the multi-modal feature data to obtain an initial personalized feature map, and perform initialization processing on the inter-modal association weights of the initial personalized feature map to obtain a weighted personalized feature map;
[0102] (2) Perform cross-modal feature aggregation processing on the weighted personality feature map to obtain aggregated personality data, and perform non-linear expression mapping processing on the aggregated personality data to obtain an expression feature vector;
[0103] (3) Perform emotional residual connection processing on the expression feature vector to obtain enhanced emotional representation data, and perform multi-head personality attention processing on the enhanced emotional representation data to obtain attention-weighted personality features;
[0104] (4) Perform scene adaptation pooling processing on the attention-weighted personality features to obtain hierarchical personality representation data, and perform global personality extraction processing on the hierarchical personality representation data to obtain a comprehensive personality vector;
[0105] (5) Perform dynamic behavior weight update processing on the comprehensive personality vector to obtain an optimized personality structure;
[0106] (6) Perform iterative personality propagation processing on the optimized personality structure to obtain dynamic personality knowledge data.
[0107] Specifically, personalized modeling of multi-modal feature data is a key step. First, perform personality feature node construction processing on the multi-modal feature data, regard each feature as a node in the graph, and construct an initial personality feature map. Assuming there is 100-dimensional feature data, a graph structure containing 100 nodes will be formed. Then, perform inter-modal association weight initialization processing on the initial personality feature map, and set the weights of the edges by calculating the Pearson correlation coefficient between the features to obtain a weighted personality feature map.
[0108] Perform cross-modal feature aggregation processing on the weighted personality feature map, and use a graph convolutional network (GCN) to aggregate the node features. The aggregation operation of GCN can be expressed as:
[0109] = ( );
[0110] Among them, is the node feature matrix of the layer, is the adjacency matrix with self-connection added, is the degree matrix, is the learnable weight matrix, is the activation function.
[0111] Perform non-linear expression mapping processing on the aggregated personality data, and use a multi-layer perceptron (MLP) to map the aggregated features to the expression space to obtain an expression feature vector. Each layer of MLP can be expressed as:
[0112] ; Among them is the weight matrix, is the bias vector, is the activation function. Perform emotional residual connection processing on the facial expression feature vector, add the original features and the mapped features to obtain enhanced emotional representation data. The residual connection can be expressed as:
[0113] F(x) = H(x) + x;
[0114] where H(x) is the mapping function and x is the input feature. Perform multi-head personality attention processing on the enhanced emotional representation data, and use the multi-head self-attention mechanism to capture the long-range dependencies between features to obtain attention-weighted personality features. The calculation of multi-head attention can be expressed as:
[0115] ;
[0116] where ;
[0117] where Q, K, and V represent the Query, Key, and Value matrices respectively. In this solution, they all originate from the same input sequence, that is, the self-attention mechanism.
[0118] is the number of attention heads, represents the th attention head, and the calculation is as follows:
[0119] ;
[0120] where: Concat is the concatenation function used to concatenate the outputs of all attention heads; is the output linear transformation matrix used to map the concatenated vector to the final output space; is the query weight matrix of the th head; is the key weight matrix of the th head; is the value weight matrix of the th head; Attention is the attention calculation function. This multi-head attention mechanism allows the model to simultaneously focus on information in different subspaces, enhancing the ability to capture complex personality features.
[0121] Then, perform scene adaptation pooling processing on the attention-weighted personality features, use graph pooling operations to reduce the complexity of the graph, and obtain hierarchical personality representation data. Perform global personality extraction processing on the hierarchical personality representation data, use graph readout operations to convert graph-level features into vector representations, and obtain comprehensive personality vectors. The graph readout uses global average pooling:
[0122] ;
[0123] where N is the number of nodes, is the node feature.
[0124] Finally, perform dynamic behavior weight update processing on the comprehensive personality vector, dynamically adjust the importance of features according to the current task and context, and obtain an optimized personality structure. The weight update uses a gated update mechanism:
[0125] ;
[0126] where is the current feature, is the context information, is the gate value, represents elementwise multiplication. Perform iterative personality propagation processing on the optimized personality structure, and obtain the final dynamic personality knowledge data by propagating personality features multiple times.
[0127] For example, assume that the initial multi-modal feature data is a 100-dimensional vector, including 30-dimensional visual features, 20-dimensional audio features, and 50-dimensional action features. First, construct an initial personality feature map containing 100 nodes. By calculating the Pearson correlation coefficient between features, a weighted personality feature map is obtained. For example, the correlation coefficient between the visual feature "eyebrow height" and the audio feature "pitch" is 0.7, and the weight of the corresponding edge is set to 0.7. Use 3-layer GCN for feature aggregation, and the output dimensions of each layer are 64, 32, and 16 respectively. In the first layer of GCN, for node i, assume it has 5 neighbor nodes and the initial feature dimension is 100, then its update process is:
[0128] ;
[0129] where is a 64×100 matrix. The non-linear expression mapping uses a 2-layer MLP, with an input of 16 dimensions and an output of 8-dimensional expression feature vectors. After emotional residual connection, the feature dimension remains 16. The multi-head attention uses 4 attention heads, and the dimension of each head is 4. The scene adaptability pooling reduces the number of nodes from 100 to 50. The global personality extraction obtains a 16-dimensional comprehensive personality vector. The dynamic behavior weight update assigns a weight between 0 and 1 to each dimension in the 16-dimensional vector according to the current scene information. Finally, through 5 iterations of propagation, stable 16-dimensional dynamic personality knowledge data is obtained. The whole process not only captures the complex relationships between multi-modal features, but also realizes the dynamic adjustment of personality features, providing rich personalized information for subsequent digital life generation.
[0130] In a specific embodiment, the process of executing step S103 may specifically include the following steps:
[0131] (1) Perform multi-modal scene feature extraction processing on the target scene information to obtain a scene feature vector;
[0132] (2) Perform cross-attention processing on the scene feature vector and the dynamic personality knowledge data to obtain personality-scene association data;
[0133] (3) Perform contextualized behavior pattern generation processing on the personality-scene association data to obtain a candidate behavior sequence;
[0134] (4) Perform filtering processing on the candidate behavior sequence based on preset behavior rules to obtain a target behavior sequence that conforms to the preset behavior rules;
[0135] (5) Perform emotional intensity parameterization processing on the target behavior sequence to obtain an emotionally parameterized behavior instruction;
[0136] (6) Perform timestamp alignment and interpolation smoothing processing on the emotionally parameterized behavior instruction to obtain contextualized behavior instruction data.
[0137] Specifically, performing intelligent matching processing on the preset target scene information and dynamic personality knowledge data to obtain contextualized behavior instruction data is a key step in the AIGC-based multi-modal digital human generation method. Performing multi-modal scene feature extraction processing on the target scene information to obtain a scene feature vector. This process uses a pre-trained multi-modal encoder, such as the CLIP model, to encode the image and text description of the scene into a unified feature space. Assuming the dimension of the scene feature vector is 128, denoted as .
[0138] Next, perform cross-attention processing on the scene feature vector and the dynamic personality knowledge data to obtain personality-scene association data. The calculation formula of the cross-attention mechanism is as follows:
[0139] ;
[0140] where, , , , is the scene feature vector, is the dynamic personality knowledge data, is the value matrix as described above, , , are learnable weight matrices, is the dimension of the key vector. The output A of this step is the personality-scene association data, with the same dimension as the dynamic personality knowledge data.
[0141] Then, perform contextualized behavior pattern generation processing on the personality-scenario association data to obtain candidate behavior sequences. This step uses a recurrent neural network (RNN) or a Transformer decoder to generate the behavior sequences. The generation process can be expressed as:
[0142] ;
[0143] where is the hidden state of the RNN, is the representation of the personality-scenario association data at time step t, is the state update function of the RNN, g is the output function, is the generated candidate behavior. Perform preset behavior rule filtering processing on the candidate behavior sequences to obtain behavior data that conforms to the rules. This step uses a predefined rule set ;
[0144] Filter the generated behaviors. Each rule is a boolean function that takes a behavior y as input and outputs whether it conforms to the rule. The filtering process can be expressed as:
[0145] ;
[0146] Perform emotional intensity parameterization processing on the behavior data that conforms to the rules to obtain emotion-parameterized behavior instructions. This step uses an emotion analysis model E to score each behavior y to obtain an emotional intensity vector e:
[0147] ;
[0148] where is the number of emotional dimensions (e.g., pleasantness, activation, etc.).
[0149] Finally, perform timestamp alignment and interpolation smoothing processing on the emotion-parameterized behavior instructions to obtain contextualized behavior instruction data. Timestamp alignment ensures that the behavior instructions are consistent with the target scenario timeline, and interpolation smoothing uses linear interpolation methods to fill the gaps between behavior instructions to ensure the continuity of behaviors.
[0150] For example, suppose there is a social scenario where the goal is to generate the behavior of a digital human at a party. The scene feature vector S is encoded multimodally to obtain a 128-dimensional vector. The dynamic personality knowledge data P is a 16-dimensional vector representing the personality characteristics of the digital human. Through cross-attention processing, 16-dimensional personality-scene association data A is obtained. A Transformer decoder is used to generate candidate behavior sequences such as "walking towards the bar", "talking to others", "dancing", etc. Preset behavior rules include "not committing illegal acts", "being polite", etc., and behaviors that do not conform to the rules are filtered out. The remaining behaviors are parameterized with emotional intensity. For example, the emotional vector of "dancing" may be [0.8, 0.7], indicating a relatively high degree of pleasure and activation. Finally, these behavior instructions are aligned to the time axis of the scene, such as "walking towards the bar at 19:30:00", "talking to others at 19:35:00", "dancing at 20:00:00", and linear interpolation is performed between these discrete time points to obtain continuous contextualized behavior instruction data. This process ensures that the generated behavior of the digital human not only meets the scene requirements but also reflects the personality characteristics, while maintaining the continuity and naturalness of the behavior.
[0151] In a specific embodiment, the process of performing step S104 may specifically include the following steps:
[0152] (1) Perform temporal decomposition processing on the contextualized behavior instruction data to obtain a single-frame behavior instruction sequence, and perform key frame extraction processing on the single-frame behavior instruction sequence to obtain key behavior node data;
[0153] (2) Perform interpolation expansion processing on the key behavior node data to obtain complete action sequence data;
[0154] (3) Perform bone mapping processing on the complete action sequence data to obtain initial bone animation data, and perform animation correction based on physical constraints on the initial bone animation data to obtain target bone animation data;
[0155] (4) Perform skinning weight calculation processing on the target bone animation data to obtain deformed mesh data, and perform texture mapping processing on the deformed mesh data to obtain initial rendering data;
[0156] (5) Perform adaptive resolution adjustment processing on the initial rendering data to obtain optimized rendering data;
[0157] (6) Perform phoneme segmentation processing on the voice part in the contextualized behavior instruction data to obtain phoneme sequence data, and perform acoustic parameter generation processing on the phoneme sequence data and the optimized rendering data to obtain multimodal digital human performance data.
[0158] Specifically, the contextual behavior instruction data is processed by temporal decomposition, discretizing continuous behavior instructions into a single-frame behavior instruction sequence. This process uses the time window sliding technique, where each window represents a time step, and the window size is determined according to the target frame rate. For example, for a target animation with 30 fps, the behavior instructions every 33.3 milliseconds are extracted as one frame.
[0159] Next, key frame extraction is performed on the single-frame behavior instruction sequence to obtain key behavior node data. This step uses an action saliency detection algorithm to calculate the difference degree between each frame and its adjacent frames. The frames with a difference degree exceeding the preset threshold are marked as key frames. The difference degree calculation can use cosine similarity:
[0160] ;
[0161] where A and B are the feature vectors of adjacent frames, and (n) is the feature dimension. Interpolation and extension processing is performed on the key behavior node data to obtain the complete action sequence data. Here, the spline interpolation method, such as cubic spline interpolation, is used to ensure that the generated action curve is smooth and continuous. The cubic spline interpolation function can be expressed as:
[0162] ;
[0163] where is the interpolation function of the i-th segment, , , , are undetermined coefficients.
[0164] Skeletal mapping processing is performed on the complete action sequence data to obtain the initial skeletal animation data. This step uses the inverse kinematics (IK) algorithm to convert the abstract action instructions into specific skeletal rotation data. The IK solution can use the Jacobian matrix method, and its iterative formula is:
[0165] ;
[0166] where is the pseudo-inverse of the Jacobian matrix, is the change in joint angle, is the change in the position of the end effector. Animation correction based on physical constraints is performed on the initial skeletal animation data to obtain the target skeletal animation data. This step introduces physical simulation and uses the spring-mass system to simulate the physical properties of muscles and joints. The spring force can be expressed as:
[0167] ; where is the spring constant, is the damping coefficient, is the spring deformation amount. The skinning weight calculation process is performed on the target skeletal animation data to obtain the deformed mesh data. Here, the linear blend skinning (LBS) algorithm is used, and the position of each vertex is affected by multiple weighted bones:
[0168] ; where is the position of the deformed vertex, is the weight, is the bone transformation matrix, is the original vertex position.
[0169] The texture mapping process is performed on the deformed mesh data to obtain the initial rendering data. In this step, the UV mapping technology is used to map the 2D texture image to the surface of the 3D model. The adaptive resolution adjustment process is performed on the initial rendering data to obtain the optimized rendering data. Here, the dynamic resolution scaling algorithm is used to dynamically adjust the rendering resolution according to the device performance and network conditions.
[0170] Finally, the phoneme segmentation process is performed on the speech part of the contextualized behavior instruction data to obtain the phoneme sequence data. In this step, the hidden Markov model (HMM) is used for speech recognition and phoneme segmentation. The acoustic parameter generation process is performed on the phoneme sequence data and the optimized rendering data to obtain the multimodal digital human performance data. Here, a deep learning model such as WaveNet is used to generate speech data synchronized with the animation.
[0171] For example, assume there is a 10-second contextualized behavior instruction data, and the goal is to generate an animation at 30 fps. First, the 10-second instruction data is decomposed into a single-frame behavior instruction sequence of 300 frames (10 seconds 30 fps). Through action saliency detection, 50 key frames are identified, with an average of one key frame every 0.2 seconds. Cubic spline interpolation is used to expand the 50 key frames into a complete action sequence of 300 frames.
[0172] In the bone mapping stage, assume the digital human model has 20 key bone points. Using the IK algorithm, the abstract action instruction of each frame is converted into the rotation data of 20 bone points, forming the initial skeletal animation data of 300×20×3 (300 frames, 20 bones, and 3 rotation angles for each bone). The physical constraint correction is iterated 1000 times with a time step of 10 ms to ensure that the action conforms to the physical laws.
[0173] In the skinning weight calculation stage, assume the 3D model has 10,000 vertices, and each vertex is affected by 4 nearest bones. Through the LBS algorithm, the deformed mesh data of 300×10,000×3 (300 frames, 10,000 vertices, and 3 coordinates for each vertex) is generated. The texture mapping uses a texture map with a resolution of 2048 2048 resolution.
[0174] Adaptive resolution adjustment dynamically adjusts the rendering resolution from 1080p to 720p according to the target device, reducing the rendering load by 30%. In terms of speech processing, 10 seconds of speech data is segmented into approximately 100 phonemes. Using the WaveNet model, speech data with a sampling rate of 44.1kHz is generated, which is precisely synchronized with 300 frames of animation. The final multi-modal digital human performance data includes 300 frames of visual data, 441,000 audio samples (10 seconds × 44.1kHz), and corresponding timestamp information, achieving a highly realistic and synchronized digital human performance.
[0175] In a specific embodiment, the process of executing step S105 may specifically include the following steps:
[0176] (1) Perform modal separation processing on the multi-modal digital human performance data to obtain visual component data, audio component data, and motion component data;
[0177] (2) Perform time window segmentation processing on the visual component data, audio component data, and motion component data to obtain a multi-modal behavior segment set, where the multi-modal behavior segment set includes: visual behavior segments, audio behavior segments, and motion behavior segments;
[0178] (3) Perform facial expression feature extraction processing on the visual behavior segments to obtain an expression feature sequence;
[0179] (4) Perform phoneme segmentation processing on the audio behavior segments to obtain phoneme time series data;
[0180] (5) Perform joint angle calculation processing on the motion behavior segments to obtain a skeletal motion sequence;
[0181] (6) Perform time series alignment processing on the expression feature sequence, phoneme time series data, and skeletal motion sequence to obtain an aligned multi-modal sequence, and perform cross-modal consistency analysis on the aligned multi-modal sequence to obtain a modal coordination degree index;
[0182] (7) Perform threshold segmentation processing on the modal coordination degree index to obtain a preliminary behavior deviation mark;
[0183] (8) Perform clustering analysis processing on the preliminary behavior deviation mark to obtain behavior deviation data, and perform user experience mapping processing on the behavior deviation data to obtain performance feedback data;
[0184] (9) Perform adaptive optimization processing on the multi-modal digital human performance data and performance feedback data through the behavior deviation data to obtain the target digital human.
[0185] Specifically, perform modal separation processing on the multi-modal digital human performance data to separate visual, audio, and action data. This process uses data parsing technology to extract data of different modalities into their respective data streams according to predefined data structures and tags.
[0186] Next, perform time window segmentation processing on the separated visual component data, audio component data, and action component data to obtain a multi-modal behavior segment set. Time window segmentation uses a sliding window technique, and the window size is usually set from 0.5 seconds to 2 seconds, with the step size being 25% to 50% of the window size. For example, for 30fps video data, a 1-second window size corresponds to 30 frames, and the step size may be 8 frames or 15 frames.
[0187] Perform facial expression feature extraction processing on the visual behavior segments to obtain an expression feature sequence. This step uses a facial feature point detection algorithm, such as an improved Active Shape Model (ASM), to extract 68 facial key points. Each key point is represented by (x, y) coordinates, so the expression feature of each frame is a 136-dimensional vector. Perform phoneme segmentation processing on the audio behavior segments to obtain phoneme time series data. Phoneme segmentation uses a Hidden Markov Model (HMM) combined with Mel Frequency Cepstral Coefficients (MFCC) features to segment continuous speech signals into discrete phoneme units. Perform joint angle calculation processing on the action behavior segments to obtain a skeletal motion sequence. Here, quaternion representation is used to calculate joint angles, and the rotation of each joint is represented by a four-dimensional vector.
[0188] Perform temporal alignment processing on the expression feature sequence, phoneme time series data, and skeletal motion sequence to obtain an aligned multi-modal sequence. Temporal alignment uses the Dynamic Time Warping (DTW) algorithm, and the distance calculation formula of DTW is:
[0189] , where X and Y are two time series, and are the feature values at the corresponding time points.
[0190] Perform cross-modal consistency analysis on the aligned multi-modal sequence to obtain a modal coordination degree index. Consistency analysis uses Mutual Information (MI) calculation, and the formula is:
[0191] ;
[0192] where p(x,y) is the joint probability distribution, and p(x) and p(y) are the marginal probability distributions.
[0193] Perform threshold segmentation processing on the modal coordination degree index to obtain preliminary behavior deviation marks. The threshold is usually set to 80% of the average coordination degree. Perform clustering analysis processing on the preliminary behavior deviation marks to obtain behavior deviation data. Here, the K-means clustering algorithm is used, and the number of clusters K is determined by the elbow method. Perform user experience mapping processing on the behavior deviation data to obtain performance feedback data. The mapping function can be a simple linear mapping or a more complex non-linear function, such as the sigmoid function:
[0194] ;
[0195] where k controls the steepness of the curve, is the midpoint.
[0196] Finally, perform adaptive optimization processing on the multi-modal digital human performance data and performance feedback data through the behavior deviation data to obtain the target digital human. The gradient descent method is used in the optimization process, and the objective function is:
[0197] ;
[0198] where, , , , are weight coefficients, , , are the loss functions of vision, audio, and action respectively, is the feedback loss.
[0199] For example, assume there is a 10 - second multi - modal digital human performance data with a video frame rate of 30 fps and an audio sampling rate of 16 kHz. First, perform modal separation to obtain 300 video frames, 160,000 audio samples, and 300 groups of action data. Use a 1 - second time window and a 0.5 - second step size for segmentation to obtain 19 multi - modal behavior segments. Extract 68 facial feature points from each visual segment to form an expression feature sequence of 30×136. The audio segments are segmented by phonemes, with an average of 10 phonemes recognized per second, resulting in a phoneme time - series data of 10×19. The action segments calculate the angles of 20 key joints to form a skeletal motion sequence of 30×80. After time - series alignment, calculate the mutual information between modalities to obtain the coordination degree index at 19 time points. Set the threshold to 0.75, and identify 3 time periods with a coordination degree lower than the threshold. Through K - means clustering (K = 2), these deviations are classified into two categories: "expression - speech mismatch" and "action - speech asynchrony". The user experience mapping converts the deviation degree into a feedback score ranging from - 0.8 to 0.2. Finally, the optimization algorithm adjusts the data in these 3 time periods. For example, it increases the smile degree of the expression from 0.3 to 0.6 and delays the action by 100 ms to match the speech, ultimately generating a target digital human performance with more coordinated behavior.
[0200] In a specific embodiment, the process of performing the step of adaptively optimizing the multi - modal digital human performance data and performance feedback data through behavior deviation data may specifically include the following steps:
[0201] (1) Perform importance scoring on the behavior deviation data to obtain a deviation priority list;
[0202] (2) Perform semantic analysis on the performance feedback data to obtain a user preference feature vector, and perform weighted fusion processing on the deviation priority list and the user preference feature vector to obtain an adjustment target matrix;
[0203] (3) Perform expression correction on the visual data in the multi - modal digital human performance data based on the adjustment target matrix to obtain a corrected expression sequence;
[0204] (4) Perform pitch adjustment on the audio data in the multi - modal digital human performance data based on the adjustment target matrix to obtain adjusted speech data;
[0205] (5) Perform pose fine - tuning on the action data in the multi - modal digital human performance data based on the adjustment target matrix to obtain an adjusted action sequence;
[0206] (6) Perform time - series alignment on the corrected expression sequence, adjusted speech data, and adjusted action sequence to obtain preliminarily adjusted multi - modal data, and perform inter - modal correlation analysis on the preliminarily adjusted multi - modal data to obtain a correlation matrix;
[0207] (7) Perform threshold screening on the correlation matrix to obtain modal coordination data, and perform multimodal fusion processing on the modal coordination data to obtain the target digital human.
[0208] Specifically, perform importance scoring on the behavior deviation data to obtain a deviation priority list. This process uses a weighted scoring model, considering the magnitude, duration, and impact on user perception of the deviation. The scoring formula can be expressed as:
[0209] ;
[0210] where, is the score of the i-th deviation, is the deviation magnitude, is the duration, is the degree of impact, , , are the corresponding weights.
[0211] Next, perform semantic analysis on the performance feedback data to obtain a user preference feature vector. This step uses natural language processing techniques, such as word embedding and sentiment analysis, to convert user feedback into a numerical feature vector. Word embedding can use the Word2Vec model, and sentiment analysis uses a deep learning model based on LSTM. The user preference feature vector can be expressed as:
[0212] ;
[0213] where, represents user preference scores in different dimensions, such as naturalness, emotional expression, action fluency, etc. Perform weighted fusion on the deviation priority list and the user preference feature vector to obtain an adjustment target matrix. The fusion process uses matrix operations, and the formula is as follows:
[0214] ;
[0215] where, is the adjustment target matrix, is the deviation priority list, is the user preference feature vector, and are balance parameters, is the identity matrix.
[0216] Based on the adjustment target matrix, correct the visual, audio, and motion data in the multi-modal digital human performance data. For the expression correction of visual data, facial action units (AUs) are used for adjustment, and the activation intensity of each AU is adjusted according to the corresponding element value of the adjustment target matrix. The pitch adjustment of audio data uses the pitchshifting algorithm, and the adjustment amplitude is determined by the adjustment target matrix. The pose fine-tuning of motion data is achieved through the inverse kinematics (IK) algorithm, and the adjustment target is specified by the adjustment target matrix.
[0217] Perform temporal alignment processing on the corrected expression sequence, speech data, and motion sequence to obtain the preliminary adjusted multi-modal data. Dynamic Time Warping (DTW) algorithm is used for temporal alignment to ensure the consistency of different modal data in the time dimension. The distance calculation formula of DTW is: ;
[0218] where X and Y are two time series, is the distance between two data points.
[0219] Perform inter-modal correlation analysis on the preliminary adjusted multi-modal data to obtain the correlation matrix. This step uses mutual information (MI) to calculate the correlation degree between different modalities, and the formula is as follows:
[0220] ;
[0221] where p(x,y) is the joint probability distribution, and p(x) and p(y) are the marginal probability distributions.
[0222] Perform threshold screening on the correlation matrix to obtain the modality-coordinated data. The Otsu method is used for threshold selection to automatically determine the optimal segmentation threshold. Finally, perform multi-modal fusion processing on the modality-coordinated data to obtain the target digital human. The attention mechanism is used in the fusion process to dynamically adjust the fusion weights according to the importance of different modalities.
[0223] For example, assume there is a 10-second multi-modal digital human performance data, including 300 frames of video (30fps), 160,000 audio samples (16kHz sampling rate), and 300 groups of motion data. Three main deviations are found through behavior deviation analysis:
[0224] 1) The expression is unnatural from 2 to 3 seconds;
[0225] 2) The speech pitch is too high from 5 to 6 seconds;
[0226] 3) The motion is rigid from 8 to 9 seconds.
[0227] Importance scores are assigned to these deviations, with scores of 0.8, 0.6, and 0.7 respectively. After semantic analysis of the user feedback, a 5-dimensional preference feature vector [0.9, 0.7, 0.8, 0.6, 0.8] is obtained, representing preferences for naturalness, emotional expression, action fluency, timbre, and synchronization respectively. Through weighted fusion, a 3x5 adjustment target matrix is obtained. Based on this matrix, the expression in the 2nd - 3rd seconds is corrected, increasing the AU activation intensity of the smile from 0.3 to 0.6; the pitch of the speech in the 5th - 6th seconds is adjusted, reducing the fundamental frequency from 280Hz to 240Hz; the action in the 8th - 9th seconds is fine-tuned, increasing the arm swing amplitude by 10%. Temporal alignment ensures that these adjustments are synchronized in time, such as adjusting the start time of the expression change from 1.95 seconds to 2.00 seconds to precisely match the audio change.
[0228] Inter-modal correlation analysis yields a 3x3 correlation matrix, indicating the degree of correlation between vision-audio, vision-action, and audio-action. Suppose the obtained matrix is:
[0229] ;
[0230] The threshold determined using the Otsu method is 0.75, and high-correlation modal combinations are retained after screening. Finally, through the attention mechanism for multi-modal fusion, the weights of different modalities are dynamically adjusted according to the correlation matrix. For example, the weight of the visual modality is increased during the expression change period, and the weight of the audio modality is increased during the speech adjustment period, ultimately generating a target digital human performance with more coordinated behavior and in line with user preferences. This optimization process not only corrects specific behavior deviations but also realizes the overall improvement of the digital human performance by considering the inter-modal correlation and user preferences.
[0231] The above described the multi-modal digital human generation method based on AIGC in the embodiments of the present application. Next, the multi-modal digital human generation system based on AIGC in the embodiments of the present application will be described. Please refer to Figure 2 , an embodiment of the multi-modal digital human generation system based on AIGC in the embodiments of the present application includes:
[0232] An enhancement module 201 for performing data enhancement processing on pre-acquired multi-modal character data to obtain an enhanced multi-modal data set, and extracting features from the enhanced multi-modal data set to obtain multi-modal feature data;
[0233] A processing module 202 for performing modeling and dynamic weight adjustment processing on the multi-modal feature data based on a graph-structured neural network to obtain dynamic personality knowledge data;
[0234] A matching module 203, configured to perform intelligent matching processing on preset target scenario information and dynamic personalized knowledge data to obtain contextualized behavior instruction data;
[0235] A generation module 204, configured to perform hierarchical action sequence generation, adaptive resolution optimization, and speech synthesis processing on a preset initial digital human according to the contextualized behavior instruction data to obtain multimodal digital human performance data;
[0236] An identification module 205, configured to perform behavior deviation identification on the multimodal digital human performance data to obtain behavior deviation data and performance feedback data, and perform adaptive optimization processing on the multimodal digital human performance data and the performance feedback data through the behavior deviation data to obtain a target digital human.
[0237] Through the collaborative cooperation of the above-mentioned various components, by performing data augmentation processing and feature extraction on the pre-acquired multi-modal human data, the quality and diversity of the original data are significantly improved, laying a solid foundation for subsequent personalized modeling. This not only enhances the realism of the generated digital human but also greatly improves the generalization ability of the model. Secondly, by adopting the modeling and dynamic weight adjustment processing method based on the graph-structured neural network, the accurate capture and dynamic expression of complex personality characteristics are realized, enabling the generated digital human to exhibit rich and diverse personalized characteristics and greatly enhancing the uniqueness and realism of the digital human. Furthermore, by performing intelligent matching processing on the target scene information and dynamic personality knowledge data, this method can dynamically generate adaptive behavior instructions according to different scene requirements, greatly improving the performance ability and naturalness of the digital human in complex and changing environments. In addition, when generating multi-modal digital human performance data, this method adopts advanced technologies such as hierarchical action sequence generation, adaptive resolution optimization, and speech synthesis, ensuring a high degree of coordination among the three modalities of vision, action, and speech and effectively solving the common modality inconsistency problem in traditional methods. It is worth noting that this method also introduces a behavior deviation recognition and adaptive optimization mechanism, which can detect and correct abnormal behaviors in the digital human performance in real time and make dynamic adjustments according to user feedback. This not only improves the performance quality of the digital human but also continuously optimizes and improves the behavior pattern of the digital human. At the same time, the modular design of this method makes the entire generation process highly flexible and scalable, and can be customized and optimized according to the requirements of different application scenarios. In addition, through the deep fusion and collaborative processing of multi-modal data, this method can generate more natural, smooth, and expressive digital humans, greatly enhancing the user's interaction experience and immersion. It is worth mentioning that the adaptive optimization mechanism of this method can not only improve the performance quality of the digital human but also continuously learn and evolve according to long-term interaction data, making the behavior and performance of the digital human increasingly close to real humans. In addition, this method performs excellently in processing long-time sequence multi-modal data, can maintain the coherence of behavior and the stability of emotional expression, solves the common long-term performance instability problem in the existing technology, and improves the efficiency and accuracy of multi-modal digital human generation based on AIGC.
[0238] This application also provides a computer-readable storage medium. The computer-readable storage medium can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the multi-modal digital human generation method based on AIGC.
[0239] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, systems, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0240] The above is the case. The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A multi-modal digital human generation method based on AIGC, characterized in that, It includes performing data augmentation processing on pre-acquired multi-modal person data to obtain an augmented multi-modal data set, and performing feature extraction on the augmented multi-modal data set to obtain multi-modal feature data; Performing modeling and dynamic weight adjustment processing on the multi-modal feature data based on a graph-structured neural network to obtain dynamic personality knowledge data; Performing intelligent matching processing on preset target scenario information and dynamic personality knowledge data to obtain contextualized behavior instruction data; Generating a hierarchical action sequence, performing adaptive resolution optimization, and performing speech synthesis processing on a preset initial digital human according to the contextualized behavior instruction data to obtain multi-modal digital human performance data; Performing behavior deviation identification on the multi-modal digital human performance data to obtain behavior deviation data and performance feedback data, and performing adaptive optimization processing on the multi-modal digital human performance data and performance feedback data through the behavior deviation data to obtain a target digital human, including performing modal separation processing on the multi-modal digital human performance data to obtain visual component data, audio component data, and action component data; performing time window segmentation processing on the visual component data, audio component data, and action component data to obtain a multi-modal behavior segment set, including visual behavior segments, audio behavior segments, and action behavior segments. The time window segmentation uses a sliding window technique, and the window size is set to 0.5 seconds to 2 seconds, and the step size is 25% to 50% of the window size; performing facial expression feature extraction processing on the visual behavior segments to obtain an expression feature sequence; performing phoneme segmentation processing on the audio behavior segments to obtain phoneme time series data; performing joint angle calculation processing on the action behavior segments to obtain a bone motion sequence; performing time series alignment processing on the expression feature sequence, phoneme time series data, and bone motion sequence to obtain an aligned multi-modal sequence, and performing cross-modal consistency analysis on the aligned multi-modal sequence to obtain a modal coordination degree index; Performing threshold segmentation processing on the modal coordination degree index to obtain preliminary behavior deviation marks; performing clustering analysis processing on the preliminary behavior deviation marks to obtain behavior deviation data, and performing user experience mapping processing on the behavior deviation data to obtain performance feedback data; performing adaptive optimization processing on the multi-modal digital human performance data and performance feedback data through the behavior deviation data to obtain a target digital human.
2. The multimodal digital human generation method based on AIGC according to claim 1, wherein Performing data augmentation processing on pre-acquired multi-modal person data to obtain an augmented multi-modal data set, and performing feature extraction on the augmented multi-modal data set to obtain multi-modal feature data, including: Performing timestamp alignment processing on multi-modal person data including video data, audio data, and motion capture data to obtain time-synchronized multi-modal raw data, and performing high-dynamic range image synthesis processing on the synchronized video data in the time-synchronized multi-modal raw data to obtain enhanced video data; Performing spectral enhancement processing on the synchronized audio data in the time-synchronized multi-modal raw data to obtain enhanced audio data, and performing joint trajectory smoothing processing on the synchronized motion capture data in the time-synchronized multi-modal raw data to obtain enhanced action data; Performing facial feature point extraction processing on the enhanced video data to obtain facial key point data; Perform Mel Frequency Cepstral Coefficient (MFCC) extraction on the enhanced audio data to obtain audio feature vectors; Perform bone angle and speed calculation on the enhanced action data to obtain action feature matrices; Perform temporal alignment on the facial key point data, audio feature vectors, and action feature matrices to obtain aligned feature data, and perform feature importance evaluation based on random forest on the aligned feature data, and select the top K features sorted by importance scores to obtain dimensionality-reduced feature data, where K is a positive integer greater than 1; Perform feature scaling on the dimensionality-reduced feature data to obtain multi-modal feature data.
3. The multi-modal digital human generation method based on AIGC according to claim 1, wherein Perform modeling and dynamic weight adjustment based on a graph-structured neural network on the multi-modal feature data to obtain dynamic personality knowledge data, including: Perform initial personality feature graph construction on the multi-modal feature data to obtain an initial personality feature graph, and perform inter-modal association weight initialization on the initial personality feature graph to obtain a weighted personality feature graph; Perform cross-modal feature aggregation on the weighted personality feature graph to obtain aggregated personality data, and perform non-linear expression mapping on the aggregated personality data to obtain expression feature vectors; Perform emotional residual connection on the expression feature vectors to obtain enhanced emotional representation data, and perform multi-head personality attention on the enhanced emotional representation data to obtain attention-weighted personality features; Perform scene adaptation pooling on the attention-weighted personality features to obtain hierarchical personality representation data, and perform global personality extraction on the hierarchical personality representation data to obtain a comprehensive personality vector; Perform dynamic behavior weight update on the comprehensive personality vector to obtain an optimized personality structure; Perform iterative personality propagation on the optimized personality structure to obtain dynamic personality knowledge data.
4. The multimodal digital human generation method based on AIGC according to claim 1, wherein, Perform intelligent matching on the preset target scene information and dynamic personality knowledge data to obtain contextualized behavior instruction data, including: perform multi-modal scene feature extraction on the target scene information to obtain scene feature vectors; perform cross-attention on the scene feature vectors and dynamic personality knowledge data to obtain personality-scene association data; Perform contextualized behavior pattern generation on the personality-scene association data to obtain candidate behavior sequences; Perform filtering based on preset behavior rules on the candidate behavior sequences to obtain target behavior sequences that conform to the preset behavior rules; Perform emotional intensity parameterization on the target behavior sequences to obtain emotionally parameterized behavior instructions; Perform timestamp alignment and interpolation smoothing on the emotionally parameterized behavior instructions to obtain contextualized behavior instruction data.
5. The multi-modal digital human generation method based on AIGC according to claim 1, characterized in that, Generate hierarchical action sequences, perform adaptive resolution optimization, and perform speech synthesis on the preset initial digital human according to the contextualized behavior instruction data to obtain multi-modal digital human performance data, including: Perform temporal decomposition on the contextualized behavior instruction data to obtain a single-frame behavior instruction sequence, and perform key frame extraction on the single-frame behavior instruction sequence to obtain key behavior node data; Perform interpolation expansion on the key behavior node data to obtain complete action sequence data; Perform skeletal mapping processing on the complete action sequence data to obtain initial skeletal animation data, and perform animation correction based on physical constraints on the initial skeletal animation data to obtain target skeletal animation data; Perform skinning weight calculation processing on the target skeletal animation data to obtain deformed mesh data, and perform texture mapping processing on the deformed mesh data to obtain initial rendering data; Perform adaptive resolution adjustment processing on the initial rendering data to obtain optimized rendering data; Perform phoneme segmentation processing on the voice part of the contextualized behavior instruction data to obtain phoneme sequence data, and perform acoustic parameter generation processing on the phoneme sequence data and the optimized rendering data to obtain multi-modal digital human performance data.
6. The multimodal digital human generation method based on AIGC according to claim 1, wherein Perform adaptive optimization processing on the multi-modal digital human performance data and the performance feedback data through the behavior deviation data to obtain the target digital human, including: performing importance scoring processing on the behavior deviation data to obtain a deviation priority list; performing semantic analysis processing on the performance feedback data to obtain a user preference feature vector, and performing weighted fusion processing on the deviation priority list and the user preference feature vector to obtain an adjustment target matrix; performing expression correction processing on the visual data in the multi-modal digital human performance data based on the adjustment target matrix to obtain a corrected expression sequence; performing pitch adjustment processing on the audio data in the multi-modal digital human performance data based on the adjustment target matrix to obtain adjusted voice data; performing pose fine-tuning processing on the action data in the multi-modal digital human performance data based on the adjustment target matrix to obtain an adjusted action sequence; performing temporal alignment processing on the corrected expression sequence, the adjusted voice data, and the adjusted action sequence to obtain preliminarily adjusted multi-modal data, and performing inter-modal correlation analysis processing on the preliminarily adjusted multi-modal data to obtain a correlation matrix; performing threshold screening processing on the correlation matrix to obtain modal coordination data, and performing multi-modal fusion processing on the modal coordination data to obtain the target digital human.
7. A multi-modal digital human generation system based on AIGC, characterized in that, For implementing the AIGC-based multi-modal digital human generation method according to any one of claims 1-6, the AIGC-based multi-modal digital human generation system includes: an enhancement module for performing data enhancement processing on the pre-acquired multi-modal human data to obtain an enhanced multi-modal data set, and performing feature extraction on the enhanced multi-modal data set to obtain multi-modal feature data; A processing module for performing modeling and dynamic weight adjustment processing on the multi-modal feature data based on a graph-structured neural network to obtain dynamic personality knowledge data; A matching module for performing intelligent matching processing on the preset target scene information and the dynamic personality knowledge data to obtain contextualized behavior instruction data; A generation module for performing hierarchical action sequence generation, adaptive resolution optimization, and speech synthesis processing on the preset initial digital human according to the contextualized behavior instruction data to obtain multi-modal digital human performance data; An identification module is used to identify behavioral deviations in multi-modal digital human performance data, obtain behavioral deviation data and performance feedback data, and perform adaptive optimization processing on the multi-modal digital human performance data and performance feedback data through the behavioral deviation data to obtain a target digital human, including performing modal separation processing on the multi-modal digital human performance data to obtain visual component data, audio component data, and motion component data; performing time window segmentation processing on the visual component data, audio component data, and motion component data to obtain a multi-modal behavior segment set, including visual behavior segments, audio behavior segments, and motion behavior segments. The time window segmentation uses a sliding window technique, and the window size is set to 0.5 seconds to 2 seconds, and the step size is 25% to 50% of the window size; performing facial expression feature extraction processing on the visual behavior segments to obtain an expression feature sequence; performing phoneme segmentation processing on the audio behavior segments to obtain phoneme time series data; performing joint angle calculation processing on the motion behavior segments to obtain a bone motion sequence; performing time series alignment processing on the expression feature sequence, phoneme time series data, and bone motion sequence to obtain an aligned multi-modal sequence, and performing cross-modal consistency analysis on the aligned multi-modal sequence to obtain a modal coordination degree index; performing threshold segmentation processing on the modal coordination degree index to obtain a preliminary behavioral deviation mark; performing clustering analysis processing on the preliminary behavioral deviation mark to obtain behavioral deviation data, and performing user experience mapping processing on the behavioral deviation data to obtain performance feedback data; performing adaptive optimization processing on the multi-modal digital human performance data and performance feedback data through the behavioral deviation data to obtain a target digital human.
8. A computer-readable storage medium having instructions stored thereon, characterized in that, When the instruction is executed by the processor, it implements the AIGC-based multi-modal digital human generation method according to any one of claims 1-6.
Citation Information
Patent Citations
Traditional Chinese medicine diagnosis and treatment data analysis system and method based on digital human cloning technology
CN119167112A
Training method and system for generating 5D digital human based on AIGC and medium
CN119378647A