A method and system for creating dynamic 3D character models and lip-syncing

CN122574218APending Publication Date: 2026-08-14BEIJING HOLOGRAPHIC JULANG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

因此,相关技术将颤音的声学波动直接转化为嘴部网格的几何形变,会导致3D人物在演唱长音时,其嘴唇和下颌跟随颤音频率发生高频的机械性开合与抖动,降低了3D人物在复杂演艺场景下口型表达的准确度

Benefits of technology

[0059]1. This application provides a method for establishing a dynamic 3D character model and lip-syncing. By configuring independent control nodes corresponding to the lip, jaw, and neck regions in the basic mesh model of the 3D character, and determining that the target physiological compensation state for generating high-frequency vibrato is in place when the periodic fluctuation frequency of the audio is detected to be greater than a preset high-frequency threshold and the phoneme sequence remains continuous, a spatial transfer matrix is ​​established to smoothly shift the deformation response center from the lips and jaw to the neck, and dynamic driving weights are assigned. This allows the model to simulate the compensatory linkage response of the neck muscles during real human vocalization based on the acoustic characteristics of high-frequency vibrato. This scheme combines the basic deformation parameters of the phoneme sequence mapping with the dynamic driving weights calculated based on the physiological compensation state to generate independent driving instructions for each region, thereby improving the accuracy of lip-syncing in complex performance scenarios for 3D characters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122574218A_ABST
    Figure CN122574218A_ABST
Patent Text Reader

Abstract

A method and system for establishing a 3D character dynamic model and lip-syncing is disclosed. In this method, independent control nodes corresponding to the lips, jaw, and neck are configured in the 3D model. Then, the phoneme sequence and acoustic energy features of the audio are extracted, and their periodic fluctuation frequencies are detected. When the frequency is greater than a preset threshold and the phoneme sequence remains unchanged, it is determined that the current state is a physiological compensation state of high-frequency vibrato. Based on this, a spatial transfer matrix is ​​established to smoothly transfer the deformation response center from the lips and jaw to the neck, and the dynamic driving weights of each region are calculated. Finally, the basic deformation parameters mapped from the phoneme sequence are weighted and calculated with the weights to generate independent driving commands, which are input to the corresponding nodes to drive the mesh vertices to perform coordinate updates and deformation rendering. This application aims to improve the accuracy of lip-syncing in complex performance scenes for 3D characters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the general field of image data processing or generation, and in particular relates to a method and system for establishing a 3D human dynamic model and lip-syncing. Background Technology

[0002] With the popularization of metaverse and digital human technology, the construction of 3D character dynamic models and voice-driven lip-syncing technology have been widely used in fields such as virtual anchors, film and television animation.

[0003] In related technologies, an end-to-end audio-driven facial deformation method based on a self-attention mechanism (Transformer) is used to achieve lip-syncing for 3D characters. A pre-trained speech feature extraction network obtains the deep acoustic feature sequence of the input audio, and a Transformer network with a sliding window captures the long and short-term temporal dependencies in the audio sequence. The acoustic features are directly mapped to the target weight sequence of facial deformation (blendshape) for the 3D character model. Finally, a temporal smoothing algorithm filters the weights of adjacent frames to drive the 3D model. This method can automatically and efficiently drive continuous mouth movements in a 3D character using only audio input.

[0004] In virtual concert scenarios, virtual idols often need to showcase vocal techniques such as sustained notes and vibrato accompanying them when singing. When a singer produces vibrato, the periodic vibration of their vocal cords causes rapid and continuous periodic fluctuations in the fundamental frequency and acoustic energy of the audio. When processing these high-frequency acoustic features frame-by-frame, the Transformer network may misinterpret them as rapid switching of phonemes, resulting in outputting drastically fluctuating deformation weights. However, the jaw and lips maintain a relatively stable opening; the sound fluctuations primarily originate from the control of the vocal cords in the throat. Therefore, directly converting the acoustic fluctuations of vibrato into geometric deformation of the mouth grid in related technologies causes the lips and jaw of 3D characters to undergo high-frequency mechanical opening and closing and shaking in response to the vibrato frequency when singing sustained notes, reducing the accuracy of lip-syncing in complex performance scenarios. Summary of the Invention

[0005] This application provides a method and system for establishing dynamic 3D character models and lip-syncing, which can improve the accuracy of lip-syncing of 3D characters in complex performing arts scenarios.

[0006] In the first aspect, this application provides a method for establishing a 3D character dynamic model and lip-syncing, which configures independent control nodes corresponding to the lip region, jaw region and neck region in the basic mesh model of the 3D character to establish a 3D character dynamic model with a layered linkage structure.

[0007] Based on a preset sliding time window, feature extraction is performed on the input audio signal to obtain a phoneme sequence that represents the pronunciation content and an acoustic energy feature that represents the sound intensity. The acoustic energy feature includes fundamental frequency feature and energy amplitude feature.

[0008] Extract the periodic fluctuation frequency and amplitude difference of acoustic energy characteristics within a preset sliding time window;

[0009] If the periodic fluctuation frequency is greater than the preset high-frequency threshold and the phoneme sequence remains continuous within the sliding time window, the current audio frame is determined to be in the target physiological compensatory state that produces high-frequency vibrato.

[0010] Based on the target physiological compensation state, a spatial transfer matrix is ​​established to smoothly transfer the deformation response center from the lip region and mandibular region to the neck region, and the transfer parameters corresponding to each region are calculated based on the spatial transfer matrix.

[0011] Dynamic driving weights for each region are assigned based on the transfer parameters;

[0012] The phoneme sequence is mapped to the basic deformation parameters, and the basic deformation parameters are weighted using dynamic driving weights to generate independent driving commands for the lip region, mandibular region and neck region.

[0013] The independent drive commands are input to the corresponding independent control nodes to drive the mesh vertices of the 3D character dynamic model to perform coordinate updates and deformation rendering.

[0014] By employing the aforementioned technical solution, independent control nodes are configured in the basic mesh model of the 3D character, corresponding to the lip, jaw, and neck regions. When the periodic fluctuation frequency of the audio is detected to be greater than a preset high-frequency threshold while the phoneme sequence remains continuous, the target physiological compensatory state for generating high-frequency vibrato is determined. A spatial transfer matrix is ​​then established to smoothly shift the deformation response center from the lips and jaw to the neck, and dynamic driving weights are assigned. This allows the model to simulate the compensatory linkage response of the neck muscles during real human vocalization based on the acoustic characteristics of high-frequency vibrato. This solution combines the basic deformation parameters of the phoneme sequence mapping with dynamic driving weights calculated based on the physiological compensatory state to generate independent driving instructions for each region, thereby improving the accuracy of lip-syncing in complex performance scenarios for 3D characters.

[0015] In conjunction with some implementation methods of the first aspect, in some implementation methods, based on the target physiological compensation state, a spatial transfer matrix is ​​established to smoothly transfer the deformation response center from the lip region and mandibular region to the neck region, and the transfer parameters corresponding to each region are calculated based on the spatial transfer matrix, specifically including:

[0016] Obtain the deformation bearing limit threshold of the lip and mandibular regions under the target physiological compensation state;

[0017] The energy amplitude characteristics in the acoustic energy features are compared with the deformation bearing limit threshold to calculate the overflow compensation energy that exceeds the deformation bearing limit threshold.

[0018] Based on the overflow compensation energy, a spatial transfer matrix is ​​constructed with the lip region and mandibular region as the starting point and the neck region as the ending point;

[0019] Extract the node mapping values ​​of the lip region, mandible region, and neck region on the transfer path from the spatial transfer matrix, and use the node mapping values ​​as the corresponding transfer parameters for each region.

[0020] By employing the aforementioned technical solution, the deformation bearing limit threshold of the lip and mandibular regions under the target physiological compensation state is obtained. The energy amplitude characteristics in the acoustic energy features are compared with this threshold to calculate the overflow compensation energy exceeding the limit. Then, based on this overflow compensation energy, a spatial transfer matrix is ​​constructed with the lips and mandible as the starting point and the neck as the ending point, and transfer parameters are extracted. This provides a precise quantitative basis for the transfer process of the deformation center. This solution binds energy transfer to the physical bearing limit of facial organs, improving the accuracy and rationality of the dynamic driving weight allocation for each region, and reducing the incidence of excessive or unnatural deformation in the lip and mandibular regions when processing high-energy audio clips.

[0021] In conjunction with some implementations of the first aspect, in some implementations, dynamic driving weights for each region are allocated according to transfer parameters, specifically including:

[0022] Based on the transfer parameters, the weight attenuation values ​​for the lip region and the mandibular region, as well as the weight increase value for the neck region, are calculated respectively.

[0023] The target driving weight of each region in the current audio frame is determined based on the weight decay value and the weight increase value.

[0024] Extract the historical driving weights of each region in the previous audio frame within the sliding time window;

[0025] The historical driving weights and the target driving weights are interpolated using a preset smooth interpolation algorithm to obtain the dynamic driving weights for each region.

[0026] By adopting the above technical solution, the target driving weight of the current audio frame is determined by calculating the weight decay values of the lip region and the mandibular region and the weight increase value of the neck region respectively according to the transfer parameters, and the historical driving weight of the previous audio frame within the sliding time window is extracted. The preset smoothing interpolation algorithm is used to interpolate and calculate the historical driving weight and the target driving weight to obtain the final dynamic driving weight, making the weight change process between adjacent audio frames smoother, improving the temporal continuity and visual smoothness of the 3D human face and neck mesh during the center transfer of the deformation response, and reducing the visual jitter or muscle twitching feeling that the model may generate under continuous speech driving.

[0027] In combination with some embodiments of the first aspect, in some embodiments, after allocating the dynamic driving weights of each region according to the transfer parameters, the method further includes:

[0028] Characterize the transfer parameter as a measure of the tension deviation of the current facial system from the natural expression manifold;

[0029] Input the tension deviation measure into a preset global correlation mapping model, which includes potential statistical association features between the main pronunciation region, the compensation region and the non-articulatory participation region;

[0030] Based on the preset global correlation mapping model, generate an induced driving weight that matches the tension deviation measure;

[0031] Apply the induced driving weight to the non-articulatory participation region to reconstruct the full facial state of the 3D human into a co-expression manifold that satisfies the potential statistical association features.

[0032] By adopting the above technical solution, by characterizing the transfer parameter as a measure of the tension deviation of the current facial system from the natural expression manifold and inputting it into a preset global correlation mapping model that includes potential statistical association features between the main pronunciation region, the compensation region and the non-articulatory participation region, generating a matching induced driving weight and applying it to the non-articulatory participation region, the non-articulatory region can generate coordinated deformations according to the local tension changes during pronunciation. Extending the local pronunciation compensation action to the reconstruction of the full facial state simulates the natural linkage reaction of other regions of the human face when speaking forcefully, thereby improving the global coordination and comprehensive expressiveness of the overall facial expression of the 3D human and reducing the visual fragmentation feeling caused by the local lip synchronization action and the immobility of other regions of the face.

[0033] In combination with some embodiments of the first aspect, in some embodiments, based on the preset global correlation mapping model, generating an induced driving weight that matches the tension deviation measure specifically includes:

[0034] Obtain the initial state features of non-speech-involved regions, which include the eye region, eyebrow region, and cheek region;

[0035] The tension deviation metric is concatenated with the initial state features to construct a multi-dimensional fused feature vector;

[0036] The fused feature vectors are input into the feature decoding network of the preset global correlation mapping model to extract the spatial linkage correlation between the main pronunciation region, the compensation region and each non-pronunciation participation region.

[0037] Based on spatial linkage correlation, attention-weighted mapping is performed on the fused feature vector to output the induction driving weights corresponding to each non-speech-participating region.

[0038] By employing the aforementioned technical solution, the initial state features of non-speech-involved areas such as the eyes, eyebrows, and cheeks are acquired and concatenated with tension deviation metrics to form a multi-dimensional fused feature vector. This enables the model to perceive the current state of each facial region. The fused feature vector is then input into a feature decoding network to extract the spatial linkage correlation between different regions, accurately quantifying the potential physiological connections between speech and non-speech areas. Based on this spatial linkage correlation, attention-weighted mapping is applied to the fused feature vector to output induction driving weights. This allows the deformation of non-speech-involved areas to adaptively adjust according to the speech state, thereby improving the naturalness and coordination of global facial expression coordination in 3D characters.

[0039] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes, before applying the induced driving weights to the non-vocalization participation region:

[0040] Obtain the first deformation displacement vector of the main sound-producing region in the current audio frame, and the second deformation displacement vector of the non-sound-producing region determined based on the induced driving weight;

[0041] Calculate the topological antagonism between the first deformation displacement vector and the second deformation displacement vector. The topological antagonism is a numerical value that characterizes the degree of reverse stretching between the main sound-producing region and the non-sound-producing region in three-dimensional space.

[0042] When the topological antagonism is greater than the preset tensile limit threshold, the first deformation displacement vector remains unchanged, and the attenuation coefficient is calculated based on the difference between the topological antagonism and the tensile limit threshold.

[0043] The induced driving weights are reduced by using the attenuation coefficient to obtain the corrected induced driving weights;

[0044] The modified induction-driven weights are applied to the non-speech-participating regions.

[0045] By employing the aforementioned technical solution, and by acquiring the first deformation displacement vector of the main vocal region and the second deformation displacement vector of the non-vocalization region, and calculating the topological antagonism of their reverse stretching in three-dimensional space, the structural tension of the facial mesh during the coordinated deformation process can be monitored in real time. When the topological antagonism exceeds a preset stretching limit threshold, while keeping the deformation displacement vector of the main vocal region unchanged, an attenuation coefficient is calculated based on the difference exceeding the threshold to reduce the induction driving weight of the non-vocalization region. This allows the system to prioritize lip-sync accuracy while dynamically suppressing excessive deformation in the non-vocalization region, reducing the risk of mesh tearing or excessive distortion when the 3D character model performs complex or exaggerated expressions, and improving the structural stability and visual smoothness of the mesh dynamic deformation process.

[0046] In conjunction with some implementations of the first aspect, in some implementations, the topological antagonism between the first deformation displacement vector and the second deformation displacement vector is calculated, specifically including:

[0047] Obtain the first grid vertex located at the boundary of the main sound-producing region, and the second grid vertex located at the boundary of the non-sound-producing region;

[0048] Calculate the initial spatial distance between the first grid vertex and the second grid vertex in the undeformed state;

[0049] Update the position of the first mesh vertex according to the first deformation displacement vector, and update the position of the second mesh vertex according to the second deformation displacement vector;

[0050] Calculate the target space distance between the updated first grid vertex and the second grid vertex;

[0051] Calculate the ratio of the target spatial distance to the initial spatial distance to obtain the deformation elongation rate;

[0052] Calculate the angle between the first deformation displacement vector and the second deformation displacement vector;

[0053] The product of the deformation elongation rate and the included angle is used as the topological antagonism.

[0054] By employing the aforementioned technical solution, and extracting the mesh vertices at the boundary between the main vocalization region and the non-vocalization region, and calculating their initial and target spatial distances before and after deformation, the linear stretching amplitude between adjacent functional regions can be accurately obtained. Further calculation of the ratio of spatial distances before and after deformation yields the deformation stretching rate. Combining this with the angle between the first and second deformation displacement vectors, multiplying the deformation stretching rate by this angle yields the topological antagonism. This allows the system to comprehensively quantify the relative tension state of the mesh from both the displacement amplitude and the divergence of the motion direction. This multi-dimensional calculation method improves the accuracy of assessing the degree of structural conflict between different facial regions, enhances the reliability of mesh tension monitoring, and thus reduces the probability of misjudging normal facial muscle coordination as overstretching.

[0055] Secondly, embodiments of this application provide a 3D character dynamic model creation and lip-syncing system, which includes: one or more processors and a memory; the memory is coupled to one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and one or more processors call the computer instructions to cause the system to perform the method as described in the first aspect and any possible implementation of the first aspect.

[0056] Thirdly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a system, cause the system to perform the method described in the first aspect and any possible implementation thereof.

[0057] Fourthly, embodiments of this application provide a computer program product that, when run on a system, causes the system to execute the method described in any possible implementation of the first aspect.

[0058] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0059] 1. This application provides a method for establishing a dynamic 3D character model and lip-syncing. By configuring independent control nodes corresponding to the lip, jaw, and neck regions in the basic mesh model of the 3D character, and determining that the target physiological compensation state for generating high-frequency vibrato is in place when the periodic fluctuation frequency of the audio is detected to be greater than a preset high-frequency threshold and the phoneme sequence remains continuous, a spatial transfer matrix is ​​established to smoothly shift the deformation response center from the lips and jaw to the neck, and dynamic driving weights are assigned. This allows the model to simulate the compensatory linkage response of the neck muscles during real human vocalization based on the acoustic characteristics of high-frequency vibrato. This scheme combines the basic deformation parameters of the phoneme sequence mapping with the dynamic driving weights calculated based on the physiological compensation state to generate independent driving instructions for each region, thereby improving the accuracy of lip-syncing in complex performance scenarios for 3D characters.

[0060] 2. This application provides a method for establishing a 3D character dynamic model and lip-syncing. By representing the transfer parameter as a tension deviation measure of the current facial system from the natural facial expression manifold, and inputting it into a pre-defined global correlation mapping model containing the potential statistical correlation features between the main vocalization region, compensatory region, and non-vocalization participation region, matching induction driving weights are generated and applied to the non-vocalization participation region. This allows the non-vocalization region to undergo coordinated deformation according to local tension changes during vocalization. Extending local vocalization compensation actions to the reconstruction of the global facial state simulates the natural coordinated response of other facial areas when humans forcefully vocalize, thereby improving the overall coordination and comprehensive expressiveness of the 3D character's facial expression and reducing the visual disconnect caused by local lip-syncing actions and the static state of other facial areas.

[0061] 3. This application provides a method for establishing a 3D character dynamic model and lip-syncing. By obtaining the first deformation displacement vector of the main vocal region and the second deformation displacement vector of the non-vocal participation region, and calculating the topological antagonism of their reverse stretching in three-dimensional space, the structural tension of the facial mesh during the linkage deformation process can be monitored in real time. When the topological antagonism exceeds a preset stretching limit threshold, while keeping the deformation displacement vector of the main vocal region unchanged, an attenuation coefficient is calculated based on the difference exceeding the threshold to reduce the induction driving weight of the non-vocal participation region. This allows the system to prioritize lip-syncing accuracy while dynamically suppressing excessive deformation of the non-vocal regions, reducing the risk of mesh tearing or excessive distortion when the 3D character model performs complex or exaggerated expressions, and improving the structural stability and visual smoothness of the mesh dynamic deformation process. Attached Figure Description

[0062] Figure 1 This is a flowchart illustrating a method for establishing a 3D character dynamic model and lip-syncing in an embodiment of this application.

[0063] Figure 2 This is another flowchart illustrating a method for establishing a 3D character dynamic model and lip-syncing in an embodiment of this application.

[0064] Figure 3 This is another flowchart illustrating a method for establishing a 3D character dynamic model and lip-syncing in an embodiment of this application.

[0065] Figure 4 This is a schematic diagram of the physical device structure of a 3D character dynamic model creation and lip-syncing system provided in an embodiment of this application. Detailed Implementation

[0066] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to any or all possible combinations including one or more of the listed items.

[0067] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.

[0068] The following example is used in conjunction with Figure 1 The present application describes a method for establishing a 3D character dynamic model and lip-syncing in an embodiment of the present application:

[0069] Please see Figure 1 This is a flowchart illustrating a method for establishing a 3D character dynamic model and lip-syncing in an embodiment of this application.

[0070] S101. Configure independent control nodes for the corresponding lip region, jaw region and neck region in the basic mesh model of the 3D character to establish a 3D character dynamic model with a layered linkage structure.

[0071] The basic mesh model of a 3D character refers to the set of polygonal topologies used in computer graphics to construct the surface appearance of a virtual character. It consists of numerous vertices with specific coordinates in three-dimensional space, edges connecting these vertices, and faces enclosed by these edges. Examples include, but are not limited to, triangular mesh models composed of triangular faces or quadrilateral mesh models composed of quadrilateral faces. The lip region, jaw region, and neck region refer to specific sets of vertices within the aforementioned basic mesh model, defined according to the anatomical structure of the human face and neck. These correspond to the surface topology of the virtual character's lips, jawbone contours, and neck muscle groups, respectively. Independent control nodes are digital drive units assigned in 3D software or rendering engines to specifically manipulate the displacement, rotation, or scaling of one or more specific mesh regions. Examples include, but are not limited to, skeletal joints, deformation target controllers, or mesh deformation lattices. A hierarchical linkage structure refers to an organizational architecture established between these independent control nodes, characterized by hierarchical dependencies and weighted transmission relationships. This allows different regions to receive commands and deform independently, while also generating natural, interconnected movements within the overall anatomical logic.

[0072] The system can achieve this step using skeletal skinning rigging technology. First, the system constructs a virtual skeleton system within the base mesh model of the 3D character. For the lip region, multiple tiny bones are created surrounding the lips; for the mandible region, pivot bones are created to simulate the opening and closing of the mandible; and for the neck region, neck bones are created to simulate the movement of the sternocleidomastoid muscle and the Adam's apple. These bones serve as independent control nodes. Subsequently, the system uses a vertex weighting rendering algorithm to calculate the weight value of each vertex in the base mesh model affected by these bones. For vertices in the lip region, the lip bones are assigned extremely high weights; for vertices in the mandible region, the mandible bones are assigned primary weights, supplemented by a small amount of lip bone weights to achieve linkage; and for vertices in the neck region, they are primarily bound to the neck bones. The system establishes a hierarchical constraint relationship between bones, for example, setting the parent node of the lip bones to the mandible bones, thereby creating a layered linkage structure. This allows the movement of the mandible to naturally drive the movement of the lips, while the neck bones exist as a relatively independent compensation layer.

[0073] The system can also use hybrid deformation technology to achieve this step. Based on the base mesh model, the system sculpts target meshes for various basic mouth shape changes in the lip region, the jaw region during opening and closing, and the neck region when muscles are tense. The system extracts the vertex displacement vector field for each region by calculating the spatial coordinate difference between the corresponding vertices of these target meshes and the base mesh model. The system encapsulates these displacement vector fields into independent deformation channels, each channel configured with a floating-point slider with a value between zero and one as an independent control node. To establish a hierarchical linkage structure, the system writes expressions or node connection logic between these control nodes, setting it so that when the value of the control node in the jaw region increases, it automatically triggers a change in the value of the relevant control node in the lip region by a specific ratio, while maintaining the independent listening state of the control node in the neck region. This achieves a relative balance between independent driving and linked deformation of the meshes in each region.

[0074] S102. Based on a preset sliding time window, feature extraction is performed on the input audio signal to obtain a phoneme sequence for representing the pronunciation content and acoustic energy features for representing the sound intensity.

[0075] The system performs feature extraction on the input audio signal based on a preset sliding time window, obtaining a phoneme sequence representing the pronunciation content and acoustic energy features representing the sound intensity. The acoustic energy features include fundamental frequency features and energy amplitude features. The preset sliding time window refers to a segment with a fixed duration that slides forward on the time axis at specific steps when processing continuous audio data streams. For example, but not limited to, a data frame segmentation window with a length of twenty milliseconds that slides forward ten milliseconds each time. The audio signal refers to a sequence of sound wave data acquired by a microphone or other sound pickup device and converted into digital format. Feature extraction refers to the process of calculating and separating data dimensions that reflect specific physical or linguistic attributes of sound from the original audio waveform using digital signal processing techniques or machine learning algorithms. A phoneme sequence refers to the ordered arrangement of the smallest phonetic units obtained after segmenting continuous speech content according to linguistic pronunciation rules, used to represent the pronunciation content, such as, but not limited to, combinations of vowels, consonants, or initials and finals. Acoustic energy features are quantitative indicators reflecting the vibration intensity and frequency characteristics of sound in physical space, used to represent the sound intensity. The fundamental frequency characteristic refers to the lowest frequency component produced when the vocal cords vibrate, usually corresponding to the pitch of the sound. The energy amplitude characteristic refers to the amplitude or sound pressure level of the audio signal within a specific time period, usually corresponding to the loudness of the sound. The preset sliding time window length and step size threshold are derived from statistical analysis of human speech pronunciation rate and auditory persistence effect, aiming to ensure that the window contains a sufficient waveform period for extracting a stable fundamental frequency, while ensuring a smooth transition between adjacent windows.

[0076] To enable 3D characters to accurately lip-sync and exhibit physiological responses, the system must function like the human brain, not only understanding "what the character said" but also perceiving "how loudly and how high-pitched" the character spoke. The system slices the input audio signal using a preset sliding time window, converting continuous sound waves into discrete time frames. Within each window, the system simultaneously extracts two core information streams: one is a phoneme sequence representing the language content, which determines the basic shape of the character's lips; the other is acoustic energy characteristics containing the fundamental frequency and energy amplitude, revealing the physical force exerted by the character when speaking.

[0077] The system can employ a combination of traditional digital signal processing and Hidden Markov Models (HMMs) to implement the refinement techniques for this step. First, the system pre-emphasizes, frames, and windows the input audio signal, using a Hamming window as a preset sliding time window to segment continuous audio into overlapping short frames. For phoneme sequence extraction, the system performs a Fast Fourier Transform (FFT) on each frame of audio, calculates the Mel-frequency cepstral coefficients, and inputs these cepstral coefficients into a pre-trained HMM. The Viterbi decoding algorithm is used to find the most probable phoneme state transition path, thereby outputting a phoneme sequence representing the phonological content. For acoustic energy feature extraction, the system uses an autocorrelation algorithm or cepstral method to periodically detect the waveform of each frame of audio, identifying the first peak position of the autocorrelation function, the reciprocal of which is the fundamental frequency feature. Simultaneously, the system calculates the sum of squares of the amplitudes of all audio sampling points within the sliding time window, and then takes the root mean square value to obtain the energy amplitude feature reflecting the loudness of the sound.

[0078] The system can also utilize an end-to-end deep learning acoustic model to refine the technical solutions for this step. The system converts the input audio signal into a Mel spectrogram and, using a preset sliding time window as the receptive field, inputs the spectrogram matrix into an acoustic model based on a convolutional recurrent neural network. This large-scale AI model achieves its functionality because its convolutional layers effectively extract local texture features from the spectrogram, while the long short-term memory network layer captures the temporal dependencies between audio frames, and the attention mechanism focuses on key pronunciation segments. The model is designed with a multi-task output architecture. Its classification head outputs the probability distribution of each phoneme corresponding to the current window through a fully connected layer and a normalized exponential function, selecting the phoneme sequence with the highest probability. Simultaneously, its regression head directly outputs continuous values ​​through a linear mapping layer, predicting the fundamental frequency feature value and energy amplitude feature value within the current window, respectively.

[0079] S103. Extract the periodic fluctuation frequency and amplitude difference of acoustic energy features within a preset sliding time window;

[0080] Periodic fluctuation frequency refers to the rate at which the acoustic energy characteristic value exhibits regular fluctuations over time within multiple consecutive sliding time windows; that is, the speed at which energy or pitch undergoes periodic oscillations, such as, but not limited to, the number of times a singer's voice vibrates per second during performance. Amplitude difference refers to the numerical difference between the highest peak and lowest trough of the acoustic energy characteristic within a complete cycle of the aforementioned regular fluctuations, representing the magnitude of the sound oscillation.

[0081] The system can implement this step using an algorithm based on time-domain envelope tracking and extremum detection. First, the system arranges the continuously extracted acoustic energy features (such as fundamental frequency or energy amplitude) within a preset sliding time window in chronological order to construct a one-dimensional time-series signal. To eliminate high-frequency noise interference, the system applies a low-pass filter to smooth the time series, extracting a low-frequency envelope reflecting the energy change trend. Next, the system runs an extremum search algorithm on the envelope, locating all local maxima (peaks) and local minima (troughs) by comparing the slope changes of adjacent data points. The system calculates the time interval between two adjacent peaks or troughs, using the reciprocal of this time interval as the periodic fluctuation frequency. Simultaneously, the system extracts the feature values ​​of adjacent peaks and troughs, subtracting the trough value from the peak value to calculate the amplitude difference within the fluctuation period.

[0082] The system can also use frequency domain transformation and spectrum analysis techniques to achieve this step. The system treats the acoustic energy feature sequence within a preset sliding time window as a low-frequency oscillation signal, performs a discrete Fourier transform on it, converting the energy features that originally fluctuated in the time domain to the frequency domain, generating an energy fluctuation spectrum. In the spectrum, the system searches for the dominant frequency component with the largest amplitude; the frequency value corresponding to this dominant frequency component is the periodic fluctuation frequency of the acoustic energy feature. Subsequently, the system uses a bandpass filter to separate this dominant frequency component from the original signal, reconstructing a clean periodic oscillation waveform. The system calculates the maximum amplitude of the positive half-cycle and the minimum amplitude of the negative half-cycle in the reconstructed waveform, and takes the sum of their absolute values ​​as the amplitude difference within the sliding time window.

[0083] S104. If the periodic fluctuation frequency is greater than the preset high-frequency threshold and the phoneme sequence remains continuous within the sliding time window, determine that the current audio frame is in the target physiological compensation state that generates high-frequency vibrato.

[0084] The preset high-frequency threshold refers to a numerical limit set by the system to define whether a sound fluctuation belongs to high-frequency vibrato. This threshold is derived from statistical analysis of a large amount of acoustic data from real humans under extreme vocal conditions. It is compared with the periodic fluctuation frequency; only when the actual frequency crosses this limit is it considered an abnormal high-frequency oscillation. The continuous invariant state refers to the phoneme sequence extracted by the system remaining consistent across multiple sliding time windows, without any switching of phoneme types, indicating that the singer is prolonging the same syllable. The target physiological compensatory state refers to the physiological phenomenon in which, when the main lip and jaw muscles cannot independently withstand the current vocal intensity or vibrato frequency, the neck muscles (such as the sternocleidomastoid muscle) are instinctively engaged to contract and maintain airflow and sound stability. High-frequency vibrato refers to a rapid and regular periodic oscillation of sound at the fundamental frequency or energy level, typically occurring during the sustained note phase of singing.

[0085] The system can implement this step using a real-time decision architecture based on logic gates and state machines. The system maintains a state machine in memory, initially in the "normal pronunciation state." The system establishes a frequency comparator module and a phoneme buffer queue. In each sliding time window, the frequency comparator receives the calculated periodic fluctuation frequency and compares it with a preset high-frequency threshold stored in a configuration file. If the frequency is greater than the threshold, a high-level signal is output. Simultaneously, the phoneme buffer queue records the phoneme sequence of the most recent windows. The system traverses the queue, checking if all elements are identical to the first element. If they are identical, it is determined to be a continuous, unchanging state, and a high-level signal is output. The system inputs these two high-level signals into a logic AND gate. Only when both are simultaneously high, the AND gate outputs a trigger signal, driving the state machine to jump from the "normal pronunciation state" to the "target physiological compensation state," and tagged the current audio frame with a high-frequency vibrato.

[0086] The system can also implement this step using an algorithm based on sliding window statistics and confidence scoring. Instead of using absolute Boolean logic, the system calculates a compensation trigger confidence level. First, the system calculates the proportion of periodic fluctuation frequencies exceeding a preset high-frequency threshold, mapping it to a frequency score between zero and one. Next, the system calculates the duration for which the phoneme sequence remains unchanged within the sliding time window, compares this duration with a preset minimum sustained note duration, and calculates a continuity score between zero and one using a non-linear mapping function. The system then performs a weighted sum of the frequency score and the continuity score to obtain the overall confidence level of the current audio frame. The system compares this overall confidence level with a preset compensation trigger confidence threshold; once the overall confidence level exceeds this threshold, the system determines that the current audio frame is in the target physiological compensation state that produces high-frequency vibrato.

[0087] S105. Based on the target physiological compensation state, establish a spatial transfer matrix that allows the deformation response center to smoothly shift from the lip region and mandibular region to the neck region, and calculate the transfer parameters corresponding to each region based on the spatial transfer matrix.

[0088] Based on the target physiological compensation state, the system establishes a spatial transfer matrix that smoothly shifts the deformation response center from the lip and mandibular regions to the neck region. Based on this spatial transfer matrix, the system calculates the corresponding transfer parameters for each region. The deformation response center refers to the mesh region in the 3D character model that currently experiences the maximum deformation driving force and the most significant displacement; for example, but not limited to, the area around the mouth during normal speech. The spatial transfer matrix is ​​a mathematical matrix model used to describe and calculate the directional flow and redistribution of deformation driving force from one region to another in three-dimensional space. Transfer parameters are specific values ​​extracted from the spatial transfer matrix to quantify the proportion of deformation each region should bear during energy transfer. The deformation bearing limit threshold refers to the maximum deformation driving energy that the mesh in the lip and mandibular regions can withstand without clipping, distortion, or violating anatomical norms. This threshold is derived from physical constraint tests of the model's topology. Overflow compensation energy refers to the excess energy after subtracting the deformation bearing limit threshold from the energy amplitude characteristic value of the current audio frame; it represents the additional sound intensity that needs to be transferred out. The node mapping value refers to the share of overflow energy allocated to each independent control node along the transfer path. The transfer path refers to the spatial topological route for energy conduction from the lips and jaw to the neck. From the inventor's perspective, once the system determines that the character has entered the target physiological compensatory state, the real challenge lies in how to transform this intangible physiological state into a visualized model deformation. The system obtains the deformation bearing limit threshold of the lips and jaw, much like measuring the capacity of a container. When the input energy amplitude characteristics are too large, exceeding this capacity, the system calculates the overflow compensatory energy.

[0089] This step can be implemented in the following way: Obtain the deformation bearing limit threshold of the lip and mandibular regions under the target physiological compensation state; compare the energy amplitude characteristics in the acoustic energy features with the deformation bearing limit threshold to calculate the overflow compensation energy exceeding the deformation bearing limit threshold; construct a spatial transfer matrix with the lip and mandibular regions as the starting point and the neck region as the ending point based on the overflow compensation energy; extract the node mapping values ​​of the lip, mandibular, and neck regions on the transfer path from the spatial transfer matrix, and use the node mapping values ​​as the corresponding transfer parameters for each region.

[0090] The system can employ an algorithm based on vertex distance fields and energy attenuation functions to implement the refinement technique in this step. First, the system reads the maximum displacement limit parameters of the control nodes in the lip and mandibular regions from the 3D character's rigging configuration file, using these as deformation bearing limit thresholds. The system subtracts the currently extracted energy amplitude feature from this limit threshold; if the result is positive, the difference is recorded as overflow compensation energy. Subsequently, the system calculates the direction vectors in 3D space, using the geometric centers of the lip and mandibular regions as the starting point and the geometric center of the neck region as the ending point, constructing a spatial transfer matrix. The system uses the vertex distance field algorithm to calculate the spatial distance from all independent control nodes on the transfer path to the starting point. The system introduces an exponentially decaying energy distribution function, using the overflow compensation energy as the initial input, and calculates the attenuation value of energy during transmission based on the distance coordinates of each node in the spatial transfer matrix. Mandibular nodes closer to the starting point absorb less energy, while neck nodes at the ending point absorb most of the attenuated overflow compensation energy. The system normalizes the calculated energy values ​​absorbed by each node and extracts them as node mapping values, i.e., the transfer parameters corresponding to each region.

[0091] The system can also use algorithms based on graph neural networks and heat conduction models to implement the refinement techniques in this step. The system abstracts the independent control nodes of the 3D character as vertices in a graph structure, and the anatomical connections between nodes as edges, thus constructing a topological graph representing the transfer path. The system sets the deformation bearing limit threshold as the upper limit of the heat capacity of the lip and mandibular nodes in the graph. The system treats energy amplitude characteristics as heat source input; when the heat source input exceeds the upper limit of heat capacity, it calculates the overflow compensation energy, i.e., the excess heat. The system uses the heat conduction equation to construct a spatial transfer matrix, simulating the diffusion process of excess heat in the topological graph, causing heat to flow directionally from the lip and mandibular nodes along the edges to the neck nodes. This graph neural network can achieve this function because it can accurately simulate the nonlinear diffusion behavior of energy in complex mesh topologies through multiple information transmission iterations. After the heat conduction simulation reaches a steady state, the system reads the final heat distribution ratios on the lip, mandibular, and neck nodes, extracts these ratio values ​​as node mapping values, and directly uses them as the corresponding transfer parameters for each region.

[0092] S106. Assign dynamic driving weights to each region according to the transfer parameters;

[0093] The system allocates dynamic driving weights to each region based on transfer parameters. Dynamic driving weights refer to floating-point multiplier factors calculated in real-time and assigned to each independent control node during model animation rendering. These factors amplify or reduce the basic deformation amplitude of the corresponding region, such as, but not limited to, the weight value controlling the degree of neck muscle expansion. Weight decay values ​​refer to the driving weight values ​​that need to be reduced in the lip and jaw regions due to the shift in the deformation response center during energy transfer. Weight increase values ​​refer to the driving weight values ​​that need to be increased in the neck region to absorb overflow energy. Target driving weights refer to the final weight values ​​that each region should ideally achieve, calculated based on the transfer parameters of the current audio frame. Historical driving weights refer to the actual driving weight values ​​used by each region in the previous audio frame immediately preceding the current frame within a preset sliding time window. The preset smooth interpolation algorithm is a mathematical transition calculation method used to generate a series of continuous intermediate values ​​between two discrete values ​​to eliminate numerical abrupt changes, such as, but not limited to, linear interpolation or Bézier curve interpolation.

[0094] This step can be implemented in the following way: Calculate the weight decay value of the lip region and the mandibular region, and the weight increase value of the neck region, based on the transfer parameters; determine the target driving weight of each region in the current audio frame based on the weight decay value and the weight increase value; extract the historical driving weight of each region in the previous audio frame within the sliding time window; and use a preset smooth interpolation algorithm to interpolate the historical driving weight and the target driving weight to obtain the dynamic driving weight of each region.

[0095] The system can implement the refinement technique in this step using linear interpolation and a buffer queue. First, the system multiplies the acquired transfer parameters by a preset global scaling factor. Negative values ​​are taken for the transfer parameters corresponding to the lip and jaw regions as weight attenuation values, while positive values ​​are taken for the transfer parameters corresponding to the neck region as weight increment values. The system performs vector addition on the base weight distribution with these attenuation and increment values ​​to calculate the target driving weight for each region in the current audio frame. Simultaneously, the system maintains a first-in-first-out buffer queue in memory to store the final weights after rendering each frame in real time. The system reads the weight data written in the previous audio frame from the tail of this buffer queue as historical driving weights. Then, the system calls a standard linear interpolation function, sets a fixed smoothing factor (e.g., 0.2), and performs linear calculations starting with the historical driving weights and ending with the target driving weights. The system uses the interpolation result as the dynamic driving weight for each region in the current frame and pushes it into the buffer queue for use in the next frame.

[0096] The system can also use an interpolation algorithm based on a spring-damped dynamics model to implement the refinement technique in this step. Similarly, based on the proportional relationship of the transfer parameters, the system calculates the weight attenuation values ​​for the lips and jaw, and the weight increase values ​​for the neck, through matrix multiplication, thereby determining the target driving weights for each region. The system not only extracts historical driving weights from the previous audio frame but also extracts the rate of change of the weights in the previous frame (i.e., the first derivative). The system treats the historical driving weights as the current position of the spring oscillator and the target driving weights as the equilibrium position of the spring. The system inputs these parameters into a preset critical damped spring physical equation for iterative calculation. By simulating Hooke's law of the spring and the viscous resistance of the damper, the system calculates the new position of the oscillator at the current time step; the value of this new position is the dynamic driving weight for each region.

[0097] S107. Map the phoneme sequence to the basic deformation parameters, and use dynamic driving weights to perform weighted calculations on the basic deformation parameters to generate independent driving commands for the lip region, mandibular region and neck region.

[0098] Basic deformation parameters refer to the initial geometric transformation data that the system maps by default, based on the input phoneme sequence, to control the opening and closing of the 3D character's mouth and facial muscle movement, without considering vocal intensity and physiological compensation. Examples include, but are not limited to, the chin descent distance and mouth corner stretching vector corresponding to the vowel "ah". Weighted calculation refers to the process of performing mathematical multiplication or tensor product operations on the above basic deformation parameters and the dynamic driving weights obtained in step S106 to adjust the deformation amplitude according to the current physiological state. Independent driving instructions refer to the final execution commands generated after weighted calculation, targeting the independent control nodes of the lip region, jaw region, and neck region. These instructions contain specific spatial coordinate offsets or rotation quaternions and can be directly read and executed by the rendering engine.

[0099] The system can implement this step using a technique based on a pre-built deformation dictionary and scalar multiplication. The system pre-builds a phoneme-to-visual-morphology mapping dictionary in local memory. The keys of the dictionary are various phoneme combinations, and the values ​​are a set of standardized basic deformation parameters (e.g., an initial floating-point array containing fifty facial blending deformation channels). Using the currently extracted phoneme sequence as an index, the system performs a hash lookup in the dictionary to retrieve the corresponding basic deformation parameters. Subsequently, the system extracts subsets of parameters belonging to the lips, jaw, and neck from the basic deformation parameters. The system performs element-wise scalar multiplication of these parameter subsets with the dynamic driving weights of the corresponding regions. For example, multiplying the basic deformation parameters of the neck by the dynamic driving weights of the neck will amplify the deformation parameters because the neck weights increase significantly during high-frequency vibrato. The system repackages the weighted parameters, formats them according to the rendering engine's communication protocol into data packets containing channel identifiers and target values, and generates independent driving instructions for each region.

[0100] The system can also implement this step using deep learning techniques based on multilayer perceptrons and tensor operations. The system converts phoneme sequences into one-hot encoded vectors and inputs them into a multilayer perceptron network trained on large-scale motion capture data. This large-scale AI model can achieve this function because its hidden layer nodes can learn the nonlinear mapping between phoneme sequences and complex facial muscle movements, thus directly outputting continuous high-dimensional fundamental deformation parameter tensors. Next, the system constructs a diagonal weight matrix for the dynamic driving weights of the lips, jaw, and neck. The system performs tensor multiplication operations in the computation graph, multiplying the fundamental deformation parameter tensors with the diagonal weight matrix to achieve global weighted calculation. Using tensor slicing operations, the system divides the multiplied tensor into three independent sub-tensors according to anatomical regions. These three sub-tensors correspond to independent driving commands for the lip, jaw, and neck regions, respectively.

[0101] S108. Input the independent drive commands to the corresponding independent control nodes to drive the mesh vertices of the 3D character dynamic model to perform coordinate updates and deformation rendering.

[0102] Mesh vertices are the basic spatial points that constitute the topological structure of a 3D character's surface. Each vertex contains spatial coordinate information in three dimensions (X, Y, and Z) as well as additional attributes such as normals and texture coordinates. Coordinate update refers to the process of recalculating and modifying the position information of each mesh vertex in the basic mesh model in the three-dimensional world coordinate system or local coordinate system based on received independent driving instructions and using specific spatial transformation algorithms. Deformation rendering refers to the process by which the graphics processing unit (GPU) receives the updated mesh vertex coordinate data, processes it through the graphics pipeline such as vertex shaders, rasterization, and fragment shaders, and finally renders the deformed 3D character surface onto a two-dimensional screen with lighting, materials, and shadow effects.

[0103] The system can implement this step using a CPU-based skeletal skinning animation pipeline. The system parses the independent drive commands for the lips, jaw, and neck into translation vectors and rotation quaternions for each independent control node (i.e., bone) in the local coordinate system. The system traverses the skeletal hierarchy of the 3D character, using matrix multiplication to multiply the local transformation matrices to calculate the global transformation matrix for each bone. Subsequently, the system traverses all mesh vertices of the base mesh model in the CPU, reading the bone index bound to each vertex and its corresponding skinning weight. The system applies a linear blending skinning algorithm, performing a weighted summation of the initial coordinates of each vertex based on the global transformation matrix of the bones and the skinning weights, thus updating the coordinates of the mesh vertices. Finally, the system uploads the updated vertex coordinate array to video memory via a graphics application programming interface (such as OpenGL or DirectX), triggering the graphics card's rendering pipeline to perform deformation rendering, drawing the 3D character with muscle compensation effects onto the screen.

[0104] The system can also implement this step using a computational shader based on the graphics processing unit (GPU) and a hybrid deformation pipeline. The system directly writes the independent drive instructions (i.e., the target weight values ​​for each deformation channel) into a unified buffer object in video memory. The system schedules a specially written computational shader program to execute in parallel on the GPU. This computational shader reads the initial vertex coordinate buffer of the base mesh model, as well as a structured buffer storing vertex offsets under various extreme deformation states of the lips, jaw, and neck. The computational shader uses the drive instructions in the unified buffer object as multipliers to perform massively parallel multiply-accumulate operations on the vertex offsets, directly updating the coordinates of all mesh vertices within video memory. The updated vertex data does not need to be sent back to the CPU; instead, it is directly used as input to the vertex shader, working in conjunction with physically based rendering materials to efficiently complete the final deformation rendering.

[0105] In the above embodiments, by configuring independent control nodes corresponding to the lip region, jaw region, and neck region in the basic mesh model of the 3D character, and determining that the target physiological compensation state for generating high-frequency vibrato is in place when the periodic fluctuation frequency of the audio is detected to be greater than a preset high-frequency threshold and the phoneme sequence remains continuous, a spatial transfer matrix is ​​established to smoothly transfer the deformation response center from the lips and jaw to the neck, and dynamic driving weights are assigned. This allows the model to simulate the neck muscle compensation response during real human vocalization based on the acoustic characteristics of high-frequency vibrato. This scheme combines the basic deformation parameters of the phoneme sequence mapping with the dynamic driving weights calculated based on the physiological compensation state to generate independent driving instructions for each region, thereby improving the accuracy of lip-syncing in complex performance scenarios for 3D characters.

[0106] In the above embodiments, by calculating transfer parameters and assigning dynamic driving weights, the system has been able to precisely control the local coordinated deformation of the main vocalization area and the neck compensation area. However, in real human vocalization, local muscle tension changes usually trigger coordinated responses in other facial areas. If only the deformation of local areas is controlled, it can easily lead to stiffness in other facial areas, resulting in a visual disconnect. To further improve the global naturalness and coordination of 3D human facial expressions, this application, based on the above embodiments, further applies the locally calculated transfer parameters to the coordinated driving of non-vocalization areas. The following is a combination of... Figure 2 Another method for establishing a 3D character dynamic model and lip-syncing in the embodiments of this application is described below:

[0107] Please see Figure 2 This is another flowchart illustrating a method for establishing a 3D character dynamic model and lip-syncing in an embodiment of this application.

[0108] S201, Characterize the transfer parameters as a measure of the tension deviation of the current facial system from the natural expression manifold;

[0109] The facial system refers to the overall structure of the mesh topology and control nodes of a 3D human head involved in facial expressions and lip movements. The natural expression manifold, in mathematical space, represents a low-dimensional data subspace comprising the deformation parameters of various facial regions in a relaxed or normal vocal state; for example, but not limited to, the set of parameters for a neutral expression or a slight smile. Tension deviation metric is a numerical indicator used to quantify the deviation of the current facial muscle tension from the aforementioned relaxed state.

[0110] The system can implement this step using a linear mapping matrix technique. The system pre-configures a tension coefficient matrix in memory, where each element represents the weight of the influence of different transfer parameters on overall facial tension. The system extracts the transfer parameters of the current audio frame and constructs a one-dimensional column vector. Using matrix multiplication, this column vector is multiplied by the tension coefficient matrix, and the resulting vector is the tension deviation measure.

[0111] The system can also use a latent space distance calculation algorithm to perform this step. The system constructs a low-dimensional latent space representing the natural facial expression manifold using principal component analysis. The system projects the current transition parameters into this latent space to obtain the current coordinate point and calculates the Euclidean distance between this coordinate point and the origin of the latent space. The system uses the calculated Euclidean distance value as a measure of the tension deviation of the current facial system from the natural facial expression manifold.

[0112] S202. Input the tension deviation measurement into the preset global correlation mapping model;

[0113] Tension deviation measurement is input into a pre-defined global correlation mapping model, which includes potential statistical correlation features between the primary vocalization region, compensatory region, and non-vocalization-involved region. The pre-defined global correlation mapping model refers to a pre-trained algorithm or neural network used to handle the linkage relationships between different facial regions. The primary vocalization region refers to the lips and jaw region directly involved in vocalization. The compensatory region refers to the neck region that shares the deformation pressure during vigorous vocalization. The non-vocalization-involved region refers to regions that do not directly produce sound but are affected by facial muscle traction, such as, but not limited to, the forehead or nose. Potential statistical correlation features are mathematical patterns extracted from a large amount of real human facial motion capture data, reflecting the interdependence and coordinated changes of the above three regions during movement.

[0114] The system can use a Gaussian mixture model to implement the refinement technique in this step. The system pre-trains the Gaussian mixture model on a full-face motion capture dataset using the expectation-maximization algorithm, allowing it to fit the potential statistical association features of the joint distribution of the main vocal region, compensatory region, and non-vocal participation region. During runtime, the system inputs the tension deviation metric as a known condition into the model and uses the conditional probability density function to calculate the optimal estimate of the state of the non-vocal participation region.

[0115] The system can also use a multilayer perceptron neural network to implement the refinement technique in this step. The system constructs a feedforward neural network containing multiple hidden layers. This large-scale artificial intelligence model can achieve the corresponding functions because its hidden layer nodes, through nonlinear activation functions, can fully learn and memorize the complex latent statistical correlation features between various facial regions in the training data. The system inputs the tension deviation metric into the input layer of this network, and calculates it layer by layer through the forward propagation algorithm, activating the correlation feature responses within the network.

[0116] S203. Based on the preset global correlation mapping model, generate induced driving weights that match the tension deviation metric;

[0117] The system generates induced driving weights that match the tension deviation metric based on a pre-defined global correlation mapping model. These induced driving weights are floating-point control parameters calculated by the system specifically to control the cooperative deformation of non-vocal areas. Initial state features refer to the current geometric displacement or control node values ​​of the non-vocal areas before applying cooperative deformation. The eye, eyebrow, and cheek regions correspond to the mesh sets of the 3D character's eye socket, eyebrow, and cheek, respectively. Feature concatenation refers to the operation of concatenating and combining data vectors from different sources in a dimensional manner. The fused feature vector is the data set containing global information formed after concatenation. The feature decoding network is the module in the model responsible for converting abstract features into specific spatial relationships. Spatial linkage correlation refers to the numerical value that quantifies the strength of mutual influence between different regions. Attention-weighted mapping is a mechanism for allocating resources and adjusting values ​​for features based on the degree of correlation.

[0118] This step can be implemented in the following way: obtain the initial state features of the non-speech-involved regions, which include the eye region, eyebrow region, and cheek region; concatenate the tension deviation measure with the initial state features to construct a multi-dimensional fused feature vector; input the fused feature vector into the feature decoding network in the preset global correlation mapping model to extract the spatial linkage correlation between the main speech region, the compensation region, and each non-speech-involved region; perform attention-weighted mapping on the fused feature vector based on the spatial linkage correlation to output the induction driving weights corresponding to each non-speech-involved region.

[0119] The system can implement this refinement technique using a transformer decoding network based on a self-attention mechanism. The system reads the current blending deformation weights of the eye, eyebrow, and cheek regions as initial state features, and concatenates them with a tension deviation metric using a vector concatenation function to form a fused feature vector. This vector is then input into the transformer decoding network. This model achieves this functionality because the self-attention mechanism adaptively captures the spatial linkage between regions by calculating the dot product of the query matrix and the key matrix. The system uses the output attention score matrix to perform a weighted summation of the value matrix, completing the attention weighted mapping. Finally, it outputs the induction driving weights for each non-speech-participating region through a fully connected layer.

[0120] The system can also use a graph convolutional network combined with a channel attention mechanism to achieve this refinement technique. The system maps initial state features and tension deviation metrics to graph node attributes for feature concatenation. The system inputs the fused feature vector into the graph convolutional network. This model achieves this function because graph convolution operations, by aggregating neighbor node information, naturally align with extracting the spatial interconnectivity of facial topology. The system then uses a squeeze-excitation network to calculate channel attention weights, performing attention-weighted mapping on node features and outputting induced driving weights.

[0121] S204. Apply induction-driven weights to non-vocal participation regions to reconstruct the global facial state of the 3D character into a co-expressive manifold that satisfies potential statistical association features.

[0122] The collaborative expression manifold refers to a data space in a mathematical model that represents the overall facial state of a 3D character, where all areas of the face are in a highly coordinated manner, conforming to the physiological linkage patterns of real humans. The system precisely applies the calculated induction driving weights to non-vocalization areas such as the eyes, eyebrows, and cheeks. Through this operation, the system forces the upper half of the face, which was originally static or only slightly moving, to make corresponding muscle contractions or relaxations based on the intense vocalization state of the lower half of the face and neck.

[0123] The system can implement this step using a vertex offset stacking technique based on a hybrid deformation pipeline. The system transmits the generated induced driving weights to the animation controller of the graphics rendering engine. The system iterates through all target deformable meshes corresponding to the eye, eyebrow, and cheek regions, performing a scalar multiplication operation between the vertex offset of each target mesh and its corresponding induced driving weight. Using vector addition, the system superimposes the weighted vertex offsets onto the current base mesh vertex coordinates of the 3D character, thereby updating the physical form of non-vocalization areas and reconstructing the global facial state into a co-expressive manifold.

[0124] The system can also use the joint rotation driving technology based on skeletal skinning animation to implement this step. The system maps the induced driving weights to the rotation angle values of the subdivision facial bones that control the trends of the eye, eyebrow, and cheek muscles. The system updates the rotation states of these facial bones in the local coordinate system using the quaternion interpolation algorithm and calculates the global transformation matrix. The system uses the linear blend skinning algorithm to drive the mesh vertices in the non-articulation participation area to displace using the updated bone matrix, thereby achieving the reconstruction of the full facial state towards the co-expression manifold.

[0125] In the above embodiments, by characterizing the transfer parameter as the tension deviation metric of the current facial system deviating from the natural expression manifold and inputting it into the preset global correlation mapping model containing the potential statistical correlation features between the main articulation area, compensation area, and non-articulation participation area, the matching induced driving weights are generated and applied to the non-articulation participation area, enabling the non-articulation area to generate cooperative deformation according to the local tension changes during articulation. Extending the local articulation compensation action to the reconstruction of the full facial state simulates the natural linkage reaction of other areas of the human face when speaking forcefully, thereby improving the global coordination and comprehensive expressiveness of the overall facial expression of the 3D character and reducing the visual fragmentation caused by the local lip-sync action and the immobility of other areas of the face.

[0126] In the above embodiments, by applying the induced driving weights to the non-articulation participation area, the system successfully achieved the co-expression of the global facial state. However, during the actual mesh driving process, especially when performing complex or exaggerated articulation actions, the severe deformation of the main articulation area and the cooperative deformation of the non-articulation participation area may cause serious reverse pulling in the three-dimensional space, resulting in mesh tearing or excessive distortion of the model. To maintain the stability of the mesh structure while ensuring the accuracy of lip-sync expression, the present application further introduces a dynamic monitoring and correction mechanism for the induced driving weights based on the above embodiments. The following is combined with Figure 3 , to describe another method for establishing a 3D character dynamic model and lip-sync in the embodiments of the present application:

[0127] Please refer to Figure 3 , which is another process schematic diagram of a method for establishing a 3D character dynamic model and lip-sync in the embodiments of the present application.

[0128] S301. Obtain the first deformation displacement vector of the main articulation area under the current audio frame and the second deformation displacement vector of the non-articulation participation area determined based on the induced driving weights;

[0129] The first deformation displacement vector refers to vector data characterizing the direction and distance of movement of grid vertices within the main articulation region in three-dimensional space, such as, but not limited to, the displacement vectors of the lip or mandibular vertices. The second deformation displacement vector refers to vector data characterizing the expected direction and distance of movement of non-articulation regions under the influence of induced driving weights.

[0130] The system can perform this step by reading vertex cache data from the graphics rendering pipeline. The system accesses the memory space of the animation controller, extracts the 3D coordinate differences between the vertices of the main sound region mesh in the current frame and the previous frame as the first deformation displacement vector, and calculates the second deformation displacement vector based on the product of the induced driving weight and the target deformation mesh.

[0131] The system can also perform this step by calling the application programming interface of the hybrid deformation engine. The system sends a query command to the hybrid deformation engine to directly obtain the local spatial offset vector corresponding to the currently active hybrid deformation channel in the main sound-producing region as the first deformation displacement vector, and obtains the expected offset vector of the non-sound-producing region under the induced driving weight mapping as the second deformation displacement vector.

[0132] S302. Calculate the topological antagonism between the first deformation displacement vector and the second deformation displacement vector;

[0133] The system calculates the topological antagonism between the first and second deformation displacement vectors. Topological antagonism is a numerical value characterizing the degree of reverse stretching between the primary vocal region and the non-vocal participation region in three-dimensional space. It quantifies the degree of reverse stretching or compression caused by the inconsistent movement directions of the primary vocal region and the non-vocal participation region in three-dimensional space. The first grid vertex refers to a three-dimensional coordinate point located on the topological structure at the edge of the primary vocal region. The second grid vertex refers to a three-dimensional coordinate point located on the topological structure at the edge of the non-vocal participation region. The initial spatial distance refers to the length span between the two types of vertices when the face is in a natural neutral state. The target spatial distance refers to the new length span between the two types of vertices after applying the displacement vector. The deformation stretching rate is a numerical value reflecting the proportion of length change before and after local grid deformation. The included angle refers to the degree difference in direction between the two displacement vectors in three-dimensional space. From the inventor's perspective, the system needs to accurately measure the severity of stretching in different areas of the face. The system selects key points on the boundaries of two regions, compares the changes in their distances before and after deformation to obtain the stretching ratio, and then combines the degree of divergence in the movement directions of the two regions to finally calculate a comprehensive value that reflects the risk of mesh tearing.

[0134] This step can be implemented as follows: Obtain the first grid vertex located at the boundary of the main sound-producing region and the second grid vertex located at the boundary of the non-sound-producing region; calculate the initial spatial distance between the first grid vertex and the second grid vertex in the undeformed state; update the position of the first grid vertex according to the first deformation displacement vector and update the position of the second grid vertex according to the second deformation displacement vector; calculate the target spatial distance between the updated first grid vertex and the second grid vertex; calculate the ratio of the target spatial distance to the initial spatial distance to obtain the deformation stretching ratio; calculate the angle between the first deformation displacement vector and the second deformation displacement vector; and use the product of the deformation stretching ratio and the angle as the topological antagonism.

[0135] The system can implement this refinement technique using Euclidean distance and inverse cosine function algorithms. The system traverses the boundary vertex indices of the main sound-producing region and the non-sound-producing region, extracts the first and second grid vertices, and calculates the initial spatial distance between them in the undeformed state using the Euclidean distance formula. The system then adds the first and second deformation displacement vectors to the corresponding vertex coordinates to update the positions, and again calculates the target spatial distance using the Euclidean distance formula. The system obtains the deformation stretching rate through division, calculates the angle between the two displacement vectors using the dot product formula and the inverse cosine function, and finally multiplies the two to obtain the topological antagonism.

[0136] The system can also implement this refinement technique using geodesic distance and cosine similarity algorithms. The system uses a shortest path algorithm to calculate the initial geodesic distance between the first and second mesh vertices along the surface contour on the facial mesh topology. After updating the vertices based on the displacement vectors, the system recalculates the target geodesic distance and divides it to obtain the deformation elongation rate. The system uses a cosine similarity algorithm to calculate the cosine values ​​of the two displacement vector directions and converts them into an angle; finally, a multiplication operation is performed to obtain the topological antagonism.

[0137] S303. When the topological antagonism is greater than the preset tensile limit threshold, keep the first deformation displacement vector unchanged, and calculate the attenuation coefficient based on the difference between the topological antagonism and the tensile limit threshold.

[0138] The preset stretch limit threshold refers to the maximum topological antagonism value allowed by the system without visual tearing or excessive distortion of the facial mesh. This threshold is derived from simulation test data of the physical properties of the facial mesh material or binding limit parameters set by the artists. The difference refers to the specific value by which the current topological antagonism exceeds the preset stretch limit threshold. The two are in a comparative relationship; the preset stretch limit threshold serves as a safety boundary, and the magnitude of the difference directly determines the strength of subsequent corrections. The attenuation coefficient refers to a floating-point value used to proportionally reduce the original driving weights, such as, but not limited to, decimals between zero and one.

[0139] The system can implement this step using a linear mapping algorithm. The system uses a comparator to determine if the topological antagonism is greater than a preset tensile limit threshold. If the condition is met, the system locks the value of the first deformation displacement vector in memory and refuses to modify it. The system uses a subtractor to calculate the difference between the topological antagonism and the preset tensile limit threshold, divides this difference by a preset maximum tolerance difference constant to obtain a normalized over-limit ratio, and subtracts this over-limit ratio from one to calculate the attenuation coefficient, which is between zero and one.

[0140] The system can also use an exponential decay function to achieve this step. After determining that the topological antagonism exceeds the limit, the system keeps the displacement parameters of the main sound region unchanged. The system calculates the difference between the topological antagonism and the preset stretching limit threshold, and substitutes this difference as the independent variable of the negative exponent of the natural constant into the exponential decay function. By calculating the value of this exponential function, the system directly outputs a decay coefficient that decreases smoothly as the difference increases.

[0141] S304. Reduce the induced driving weight using the attenuation coefficient to obtain the corrected induced driving weight;

[0142] The revised induced drive weights refer to the final control parameters, after numerical reduction, that ensure the safety of the facial mesh structure without causing excessive stretching. The system needs to implement the newly calculated penalty mechanism into specific control parameters.

[0143] The system can implement this step using scalar multiplication. The system reads the previously generated induced driving weight floating-point value and the attenuation coefficient calculated in the previous step from memory. The system calls the floating-point unit of the central processing unit to perform a direct multiplication operation between the induced driving weight and the attenuation coefficient. Taking advantage of the fact that the attenuation coefficient is less than one, the original weight value is proportionally reduced, and the product result is overwritten to the original memory address as the corrected induced driving weight.

[0144] The system can also use a percentage-based subtraction compensation algorithm to implement this step. The system reads the attenuation coefficient and subtracts it from the original value to obtain the percentage reduction required. The system multiplies the induced driving weight by this percentage reduction to calculate the specific weight value to be deducted. The system then subtracts this deducted value from the original induced driving weight and outputs the final difference as the corrected induced driving weight.

[0145] S305. Apply the modified induction-driven weights to the non-speech-participating regions.

[0146] The system sends the driver commands, which have undergone security degradation processing, to the control nodes in non-vocalization areas. Upon receiving these commands, the non-vocalization areas coordinate with the actions of the main vocalization area, thus preserving the natural feel of global linkage while reducing the risk of mesh tearing.

[0147] The system can achieve this step by updating the blending deformation channel weights. The system encapsulates the corrected induced-drive weights into a specific data packet and sends it to the facial animation system of the 3D rendering engine. Based on the vertex group identifiers of the non-vocalization participating regions, the system locates the corresponding blending deformation target channel, replaces the current value of that channel with the corrected induced-drive weights, and triggers the rendering pipeline to recalculate the final positions of the mesh vertices.

[0148] The system can also achieve this step by adjusting the local transformation matrix of the facial bones. The system transforms the corrected induced driving weights into displacement and rotation parameters of the bottom-level driving bones in the non-vocalization region. The system uses these parameters to update the transformation matrix of the corresponding bones in the local coordinate system and calculates the global transformation matrix through the bone hierarchy structure. Finally, the system uses skinning weights to apply the bone transformation to the mesh vertices in the non-vocalization region, achieving the corrected cooperative deformation.

[0149] In the above embodiments, by acquiring the first deformation displacement vector of the main vocal region and the second deformation displacement vector of the non-vocalization region, and calculating the topological antagonism of their reverse stretching in three-dimensional space, the structural tension of the facial mesh during the linkage deformation process can be monitored in real time. When the topological antagonism exceeds a preset stretching limit threshold, while keeping the deformation displacement vector of the main vocal region unchanged, an attenuation coefficient is calculated based on the difference exceeding the threshold to reduce the induction driving weight of the non-vocalization region. This allows the system to prioritize lip-sync accuracy while dynamically suppressing excessive deformation of the non-vocalization region, reducing the risk of mesh tearing or excessive distortion when the 3D character model performs complex or exaggerated expressions, and improving the structural stability and visual smoothness of the mesh dynamic deformation process.

[0150] The system in the embodiments of this invention is described below from the perspective of hardware processing. Please refer to [link / reference needed]. Figure 4 This is a schematic diagram of the physical device structure of a 3D character dynamic model creation and lip-syncing system provided in an embodiment of this application.

[0151] It should be noted that, Figure 4 The structure of the system shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.

[0152] like Figure 4As shown, the system includes a Central Processing Unit (CPU) 401, which can perform various appropriate actions and processes based on a program stored in Read-Only Memory (ROM) 402 or a program loaded from storage portion 408 into Random Access Memory (RAM) 403, such as executing the methods described in the above embodiments. The RAM 403 also stores various programs and data required for system operation. The CPU 401, ROM 402, and RAM 403 are interconnected via a bus 404. An Input / Output (I / O) interface 405 is also connected to the bus 404.

[0153] The following components are connected to I / O interface 405: input section 406 including a camera, infrared sensor, etc.; output section 407 including a liquid crystal display (LCD) and speakers, etc.; storage section 408 including a hard disk, etc.; and communication section 409 including a network interface card such as a LAN (Local Area Network) card and a modem, etc. Communication section 409 performs communication processing via a network such as the Internet. Drive 410 is also connected to I / O interface 405 as needed. Removable media 411, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 410 as needed so that computer programs read from it can be installed into storage section 408 as needed.

[0154] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 409, and / or installed from removable medium 411. When the computer program is executed by central processing unit (CPU) 401, it performs the various functions defined in the present invention.

[0155] It should be noted that the computer-readable medium shown in the embodiments of the present invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, wherein a computer-readable computer program is carried. The transmitted data signal can take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof.

[0156] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0157] In another aspect, the present invention also provides a computer-readable storage medium, which may be included in the system described in the above embodiments; or it may exist independently and not assembled into the system. The storage medium carries one or more computer programs that, when executed by a processor of a system, cause the system to implement the methods provided in the above embodiments.

[0158] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0159] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".

[0160] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0161] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A method for establishing a 3D character dynamic model and lip-syncing, characterized in that, include: Configure independent control nodes for the lip, jaw and neck regions in the basic mesh model of the 3D character to create a dynamic 3D character model with a layered linkage structure. Based on a preset sliding time window, feature extraction is performed on the input audio signal to obtain a phoneme sequence for characterizing the pronunciation content and an acoustic energy feature for characterizing the sound intensity. The acoustic energy feature includes fundamental frequency feature and energy amplitude feature. Extract the periodic fluctuation frequency and amplitude difference of the acoustic energy characteristics within the preset sliding time window; If the periodic fluctuation frequency is greater than a preset high-frequency threshold and the phoneme sequence remains continuous within the sliding time window, the current audio frame is determined to be in the target physiological compensatory state that generates high-frequency vibrato. Based on the target physiological compensation state, a spatial transfer matrix is ​​established to smoothly transfer the deformation response center from the lip region and the mandibular region to the neck region, and the transfer parameters corresponding to each region are calculated based on the spatial transfer matrix. Dynamic driving weights are assigned to each region based on the transfer parameters; The phoneme sequence is mapped to basic deformation parameters, and the basic deformation parameters are weighted using the dynamic driving weights to generate independent driving commands for the lip region, the jaw region, and the neck region. The independent driving instructions are input to the corresponding independent control nodes to drive the mesh vertices of the 3D character dynamic model to perform coordinate updates and deformation rendering.

2. The method according to claim 1, characterized in that, Based on the target physiological compensation state, a spatial transfer matrix is ​​established to smoothly transfer the deformation response center from the lip region and the mandibular region to the neck region, and the transfer parameters corresponding to each region are calculated based on the spatial transfer matrix, specifically including: Obtain the deformation bearing limit threshold of the lip region and the mandibular region under the target physiological compensation state; The energy amplitude characteristics in the acoustic energy features are compared with the deformation bearing limit threshold to calculate the overflow compensation energy that exceeds the deformation bearing limit threshold. Based on the overflow compensation energy, a spatial transfer matrix is ​​constructed with the lip region and the mandibular region as the starting point and the neck region as the ending point; Extract the node mapping values ​​of the lip region, the mandibular region, and the neck region on the transfer path from the spatial transfer matrix, and use the node mapping values ​​as the transfer parameters corresponding to each region.

3. The method according to claim 1, characterized in that, The allocation of dynamic driving weights for each region based on the transfer parameters specifically includes: Based on the transfer parameters, the weight attenuation values ​​for the lip region and the mandibular region, and the weight increase value for the neck region are calculated respectively. Based on the weight decay value and the weight increase value, the target driving weight of each region in the current audio frame is determined; Extract the historical driving weights of each region in the previous audio frame within the sliding time window; The historical driving weights and the target driving weights are interpolated using a preset smooth interpolation algorithm to obtain the dynamic driving weights for each region.

4. The method according to claim 1, characterized in that, After allocating the dynamic driving weights of each region according to the transfer parameters, the method further includes: The transfer parameters are characterized as a measure of the tension deviation of the current facial system from the natural expression manifold; The tension deviation metric is input into a preset global correlation mapping model, which includes potential statistical correlation features between the primary vocal region, the compensatory region, and the non-vocal participation region. Based on the preset global correlation mapping model, an induced driving weight that matches the tension deviation metric is generated; The induced driving weights are applied to the non-vocal participation region to reconstruct the global facial state of the 3D character into a co-expressive manifold that satisfies the latent statistical association features.

5. The method according to claim 4, characterized in that, The step of generating induced driving weights that match the tension deviation metric based on the preset global correlation mapping model specifically includes: Obtain the initial state features of the non-speech-involved region, which includes the eye region, eyebrow region, and cheek region; The tension deviation metric and the initial state features are concatenated to construct a multi-dimensional fused feature vector; The fused feature vector is input into the feature decoding network in the preset global correlation mapping model to extract the spatial linkage correlation between the main pronunciation region, the compensation region and each of the non-pronunciation participation regions. Based on the spatial linkage correlation, the fused feature vector is subjected to attention-weighted mapping, and the induction driving weights corresponding to each of the non-speech participation regions are output.

6. The method according to claim 4, characterized in that, Prior to applying the induced driving weights to the non-vocal participation region, the method further includes: Obtain the first deformation displacement vector of the main sound-producing region in the current audio frame, and the second deformation displacement vector of the non-sound-producing region determined based on the induced driving weight; Calculate the topological antagonism between the first deformation displacement vector and the second deformation displacement vector, wherein the topological antagonism is a numerical value characterizing the degree of reverse stretching between the main sound-producing region and the non-sound-producing region in three-dimensional space; When the topological antagonism is greater than the preset tensile limit threshold, the first deformation displacement vector remains unchanged, and the attenuation coefficient is calculated based on the difference between the topological antagonism and the tensile limit threshold. The induced driving weight is reduced by the attenuation coefficient to obtain the corrected induced driving weight. The modified induction-driven weights are applied to the non-vocalization-participating region.

7. The method according to claim 6, characterized in that, The calculation of the topological antagonism between the first deformation displacement vector and the second deformation displacement vector specifically includes: Obtain the first grid vertex located at the boundary of the main sound-producing region, and the second grid vertex located at the boundary of the non-sound-producing region; Calculate the initial spatial distance between the first mesh vertex and the second mesh vertex in the undeformed state; The positions of the first mesh vertices are updated according to the first deformation displacement vector, and the positions of the second mesh vertices are updated according to the second deformation displacement vector; Calculate the updated target space distance between the first and second grid vertices; Calculate the ratio of the target spatial distance to the initial spatial distance to obtain the deformation elongation rate; Calculate the angle between the first deformation displacement vector and the second deformation displacement vector; The product of the deformation elongation rate and the included angle is used as the topological antagonism.

8. A 3D character dynamic model creation and lip-syncing system, characterized in that, The system includes: One or more processors and a memory; the memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the system to perform the method as described in any one of claims 1-7.

9. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are executed on the system, the system performs the method as described in any one of claims 1-7.

10. A computer program product, characterized in that, When the computer program product is run on the system, the system performs the method as described in any one of claims 1-7.