Character large model driven digital avatar lip micro-decoupling synchronous generation method and product

CN122618052APending Publication Date: 2026-08-21LIANGSHENG DIGITAL CREATIVE DESIGN (HANGZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610700139.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

[0006]本发明提供一种角色大模型驱动数字分身唇微解耦同步生成方法及产品,以解决相关技术中数字分身缺少面部其他区域的微表情变化,导致数字分身的自然度与感染力不佳的问题

Benefits of technology

[0012]本发明实施例提供的角色大模型驱动数字分身唇微解耦同步生成方法及产品中,通过将音频特征序列与角色画像向量输入角色大模型进行角色情感模拟,生成表情控制信息;并基于表情控制信息作为情感约束,分别生成与音频时序对齐的唇形动作参数序列和微表情动作参数序列,进而通过对重叠区域进行冲突消解的融合处理,得到全脸表情参数序列并进行预设动作参数叠加与平滑处理,以输出全脸表情参数序列用于驱动数字分身实现唇形与微表情的同步。相较于依赖随机或模板化方式的微表情生成方案,本发明实施例通过引入角色大模型进行角色情感模拟,使生成的表情控制信息能够综合反映语音内容、韵律特征及角色属性,并对微表情生成过程进行条件化调制,从而使面部微表情能够随语音内容及韵律变化产生自适应响应,同时结合角色特征实现差异化表达,提高微表情在情绪语气与角色风格上的一致性。在此基础上,通过对唇形动作参数与微表情动作参数进行解耦建模并融合生成全脸表情参数序列,并在融合过程中针对唇形动作与微表情动作在重叠区域进行冲突消解处理,有效减少因动作叠加引起的局部形变冲突及表情失真问题;此外,还在全脸表情基础上进行预设动作叠加和平滑处理,提升面部微表情的丰富度、时序一致性和连贯性。本方案在保证唇形与音频严格同步的同时,使数字分身的微表情随情绪语气与角色风格联动变化,实现更加自然连贯的情感表达,从而在表情丰富性、唇形同步精度及整体自然度之间实现了有效平衡,显著提升了数字分身面部表现的真实感与感染力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122618052A_ABST
    Figure CN122618052A_ABST
Patent Text Reader

Abstract

The present application belongs to the field of digital avatar facial animation generation and human-computer interaction, and provides a character large model driven digital avatar lip micro-decoupling synchronous generation method and product. The method comprises: inputting an audio feature sequence and a character portrait vector into a character large model to simulate character emotions, generating expression control information as emotional constraints, respectively generating a lip shape motion parameter sequence and a micro-expression motion parameter sequence aligned with the audio time sequence, then performing conflict resolution fusion processing on the parameters of the overlapping area, obtaining a full-face expression parameter sequence and performing preset motion superposition and parameter smoothing processing, obtaining a target expression parameter sequence and driving digital avatar rendering output, realizing the synchronization of lip shape and micro-expression, and triggering expression fault tolerance rollback when the audio signal is abnormal. While ensuring strict synchronization of lip shape and audio, the micro-expression changes with the emotional tone and character style, improving the naturalness, stability and real-time interaction experience of digital avatar expressions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence technology and computer graphics, particularly to the fields of digital avatar facial animation generation and human-computer interaction technology. Specifically, it relates to a method for the synchronous generation of digital avatar lips driven by a large character model, as well as the supporting devices, computer equipment, computer-readable storage media, and computer program products. Background Technology

[0002] With the development of virtual reality, artificial intelligence, and digital content generation technologies, digital avatars (i.e., virtual characters) have been widely used in film and television production, game interaction, virtual anchors, and intelligent customer service. In these applications, it is usually necessary to generate facial animations for the digital avatar through audio-driven methods, enabling the digital avatar to change its lip movements in response to speech and complete basic expressions.

[0003] Currently, mainstream audio-driven digital avatar facial animation solutions can be divided into two categories: one is a single-region modeling solution that focuses on lip-sync. This type of solution mainly extracts audio features and directly maps them to lip movement parameters to achieve precise alignment between lip shape and audio. However, it only models the lip area and completely lacks micro-expression changes in other areas of the face. The resulting virtual human face is stiff, lacks expressiveness, and cannot achieve natural emotional expression.

[0004] Another type is the unified facial expression prediction scheme. This type of scheme extracts emotional features from audio and directly predicts facial expression parameters, generating full-face expressions with emotion. However, it has three major drawbacks: First, lip shape and micro-expression are not decoupled and modeled. The modulation of emotional features will interfere with the synchronization accuracy between lip shape and audio, easily leading to problems such as mouth shape deviation and incomplete closure of closed sounds. Second, it does not handle conflict resolution for overlapping areas such as the corners of the lips and the edge of the upper lip, which are involved in both pronunciation and emotional expression. The superposition of the two types of action parameters is prone to problems such as local deformation distortion and mouth tremors. Third, it does not set up a fault-tolerant fallback mechanism for abnormal audio scenarios. When the audio is interrupted, muted, or the signal is distorted, abnormalities such as abrupt changes in expression and freezing may occur, making it unable to adapt to the stability requirements of real-time interactive scenarios.

[0005] In summary, current audio-driven facial animation solutions for virtual characters cannot simultaneously achieve high lip-sync accuracy, consistent micro-expression emotion, and stable real-time interaction. This is especially true in close-up scenarios requiring high realism (such as close-up shots of virtual anchors), where the lack of synchronized micro-expressions significantly reduces the realism and emotional impact of the virtual character. Therefore, there is an urgent need for an improved technical solution that enables digital avatars to synchronize their facial expressions with the speech content and rhythm, resulting in more vivid and expressive performance, while ensuring accurate lip-sync, natural expressions, and stable operation. Summary of the Invention

[0006] This invention provides a method and product for the decoupled synchronous generation of lips in digital clones driven by a large character model, in order to solve the problem in related technologies where digital clones lack micro-expression changes in other areas of the face, resulting in poor naturalness and appeal of the digital clones.

[0007] In a first aspect, embodiments of the present invention provide a method for the decoupled and synchronous generation of a large character model-driven digital clone, comprising: The audio feature sequence obtained by extracting features from the input audio data and the character portrait vector obtained by extracting features from the character portrait information of the digital clone are obtained. The character portrait vector is used to represent the expression style of the digital clone. Audio feature sequences and character portrait vectors are input into a pre-trained large-scale character model to simulate character emotions and obtain facial expression control information for the digital clone. The facial expression control information includes at least emotion labels and / or facial expression control vector sequences. The large-scale character model and the parameter generation model are obtained by deep training the corresponding neural networks based on a multimodal facial animation dataset. Using facial expression control information as an emotional condition constraint, a pre-trained parameter generation model is used to simulate lip movements and facial micro-expressions on audio feature sequences, generating lip movement parameter sequences and micro-expression movement parameter sequences that are time-aligned with the audio feature sequences. Based on the lip movement parameter sequence and the micro-expression parameter sequence, the lip movement and micro-expression are fused in the overlapping area to resolve conflicts and generate the full-face expression parameter sequence of the digital clone. The full-face expression parameter sequence is superimposed with preset motion parameters and smoothed to obtain the target expression parameter sequence. The target expression parameter sequence is then mapped to the digital clone's facial driving parameters and driven to render the output, so as to achieve synchronization of the digital clone's lip shape and micro-expression.

[0008] Secondly, embodiments of the present invention provide a device for the micro-decoupling and synchronous generation of a large character model-driven digital clone, comprising: The audio feature extraction module is used to extract audio feature sequences from the input audio data. The character portrait acquisition module is used to extract features from the character portrait information of the digital clone to obtain the character portrait vector, which is used to represent the expression style of the digital clone. The character large model inference module is used to input audio feature sequences and character portrait vectors into a pre-trained character large model to simulate character emotions and obtain the expression control information of the digital clone; the expression control information includes at least emotion labels and / or expression control vector sequences; the character large model and parameter generation model are obtained by deep training the corresponding neural networks based on a multimodal facial animation dataset; The lip shape parameter generation module is used to generate a sequence of lip shape action parameters that are temporally aligned with the audio feature sequence based on the audio feature sequence. The micro-expression parameter generation module is used to generate a sequence of micro-expression action parameters that are time-aligned with the audio feature sequence, based on the audio feature sequence and using facial expression control information as an emotional condition constraint. The expression fusion module is used to perform conflict resolution fusion processing on the overlapping areas of lip movements and micro-expression movements based on the lip movement parameter sequence and micro-expression movement parameter sequence, and generate the full-face expression parameter sequence of the digital clone. The facial expression post-processing module is used to overlay preset motion parameters and smooth the full-face facial expression parameter sequence to obtain the target facial expression parameter sequence. The face-driven rendering module is used to map the target expression parameter sequence to the digital clone's face driving parameters and drive the rendering output to achieve synchronization of the digital clone's lip shape and micro-expressions.

[0009] Thirdly, embodiments of the present invention provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.

[0010] Fourthly, embodiments of the present invention provide a readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.

[0011] Fifthly, embodiments of the present invention provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the above-described method.

[0012] The method and product for decoupled synchronous generation of lip and micro-expressions in a digital avatar driven by a large character model provided in this invention simulates character emotions by inputting audio feature sequences and character portrait vectors into a large character model to generate facial expression control information. Based on the facial expression control information as emotional constraints, lip movement parameter sequences and micro-expression movement parameter sequences aligned with the audio time sequence are generated respectively. Then, through conflict resolution fusion processing of overlapping areas, a full-face expression parameter sequence is obtained and subjected to preset movement parameter superposition and smoothing processing to output a full-face expression parameter sequence for driving the digital avatar to achieve synchronization of lip movements and micro-expressions. Compared with micro-expression generation schemes that rely on random or template-based methods, this invention introduces a large character model for character emotion simulation, enabling the generated facial expression control information to comprehensively reflect speech content, prosodic features, and character attributes, and conditionally modulates the micro-expression generation process. This allows facial micro-expressions to adaptively respond to changes in speech content and prosodic features, while also achieving differentiated expression by combining character characteristics, improving the consistency of micro-expressions in emotional tone and character style. Building upon this foundation, a full-face expression parameter sequence is generated by decoupling and modeling lip-shape motion parameters and micro-expression motion parameters. During the fusion process, conflict resolution is implemented in overlapping areas of lip-shape and micro-expression motions, effectively reducing local deformation conflicts and expression distortion caused by motion overlap. Furthermore, preset motion overlays and smoothing are applied to the full-face expression, enhancing the richness, temporal consistency, and coherence of facial micro-expressions. This solution ensures strict synchronization between lip shape and audio while enabling the digital avatar's micro-expressions to change in sync with emotional tone and character style, achieving a more natural and coherent emotional expression. This results in an effective balance between expressive richness, lip-shape synchronization accuracy, and overall naturalness, significantly improving the realism and appeal of the digital avatar's facial performance.

[0013] Compared with the approach of simultaneously inputting facial expression control information into the lip shape parameter generation model and the micro-expression parameter generation model, in the embodiment of this application, facial expression control information is only input into the micro-expression parameter generation model, so that the lip shape parameter generation process is directly driven by audio features, and is improved in at least one of the indicators of closed-mouth closure integrity, lip shape synchronization error and / or mouth corner tremor amplitude, thereby reducing the interference of emotion modulation on lip shape synchronization. Attached Figure Description

[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1This is a schematic diagram of the structure of an expression generation system according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating a method for the decoupling and synchronous generation of a large character model-driven digital clone in one embodiment of the present invention. Figure 3 This is a schematic diagram illustrating the generation process of the full-face expression parameter sequence in one embodiment of the present invention; Figure 4 yes Figure 2 A schematic diagram of the implementation process of step S30; Figure 5 yes Figure 2 A schematic diagram of the implementation process of step S40; Figure 6 yes Figure 2 A schematic diagram of the implementation process of step S50; Figure 7 This is a schematic flowchart of another method including audio anomaly fault tolerance rollback in one embodiment of the present invention; Figure 8 This is a schematic diagram of the data processing pipeline in one embodiment of the present invention; Figure 9 This is a schematic diagram of the generating device in one embodiment of the present invention; Figure 10 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof. It should also be understood that, as used in this specification and the appended claims, the term "and / or" refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0018] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0019] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0020] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0021] To illustrate the technical solution of the present invention, specific embodiments are described below.

[0022] For ease of understanding, the following terms are defined: A digital avatar is a two-dimensional or three-dimensional virtual character that can be rendered using facial driving parameters. It can simulate a specific role or character and possess the ability to synchronize lip movements with speech and express emotional facial expressions. The digital avatar can be rendered using its facial driving parameters to achieve facial animation output. These facial driving parameters include, but are not limited to, facial action unit (AU) parameters, blendshape coefficients (i.e., deformation coefficients), and / or facial bone pose parameters. Character profile information is a set of information representing the identity attributes and expression style preferences of the target digital clone. The character profile vector is a vector representation obtained by feature extraction of the character profile information, which includes at least the character identity embedding vector and / or expression style parameter vector. The character big model refers to a sequence reasoning model trained on a deep neural network, which is used to combine audio feature sequences with character portrait vectors to output facial expression control information. Lip micro-decoupling refers to the decoupling modeling of lip movements and facial micro-expressions in digital clones. It generates time-aligned lip movement parameter sequences and micro-expression parameter sequences through independent generative models, and then achieves coordinated synchronization of the two by resolving conflicts in overlapping areas, taking into account both the synchronization accuracy of lip movements and audio and the naturalness of micro-expressions.

[0023] Facial expression control information is the structured control information output by the large character model, including at least emotion labels and / or facial expression control vector sequences, and optionally semantic labels, intention labels, and / or prosodic event labels; lip movement parameters refer to a set of parameters related to pronunciation and directly corresponding to changes in mouth / jaw shape, including at least one or more of jaw opening and closing, lip opening and closing, lip rounding, and lip stretching; micro-expression movement parameters refer to a set of subtle facial expression parameters other than lip movement parameters, including at least movement parameters of areas such as eyebrows, eyes, and cheeks, and may also include subtle movement parameters of the mouth edge area; A parameter generation model is a deep neural network sequence generation model trained on a multimodal facial animation dataset. It is used to generate corresponding action parameter sequences based on audio feature sequences and facial expression control information, including independent lip shape parameter generation models and micro-expression parameter generation models.

[0024] The overlapping region refers to the mouth edge region in the facial driving parameter space that may participate in both lip shape deformation for pronunciation and emotional expression, including at least lip corner stretching / contraction and upper lip lifting; The parameter partitioning rule refers to the rule of dividing the full-face expression parameters into a set of lip shape parameters, a set of micro-expression parameters, and a set of overlapping parameters according to function or region, and using it for conflict resolution and smooth transition in the fusion stage; The full-face expression parameter sequence refers to the parameter sequence obtained by fusing the lip movement parameter sequence and the micro-expression movement parameter sequence; The target facial expression parameter sequence refers to the final output sequence obtained by performing rule-based action superposition, subparameter temporal smoothing, and amplitude / velocity constraints on the full-face facial expression parameter sequence; Neutral expression refers to the baseline expression of a character with no facial micro-expressions and closed lips. The parameters of neutral expression are the full-face expression parameters corresponding to neutral expression. Fault-tolerant rollback refers to a processing mechanism that smoothly rolls back the output expression to a neutral expression within a preset transition time when an audio signal interruption, signal abnormality, or prolonged silence is detected, and then performs a transition fusion after recovery.

[0025] Most common audio-driven digital avatar methods focus on motion modeling of the lip area. The resulting animations typically only show the mouth opening and closing in response to speech, lacking micro-expression changes in other facial areas, making the avatars stiff and lifeless. In reality, when humans converse, their lip movements not only need to be accurately aligned with the audio, but they also produce subtle facial expressions accompanying changes in speech content and tone, such as blinking, eyebrow raising, eye opening and closing, and cheek muscle changes, to convey appropriate emotions. Since audio signals do not directly contain information about upper facial expressions such as blinking and eyebrow movement, audio-driven animation methods, while capable of generating high-quality lip-synchronized animations, lack rich information about facial expressions during the modeling process. This often results in static or emotionless facial animations, leading to virtual avatars with dull eyes, lack of eyebrow and eye movement, and an overall stiff and unnatural expression, lacking naturalness and appeal.

[0026] To address this, one could try introducing random emotional signals or preset emotion-driven signals to enrich expressions. However, this might compromise the alignment between lip movements and audio, leading to decreased lip shape accuracy and inconsistencies between lip shape and facial expression. Simultaneously, controlling the amplitude of facial expressions is a significant challenge: if the micro-expression amplitude is too large, the virtual clone's expression appears exaggerated and unrealistic; if it's too small, it fails to achieve the desired effect. Micro-expression generation schemes relying on random parameters or template-based methods struggle to achieve a balance between richness of expression and accuracy and naturalness of lip shape.

[0027] Furthermore, the inventors discovered overlapping areas on the face that are involved in both lip shape deformation related to pronunciation and facial expression changes related to emotion, such as the mouth edge area (including the corners of the lips and the edge of the upper lip). Deformations in these overlapping areas can easily affect the expression of lip shape and other facial features. When using the same model to uniformly predict all facial parameters, or when predicting lip shape parameters and other micro-expression parameters separately and fusing them to obtain all facial parameters, the parameters in the overlapping areas are prone to conflict and overlap. This manifests as difficulty in completely closing the lips when pronouncing closed sounds, sharp tremors at the corners of the mouth, or "leakage interference" from emotional expressions, thus simultaneously impairing both lip shape accuracy and facial expression naturalness. Moreover, these conflicts are more likely to accumulate and amplify in real-time dialogue scenarios, significantly reducing the interactive experience.

[0028] It is evident that current audio-driven facial animation solutions for virtual characters are insufficient to express the emotions and tone of voice conveyed in the audio. Digital avatars often give the impression of being "stiff" and "lacking eye contact" during interaction. Especially in close-up scenarios requiring high realism (such as close-up shots of virtual anchors), the lack of synchronized micro-expressions significantly reduces the realism and appeal of the virtual character. Even some solutions that use pre-trained models to directly predict facial expressions from audio fail to simultaneously meet the requirements of real-time audio playback and precise micro-expression synchronization, exhibiting shortcomings in terms of naturalness and robustness. Therefore, there is an urgent need for an improved technical solution that enables digital avatars to achieve more vivid and expressive facial expressions that change synchronously with the speech content and rhythm, while ensuring accurate lip movements and natural expressions.

[0029] To address the above issues, this invention provides a method, apparatus, system, device, medium, and program product for generating decoupled and synchronized lip and micro-expression movements of a digital clone driven by a large character model. By inputting an audio feature sequence and a character portrait vector into a large character model to simulate character emotions, facial expression control information is generated. Using this information as a constraint, lip movement parameter sequences and micro-expression movement parameter sequences aligned with the audio timing are generated respectively. Then, a full-face expression parameter sequence is obtained through fusion processing, and preset movement parameters are superimposed and smoothed to output the full-face expression parameter sequence used to drive the digital clone to achieve lip and micro-expression synchronization. This achieves the following effects: (1) In the process of generating facial expressions, based on the audio feature sequence and the character profile vector, the large character model outputs emotion labels and / or expression control vector sequences to conditionally modulate the generation process of facial micro-expressions. Through the cross-modal alignment mechanism of the expression condition control output by the large character model and the parameter generation model, the detailed expressions of the upper half of the face can be generated synchronously with the changes in audio prosody and emotional tone. Compared with micro-expression generation schemes that rely on random parameters or templates, this embodiment of the invention introduces a large character model to simulate character emotions. The generated expression control information can comprehensively reflect the speech content, prosodic features and character attributes, and conditionally modulate the micro-expression generation process accordingly, so that the changes in facial micro-expressions respond reasonably to the speech content and prosody, and at the same time, combine with character features to achieve differentiated expression, thereby ensuring that the output micro-expression actions are consistent in emotional tone and character profile style. Building upon this foundation, the full-face expression parameter sequence obtained by fusing micro-expression and lip movement parameters can enhance the synchronization richness of other facial micro-expressions such as eyebrows and eyes, while ensuring strict synchronization between the digital clone's lip movements and audio. This allows the digital clone to present a more natural, coherent, and emotionally expressive facial performance during speech. Therefore, the digital clone driving process in this embodiment of the invention achieves an effective balance between expression richness and lip shape accuracy and naturalness, significantly improving the naturalness and appeal of the digital clone's facial expressions.

[0030] (2) During the parameter generation process, a lip shape parameter generation model is used to generate a sequence of lip shape action parameters, and a micro-expression parameter generation model is used to generate a sequence of micro-expression action parameters. The lip shape generation branch focuses on pronunciation-related movements, while the micro-expression generation branch focuses on subtle expression changes. By decoupling generation and partitioning fusion, the interference between lip shape and other facial expressions is reduced, improving the accuracy of lip shape and micro-expression parameter generation. This ensures lip shape alignment while enhancing the expression of details such as eyebrows and eyes. Then, the two are fused according to the parameter partitioning rules. During fusion, a weighted transition is applied to overlapping areas such as the corners of the lips to resolve conflicts and improve the naturalness of the transition, thereby improving the overall naturalness of the digital clone.

[0031] (3) By superimposing rules into actions and using a parameter smoothing mechanism, the accuracy of the final generated facial expression parameters is further improved, thereby enhancing the naturalness (such as realism) and stability of the digital avatar. Introducing regular actions such as blinking and nodding on the basis of parameter fusion output can simulate the physiological details and communication habits of real people or virtual characters. For the generated facial expression parameter sequence, different filtering intensities are applied to different types of parameters for temporal smoothing, amplitude, or speed constraints. This can suppress jitter and ensure smooth transitions of action parameters, improving the continuity and realism of the output parameter sequence and enhancing the overall consistency of the digital avatar's appearance.

[0032] (4) Provides real-time parameter output based on audio and parameter rollback mechanism based on audio anomalies, enabling the digital clone driving system to have real-time output and fault-tolerant rollback capabilities. The parallel pipeline mode of audio input and parameter synchronous generation can reduce end-to-end signal delay during digital clone driving. When the input audio signal is interrupted, abnormal, or silent for a long time, the expression parameter is triggered to smoothly roll back to the neutral expression parameter, and the synchronous output continues after the signal returns to normal. This can maintain output continuity, enhance the stability of the system under complex input conditions, and is suitable for real-time interaction and online deployment scenarios.

[0033] The method for decoupling and synchronizing the generation of large character models-driven digital clones provided in this invention can be applied to, for example... Figure 1 The expression generation system shown includes a character large model-driven digital clone lip micro-decoupling synchronous generation device (hereinafter referred to as the generation device) and a user terminal. The user terminal communicates with the character large model-driven digital clone lip micro-decoupling synchronous generation device via a network.

[0034] Among them, the user terminal refers to the terminal device used by the user. Digital avatars include virtual digital images built based on real people, as well as virtual digital images built based on virtual characters (such as anime characters). Different digital avatars correspond to different basic attributes, styles, and preferences, and their corresponding character profile information is different.

[0035] When there is a need for generating facial expression parameters and displaying digital avatars, the user sends a command to the generation device to generate facial expression parameters via a user terminal. The generation device collects audio data or receives audio data sent by the user terminal, and extracts features from the input audio data to obtain an audio feature sequence. Simultaneously, it acquires the character profile information of the specified digital avatar, and extracts features from the character profile information of the digital avatar to obtain a character profile vector. The character profile vector is used to represent the facial expression style information, identity information, etc. of the digital avatar.

[0036] Then, the generation device inputs the audio feature sequence and the character portrait vector into a pre-trained large-scale character model to simulate character emotions, obtaining the expression control information of the digital clone. Using this expression control information as emotional constraints, a pre-trained parameter generation model simulates lip movements and facial micro-expressions on the audio feature sequence, generating lip movement parameter sequences and micro-expression parameter sequences that are time-aligned with the audio feature sequence. Next, expression fusion processing is performed on the lip movement parameter sequences and micro-expression parameter sequences to generate the full-face expression parameter sequence of the digital clone. Subsequently, post-processing of the full-face expression parameter sequence is performed, including motion overlay, partition smoothing, and fault-tolerant rollback processing. Based on the processed expression parameters, facial-driven rendering is performed on the digital clone to generate facial-driven animation, which is then sent to the user terminal for output display, achieving synchronization of lip movements and micro-expressions. The large-scale character model and the parameter generation model are respectively obtained by deep training of corresponding neural networks based on a multimodal facial animation dataset.

[0037] In this embodiment, compared with the micro-expression generation scheme that relies on random or templated methods, in the embodiment of the present invention, by introducing a character large model for character emotion simulation, the generated expression control information can comprehensively reflect the speech content, prosody, and character attributes, and conditionally modulate the micro-expression generation process based on this, so that the facial micro-expression changes respond reasonably to the speech content and prosody. At the same time, combined with character features, differential expression is achieved, thereby ensuring that the output micro-expression actions are consistent in terms of emotional tone and character portrait style. On this basis, full-face parameter fusion is performed based on micro-expression and lip movement parameters, and conflict resolution processing is carried out for the lip movement and micro-expression actions in the overlapping area during the fusion stage to reduce local deformation conflicts and expression distortion problems caused by the superposition of lip shape and expression; in addition, preset action superposition and smoothing processing are also performed on the full-face expression to improve the richness, temporal consistency, and coherence of facial micro-expressions. While ensuring the strict synchronization of the lip shape and audio of the digital avatar, this solution improves the synchronization richness of other facial micro-expressions to achieve the联动变化 of micro-expressions with emotional tone and character style, making the digital avatar present a more natural, coherent, and emotional facial performance during speech. Thus, in the digital avatar driving process in the embodiment of the present invention, an effective balance is achieved between expression richness, lip shape accuracy, and naturalness, significantly enhancing the naturalness and appeal of the digital avatar's facial expressions.

[0038] Among them, the device for synchronously generating the decoupled lip of the digital avatar driven by the character large model can be a server or a terminal device. Terminal devices include, but are not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. Portable wearable devices can include, but are not limited to, AR glasses, VR glasses, smart watches, exoskeletons, etc. The server can be implemented by an independent server or a server cluster composed of multiple servers.

[0039] The above method can be deployed to a terminal device (i.e., a user terminal), where the entire expression generation system runs independently, achieving local terminal deployment. This terminal device needs sufficient computing power, such as including a GPU or NPU accelerator, to meet the real-time inference requirements of the aforementioned neural network model. The terminal device contains multiple functional modules, each working collaboratively as a local application component to perform tasks such as audio data acquisition, feature extraction, expression generation, and digital avatar rendering. The advantages of local terminal deployment are high real-time performance, independence from network connections, and the elimination of the need to upload input audio data to a cloud server, ensuring data privacy and security. To address the limited computing power of mobile terminal devices, lightweight optimization can be performed on various application neural network models, such as through model quantization, pruning, or using a smaller network architecture, to maintain smooth operation under limited computing resources. After lightweight optimization, the method provided in this embodiment can achieve near real-time processing speeds on high-performance terminal chips (such as mobile phone chips), for example, achieving expression generation and rendering speeds of over 20 frames per second, enabling synchronized display of digital avatar facial expressions and audio, meeting real-time interaction requirements.

[0040] Furthermore, the above methods are deployed to servers to achieve cloud deployment. An expression generation system is deployed on a cloud server or backend server, where the server performs the computationally intensive expression generation task, while the user terminal is only responsible for audio data acquisition and task result display. Specifically: the user terminal acquires audio data and uploads it to the server in real time via the network. The server extracts features from the audio data, generates an expression parameter sequence based on the extracted audio feature sequence, and performs animation rendering on the digital avatar based on the expression parameter sequence to obtain video data. The generated expression parameter sequence and / or the rendered video data are then sent to the user terminal as the task result. After receiving the task result, the user terminal uses it to present the digital avatar. Servers are typically equipped with high-performance GPUs / TPUs, capable of serving multiple user requests in parallel, achieving efficient batch processing. Server deployment fully utilizes powerful computing capabilities, supporting more complex or higher-precision models, thus achieving optimal results in scenarios with extremely high requirements for expression details. In this mode, network latency and bandwidth limitations need to be considered. Therefore, the method provided in this embodiment of the invention can transmit only a compact expression parameter sequence instead of complete video data to reduce bandwidth consumption. The animation is then reconstructed locally on the user terminal based on the received sequence of facial expression parameters, thereby minimizing latency and ensuring that audio and facial expressions are synchronized on the user terminal.

[0041] In summary, the method for decoupling and synchronously generating digital avatars based on large character models can be deployed on servers or on terminal devices with certain computing capabilities. Regardless of whether terminal device or server deployment is used, the method in this embodiment can be implemented in software and run on general-purpose computing devices or specific hardware, providing a flexible integration solution for various digital avatar applications.

[0042] In one embodiment, such as Figure 2 As shown, a method for decoupling and synchronizing the generation of digital clone lips driven by a large character model is provided, and this method is applied to... Figure 1 Taking the generating device in the example, the following steps are included: S10: Obtain the audio feature sequence obtained by feature extraction from the input audio data, and the character portrait vector obtained by feature extraction from the character portrait information of the digital clone.

[0043] Upon receiving an instruction to generate facial expression parameters for a specific digital clone, the generation device collects audio data from the external environment or receives audio data sent by the user through a user terminal to obtain input audio data. The generation device then extracts audio features from the input audio data to extract prosodic features and / or content-related features (such as semantic features), resulting in an audio feature sequence. That is, the audio feature sequence includes a sequence formed by at least one of prosodic features and content-related features.

[0044] Simultaneously, the generation device also acquires the character portrait information of the digital clone and extracts features from it to obtain a character portrait vector. For example, it can generate the digital clone's identity identifier (such as ID) based on the expression parameters in the command, retrieve the digital clone information corresponding to that identifier from the database, extract the character portrait information of the digital clone from the digital clone information, and then perform feature extraction to obtain a character portrait vector. The character portrait vector is used to represent the digital clone's identity, expression style, and preference information.

[0045] The digital avatar information includes the avatar's character profile information and facial driving parameters. The character profile information is a conditional input representing the digital avatar's personalized characteristics, and includes, but is not limited to, the character's identity ID, basic attribute information, expression style, and preference information. Basic attribute information includes, but is not limited to, gender, age, occupation, and appearance. Correspondingly, the character profile vector includes, but is not limited to, the digital avatar's identity ID embedding vector, expression style parameters, and expression preference parameters.

[0046] Among them, facial driving parameters are a set of all adjustable variables used to control the facial movements of the digital avatar. Each parameter value corresponds to a specific local deformation (such as opening the mouth, smiling, or frowning). By adjusting the values ​​of these parameters, arbitrary facial expressions can be linearly combined to achieve precise control of the digital avatar's facial movements. Facial driving parameters include, but are not limited to, facial motion unit parameters, BlendShape coefficients and facial bone posture parameters, facial key point coordinates, and muscle driving parameters.

[0047] Upon receiving instructions to generate facial expression parameters for a specific digital clone, the generation device can also acquire or invoke pre-trained character model and parameter generation model. The character model and parameter generation model are obtained by deep training the corresponding neural networks based on a multimodal facial animation dataset.

[0048] S20: Input the audio feature sequence and the character portrait vector into the pre-trained large character model to simulate the character's emotions and obtain the expression control information of the digital clone.

[0049] The generation device inputs the extracted audio feature sequence and character profile vector into a pre-trained character model to simulate character emotions, thereby obtaining the facial expression control information of the digital avatar. The output of the character model (i.e., the facial expression control information) may include at least one of a facial expression control vector sequence and facial expression labels. In other embodiments, the output of the character model may also include at least one of a semantic label, an intent label, and a prosodic event label corresponding to the audio feature sequence. The label information is output through the classifier of the character model, and the facial expression control vector sequence is output through the linear projection layer of the character model.

[0050] Specifically, the large-scale character model can be a sequence model obtained by deep training a pre-defined neural network structure using audio-animation sample pairs extracted from a multimodal facial animation dataset. The pre-defined neural network structure (i.e., the large-scale character model) can employ a multi-branch, multimodal fusion network structure, meaning the large-scale character model includes an audio encoding layer, a character portrait encoding layer, a cross-modal fusion layer, a temporal modeling layer, and a multi-task output module. The multi-task output module includes a sequence generation layer (i.e., a linear projection layer) and one or more classifiers. The sequence generation layer is responsible for generating expression control vector sequences, and each classifier generates one type of label information.

[0051] The generation device inputs the audio feature sequence and the character portrait vector into a pre-trained large-scale character model. The large-scale model first performs dimensional alignment and semantic mapping on the audio feature sequence through an audio coding layer, and simultaneously performs dimensional alignment and semantic mapping on the character portrait vector through a character portrait coding layer. Subsequently, the outputs of the audio coding layer and the character portrait coding layer are fused in a cross-modal fusion layer, imbuing the audio features with character style information. Then, a temporal modeling layer dynamically models the fused features to obtain temporal semantic features reflecting speech content, tone, and character style characteristics. Finally, a multi-task output module generates expression control vector sequences, expression labels, semantic labels, intent labels, and prosodic event labels, thereby obtaining the expression control information used to drive the generation of facial expressions in the digital avatar.

[0052] In this scheme, the large-scale character model aligns and semantically maps audio feature sequences and character profile vectors, and then fuses these features in a cross-modal fusion layer, imbuing the audio representation with character style information. Furthermore, it dynamically models the fused features using temporal modeling to obtain temporal semantic features that simultaneously reflect speech content, tone, and character characteristics. A multi-task output module generates expression control vector sequences, along with expression labels, semantic labels, intent labels, and prosodic event labels, forming structured and multi-layered expression control information. This transforms the expression control information generation process from solely relying on audio features to a multi-modal driven approach that integrates speech and character semantics. This ensures that the generated micro-expressions remain consistent with speech content, tone changes, and character characteristics. Simultaneously, multi-type label information provides fine-grained characterization of expression categories, intensity, and timing of changes, offering a unified and controllable driving basis for subsequent facial animation generation. This enhances the naturalness, coherence, and emotional expressiveness of the digital avatar's facial expressions.

[0053] S30: Using facial expression control information as an emotional condition constraint, a pre-trained parameter generation model is used to simulate lip movements and facial micro-expressions on the audio feature sequence, generating lip movement parameter sequences and micro-expression movement parameter sequences that are time-aligned with the audio feature sequence.

[0054] The generation device uses the facial expression control information obtained in step S20 as a conditional constraint for micro-expression generation, and generates lip movement parameter sequences and micro-expression movement parameter sequences aligned with the audio time sequence based on the audio feature sequence, thus providing two inputs for subsequent fusion to generate a full-face expression parameter sequence. Lip movement parameters refer to the movement parameters of the lips and their edge regions; micro-expression movement parameters refer to the detailed facial expression movement parameters of other facial regions besides the lips, such as the movement parameters of the lip edges (e.g., corners of the lips, upper and lower lip edges), eyebrows, eyes, nose, cheeks, and chin. In other words, micro-expression movement parameters are used to characterize subtle facial expression changes in areas such as the eyebrows, eyes, cheeks, and lip edges.

[0055] The parameter generation model can be a sequence generation model obtained by deep training a sequence generation network based on audio-animation sample pairs extracted from a multimodal facial animation dataset. The parameter generation model can include, in sequence, an input alignment and feature transformation module, an attention-based conditional modulation module, a shared temporal modeling module, a region decoupling branch module (including lip shape branch and micro-expression branch), and a parameter output module. The region decoupling branch module includes a lip shape branch (i.e., a lip shape decoder) and a micro-expression branch (i.e., a micro-expression decoder).

[0056] Specifically, after the generation device inputs facial expression control information and audio feature sequences into the parameter generation model, the model first performs feature alignment and mapping processing on the input audio feature sequences and facial expression control information. Then, through a conditional modulation module, the facial expression control information is injected into the audio features to obtain a controlled feature representation. Subsequently, a shared temporal modeling module performs temporal modeling on this controlled feature to obtain a high-level temporal feature that integrates speech information and facial expression constraints. Based on this, a lip-related feature sequence is modeled through a lip-shaped branch in the region decoupling branch module to obtain a lip-shaped latent feature sequence, and a micro-expression latent feature sequence is obtained by modeling non-lip facial features through a micro-expression branch. Finally, a parameter output module generates lip-shaped action parameter sequences and micro-expression action parameter sequences that are temporally aligned with the audio feature sequences. By introducing a conditional modulation mechanism and a region decoupling structure, the parameter generation model performs temporal modeling on audio features under the constraints of facial expression control information and generates lip-shaped action parameter sequences and micro-expression action parameter sequences respectively, thereby achieving the collaborative generation of lip movements and facial micro-expressions driven by speech.

[0057] In other embodiments, the parameter generation model includes two independent models: a micro-expression parameter generation model and a lip shape parameter generation model. This decouples lip movements from facial micro-expressions, reducing mutual interference. The lip shape parameter generation model is a neural network model trained on multiple pairs of synchronized speech lip-shape samples. The micro-expression parameter generation model is also a neural network model trained on multiple pairs of synchronized speech micro-expression samples. Using two independent parameter generation models to generate lip shape and micro-expression parameters separately decouples them, reducing the likelihood of inaccurate parameters due to mutual interference. Preferably, facial expression control information is only used as a conditional input to the micro-expression parameter generation model and not directly involved in the inference of the lip shape parameter generation model, to avoid interference from emotion modulation on lip-shape synchronization.

[0058] Specifically, the generation device inputs the audio feature sequence into a pre-trained lip shape parameter generation model to simulate lip movements, generating a lip movement parameter sequence that is temporally aligned with the audio feature sequence. These lip movement parameters characterize the primary movements of the jaw and lips directly corresponding to the pronunciation content, including at least one or more of jaw opening and closing, lip opening and closing, lip rounding, and lip stretching. Content-related features in the audio feature sequence drive the correspondence between mouth shape and phonemes, while prosodic features adjust the opening and closing rhythm and amplitude variations, thereby ensuring the synchronization accuracy between lip shape and audio.

[0059] The generation device inputs the audio feature sequence and facial expression control information into the micro-expression parameter generation model to simulate micro-expression movements, generating a sequence of micro-expression movement parameters that is temporally aligned with the audio feature sequence. These micro-expression movement parameters are used to characterize subtle facial expression changes in areas such as the eyebrows, eyes, and cheeks, excluding the main lip movements. Optional parameters may include subtle movement parameters of the mouth edge region to enhance naturalness. The micro-expression parameter generation model injects facial expression control information into the audio feature representation through feature modulation and / or cross-attention mechanisms. This allows the type, intensity, and temporal distribution of micro-expressions to be constrained by emotion labels, facial expression control vector sequences, and optional prosodic event labels, thereby ensuring that micro-expressions change with emotional tone and remain synchronously associated with prosodic events such as pauses and accents.

[0060] It should be noted that the lip edge area (such as lip corner stretching / contraction, upper lip elevation, etc.) is an overlapping area that participates in both pronunciation and emotional expression. To ensure that lip shape synchronization is prioritized during the key pronunciation stages, the parameters of the overlapping area are not simply determined by a single branch in stage S30, but are resolved and naturally integrated in the subsequent step S40 through parameter partitioning and weighted transition.

[0061] The parameter generation model described above can employ the Transformer architecture. In other embodiments, the parameter generation model can also be other types of sequence models, such as a network architecture consisting of a bidirectional long short-term memory (LSTM) encoder and a feedforward decoder, or a temporal recurrent neural network (Temporal CNN). Furthermore, in offline scenarios or scenarios where real-time performance requirements are not high, the parameter generation model can also be a more complex generative model such as a diffusion model or a variational autoencoder (VAE) model to obtain more refined facial expression details.

[0062] In a preferred embodiment, the parameter generation model includes a lip shape parameter generation model and a micro-expression parameter generation model deployed independently of each other. The input of the lip shape parameter generation model includes only an audio feature sequence and does not receive expression control information; the input of the micro-expression parameter generation model includes an audio feature sequence and expression control information, and the expression control information is injected into the audio feature representation through a conditional modulation unit and / or a cross-attention unit to generate a micro-expression action parameter sequence.

[0063] In this application, "expression control information does not participate in the inference of the lip shape parameter generation model" means that expression control information is not used as input to the lip shape parameter generation model, is not used for feature modulation of the intermediate layer of the model, and is not used to control the lip shape parameter regression results of the output layer of the model.

[0064] S40: Based on the lip movement parameter sequence and the micro-expression parameter sequence, the fusion processing of conflict resolution between the lip movement and micro-expression in the overlapping area is performed to generate the full-face expression parameter sequence of the digital clone.

[0065] In the process of fusing lip movement parameter sequences and micro-expression parameter sequences, the two types of parameters are first mapped to a unified full-face parameter space, and then divided according to the facial regions of the digital avatar: for non-lip regions, micro-expression parameters are used as the corresponding region's movement parameters; for lip regions, lip movement parameters are used as the corresponding region's movement parameters; for overlapping regions, lip movement parameters and micro-expression parameters are weighted and transitionally fused according to their respective weight coefficients (such as preset weights) to serve as the movement parameters for the overlapping region, thereby resolving conflicts between lip movement and micro-expression movement parameters in the overlapping region; additionally, the weights of the two types of parameters at different times can be adaptively determined based on prosodic features and / or detected prosodic events in the audio sequence, and then the overlapping region parameters are weighted and transitionally fused based on the determined adaptive weights to achieve conflict resolution and ensure that lip movement parameters dominate at key pronunciation moments. Then, the movement parameters of each region are aligned temporally and concatenated to obtain the full-face expression parameter sequence. The overlapping region is the facial region that affects lip movement and micro-expression, such as the lip edge region.

[0066] This fusion step integrates lip shape-driven information and micro-expression information into a unified parameter space through parameter selection and weighted fusion mechanisms based on facial regions. While ensuring lip shape accuracy, it introduces subtle facial variations, achieving natural and coordinated full-face expression generation. Simultaneously, the fusion stage identifies overlapping areas between lip shape and micro-expressions, performing conflict resolution processing based on weighted fusion or interpolation calculations in these areas. This prevents overlapping conflicts in areas such as the corners of the lips and upper lip from being simultaneously affected by both types of parameters, suppressing non-physical deformations caused by parameter overlap. This effectively avoids problems such as incomplete closure of the mouth, trembling corners of the mouth, and leakage of facial expressions, thereby improving the stability and realism of local facial regions.

[0067] like Figure 3 The diagram illustrates the various models and their data processing steps in the above method. The extraction module of the generation device inputs the extracted character feature vectors and audio feature sequences into the character large-scale model inference module. This module uses the character large-scale model to simulate character emotions and outputs facial expression control information. The character large-scale model inference module can also input the audio feature sequences into an audio encoder for feature transformation, mapping the audio feature sequences to the feature space of the facial expression control information for easier subsequent processing. The generation module inputs the facial expression control information and the transformed audio feature sequences into a parameter generation model. The lip shape parameter generation model (or lip shape decoder) within this model simulates lip shapes, generating a lip shape action parameter sequence. Similarly, the micro-expression parameter generation model (or micro-expression decoder) within this model simulates micro-expressions in other facial areas, generating a micro-expression action parameter sequence. Finally, the lip shape action parameter sequence and the micro-expression action parameter sequence are input into a fusion module for parameter fusion, outputting a full-face expression parameter sequence.

[0068] In one embodiment, the fused full-face expression parameter sequence can be smoothed to obtain a continuously changing facial expression parameter sequence, which can then be used to drive the digital clone, enabling the synchronous presentation of lip movements and facial micro-expressions, while ensuring the coherence and stability of the presentation effect.

[0069] S50: The full-face expression parameter sequence is superimposed with preset motion parameters and smoothed to obtain the target expression parameter sequence. The target expression parameter sequence is then mapped to the facial driving parameters of the digital clone and driven to render the output, so as to achieve synchronization of the digital clone's lip shape and micro-expression.

[0070] The generation device superimposes preset action parameters onto the full-face expression parameter sequence, such as superimposing preset blinking action parameters and / or nodding action parameters; it performs temporal smoothing processing on the superimposed full-face expression parameter sequence, such as temporal balancing for parameters of different regions, to obtain the target expression parameter sequence, so that facial expression movements conform to physiological habits and interaction habits, and improve the richness, temporal consistency and coherence of facial micro-expressions.

[0071] Subsequently, the generation device can map the target expression parameter sequence to the facial driving parameters of the digital avatar, and input the facial driving parameters into the facial driving system of the digital avatar for rendering output, so as to achieve synchronization of lip shape and micro-expression in the digital avatar. Specifically, if the facial driving system is used to acquire BlendShape driving, the target expression parameter sequence is converted into corresponding shape key weights and superimposed on the neutral face model; if the skeleton driving is used, the relevant parameters in the target expression parameter sequence are converted into the rotation or displacement of facial bones to drive the facial animation display, and the parameters of other parts are applied to the corresponding bones, such as the head posture parameters such as nodding are applied to the neck or head bones, so that the digital avatar can present lip shape and micro-expression in sync with the audio.

[0072] In this embodiment, by inputting the audio feature sequence and the character portrait vector into the large character model to generate expression control information, and using the expression control information as emotional condition constraints, lip movements and facial micro-expressions are decoupled, modeled, and generated separately. At the same time, in the fusion stage, conflict resolution processing is performed on the overlapping areas of lip movements and micro-expression movements. Then, the output full-face expression parameter sequence is superimposed and smoothed with preset action parameters, so that the generated target expression parameter sequence can achieve temporal synchronization and consistent emotional expression of micro-expressions while ensuring accurate alignment between lip shape and audio. This effectively improves the naturalness, coherence, and emotional expressiveness of the digital clone's facial expressions, and avoids the problems of local deformation conflict and expression distortion caused by the superposition of lip shape and expression.

[0073] In one embodiment, before feature extraction from the audio data, the acquired audio data can be digitized into a processable audio data stream. Then, audio signal preprocessing operations are performed on the digitized audio data to obtain clean audio data. These preprocessing operations include, but are not limited to, noise reduction filtering, dereverberation, volume normalization, and silence detection. Subsequently, the clean audio data is segmented into frames with a fixed frame length and frame shift, forming a continuous audio frame sequence. Feature extraction is then performed on each frame in this continuous audio frame sequence to obtain an audio feature sequence. For cases where the audio data is excessively long, such as long speech audio, streaming or segmented processing can be used to adapt and avoid the accumulation of errors in the model over time. For example, before preprocessing, long audio can be segmented into paragraphs or fixed lengths to obtain multiple audio segments. Sentiment analysis and facial expression generation are performed on each segment separately, and facial expression parameters are smoothly connected at the segment boundaries to ensure natural facial expression transitions during long, continuous speech.

[0074] In this embodiment, the acquired audio data is digitized, and preprocessing operations such as noise reduction filtering, reverberation removal, volume normalization, and silence detection are performed on the digitized audio data. This effectively reduces the interference of environmental noise and reverberation on the audio signal, eliminates volume differences between different audio sources, and removes invalid silence segments, thereby improving the purity and stability of the audio signal. Based on this, the audio is segmented into frames according to fixed frame lengths and frame shifts to construct a temporally continuous and structurally regular audio frame sequence, ensuring consistency and alignment of the audio data in the temporal dimension. This provides high-quality, standardized input data for subsequent audio feature extraction, enhancing the ability of the extracted features to represent speech content and prosodic information, thereby improving the accuracy, stability, and temporal consistency of subsequent lip-sync and facial micro-expression generation.

[0075] In the process of extracting audio feature sequences, frame-level acoustic features can be calculated for each frame in a continuous audio frame sequence. These frame-level acoustic features include prosodic features such as fundamental frequency (F0), energy envelope, and speech rate features, as well as content-related features such as Mel spectrum, Mel frequency cepstral coefficients (MFCC), and phoneme-related acoustic representations. Then, contextual modeling can be performed on the frame-level acoustic features of multiple audio frames to obtain the audio feature vector of each frame, and these vectors can be arranged in temporal order to generate an audio feature sequence.

[0076] Among them, prosodic features are used to characterize the rhythm and expression of audio signals in the time dimension. Content-related features are used to characterize speech content, and phoneme-related acoustic representation refers to the representation method that establishes a relationship between phonemes in audio data and their corresponding acoustic features (such as spectrum, energy, duration, etc.).

[0077] In this embodiment, frame-by-frame acoustic feature extraction of audio data yields frame-level acoustic features containing information such as Mel spectrum, MFCC, fundamental frequency F0, and energy envelope. This allows for fine-grained characterization of the speech's spectral structure, pitch variations, and energy distribution. Furthermore, a deep encoder is used to perform contextual modeling of the frame-level features, establishing temporal relationships between adjacent frames to obtain a continuous sequence of audio feature vectors. This enables the audio features to not only reflect the characteristics of a single frame of speech but also to characterize the dynamic changes of speech over time. This enhances the ability of audio features to express speech content and prosodic information, providing high-quality input with temporal consistency and semantic continuity for subsequent lip movements and facial micro-expression generation, thereby improving the accuracy and naturalness of animation generation.

[0078] In one embodiment, the character profile vector is a vector representation used to characterize the identity, style, and facial expression preferences of the target digital avatar. The character profile vector includes basic attribute information such as character identity ID, character type, and personality tags, as well as historical facial expression data (including preference information) and preset style parameters.

[0079] In the character portrait vector extraction process, the generation device inputs the character ID into a pre-built embedding mapping model for mapping, obtaining a character ID embedding vector to represent the identity differences between different characters. The embedding mapping model can be an embedding lookup table or a mapping function trained through a neural network. Then, based on historical expression data or preset style parameters, expression style-related parameters are calculated, including but not limited to: an expression exaggeration coefficient, used to represent the range of expression changes; an emotion baseline bias, used to represent the character's basic expression tendency in a neutral state; and a personalized expression tremor suppression coefficient, used to represent the smoothness of short-term expression fluctuations. Subsequently, the character ID embedding vector is concatenated or fused with the expression style parameters to obtain a unified character portrait vector, thereby achieving consistent expression generation style control for each character.

[0080] In this embodiment, through the aforementioned character profile vector extraction process, character identity information and facial expression style parameters are modeled in a unified manner. This allows the character profile vector to not only distinguish the identity characteristics of different digital clones but also to depict their facial expression preferences. Therefore, in the subsequent facial expression generation process, style modulation can be applied to facial expression generation based on the character profile vector, enabling different characters to exhibit differentiated facial expressions under the same voice input. Simultaneously, the facial exaggeration coefficient is used to adjust the facial expression amplitude, the emotional baseline bias is used to control the overall facial expression tendency, and the jitter suppression coefficient is used to reduce short-term unstable changes. This achieves style consistency and stability control in the facial expression generation process, enhancing the personalization and naturalness of the digital clone's facial expressions.

[0081] In summary, the character profile vector includes at least one of an expression style parameter vector and a character identity embedding vector used to map character profile information; the expression style parameter vector includes a vector of at least one of the following parameters: expression exaggeration coefficient, emotion baseline bias, and expression jiggling suppression coefficient. The audio feature sequence includes at least one of prosodic features and content-related features. Prosodic features include at least one of fundamental frequency, energy, and speech rate features. Content-related features include at least one of semantic features, Mel spectrum, Mel frequency cepstral coefficients, and phoneme-related acoustic representations. Phoneme-related acoustic representations refer to the representation method that establishes a relationship between phonemes in audio data and their corresponding acoustic features (such as spectrum, energy, duration, etc.). Expression control information includes at least one of emotion labels and expression control vector sequences, as well as at least one of prosodic event labels, semantic labels, and intent labels corresponding to the audio feature sequence. Prosodic event labels include at least one of pause events and accent events.

[0082] Based on the above, in one embodiment, step S20, the process of generating facial expression control information includes the following steps: embedding prosodic features and / or content-related features, as well as facial expression style parameter vectors and / or role identity identifiers into a vector input role model for emotion simulation, thereby generating facial expression control information. Specifically, the following process is implemented: An audio feature representation is obtained by dimensional alignment and semantic mapping of prosodic and content-related features in an audio feature sequence through an audio feature encoding layer. Specifically, fundamental frequency and energy features in the prosodic features characterize the intensity variations and stress trends of speech, speech rate features reflect changes in speech rhythm, and content-related features represent phonemes and semantic information. Simultaneously, a character profile encoding layer dimensionally aligns and semantically maps the facial expression style parameter vector and the embedding vector of the character identity ID in the character profile vector to obtain character features. Specifically, the facial exaggeration coefficient adjusts the range of facial expression changes, the emotion baseline bias sets the overall facial expression tendency, and the facial tremor suppression coefficient suppresses short-term unstable fluctuations.

[0083] Subsequently, character features are incorporated into the audio feature representation through a cross-modal fusion layer. Conditional modulation of the audio feature representation allows prosodic features-driven temporal variations and character style parameters to jointly influence the expression generation process, resulting in fused features. Specifically, stress and energy variations enhance the intensity of expressions at corresponding moments, while pauses and corresponding speech rate variations trigger expression transitions or softening. A temporal modeling layer dynamically models the fused features, extracting semantic information and expressive intent from content-related features. Based on prosodic features, pause and stress events are identified, forming temporal semantic features that reflect speech content, tone changes, and character characteristics. Finally, a multi-task output module processes the temporal semantic features to generate expression control information. This information collaboratively constrains the amplitude, timing, and type of expression changes, achieving facial expression generation consistent with speech content, tone, and character characteristics.

[0084] Among them, the expression control information is used to constrain the expression type, intensity and temporal distribution during the micro-expression generation process; the expression control vector sequence in the expression control information is used to depict the continuous changes of expression over time; the emotion label is used to indicate the overall emotion type; the prosodic event label is used to identify the timing of expression changes; the semantic label and intention label are used to limit the semantic category and expression purpose of the expression, providing a control basis for subsequent facial animation generation that is simultaneously constrained by the phonological prosody, semantic content and character style.

[0085] In this embodiment, by introducing expression style parameter vectors and / or character identity embedding vectors into the character profile vectors, the character profile vectors can represent the expression tendencies of different characters. Furthermore, by utilizing prosodic features and content-related features contained in the audio feature sequence, the rhythm, intensity, and semantic information of the speech are comprehensively depicted. Based on this, a large character model is used to generate expression control information including emotion tags, expression control vector sequences, prosodic event tags, semantic tags, and intent tags, thereby unifying the modeling of speech content, tone changes, and character features. This allows the expression generation process to not only respond to the pronunciation information of the audio but also to finely control the timing and type of expression changes based on prosodic events such as pauses and stresses in the speech, as well as semantic and intent information. Simultaneously, by combining expression exaggeration coefficients, emotion baseline bias, and expression tremor suppression coefficients, the amplitude and stability of expressions are adjusted, resulting in differentiated styles of generated facial micro-expressions among different characters, and a more coherent and natural temporal sequence. This effectively improves the realism, stability, and emotional expression capabilities of the digital avatar's facial expressions.

[0086] In one embodiment, the parameter generation model includes a micro-expression parameter generation model and a lip shape parameter generation model, which are used to generate facial micro-expression action parameter sequences and lip shape action parameter sequences, respectively, to achieve decoupled modeling of different facial regions and reduce mutual interference. Specifically, the micro-expression parameter generation model processes the audio feature sequence under the constraint of expression control information to generate a micro-expression action parameter sequence that is temporally aligned with the audio feature sequence. The lip shape parameter generation model simulates lip movements based on the audio feature sequence to generate a lip shape action parameter sequence that is temporally aligned with the audio feature sequence.

[0087] Specifically, multiple audio-lip action sample pairs can be extracted from a multimodal facial animation dataset, and a sequence generation network can be trained based on these pairs to obtain a lip shape parameter generation model. Simultaneously, multiple audio-facial micro-expression sample pairs can be extracted from the same dataset, and a micro-expression parameter generation model can be trained based on these pairs. The lip action samples are sequences of lip parameters extracted frame-by-frame, and the facial micro-expression samples are sequences of micro-expression parameters extracted frame-by-frame. Both serve as supervision signals for training and can be AU parameters or BlendShape coefficients extracted for each frame using expression capture or 3D reconstruction tools.

[0088] In the training process of the lip shape parameter generation model and the micro-expression parameter generation model, a Generative Adversarial Network (GAN) training framework can be introduced. A discriminator network is introduced at the expression generation end, and an optimization strategy combining regression loss and adversarial loss is adopted during training. That is, the weighted sum of the regression loss and adversarial loss is used as the total training loss. When the total loss meets certain conditions, such as the total loss being less than a threshold or the training iterations reaching a preset number and the regression loss being minimized, the model is considered converged, and the converged model is output as the corresponding parameter generation model. The total loss of the above parameter generation model is represented by the following function: ; in, This represents the total loss of the micro-expression parameter generation model, or the total loss of the lip shape parameter generation model; This represents the regression loss of the micro-expression parameter generation model or the regression loss of the lip shape parameter generation model. This represents the adversarial loss of the micro-expression parameter generation model or the adversarial loss of the lip shape parameter generation model. This represents a constant value, which is a coefficient used to adjust the relative weights of the two losses.

[0089] The regression loss measures the difference between the model output and facial micro-expression samples (or lip movement samples). Minimizing the regression loss improves the numerical accuracy of expression prediction. The regression loss is represented by the following function: ; in, Indicates regression loss; Indicates the first input audio sample The actual facial expression parameters corresponding to the frame; The representation model generates the first [audio sample] based on the audio sample. The predicted values ​​of facial expression parameters for the frame, i.e., the model output; This represents the square of the vector norm; This indicates the total number of frames in the input audio samples.

[0090] In this model, the discriminator takes both real and predicted facial expression parameters as input and learns to distinguish between them. The generative model corresponding to the parameters aims to make it difficult for the discriminator to differentiate between real and fake expressions. By introducing adversarial training, the model not only fits the average error but also learns facial expression sequences that are closer to the true distribution, thereby reducing the problem of stiff output or deviation from the true distribution.

[0091] During training, the micro-expression parameter generation model and the lip shape parameter generation model can be trained independently. Alternatively, the lip shape parameter generation model can be trained first, and then loaded into the micro-expression parameter generation model for joint training. During the joint training phase, lip-related parameters are fixed or slightly adjusted, with a focus on training micro-expression related parameters to avoid compromising lip shape accuracy. Simultaneously, higher weights can be assigned to micro-expression parameters such as eyebrows and eyes in the loss function to improve the learning of details.

[0092] The multimodal facial animation data can be from real people, used to train a parameter generation model for creating digital avatars of real people. Alternatively, it can be from cartoon characters, used to train a parameter generation model for creating digital avatars of cartoon characters, ensuring that the expression parameters generated by the model are compatible with the character type simulated by the digital avatar. Since the facial proportions and exaggerated movements of digital avatars of cartoon characters typically differ from those of real people, the variation range of parameters such as eyebrows and eyes can be appropriately amplified during training to match exaggerated cartoon expressions. Furthermore, other parameters, such as blinking frequency or movement amplitude, can be adjusted for the cartoon character to match its overall style. Through these adaptations, this invention can be widely used for various realistic or cartoon digital avatars, meeting the expression-driven needs of virtual characters with different styles.

[0093] Furthermore, the method provided in this embodiment of the invention does not rely on a specific language. As long as there is audio-facial expression training data in the corresponding language, it can support multilingual model training, thereby supporting multilingual audio-driven expression generation. For example, by adding English, Japanese, and other speech and expression samples to the training set, the large character model and the corresponding parameter generation model can learn the phonetic prosodic features of different languages, enabling them to correctly synchronize facial expression changes when faced with multilingual input. For real-time speech translation scenarios, the expression can be driven after translation, or the expression output can be adjusted according to the translated emotion to display the facial expression habits of the corresponding language content.

[0094] In one embodiment, such as Figure 4 As shown, in step S30, independent micro-expression parameter generation models and lip shape parameter generation models can be used to generate lip shape action parameter sequences and micro-expression action parameter sequences respectively, thereby achieving decoupled modeling of lip shape actions and facial micro-expressions. Specifically, this process includes the following steps: S31: Input the audio feature sequence into the lip shape parameter generation model to simulate lip movements and generate a lip movement parameter sequence that is time-aligned with the audio feature sequence.

[0095] The lip shape parameter generation model may include an audio feature conversion module, a temporal modeling module, and a lip shape parameter output module connected in sequence. The generation device inputs the audio feature sequence into the lip shape parameter generation model to perform the following steps: First, the audio feature transformation module performs feature transformation on the prosodic and content-related features in the audio feature sequence, mapping them to the same feature space to obtain a multimodal audio feature representation in a unified feature space. Then, the transformed multimodal audio features are input into the temporal modeling module to model the continuous changes of the audio features over time, capturing phoneme transitions and speech rhythm changes, thereby obtaining a temporal feature representation to drive lip movements. Finally, the lip movement parameter output module regresses the temporal features to generate a lip movement parameter sequence that is temporally aligned with the audio feature sequence.

[0096] Content-related features are used to characterize phoneme information to drive the correspondence between lip movements and pronunciation content; prosodic features are used to help depict pronunciation rhythm and changes in mouth opening amplitude. The lip movement parameter sequence is used to characterize the movement changes of the lips and their edge regions, including parameters such as lip opening and closing, lip corner movement, and upper lip elevation.

[0097] S32: Input facial expression control information and audio feature sequence into the micro-expression parameter generation model, use facial expression control information as emotional condition input, and modulate the type, intensity and temporal distribution of facial micro-expression movements through feature modulation and / or cross-attention weighting to generate a micro-expression movement parameter sequence that is temporally aligned with the audio feature sequence.

[0098] The micro-expression parameter generation model may include a feature transformation module, a conditional modulation module, a temporal modeling module, and a parameter output module connected in sequence. The generation device inputs the audio feature sequence and expression control information into the micro-expression parameter generation model and executes the following process: First, a feature transformation module performs feature transformation on the prosodic features and content-related features in the audio feature sequence, and embeds and maps the facial expression control vector sequence and / or various label information in the facial expression control information, thereby obtaining a multimodal feature representation in a unified feature space. This multimodal feature representation includes audio feature representation and facial expression feature representation.

[0099] Subsequently, the multimodal features after feature mapping are input into the conditional modulation module. Using facial expression control information as emotional constraints, the module modulates and / or weights the multimodal audio feature representations based on a cross-attention mechanism, resulting in modulated features. Specifically, the facial expression control vector sequence is used to adjust the continuous change trend and intensity distribution of micro-expressions; emotional labels define the overall emotional type of the micro-expressions; prosodic event labels indicate the temporal position of micro-expression changes, with pause events triggering expression transitions or softening, and accent events enhancing the intensity of the expression at the corresponding moment; semantic and intention labels define the expression type and purpose of the micro-expressions. Through this modulation process, the audio features simultaneously integrate speech rhythm information and facial expression semantic constraint information in the temporal dimension.

[0100] Then, the modulated features are input into the temporal modeling module to perform temporal dependency modeling on the modulated features, thereby obtaining a temporal feature representation that reflects speech content, tone changes, and semantic constraints of facial expressions. Finally, the temporal feature representation is regressed through the parameter output module to generate a micro-expression action parameter sequence that is temporally aligned with the audio feature sequence. This micro-expression action parameter sequence is used to characterize detailed facial expression changes in areas other than the lips, including action parameters of areas such as eyebrows, eyes, cheeks, nose, and chin.

[0101] In this embodiment, the micro-expression parameter generation model and the lip shape parameter generation model process the input information in a targeted manner: content-related features in the audio feature sequence mainly drive the lip shape parameter generation model to generate lip movements consistent with pronunciation; prosodic features in the audio feature sequence act on both models simultaneously, adjusting the lip opening and closing rhythm in the lip shape parameter generation model and triggering changes in the timing and intensity of expression changes in the micro-expression parameter generation model; expression control information acts only on the micro-expression parameter generation model, constraining the type, intensity, and temporal distribution of micro-expressions to avoid interfering with the lip shape generation process. Thus, by structurally decoupling lip movement generation from micro-expression generation and constraining the micro-expression generation process with expression control information, lip movements can accurately align with speech content, while other facial areas can undergo reasonable changes based on speech prosody and semantic information, thereby achieving the collaborative generation of lip shapes and micro-expressions. Therefore, by introducing a decoupled generation mechanism with emotional condition constraints and combining it with a conflict resolution and fusion strategy for overlapping areas, the synergistic optimization of lip shape accuracy and the naturalness of micro-expressions is achieved, thereby improving the realism and stability of digital avatar expressions.

[0102] In one embodiment, the micro-expression parameter generation model adopts the Transformer architecture. In step S32, using expression control information as constraints, a cross-attention mechanism is used to modulate the type, intensity, and temporal distribution of facial micro-expression movements, generating a micro-expression movement parameter sequence that is temporally aligned with the audio feature sequence.

[0103] The micro-expression parameter generation model employs a Transformer encoder-decoder architecture. The encoder performs feature transformation, while the decoder incorporates expression control information through a cross-attention mechanism to achieve conditional modulation, and combines this with a self-attention mechanism to complete temporal modeling and expression parameter generation. Specifically, the feature transformation module includes a Transformer encoder for contextual modeling of the audio feature sequence; the conditional modulation module, implemented by the cross-attention substructure in the Transformer decoder, modulates the audio features under the constraints of expression control information; the temporal modeling module, implemented by the self-attention substructure in the Transformer decoder, characterizes the temporal dependencies of the micro-expression parameters; and the parameter output module, implemented by the feedforward network and output layer in the decoder, generates the micro-expression action parameter sequence.

[0104] Specifically, the audio feature sequence and facial expression control vector are input into the micro-expression parameter generation model, and the following process is completed through the collaboration of various modules: mapping the audio feature sequence to the micro-expression action parameter sequence. S321: Input the audio feature sequence into the feature transformation module, encode it through the Transformer encoder, and calculate the key vector and value vector of each audio feature in the audio feature sequence.

[0105] Specifically, the Transformer encoder in the feature transformation module performs context modeling on the features in the audio feature sequence to obtain the audio context representation corresponding to the audio feature sequence. To facilitate subsequent attention calculations, for each audio feature in the audio context representation, a linear mapping can be performed on the encoder output audio feature based on key weights and value weights to generate the key vector and value vector of that audio feature, thus obtaining the key vector and value vector of each audio feature in the audio context representation. The audio context representation is represented by the following sequence: ; in, This represents the audio features after fusing temporal information before and after fusion, i.e., the first... Audio context representation of time frames; This represents the length of the audio context representation, i.e., the total number of frames in the audio feature sequence. The key vector is the key weight and... The product of the values; the value vector is the product of the value weights and... The product of. It can characterize prosodic features in audio data (including pitch fluctuations, accent positions, and pause boundaries), providing a basis for subsequent decoding.

[0106] S322: Input the key vector and value vector of each audio feature in the facial expression control information and audio context representation into the conditional modulation module, and perform feature fusion processing through the cross-attention substructure to obtain the audio context vector corresponding to the audio feature sequence.

[0107] The facial expression control information can include facial expression control vectors or embedded vector representations of facial expression labels. For each time step, the cross-attention substructure can obtain the decoding state from the previous time step, fusing the facial expression control information with the previous decoding state to obtain the current input state of the decoder. Emotional constraints are introduced at the initial stage of sequence generation, making subsequent feature selection subject to emotional modulation. This can be achieved by directly concatenating or adding the facial expression control information with the previous decoding state. In other embodiments, emotional modulation can be applied to some facial expression parameters at the decoder's output layer to achieve data fusion; certain parameters can be scaled according to the emotional intensity in the facial expression control information: for example, the eyebrow raising amplitude parameter is multiplied by a coefficient greater than 1 when the emotion is "surprise," and multiplied by a coefficient less than 1 when the emotion is "calm," to strengthen or weaken specific expressions. Furthermore, facial expression control information can be directly input into the network as an emotional condition, and the model automatically learns the influence of emotion on each facial expression parameter through training, without the need for manually specifying coefficients.

[0108] This scheme combines the expression control vector with the decoding state in the first layer of the decoder, allowing the decoder to incorporate corresponding emotional information from the initial generation stage. Experiments show that this method can make the facial expressions generated under different emotions exhibit significant differences in overall style: for example, the facial expressions corresponding to happy voices have more smiling and raised eyebrows, while the expressions corresponding to sad voices are characterized by lowered eyebrows and downcast eyes. Through the fusion of expression control vectors (i.e., emotion modulation), the model can globally adjust the expression generation according to the emotional tone of the input audio, enhancing the emotional realism conveyed by the final digital avatar's expressions.

[0109] Then, in any layer of the cross-attention substructure, the current input state of the decoder is linearly transformed into a query vector, and the attention weight of the audio feature at the current time step is calculated based on the key vector of the audio feature. The attention weight of the audio feature is calculated using the following formula: ; in, Indicates the output time Audio features in audio context representation Attention weights are used to characterize the degree of attention given to each audio frame by the expression generation at the current moment; Represents the cross-attention function; Indicates the decoder at the output time The query vector; Representing audio features The key vector, , This represents the total number of frames in the audio feature sequence; This represents the scaling factor, which can be a constant value.

[0110] After calculating the attention weights of each audio feature at the current moment, the value vectors of all audio features are weighted and summed based on their attention weights to obtain the audio context vector at the current moment. In other words, this yields the audio context information most relevant to the generation of the current micro-expression.

[0112] Next, based on a preset modulation function, the current audio context vector is fused with facial expression control information to obtain modulated features. The modulation function can be a gating or affine transformation function. This process enables control over facial expression type, intensity, and temporal distribution.

[0113] S323: Input the modulated features into the temporal modeling module, perform temporal dependency modeling on the modulated features through a self-attention substructure, and input the output temporal feature representation into the parameter output module for regression processing to generate a micro-expression action parameter sequence that is temporally aligned with the audio feature sequence.

[0114] A self-attention substructure is used to model the current and historical states of the decoder, characterizing the temporal dependencies of micro-expression parameters and ensuring the temporal continuity of the generated sequence. This yields the updated hidden representation of the decoder, which is then output as a temporal feature representation. This temporal feature representation reflects speech content, tone variations, and semantic constraints of facial expressions. Then, a parameter output module regresses the temporal feature representation to generate a sequence of micro-expression action parameters that is temporally aligned with the audio feature sequence.

[0115] In this embodiment, by introducing auxiliary structures such as input embedding, feature projection, and conditional modulation into the Transformer structure, and fusing audio context representation and facial expression control information based on the cross-attention mechanism, the model can dynamically select the most relevant audio features at the current moment when generating micro-expression parameters at each moment. At the same time, it is controlled by the facial expression control information to modulate the expression type, intensity, and temporal distribution, thereby generating a micro-expression action parameter sequence that is temporally aligned with the audio feature sequence and has continuous changes, thus improving the naturalness and expressiveness of the digital clone's facial expressions.

[0116] In one embodiment, step S31, namely the process of generating the lip movement parameter sequence, specifically includes the following steps: S311: Input prosodic features and content-related features into the lip shape parameter generation model to simulate lip movements and generate a lip movement parameter sequence that is aligned with the temporal, prosodic, and phonemic sequence of the audio feature sequence.

[0117] Among them, content-related features include at least one of Mel spectrum, Mel frequency cepstral coefficients, and phoneme-related acoustic representations, used to characterize the phoneme composition and pronunciation of speech; prosodic features include at least one of fundamental frequency, energy, and speech rate features, used to characterize the rhythmic changes and pronunciation intensity of speech.

[0118] The generation device inputs prosodic features and content-related features into the lip shape parameter generation model, and then performs the following lip movement simulation and parameter generation process, whereby the lip movements include at least one of the following: jaw opening and closing, lip opening and closing, lip rounding, and lip stretching: First, the audio feature transformation module performs feature transformation and fusion on prosodic features and content-related features to obtain a multimodal audio feature representation in a unified feature space. In this module, content-related features are mapped to highlight phoneme differences, allowing different phonemes to correspond to different mouth shapes; prosodic features are mapped to extract rhythm and intensity variations in speech; and these two types of features are then fused to form a unified audio-driven feature representation.

[0119] Subsequently, the fused multimodal audio features are input into the temporal modeling module for temporal modeling to capture the continuous changes in audio over time. In this process: based on content-related features, the transition relationships between phonemes are modeled, allowing lip shape changes to smoothly transition with the phoneme sequence, thus achieving phoneme-level alignment; based on speech rate information in prosodic features, the time scale of lip shape changes is adjusted to ensure lip movements are consistent with speech rhythm; based on fundamental frequency and energy changes, the amplitude of lip opening and closing and the amplitude of jaw movement are adjusted to make stress positions correspond to more obvious mouth shape changes. Thus, a temporal feature representation containing both phoneme and rhythmic information is obtained.

[0120] Finally, the temporal features containing phoneme and rhythm information are input into the lip shape parameter output module, and a lip shape action parameter sequence is generated through regression. This lip shape action parameter sequence includes at least one of the following parameters: mandibular opening and closing parameters, used to characterize the overall degree of mouth opening; lip opening and closing parameters, used to characterize the distance change between the upper and lower lips; rounded lip parameters, used to characterize the protruding and contracting shape of the lips (such as when pronouncing the "u" sound); and lip stretching parameters, used to characterize the lateral stretching of the corners of the mouth (such as when pronouncing the "i" sound).

[0121] Among them, content-related features mainly determine the shape type of the lips (such as rounded or stretched); energy and fundamental frequency in prosodic features are used to adjust the opening and closing amplitude of the lips; speech rate in prosodic features is used to control the temporal rhythm of lip shape changes. Finally, a sequence of lip movement parameters is generated that is temporally aligned frame by frame with the audio feature sequence, so that the lip movements can simultaneously meet the requirements of phoneme correspondence, speech rhythm consistency, and pronunciation intensity variation.

[0122] In this embodiment, the lip shape parameter generation process is modeled by combining content-related features and prosodic features. This enables the generated lip shape action parameter sequence to not only accurately reflect the phoneme information in speech, but also to dynamically adjust the amplitude and speed of lip shape action according to the rhythm and intensity of speech, and to achieve a continuous and smooth transition in the time dimension. This ensures that the lip shape action is consistent with the speech content, prosody and time, thereby improving the accuracy, naturalness and coherence of lip shape animation.

[0123] In one embodiment, such as Figure 5 As shown, step S40, which involves fusing the lip movement parameter sequence and micro-expression parameter sequence to resolve conflicts in the overlapping area, includes the following steps: S41: Perform partition parameter identification on the lip movement parameter sequence to obtain the lip movement parameters and lip parameter set of the overlapping region; S42: Perform partition parameter recognition on the micro-expression action parameter sequence to obtain the micro-expression action parameters and micro-expression parameter set of the overlapping area; S43: Perform weighted transition fusion on the lip movement parameters and micro-expression movement parameters at the same position to generate an overlapping parameter set that is temporally aligned with the audio feature sequence. Then, temporally align the lip movement parameter set, the micro-expression parameter set, and the overlapping parameter set to obtain a full-face expression parameter sequence.

[0124] Specifically, the overlapping lip movement parameters in the lip movement parameter sequence and the overlapping micro-expression movement parameters in the micro-expression movement parameter sequence are identified to obtain a temporally aligned overlapping region parameter sequence.

[0125] Specifically, the generation device performs parameter space analysis on the lip movement parameter sequence and the micro-expression movement parameter sequence. Based on preset facial region segmentation rules, it identifies overlapping regions that simultaneously participate in lip movements and micro-expression changes. These overlapping regions include, but are not limited to, the corners of the lips, the upper lip edge, and the facial muscle regions around the mouth. Subsequently, the two types of parameter sequences are aligned in the temporal dimension, and parameters belonging to the overlapping regions in the corresponding frames are extracted to obtain a temporally aligned overlapping region parameter sequence. This overlapping region parameter sequence includes a lip-side overlapping parameter sequence and a micro-expression-side overlapping parameter sequence, thus providing a unified input basis for subsequent fusion.

[0126] Then, the lip movement parameter sequence after removing overlapping regions is output as a lip movement parameter set, and the micro-expression parameter sequence after removing overlapping regions is output as a micro-expression parameter set. That is, after identifying overlapping regions, the generation device removes parameters belonging to overlapping regions from the lip movement parameter sequence, retaining only parameters that act solely on pure lip movements, forming the lip movement parameter set. Simultaneously, it removes parameters belonging to overlapping regions from the micro-expression parameter sequence, retaining only parameters that act solely on non-lip facial areas (such as eyebrows, eyes, cheeks, etc.), forming the micro-expression parameter set. Through this step, lip movement parameters are only responsible for pronunciation-related areas, and micro-expression parameters are only responsible for non-lip expression areas, thus achieving region decoupling at the parameter level.

[0127] Subsequently, a weighted transition fusion is performed on the lip movement parameters and micro-expression parameters at the same location in the overlapping region parameter sequence to generate an overlapping parameter set that is temporally aligned with the audio feature sequence. Here, "the same location" refers to data that is identical in both the time dimension and the specific location of the facial overlapping region. Specifically, for the overlapping region parameter sequence obtained in step S41, a weighted transition fusion is performed on the lip movement parameters and micro-expression parameters at the same location within the overlapping region in the same time frame to obtain the overlapping parameters at that location within the overlapping region in that time frame; the overlapping parameters at different facial locations are then temporally aligned to obtain an overlapping parameter set that is temporally aligned with the audio feature sequence.

[0128] In the weighted processing, the weight coefficient of the lip movement parameter is greater than that of the micro-expression parameter. The weight coefficients of the lip movement parameter and the micro-expression parameter can be preset fixed weights. For example, the weight coefficient of the lip movement parameter can range from 0.7 to 0.9; the weight coefficient of the micro-expression parameter can range from 0.1 to 0.3. In other embodiments, the weight coefficients of the lip movement parameter or the micro-expression parameter can be adaptively and dynamically adjusted based on audio prosodic features (such as energy and stress) and / or detected prosodic events, so that the lip movement parameter has a higher weight at stressed moments, thus giving dominance to the lip movement parameter at key moments of pronunciation.

[0129] Finally, the generation device integrates the obtained lip shape parameter set, micro-expression parameter set, and overlap parameter set, and aligns them temporally according to the time sequence of the audio feature sequence to reconstruct a complete full-face expression parameter sequence. Specifically, lip shape parameters fill the lip region, micro-expression parameters fill the non-lip region, and overlap parameters fill the overlapping region, ultimately forming a complete, frame-by-frame corresponding full-face expression parameter sequence to drive the digital clone in facial animation generation.

[0130] In this embodiment, by dividing the lip movement parameter sequence and the micro-expression parameter sequence into regions, separating the parameters, and weighting and fusing overlapping regions, the collaborative modeling of lip movements and micro-expressions is achieved while avoiding conflicts between different parameters in the same region. This enables lip movements to accurately align with speech pronunciation, while other facial regions can present natural facial expression changes. Furthermore, a full-face expression parameter sequence is generated through unified temporal alignment, thereby improving the naturalness, coordination, and realism of the digital avatar's facial animation.

[0131] In one embodiment, step S43 involves performing a weighted transition fusion on the lip movement parameters and micro-expression movement parameters at the same position in the overlapping region parameter sequence to generate an overlapping parameter set that is temporally aligned with the audio feature sequence. This specifically includes the following steps: S431: Determine the first weight of the lip movement parameters at different times based on the prosodic features in the audio feature sequence, and determine the second weight of the micro-expression movement parameters at the same time based on the first weight.

[0132] For each time frame in the audio feature sequence, a first weight corresponding to the lip movement parameters in that time frame is determined, and a second weight corresponding to the micro-expression movement parameters in the same time frame is determined based on this first weight. For example, the second weight can be the difference between 1 and the first weight. The first weight of each time frame in the audio feature sequence can be determined based on the prosodic features of the audio feature sequence, such as changes in prosodic features and / or prosodic events detected based on prosodic features. This ensures that the first weight of the lip movement parameters at different times adaptively changes with the temporal position or prosodic of the audio data; that is, the weight coefficients of different time frames adaptively change with the temporal position or prosodic of the audio data.

[0133] In one embodiment, the first weight can be adaptively changed based on the prosody of the audio data. Specifically, the first weight in each time frame is determined according to the prosodic features of the audio data, such as the fundamental frequency (F0), energy, and speech rate. For example, if the energy (emphasis) of a time frame in the audio data is high, the first weight increases, becoming greater than the first weight of the previous time frame, to emphasize lip movements. If the energy of a time frame in the audio data is low or the speech rate is slow, the first weight decreases, becoming less than the first weight of the previous time frame, to increase the proportion of micro-expression movements. If the speech rate of a time frame in the audio data is fast, the first weight increases, becoming greater than the first weight of the previous time frame, to ensure that the lip movements quickly follow the pronunciation.

[0134] Furthermore, a first weight can be determined based on the type of prosodic event detected by prosodic features. Different types of prosodic events correspond to different first weights, enabling dynamic adjustment of lip movement parameters and micro-expression parameters in the overlapping region. Prosodic events can include stress events, pause events, and speech rate change events. For example, when a local peak in audio energy or a significant increase in fundamental frequency is detected, it can be identified as a stress event; when audio energy remains below a preset threshold and speech rate decreases, it can be identified as a pause event; when speech rate or rhythm changes significantly, it can be identified as a speech rate change event. When it is a stress event, a larger first weight is determined to enhance the dominant role of lip movement in the overlapping region, thereby ensuring the accuracy of lip expression at key moments of pronunciation; when it is a pause event, a smaller first weight is determined to increase the proportion of micro-expression movements, making facial expressions more natural and rich during speech pauses; when the speech rate is fast or the rhythm is compact, the first weight can be appropriately increased so that lip movement can quickly follow speech changes.

[0135] In other embodiments, the first weight can adaptively change based on the temporal position of the audio data. For example, during the pronunciation segment of the audio data, the first weight decreases as time increases; during the transition or silence segment of the audio data, the first weight remains unchanged, and the first weight of the pronunciation segment is greater than the first weight of the transition or silence segment, so that the lip shape dominates during the pronunciation segment and the facial expression is enhanced during the non-pronouncing segment.

[0136] S432: Based on the first weight and the corresponding second weight, the lip movement parameters and micro-expression movement parameters at the same time and position in the overlapping region parameter sequence are weighted and summed to obtain the overlapping parameter set that is temporally aligned with the audio feature sequence.

[0137] After determining the first weight and the second weight, based on the first weight and the second weight of the same time frame, the lip movement parameters and micro-expression movement parameters at the same overlapping region position in the overlapping region parameter sequence are weighted and summed to obtain the overlapping parameters corresponding to that time frame. After performing the above calculation on all time frames and all overlapping region parameters, the overlapping parameter set is obtained. This overlapping parameter set is aligned with the audio feature sequence frame by frame in the time dimension.

[0138] Specifically, for motion parameters at a certain location within an overlapping region in the same time frame (i.e., at the same moment), the fusion result can be calculated in the following way: ; in, This represents the fusion result, i.e., the first... Motion parameters at a certain location within the overlapping region of a time frame The overlap parameter; Indicates the first Motion parameters at a certain location within the overlapping region of a time frame lip movement parameters in the text; Indicates the first Motion parameters at a certain location within the overlapping region of a time frame Micro-expression motion parameters; Indicates the first The weighting coefficients of the lip movement parameters in the time frame, i.e., the first... The first weight of the time frame. Indicates the first The second weight of the time frame, i.e., the first Weighting coefficients for micro-expression motion parameters in time frames.

[0139] In this embodiment, the weights of lip movement parameters and micro-expression parameters are adaptively determined based on the temporal position and prosodic features of the audio data, and the parameters in the overlapping area are weighted and transitionally fused. This allows lip movement and micro-expression to have different degrees of dominance in different time periods. While ensuring the accuracy of lip movement during the key stages of pronunciation, the expression is enhanced during non-key stages. This achieves dynamic coordination and collaborative control of lip movement and micro-expression, thereby improving the accuracy, naturalness, and rhythmic consistency of the generated facial animation.

[0140] In one embodiment, such as Figure 6 As shown, the above-mentioned process involves superimposing and smoothing preset motion parameters on the full-face expression parameter sequence to obtain the target expression parameter sequence. This target expression parameter sequence is then mapped to digital avatar facial driving parameters and used to drive the rendering output. Specifically, the process includes the following steps: S51: Perform a superposition operation of preset action parameters on the full-face expression parameter sequence to generate an optimized full-face expression parameter sequence. The preset action parameters include at least one of blinking action parameters and nodding action parameters.

[0141] The generation device can perform a superposition operation of preset action parameters on a full-face expression parameter sequence based on preset action superposition rules to generate an optimized full-face expression parameter sequence. The preset action parameters include at least one of blinking action parameters (such as the amplitude of eyelid opening and closing and the duration of its closure) and nodding action parameters (such as the chin, changes in head tilt angle, etc.).

[0142] The preset action parameters can be obtained in the following ways: generating blinking or nodding actions according to preset time intervals or probabilities; triggering actions based on pause events, speech rate changes, or stress positions in the audio feature sequence; and triggering corresponding actions based on intention tags or emotion tags in the facial expression control information.

[0143] The preset action overlay rules can include the following: region constraint, that is, only overlaying relevant action parameters in the corresponding facial region, such as blinking only affecting the eyes; amplitude limitation, that is, constraining the amplitude of the overlay parameters to avoid excessive movement that would look unnatural; time window control, that is, limiting the duration and transition range of the action so that blinking or nodding actions have a start, peak and recovery phase; priority control, that is, when multiple actions exist at the same time, they are selected or weighted overlay according to priority.

[0144] In this process, a preset action parameter superposition operation is performed on the full-face expression parameter sequence. The superposition method can be to directly insert action parameters for the corresponding region at the corresponding time frame position, such as directly inserting blinking action parameters (eyelid opening and closing amplitude). In other embodiments, parameter weighted fusion can also be used for parameter superposition. For example, for each facial region position in each time frame where preset action parameters need to be superimposed, based on the corresponding weights, the original action parameters at that position in that time frame of the full-face expression parameter sequence, and the preset action parameters to be superimposed, are weighted and fused to obtain the superimposed action parameters; the superimposed action parameters are then updated to the position of the original action parameters to obtain the optimized full-face expression parameter sequence.

[0145] S52: Perform time-series smoothing on the optimized full-face expression parameter sequence to generate the target expression parameter sequence, which is then mapped to the digital clone face driving parameters and used to drive the rendering output.

[0146] After obtaining the optimized full-face expression parameter sequence, the parameters can be classified based on preset classification rules, such as distinguishing them into action parameters of different detail regions, or classifying action parameters of different prosodic events according to the prosodic feature changes in the audio feature sequence. Then, moving average filtering and exponential smoothing are used to perform temporal smoothing on different types of action parameters. Alternatively, a temporal filtering network or Kalman filtering can be used to perform temporal smoothing on the optimized full-face expression parameter sequence to generate the target expression parameter sequence. After smoothing, the target expression parameter sequence is obtained and mapped to digital avatar facial driving parameters, which drive rendering to generate facial animation output, thereby achieving synchronized expression of lip shapes and micro-expressions. Through smoothing, parameter abrupt changes in the expression parameter sequence are eliminated, high-frequency jitter is reduced, and the continuity of movement is maintained.

[0147] In this embodiment, by introducing preset action parameters into the full-face expression parameter sequence and superimposing them with constraints, the digital clone can express natural behaviors such as blinking and nodding on the basis of lip shape and micro-expression synchronization. By performing temporal smoothing processing on the superimposed parameter sequence, parameter abrupt changes and high-frequency jitter are eliminated, making facial movements more continuous and stable in the time dimension, thereby improving the naturalness, coherence and realistic performance of the digital clone's facial animation.

[0148] In other embodiments, the preset motion parameters also include other subtle motion parameters that characterize human interaction habits, such as a brief shift in gaze direction to simulate the occasional deviation of the gaze from the other person during conversation and its return. Additionally, parameters for changes in the corners of the mouth may be included, such as a slight upward or downward tilt of the corners of the mouth to indicate a slightly confident or doubtful expression. By adding parameters for these subtle movements commonly found in real-life communication, the digital avatar can be driven to perform corresponding actions during interaction, helping to enhance the realism or naturalness of the digital avatar's performance.

[0149] In one embodiment, step S51, which involves performing a superposition operation of preset action parameters on the full-face expression parameter sequence, specifically includes at least one of the following steps: S511: When no blinking event is detected within the first time period based on the micro-expression action parameter sequence, blinking action parameters are inserted at the corresponding position in the full-face expression parameter sequence.

[0150] The generation device first analyzes the eye-related parameters in the micro-expression action parameter sequence to detect blinking events. The blinking event can be determined as follows: a blinking event is detected when the eyelid opening / closing parameter changes from a high value to a low value and then recovers within a short period; or, a blinking event is detected when the degree of eye closure exceeds a preset threshold, indicating a closed eye state. A timer or frame counter is maintained to record the time interval since the last blinking event, obtaining the duration of no blinking. When the duration of no blinking exceeds a preset first duration, it is determined that the current full-face expression parameter sequence lacks natural blinking behavior, thus triggering blinking action overlay to insert blinking action parameters at the corresponding time position (i.e., time frame) in the full-face expression parameter sequence.

[0151] The first duration is a random threshold value taken within a preset time range (e.g., 3-5 seconds). The first duration is evenly distributed within the preset time range, and a new, randomly selected threshold value is used as the first duration after each blink event is detected. In other embodiments, the first duration can also be a fixed threshold. By randomly selecting a first duration value within a certain range, blinking actions can be inserted at appropriate locations, increasing the randomness of blinking and reducing the stiffness caused by blinking at a fixed frequency, thereby improving the overall naturalness of the digital avatar.

[0152] S512: When a pause event is detected based on the lip movement parameter sequence, blinking action parameters are inserted at the corresponding position in the full-face expression parameter sequence.

[0153] The generation device first analyzes the lip movement parameter sequence to detect pause events in speech (i.e., speech pause events), such as the end of a sentence or the appearance of a noticeable silence. The pause event can be determined as follows: a pause event is detected when the lip opening / closing parameter is consistently below a preset threshold and the jaw opening / closing amplitude is small; or, based on the energy or speech rate characteristics of each feature in the audio feature sequence, a decrease in speech energy or a silence segment is detected, thus determining that a pause event has been detected. When a pause event is detected, it is determined that the current time position is suitable for inserting a natural behavioral action, thereby triggering a blinking action overlay. The blinking action parameter is then inserted into the corresponding time position in the full-face expression parameter sequence. Furthermore, if no blinking event is detected within a certain duration (this value is less than or equal to a first duration) after a pause event is detected, a blinking action overlay is triggered, and the blinking action parameter is inserted into the corresponding position in the full-face expression parameter sequence.

[0154] The blinking parameters can be generated using preset functions or templates. For example, a time window function can be used to model the eyelid opening and closing parameters, giving the blinking action a natural speed variation. The blinking parameters include parameters for three stages: the closing phase (rapid eyelid closure), the holding phase (short eyelid closure), and the opening phase (eyelid reopening). Alternatively, the blinking parameters can also include parameters for two stages: the closing phase (rapid eyelid closure) and the opening phase (gradual eyelid reopening). This can be achieved by rapidly increasing the eyelid closure parameters to near 1 (eyelid closure) in a single frame, and then gradually decreasing them back to 0 (eyes fully open) over the next few frames (e.g., 100 milliseconds), thus completing a full blink.

[0155] The process involves inserting the generated blinking motion parameters into the corresponding time positions of the full-face expression parameter sequence. Specifically, if the blinking is triggered by the duration of non-blinking, it is inserted at or near the current time point; if it is triggered by a pause event, it is inserted at an appropriate position within the pause interval. During the insertion of blinking motion parameters, the following processing can also be performed: superimposing or replacing existing eye parameters; and smoothing the parameter transition at the boundaries of the inserted blinking motion parameters (i.e., the closing and opening phases) to avoid abrupt changes. This results in an optimized full-face expression parameter sequence that includes natural blinking behavior. For example, the parameters of several consecutive frames within the boundaries of the blinking motion parameters can be subtracted to make the eyelid opening and closing movements smoother and more natural, reducing overly mechanical blinking.

[0156] In this embodiment, by detecting the duration of non-blinking based on micro-expression action parameter sequences and speech pause events based on lip-shape action parameter sequences, blinking action parameters are adaptively inserted into the full-face expression parameter sequence when preset conditions are met. This allows the digital clone to blink periodically at a frequency consistent with human eye physiology, avoiding the unnaturalness caused by prolonged absence of blinks and coordinating the blinking action with the speech rhythm. Simultaneously, a smooth insertion method ensures the continuity of action transitions, thereby improving the naturalness, coherence, and realism of the digital clone's facial expressions.

[0157] In one embodiment, step S51, which involves performing a superposition operation of preset action parameters on the full-face expression parameter sequence, specifically includes at least one of the following steps: S513: Based on the detection of accent events in the audio feature sequence, nodding action parameters are inserted at the corresponding positions in the full-face expression parameter sequence.

[0158] Nodding is a common form of body language in human communication; a slight nod can be used to emphasize a point or match the rhythm of a conversation. Introducing nodding into digital avatar facial expression synthesis can enhance the interactivity of the character. The generation device first analyzes the audio feature sequence to detect stress events in the speech. Stress events can be determined based on at least one of the following features: local peaks in energy features; significant increases or changes in the fundamental frequency (F0); and the detection of emphasized syllables in conjunction with speech rhythm information.

[0159] When an accented event is detected at a specific time frame, that time frame should be identified as the trigger point for a nodding motion. Subsequently, a corresponding nodding motion parameter sequence is generated to overlay a small-amplitude head posture change upon detecting the accented event. Finally, the nodding motion parameters are inserted into the full-face expression parameter sequence at the time frame corresponding to the accented event, and the insertion boundaries are smoothed to avoid abrupt changes in motion.

[0160] S514: When a single prosodic event lasting a second duration is detected based on the audio feature sequence, a nodding motion parameter is inserted at the corresponding position in the full-face expression parameter sequence.

[0161] The generating device can also perform continuous analysis on the audio feature sequence to detect single prosodic events with a duration reaching a second duration. A single prosodic event can include at least one of the following: the fundamental frequency remains consistently within a relatively stable range; the energy remains consistently at a high or low level; the speech rate remains substantially constant over a certain period of time.

[0162] When the duration of the aforementioned single prosodic state exceeds a preset second duration, the current speech is determined to be in a relatively monotonous or continuous expression state. In this case, an appropriate position within the corresponding time interval is determined as the insertion point for the nodding action. For example, the nodding action can be inserted in the middle or at the end of a long sentence to ensure that the synchronization between audio and lip movement is not disrupted. Subsequently, a nodding action parameter sequence is generated. Finally, the nodding action parameters are inserted into the corresponding time position of the full-face expression parameter sequence, and a smooth transition is performed to make the action flow naturally. During the generation of the nodding action parameter sequence, the amplitude, duration, and minimum trigger interval of the action can be appropriately adjusted according to the duration of the single prosodic event to avoid abrupt actions.

[0163] For example, when the audio feature sequence is detected to have a stable and monotonous tone and energy over a prolonged period, lacking variation, the system determines that the current speech segment may be too monotonous. Detecting a prolonged period of single-rhythmic events, to avoid the digital avatar appearing stiff, a subtle nodding motion is triggered at an appropriate moment. Specifically, the nodding motion is achieved by controlling the digital avatar's head to rotate slightly along the left and right ear axes; that is, the head tilts slightly forward and downward for a short time before returning to its original position. A single nod can be broken down into a continuous sequence of three frames: downward—brief pause—upward, each lasting several frames, forming a gentle pitching motion. In other words, the nodding motion parameters can include parameters for changes in head pitch angle, as well as parameters for jaw coordination. The nodding motion parameters can be constructed according to a preset time window, including motion parameters for the initial stage (i.e., head moving downward), the peak stage (i.e., the head reaching its maximum pitch angle and pausing briefly), and the recovery stage (i.e., the head lifting up, returning to the initial posture before stacking).

[0164] To avoid interfering with other facial expression parameters being output, the nodding motion parameter is added as an additional head posture parameter offset during the post-processing stage of facial expression generation. The angle change of the nodding motion should not be too large, generally within a few degrees, to avoid appearing abrupt. When there are stressed or emphasized words, a smaller, faster nod can be selectively triggered to strengthen the tone. Through the application of the nodding motion superposition rule, the digital avatar will occasionally exhibit slight head movements when speaking for long periods or with a flat tone, as if consciously interacting with the rhythm of the speech, enhancing the overall expressiveness.

[0165] In this embodiment, by detecting stress events and continuous single prosodic events based on audio feature sequences, head nodding motion parameters are adaptively inserted at the corresponding time positions. This allows the head nodding behavior to match the emphasis information and rhythmic changes in the speech, enhancing the speech expression while avoiding monotonous facial movements. Furthermore, smooth insertion ensures the naturalness of the motion transition, thereby improving the expressiveness, rhythm, and overall coordination of the digital avatar's facial animation.

[0166] In one embodiment, step S52, which involves performing temporal smoothing on the optimized full-face expression parameter sequence to generate the target expression parameter sequence, specifically includes the following steps: S521: Classify the motion parameters in the optimized full-face expression parameter sequence to obtain multiple parameter sequence groups of different types.

[0167] The generation device can classify the motion parameters in the optimized full-face expression parameter sequence based on preset parameter classification rules, resulting in multiple parameter sequence groups of different types. Specifically, for any parameter dimension, its type can be determined as follows: calculate the duration or rate of change of the parameter in the time dimension; when the duration of parameter change is less than or equal to a preset threshold, or the rate of change is greater than or equal to a preset threshold, it is classified as an instantaneous behavior parameter; when the duration of parameter change is greater than a preset threshold, or the rate of change is less than a preset threshold, it is classified as a slowly changing behavior parameter.

[0168] Based on the above classification rules, the optimized full-face expression parameter sequence is divided into multiple parameter sequence groups. For example, the preset parameter classification rules include rules for classifying dynamic actions based on the duration of action changes, so as to classify the action parameters in the optimized full-face expression parameter sequence into different types. For example, dynamic actions include instantaneous behaviors and gradual changes, that is, multiple parameter sequence groups can include instantaneous behavior parameter sequence groups and gradual change behavior parameter sequence groups. Among them, instantaneous behaviors refer to actions with short duration and rapid changes, such as blinking, nodding quickly, etc.; gradual changes refer to actions with long duration and gentle changes, such as raising eyebrows, raising cheeks, etc.

[0169] S522: Based on the type corresponding to each parameter sequence group, determine the filtering strength of each parameter sequence group, and perform smoothing filtering on each parameter sequence group based on the corresponding filtering strength to generate the target expression parameter sequence.

[0170] Specifically, the filtering intensity for instantaneous behavioral parameters is less than that for slowly varying behavioral parameters. For instantaneous behavioral parameter sequences, a lower filtering intensity (such as a smaller smoothing window) or no filtering is used to preserve the rapid change characteristics of the action; for slowly varying behavioral parameter sequences, a higher filtering intensity (such as a larger smoothing window) is used for low-pass filtering, moving average, or exponential smoothing to suppress high-frequency jitter and improve the smoothness and stability of the overall change.

[0171] Then, based on the corresponding filtering intensities, after smoothing each parameter sequence group, the smoothing results of different groups are recombined according to the original parameter structure to generate the target expression parameter sequence, which is used to drive the digital clone.

[0172] In this embodiment, facial expression parameters are classified based on the duration of action changes, and filtering intensity is determined for different types of parameters for smoothing. This improves the smoothness and stability of gradual changes while preserving the dynamic characteristics of instantaneous behavior, avoiding the loss of detail and jitter caused by uniform smoothing. As a result, the generated facial animation achieves a balance between action clarity and overall naturalness, enhancing the realism and coordination of the digital avatar's facial expressions.

[0173] In other embodiments, the classification method for motion parameters in the optimized full-face expression parameter sequence can also be different. For example, the optimized full-face expression parameter sequence can be spatially divided based on preset facial region division rules to obtain parameter sequence groups for different region locations. Facial regions may include, but are not limited to: lip region (including upper and lower lips and corners of the lips), eye region (including eyelids and periorbital muscles), eyebrow region, cheek region, chin region, and nose region. Parameter sequence groups for different region locations may include lip parameter sequence groups, eye parameter sequence groups, eyebrow parameter sequence groups, cheek parameter sequence groups, etc.

[0174] After the regions are divided, the filtering intensity of each parameter sequence group can be differentiated according to the motion characteristics of different regions: for the lip region parameters, a lower filtering intensity is used to ensure the fine changes in lip movements and the accuracy of pronunciation; for the eye region parameters, the filtering intensity is reduced during instantaneous actions such as blinking and increased during non-action stages; for the cheek, eyebrow, and other region parameters, a higher filtering intensity is used to enhance the naturalness of facial expression transitions.

[0175] In this embodiment, facial expression parameters are classified based on facial regions, and the movement characteristics of different regions are differentiated. This allows the lips, eyes, and other facial regions to be independently optimized according to their respective expressive functions. While ensuring the accuracy of lip movements, the naturalness of micro-expressions is enhanced. Furthermore, coordination and unity between regions are achieved through regional weights and constraint control, thereby improving the precision, naturalness, and overall consistency of the digital avatar's facial expressions.

[0176] In one embodiment, such as Figure 7 As shown, the method also includes a digital clone's expression error correction and rollback process, which includes the following steps: S61: Real-time detection of the signal status of input audio data.

[0177] During the generation of a full-face expression parameter sequence or the driving of a digital clone, the generation device can also acquire input audio data in real time and detect its signal status to determine whether there is signal interruption, signal abnormality, or prolonged silence. Signal abnormalities can include audio data loss, data abrupt changes or distortion, input stream interruption, etc. Prolonged silence indicates that the detected audio silence duration exceeds a certain preset time threshold.

[0178] The audio signal status detection includes: determining that silence or signal interruption is detected when the energy in the audio feature sequence is lower than a preset energy threshold; determining that signal abnormality is detected when several consecutive frames of audio features in the audio feature sequence are missing or abnormal; and determining that long-term silence or abnormal interruption is detected when the duration of silence or signal interruption exceeds a preset time threshold.

[0179] S62: When an interruption, abnormality, or prolonged silence is detected in the input audio data, the sequence of facial expression parameters from the most recent frame or a recent period is used as the current output to drive the digital clone.

[0180] S63: Within a preset transition time, the currently output expression parameter sequence is gradually rolled back to neutral expression parameters to drive the digital clone's facial expression to gradually change to a neutral expression.

[0181] When the input audio data is detected to have signal interruption, signal abnormality, or long-term silence, the generating device can directly use the full-face expression parameters (or target expression parameters) of the most recent frame as the current output to drive the digital clone; or it can use the average of the full-face expression parameters of the most recent several frames as the current output to drive the digital clone, so that the digital clone can maintain the current expression in the early stage of audio interruption and avoid sudden changes.

[0182] Subsequently, within a preset transition time, a neutral constraint is introduced during the generation process, causing the generated expression parameters to gradually converge towards neutral expression parameters, generating a continuous and smooth sequence of expression parameters that transitions to a neutral state. This means the current expression parameters gradually transition to neutral expression parameters. The transitioned sequence of expression parameters is then used as the driving input to control the digital clone's facial expressions to gradually change to a neutral expression state.

[0183] The neutral expression is the baseline expression for a character with no facial micro-expressions and closed lips. The neutral expression parameters are the full-face expression parameters corresponding to the neutral expression. Smooth transitions can be achieved using interpolation.

[0184] When no signal interruption, signal abnormality, or prolonged silence is detected in the input audio data, i.e. when the signal is normal, the expression parameter generation process is executed normally to output the generated expression parameters to drive the digital clone.

[0185] In this embodiment, the audio signal status is detected in real time during the generation or driving of the facial expression parameter sequence. When signal interruption, abnormality or long-term silence is detected, the most recent facial expression parameter is used for transition driving. The expression is gradually adjusted to a neutral expression within a preset time. This prevents the digital clone from abruptly changing its expression in the case of audio abnormality, and achieves a natural and smooth transition process, thereby improving the robustness, stability and realistic facial expression performance of the digital clone.

[0186] In one embodiment, the facial expression error correction rollback process further includes the following steps: S64: During the process of gradually reverting the expression parameter sequence to neutral expression parameters, if the signal of the input audio data is detected to return to normal, record the current expression parameter sequence corresponding to the recovery time.

[0187] During the process of gradually adjusting the expression parameter sequence to neutral expression parameters, the generation device can continuously monitor the input audio data in real time. When the audio signal is detected to have recovered from an interrupted, abnormal, or silent state to a normal state, the moment of audio signal recovery is determined, and the current expression parameter sequence corresponding to that moment is recorded. This expression parameter sequence can be either the expression parameters currently transitioning to a neutral expression, or the expression parameters that have recently stabilized.

[0188] S65: Generate the restored target facial expression parameter sequence based on the restored input audio data.

[0189] Meanwhile, based on the recovered input audio data, the above expression parameter generation process is re-executed to obtain a full-face expression parameter sequence or target expression parameter sequence corresponding to the recovered audio. This parameter sequence is aligned with the recovered audio data in the time dimension.

[0190] S66: The current expression parameter sequence corresponding to the recovery time is transitionally fused with the recovered target expression parameter sequence to generate a continuously output expression parameter sequence, and the digital clone is driven according to the continuously output expression parameter sequence to avoid abrupt changes in output.

[0191] Specifically, a preset transition duration can be set. Within this duration, a smooth transition is achieved from the expression parameter sequence corresponding to the recovery time to the recovered full-face expression parameter sequence, generating a continuously output expression parameter sequence. For each time frame within the preset transition duration, interpolation is used to fuse the full-face expression parameters and the expression parameter sequence corresponding to the recovery time within that time frame, resulting in a fused expression parameter sequence for the preset transition duration. Then, the expression parameter sequence corresponding to the recovery time, the fused expression parameter sequence for the preset transition duration, and the recovered full-face expression parameter sequence are sequentially concatenated to obtain a continuously output expression parameter sequence. Finally, the digital clone is driven according to the continuously output expression parameter sequence.

[0192] In this embodiment, by recording the facial expression parameters corresponding to the recovery time when the audio signal is recovered, and generating a new sequence of facial expression parameters based on the recovered audio, and simultaneously performing transition fusion on the two, the facial expression smoothly transitions from the abnormal state to the normal driving state, avoiding abrupt changes or discontinuities in facial expression, thereby improving the stability, continuity, and natural performance of digital clone facial animation in signal recovery scenarios.

[0193] In one embodiment, the generation device can also monitor the amplitude of changes in the output full-face or target expression parameter sequence. When an abnormal, drastic change in the expression parameter is detected (e.g., an abnormal peak in a parameter), a protection mechanism is triggered to prune or moderate the parameter, preventing the output of unreasonable extreme expressions. Furthermore, for uncertainties in the emotion recognition results (i.e., expression control information), threshold filtering or conservative strategies can be set, such as outputting a neutral expression when the confidence of the emotion classification label is too low, to improve the reliability of expression generation. Through these fault-tolerant measures, even in noisy environments, with network latency, or with abnormal input, the digital avatar's facial expressions can remain coherent and natural, without significant distortion or discrepancies with the speech content.

[0194] In a specific embodiment, such as Figure 8 As shown, the method provided in this embodiment of the invention can generate a final target expression parameter sequence through the following data processing flow, so as to generate a rendering animation of a digital clone for output display: Step S1: Audio data input; Step S2: Audio preprocessing, which involves preprocessing the audio data, including noise reduction, normalization, and frame segmentation; Step S3: Extract features from the preprocessed audio data to generate an audio feature sequence; Step S4: Perform character large model inference, that is, input the audio feature sequence and the extracted digital clone character portrait vector into the character large model to simulate character emotions, and output facial expression control information that is aligned with the audio feature sequence in time. Step S5: Micro-expression parameter generation, that is, inputting the audio feature sequence and expression control information into the micro-expression parameter generation model to generate facial action parameters, and outputting a micro-expression action parameter sequence that is temporally aligned with the audio feature sequence; Step S6: Lip shape parameter generation, that is, inputting the audio feature sequence into the lip shape parameter generation model to generate lip shape action parameters, and outputting a lip shape action parameter sequence that is temporally aligned with the audio feature sequence; Step S7: Parameter fusion, which involves dividing the lip movement parameter sequence and the micro-expression movement parameter sequence into parameter partitions and weighted fusion of parameters in overlapping areas to obtain the full-face expression parameter sequence; Step S8: Perform facial expression post-processing on the full-face expression parameter sequence, including superimposing preset motion parameters, partition smoothing, and error-tolerant rollback, to generate the final target expression parameter sequence; Step S9: Drive rendering output, that is, map the target expression parameter sequence to the facial driving parameters of the digital clone, generate the facial animation of the digital clone and output it for display.

[0195] In this embodiment, the specific implementation process of each step can be found above, and will not be repeated here.

[0196] It should be noted that the above steps can be executed continuously in parallel and pipelined modes to support real-time applications. That is, as subsequent audio data is continuously input, the generation device can extract features and generate parameters from the newly input audio data, while simultaneously rendering and outputting the already generated facial expression parameters in digital clone form, thereby reducing end-to-end latency and maintaining audiovisual synchronization. Furthermore, upon detecting audio data interruption, abnormality, or prolonged silence, a fault-tolerant fallback mechanism can be triggered, causing the output facial expression to gradually transition to a neutral state within a preset transition time to ensure continuity.

[0197] This solution utilizes a cross-modal alignment mechanism between the character's large model conditional control and the micro-expression generation model to enable the generation of detailed facial expressions on the upper half of the face synchronously with changes in rhythm and emotional tone. While ensuring alignment between lip shape and audio timing, it enhances the synchronous richness of micro-expressions such as eyebrows and eyes. In this embodiment, the lip shape branch focuses on pronunciation-related movements, while the micro-expression branch focuses on subtle facial changes. Decoupling and partitioning these two types of parameters reduces interference between lip shape and facial expression. The fusion stage uses partitioning rules and weighted transitions of overlapping areas to resolve conflicts, thereby improving overall naturalness. Overlaying regular actions such as blinking and nodding into the full-face expression parameter sequence adds details that conform to physiological and communication habits. The temporal smoothing and constraints of the sub-parameters suppress jitter and ensure smooth transitions, improving visual consistency. Thus, through regular actions and sub-parameter smoothing, the naturalness and stability of the digital avatar's performance are enhanced.

[0198] In addition, this digital clone also has real-time output and fault-tolerant rollback capabilities, reducing end-to-end latency through parallel and pipelined processing; it triggers smooth rollback of facial expressions to a neutral state when audio is interrupted, abnormal, or when there is a long period of silence, and continues to output synchronously after the abnormality is recovered, thereby enhancing the stability of the system under complex input conditions and making it suitable for real-time interaction and online deployment scenarios.

[0199] The method provided in this embodiment of the invention was used as the target solution, and the test results were compared with two other solutions. The first comparison solution only generated a lip shape parameter sequence without generating other micro-expression parameters; the second comparison solution, based on the first, simply superimposed simple blinking rules at a fixed frequency. Based on the same test audio data, the target solution and the two comparison solutions generated digital avatar facial animations, and subjective evaluations were performed using lip shape synchronization, expression naturalness, and emotional consistency as indicators. Objective quantities such as blinking frequency and output sequence jitter amplitude were also statistically analyzed. Test results show that: compared to the first comparison solution, the target solution provided in this embodiment significantly improves the richness of subtle expressions in areas such as the eyebrows and eyes without reducing lip shape synchronization; compared to the second comparison solution, the target solution provided in this embodiment, through large character model control and subparameter smoothing strategies, makes expression transitions more coherent, blink triggers more consistent with the distribution of pause events, and suppresses output jitter, thereby improving overall naturalness and stability.

[0200] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0201] In one embodiment, a generation device is provided, which corresponds one-to-one with the method for generating digital clones based on a large character model driven by micro-decoupling and synchronization, and related product methods described in the above embodiments. For example... Figure 9 As shown, the generation device includes an audio input and preprocessing module 901, an audio feature extraction module 902, a character portrait acquisition module 903, a character large model reasoning module 904, a lip shape parameter generation module 905, a micro-expression parameter generation module 906, an expression fusion module 907, an expression post-processing module 908, and a face-driven rendering module 909.

[0202] The audio input and preprocessing module 901 is used to acquire the input audio data and perform noise reduction, normalization, silence detection and frame segmentation on it, and output an audio frame sequence. The audio feature extraction module 902 is used to extract prosodic features and / or content-related features from the audio frame sequence and form an audio feature sequence. The character profile acquisition module 903 is used to acquire a character profile vector obtained by feature extraction of the character profile information of the target digital clone. The character profile vector is used to represent the character's identity, expression style and preferences. The character large model inference module 904 is used to input the audio feature sequence and the character portrait vector into the pre-trained character large model to simulate character emotions and obtain facial expression control information. The lip shape parameter generation module 905 is used to generate a sequence of lip shape action parameters aligned with the audio time sequence based on the audio feature sequence. The micro-expression parameter generation module 906 is used to generate a sequence of micro-expression action parameters aligned with the audio time sequence based on the audio feature sequence and expression control information. The expression fusion module 907 is used to partition and fuse the lip movement parameter sequence and the micro-expression movement parameter sequence according to the parameter partitioning rules to obtain the full-face expression parameter sequence. The overlapping area parameters obtained by partitioning are weighted transition fusion to resolve conflicts. The facial expression post-processing module 908 is used to perform the superposition of preset action parameters and the temporal smoothing of sub-parameters on the full-face facial expression parameter sequence to obtain the target facial expression parameter sequence, and to trigger facial expression fault tolerance backoff to ensure continuous output when audio interruption, signal abnormality or long-term silence is detected. The face-driven rendering module 909 is used to map the target expression parameter sequence to the face-driven parameters of the digital clone, and drive the digital clone to render and output, so as to realize the synchronous presentation of lip shape and micro-expression.

[0203] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of the present invention. Their specific functions and the resulting technical effects, such as the superposition of preset action parameters, the temporal smoothing of sub-parameters, and the specific implementation process and effects of expression error tolerance rollback processing executed by the expression post-processing module 908, as well as the specific implementation process and effects of the functions of other modules, can be found in the method embodiments section, and will not be repeated here.

[0204] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the functions described above can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0205] This invention also provides a computer device, such as... Figure 10 As shown, the computer device includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the processor executes the computer program, it implements the steps in any of the above-described method embodiments, or when the processor executes the computer program, it implements the functions of each module / unit in the above-described device embodiments. For example, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the computer device.

[0206] Those skilled in the art will understand that Figure 10 The computer device described is merely an example and does not constitute a limitation on the computer device. It may include more or fewer components than shown, or combine certain components, or different components. For example, the computer device may also include input / output devices, network access devices, buses, etc.

[0207] The processor mentioned above can be a central processing unit, or it can be other general-purpose processors, digital signal processors, application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0208] The memory can be an internal storage unit of the computer device, such as a hard drive or RAM. The memory can also be an external storage device of the computer device, such as a plug-in hard drive, smart memory card, secure digital card, flash memory card, etc. Furthermore, the memory can include both internal and external storage units of the computer device.

[0209] This invention also provides a readable storage medium storing a computer program, which, when executed by a processor, can implement the steps in the above-described method embodiments.

[0210] This invention provides a computer program product that, when run on an electronic device, enables the electronic device to implement the steps described in the various method embodiments above.

[0211] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: a medium capable of carrying the computer program code, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0212] In the above embodiments, the descriptions of each embodiment have their own emphasis. Parts not described in detail or in a particular embodiment can be referred to in the relevant descriptions of other embodiments. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0213] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / device and method can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0214] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0215] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for decoupled and synchronously generating digital clones using a large character model, characterized in that... include: The system obtains an audio feature sequence obtained by extracting features from the input audio data, and a character profile vector obtained by extracting features from the character profile information of the digital clone. The audio feature sequence includes prosodic features and optionally includes content-related features. The character profile vector is used to characterize the expression style of the digital clone. The audio feature sequence and the character portrait vector are input into a pre-trained large character model to simulate character emotions, thereby obtaining the facial expression control information of the digital clone; The facial expression control information includes at least emotion tags and / or facial expression control vector sequences; Decoupled modeling is performed using independent lip shape parameter generation models and micro-expression parameter generation models. Specifically, the audio feature sequence is input into the lip shape parameter generation model to simulate lip movements, generating a lip movement parameter sequence that is temporally aligned with the audio feature sequence. The expression control information and the audio feature sequence are input into the micro-expression parameter generation model, using the expression control information as an emotional condition input. The type, intensity, and temporal distribution of facial micro-expression movements are constrained through feature modulation and / or cross-attention weighting, generating a micro-expression movement parameter sequence that is temporally aligned with the audio feature sequence. Based on preset parameter partitioning rules, the parameters of overlapping regions in the lip movement parameter sequence and the micro-expression parameter sequence are identified. The lip movement parameter sequence after removing the parameters of the overlapping regions is formed into a lip movement parameter set, and the micro-expression parameter sequence after removing the parameters of the overlapping regions is formed into a micro-expression parameter set. A first weight of the lip movement parameters in the overlapping regions is determined based on the prosodic features in the audio feature sequence, and a second weight of the micro-expression parameters at the same time is determined based on the first weight. The lip movement parameters and micro-expression parameters at the same time and position are weighted and transitionally fused to generate an overlapping parameter set that is temporally aligned with the audio feature sequence. The lip movement parameter set, the micro-expression parameter set, and the overlapping parameter set are temporally aligned to generate the full-face expression parameter sequence of the digital clone, wherein the overlapping region is the facial region that simultaneously participates in lip movement and micro-expression movements. The full-face expression parameter sequence is superimposed with preset motion parameters and smoothed to obtain the target expression parameter sequence. The target expression parameter sequence is then mapped to the facial driving parameters of the digital clone and driven to render and output, so as to achieve synchronization of the lip shape and micro-expression of the digital clone.

2. The method according to claim 1, characterized in that, The input to the lip shape parameter generation model includes only the audio feature sequence, while the input to the micro-expression parameter generation model includes the audio feature sequence and the expression control information. The expression control information is injected into the audio feature representation through a conditional modulation unit and / or a cross-attention unit to generate the micro-expression action parameter sequence.

3. The method according to claim 1, characterized in that, The audio features include prosodic features and content-related features. The prosodic features include at least one of fundamental frequency, energy, and speech rate features. The content-related features include at least one of Mel spectrum, Mel frequency cepstral coefficients, and phoneme-related acoustic characterization. The method for generating the lip movement parameter sequence includes: inputting the prosodic features and the content-related features into the lip movement parameter generation model to simulate lip movements, generating the lip movement parameter sequence that is aligned with the temporal, prosodic, and phonemic sequence of the audio feature sequence, wherein the lip movements include at least one of jaw opening and closing, lip opening and closing, lip rounding, and lip stretching.

4. The method according to claim 1, characterized in that, The character profile vector includes at least one of an expression style parameter vector and a character identity embedding vector used to map the character profile information. The expression style parameter vector includes at least one of an expression exaggeration coefficient, an emotion baseline bias, and an expression tremor suppression coefficient.

5. The method according to claim 1, characterized in that, The facial expression control information also includes at least one of semantic tags, intent tags, and prosodic event tags. The prosodic event tags include at least pause events and accent events. The facial expression control information is used to constrain the type, intensity, and temporal distribution of facial expressions during the micro-expression generation process.

6. The method according to claim 1, characterized in that, The weighted transition fusion of lip movement parameters and micro-expression movement parameters at the same time and location includes: Weighted transition fusion is performed on lip movement parameters and micro-expression movement parameters at the same time and location; The first weight of the lip movement parameters at different times is determined based on the prosodic features in the audio feature sequence, and the second weight of the micro-expression movement parameters at the same time is determined based on the first weight. Based on the first weight and the corresponding second weight, the lip movement parameters and micro-expression movement parameters at the same time and position in the overlapping region are weighted and summed to obtain the overlapping parameter set.

7. The method according to claim 6, characterized in that, The first weight is determined based on at least one of the fundamental frequency, energy, and speech rate features in the audio feature sequence and / or based on prosodic events detected by the prosodic features; the first weight is increased when an accent event is detected and decreased when a pause event is detected.

8. The method according to claim 1, characterized in that, The step of superimposing and smoothing the full-face expression parameter sequence with preset action parameters includes: A preset action parameter superposition operation is performed on the full-face expression parameter sequence to generate an optimized full-face expression parameter sequence. The preset action parameters include at least one of blinking action parameters and nodding action parameters. The optimized full-face expression parameter sequence is subjected to parameter-based temporal smoothing to generate the target expression parameter sequence, which is then mapped to the facial driving parameters of the digital clone and used to drive the rendering output.

9. The method according to claim 8, characterized in that, The superposition operation of the preset action parameters includes: when no blinking event is detected within a first duration based on the micro-expression action parameter sequence, and / or when a pause event is detected based on the lip movement parameter sequence, inserting blinking action parameters at the corresponding position in the full-face expression parameter sequence.

10. The method according to claim 8, characterized in that, The superposition operation of the preset action parameters includes: when an accent event is detected based on the audio feature sequence, and / or a single rhythmic event lasting a second duration is detected, inserting a nodding action parameter at the corresponding position in the full-face expression parameter sequence.

11. The method according to claim 8, characterized in that, The step of performing time-series smoothing of the optimized full-face expression parameter sequence by sub-parameters includes: The motion parameters in the optimized full-face expression parameter sequence are classified to obtain multiple parameter sequence groups of different types; Based on the type corresponding to each parameter sequence group, the filtering intensity of each parameter sequence group is determined, and based on the corresponding filtering intensity, a smoothing filtering process is performed on each parameter sequence group to generate the target expression parameter sequence.

12. The method according to any one of claims 1 to 11, characterized in that, The method also includes facial expression error correction and rollback processing, which includes: Real-time detection of the signal status of input audio data; When signal interruption, signal abnormality, or prolonged silence is detected in the input audio data, the expression parameter sequence of the most recent frame or the most recent period is used as the current output to drive the digital clone, and the currently output expression parameter sequence is gradually rolled back to neutral expression parameters within a preset transition time. The neutral expression is the character's baseline expression with no facial micro-expressions and closed lips, and the neutral expression parameters are the full-face expression parameters corresponding to the neutral expression. During the process of gradually reverting the expression parameter sequence to neutral expression parameters, if the signal of the input audio data is detected to return to normal, the current expression parameter sequence corresponding to the recovery time is recorded, and a recovered target expression parameter sequence is generated based on the recovered input audio data. The current expression parameter sequence corresponding to the recovery time and the recovered target expression parameter sequence are then transitionally fused to generate a continuously output expression parameter sequence, and the digital clone is driven according to the continuously output expression parameter sequence to avoid abrupt changes in output.

13. A device for the micro-decoupling and synchronous generation of a large character model-driven digital clone, characterized in that, include: The audio feature extraction module is used to extract features from the input audio data to obtain an audio feature sequence that includes prosodic features and optionally content-related features. The character portrait acquisition module is used to acquire the character portrait vector obtained by extracting features from the character portrait information of the digital clone. The character portrait vector is used to represent the expression style of the digital clone. The character large model inference module is used to input the audio feature sequence and the character portrait vector into a pre-trained character large model to simulate character emotions and obtain the expression control information of the digital clone. The expression control information includes at least emotion tags and / or expression control vector sequences. The lip shape parameter generation module is used to input the audio feature sequence into a pre-trained lip shape parameter generation model to simulate lip movements and generate a lip movement parameter sequence that is temporally aligned with the audio feature sequence. The micro-expression parameter generation module is used to input the expression control information and the audio feature sequence into a pre-trained micro-expression parameter generation model, using the expression control information as the emotional condition input, and constraining the type, intensity and temporal distribution of facial micro-expression movements through feature modulation and / or cross-attention weighting, to generate a micro-expression movement parameter sequence that is temporally aligned with the audio feature sequence. The expression fusion module is used to identify parameters in overlapping regions of the lip movement parameter sequence and the micro-expression action parameter sequence based on preset parameter partitioning rules; to form a lip movement parameter set by removing the parameters in the overlapping regions from the lip movement parameter sequence; and to form a micro-expression parameter set by removing the parameters in the overlapping regions from the micro-expression action parameter sequence. Based on the prosodic features in the audio feature sequence, a first weight of the lip movement parameters in the overlapping regions is determined, and based on the first weight, a second weight of the micro-expression action parameters at the same time is determined. The lip movement parameters and micro-expression action parameters at the same time and position are then weighted and transitionally fused to generate an overlapping parameter set that is temporally aligned with the audio feature sequence. The lip shape parameter set, the micro-expression parameter set, and the overlapping parameter set are time-aligned to generate the full-face expression parameter sequence of the digital clone; The facial expression post-processing module is used to perform preset action parameter superposition and smoothing on the full-face facial expression parameter sequence to obtain the target facial expression parameter sequence; The facial-driven rendering module is used to map the target expression parameter sequence to the facial driving parameters of the digital clone and drive the rendering output to achieve synchronization of the lip shape and micro-expression of the digital clone.

14. The generating apparatus according to claim 13, characterized in that, The expression post-processing module is also used to perform expression error correction and rollback processing, which includes: real-time detection of the signal status of the input audio data; when the input audio data is detected to have signal interruption, signal abnormality, or long-term silence, the expression parameter sequence of the most recent frame or the most recent period is used as the current output to drive the digital clone; within a preset transition time, the currently output expression parameter sequence is gradually rolled back to neutral expression parameters to drive the facial expression of the digital clone to gradually change to a neutral expression, wherein the neutral expression is a character baseline expression with no facial micro-expressions and closed lips, and the neutral expression parameters are the full-face expression parameters corresponding to the neutral expression.

15. The generating apparatus according to claim 14, characterized in that, The expression fault tolerance rollback processing performed by the expression post-processing module further includes: during the process of gradually rolling back the expression parameter sequence to neutral expression parameters, if the signal of the input audio data is detected to have returned to normal, the current expression parameter sequence corresponding to the recovery time is recorded, and the recovered target expression parameter sequence generated based on the recovered input audio data is obtained; the current expression parameter sequence corresponding to the recovery time and the recovered target expression parameter sequence are transitionally fused to generate a continuously output expression parameter sequence, and the digital clone is driven according to the continuously output expression parameter sequence to avoid abrupt output jumps.

16. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the character large model-driven digital clone lip micro-decoupling synchronous generation method as described in any one of claims 1 to 12.

17. A readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the character large model-driven digital clone lip micro-decoupling synchronous generation method as described in any one of claims 1 to 12.

18. A computer program product, characterized in that, The method includes a computer program that, when executed by a processor, implements the character large model-driven digital clone lip micro-decoupling synchronous generation method as described in any one of claims 1 to 12.