Personality-Driven Role Model Multimodal Digital Avatar Interaction Method and System

By using emotion recognition and multimodal synchronization control driven by a large role model, the problem of personality consistency in the multimodal output of digital avatars is solved, improving the naturalness of interaction and user satisfaction.

CN121745256BActive Publication Date: 2026-07-17LIANGSHENG DIGITAL CREATIVE DESIGN (HANGZHOU) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
LIANGSHENG DIGITAL CREATIVE DESIGN (HANGZHOU) CO LTD
Filing Date
2026-02-28
Publication Date
2026-07-17

Smart Images

  • Figure CN121745256B_ABST
    Figure CN121745256B_ABST
Patent Text Reader

Abstract

This invention discloses a personality-driven, large-scale role model-based multimodal digital avatar interaction method and system. The method includes: performing emotion recognition on user input information to determine emotion tags; determining the current personality vector based on the emotion tags and basic personality vectors; processing the role setting fields, dialogue context, and current personality vector using a large-scale role model to determine the target response text; determining a speech style vector based on the current personality vector, injecting the speech style vector into a text-to-speech model, and performing speech synthesis on the target response text to determine the target response speech; determining an animation style vector based on the current personality vector, injecting the animation style vector into an animation generation model, and processing the target response text and target response speech to determine the target facial expression animation; and performing multimodal synchronization and consistency checks on the target response text, target response speech, and target facial expression animation to reduce cross-modal personality conflicts and ensure personality consistency in the output results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and human-computer interaction technology, and relates to character large model, digital avatar generation and multimodal interaction control technology, specifically to a personality-driven character large model multimodal digital avatar interaction method and system. Background Technology

[0002] With the large-scale application of digital avatars in fields such as intelligent assistants, virtual anchors, online education, and rehabilitation companionship, users' demands for interactive immersion are upgrading from "modal realism" to "personality consistency." Existing technologies suffer from the following main pain points:

[0003] Cross-modal style inconsistency: Due to the lack of unified personality control, different modalities often exhibit disjointed styles. For example, in a virtual customer service scenario, while the target's reply text may be rigorous and professional, the tone of voice sounds stiff and the facial expressions lack patience, leading to decreased user satisfaction. In a virtual broadcaster scenario, the voice may sound enthusiastic and energetic, but the body language is too restrained, resulting in an unnatural phenomenon of "voice detachment and stiff movements," which reduces viewer retention.

[0004] The lack of a quantitative mechanism for personality control: Existing solutions mostly control a single modality by pre-setting "style tags," failing to transform abstract personality styles into controllable multi-dimensional parameters, leading to unstable personality performance. For example, the same "composed" digital clone may exhibit significant fluctuations in speech rate across different conversation rounds, making it difficult to maintain a consistent style.

[0005] Cross-modal consistency lacks closed-loop verification: Current multimodal synchronization is limited to the temporal alignment of speech and lip movements, and cannot detect the consistency of speech, text, and facial expressions in terms of emotional style. According to statistics, existing digital avatar multimodal output has a high rate of style conflict. For example, when expressing apology in text, the speech lacks the proper falling intonation, and the facial expression does not show remorse. These conflicts reduce the credibility and immersion of the interaction. Summary of the Invention

[0006] This invention provides a personality-driven, multimodal digital avatar interaction method and system to solve the problem of lack of personality consistency in the multimodal output of digital avatars in the prior art.

[0007] A personality-driven, large-scale role model, multimodal digital avatar interaction method, including:

[0008] Emotion recognition is performed on user input information collected during the interaction between the digital avatar and the user to determine emotion tags;

[0009] Based on the emotion tag, an emotion offset is determined; based on the emotion offset and the base personality vector corresponding to the digital clone, the current personality vector corresponding to the digital clone is determined.

[0010] A large role model is used to process the role setting fields, dialogue context, and the current personality vector to determine the target response text that represents personality style. The large role model is injected with the role setting fields and the current personality vector during the inference period to constrain the target response text to maintain role consistency and personality consistency in multiple rounds of dialogue.

[0011] Based on the current personality vector, a target value for at least one voice control parameter is determined. Based on the target value of at least one voice control parameter, a voice style vector is determined. The voice style vector is injected into a text-to-speech model. The text-to-speech model is used to synthesize the target response text into speech, and a target response speech representing the personality style is determined.

[0012] Based on the current personality vector, a target value for at least one animation control parameter is determined. Based on the target value of at least one animation control parameter, an animation style vector is determined. The animation style vector is injected into an animation generation model. The animation generation model is used to process the target response text and the target response speech to determine the target facial expression animation that represents the personality style.

[0013] Perform multimodal synchronization and consistency checks on the target response text, the target response voice, and the target facial expression animation, and output the target response text, target response voice, and target facial expression animation that pass the checks.

[0014] A personality-driven, large-scale, multimodal digital avatar interaction system, comprising:

[0015] The emotion recognition module is used to identify emotions from user input information collected during the interaction between the digital avatar and the user, and to determine emotion tags.

[0016] The personality center module is used to determine the emotional offset based on the emotional label, and to determine the current personality vector corresponding to the digital clone based on the emotional offset and the basic personality vector corresponding to the digital clone.

[0017] The character big model text generation module is used to process the character setting fields, dialogue context and the current personality vector using the character big model to determine the target response text that represents the personality style. The character big model is injected with the character setting fields and the current personality vector during the inference period to constrain the target response text to maintain character consistency and personality consistency in multiple rounds of dialogue.

[0018] The speech synthesis module is used to determine the target value of at least one speech control parameter based on the current personality vector, determine the speech style vector based on the target value of at least one speech control parameter, inject the speech style vector into the text-to-speech model, use the text-to-speech model to perform speech synthesis on the target response text, and determine the target response speech that represents the personality style.

[0019] The facial expression animation module is used to determine the target value of at least one animation control parameter based on the current personality vector, determine the animation style vector based on the target value of at least one animation control parameter, inject the animation style vector into the animation generation model, and use the animation generation model to process the target response text and the target response speech to determine the target facial expression animation representing the personality style.

[0020] The multimodal synchronization module is used to perform multimodal synchronization and consistency verification on the target reply text, the target reply voice, and the target facial expression animation, and output the target reply text, target reply voice, and target facial expression animation that pass the verification.

[0021] The aforementioned personality-driven, large-scale model-based multimodal digital avatar interaction method and system avoids personality fragmentation caused by separate generation of multiple modalities by integrating personality parameters such as basic personality vectors, emotional offsets, and current personality vectors throughout the entire process of text, voice, facial expressions, and actions. It employs a large-scale model to process character setting fields, dialogue context, and the current personality vector, ensuring consistency in character and multimodal expression style across multiple rounds of interaction. Furthermore, it achieves personality evolution, traceability, and controllability through long-term personality optimization and version rollback mechanisms. By performing multimodal synchronization and consistency verification on the target response text, target response voice, and target facial expression animation, it enables threshold-based triggering of fine-tuning or regeneration, forming an executable "generation-verification-adjustment" closed-loop control. Attached Figure Description

[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort:

[0023] Figure 1 This is a flowchart of a personality-driven, multimodal digital avatar interaction method for a large role model in this invention.

[0024] Figure 2 yes Figure 1 A flowchart following step S106;

[0025] Figure 3 yes Figure 1 A flowchart of step S103;

[0026] Figure 4 yes Figure 1 A flowchart of step S106;

[0027] Figure 5 yes Figure 4 A flowchart of step S403;

[0028] Figure 6 This is a schematic diagram of a personality-driven, multimodal digital avatar interaction system based on a large role model, as described in an embodiment of the present invention. Detailed Implementation

[0029] To make the technical problems solved, the technical solutions, and the beneficial effects of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0030] This invention provides a personality-driven, multimodal digital avatar interaction method based on a large-scale character model. The method is illustrated using mobile phones, computers, or other electronic devices as examples. Figure 1 As shown, the method includes:

[0031] S101: Perform emotion recognition on user input information collected during the interaction between the digital clone and the user, and determine the emotion label;

[0032] S102: Determine the emotion offset based on the emotion tag, and determine the current personality vector corresponding to the digital clone based on the emotion offset and the basic personality vector corresponding to the digital clone;

[0033] S103: The role model is used to process the role setting fields, dialogue context and the current personality vector to determine the target response text that represents the personality style. The role model is injected with the role setting fields and the current personality vector during the inference period to constrain the target response text to maintain role consistency and personality consistency in multiple rounds of dialogue.

[0034] S104: Determine the target value of at least one voice control parameter based on the current personality vector, determine the voice style vector based on the target value of at least one voice control parameter, inject the voice style vector into the text-to-speech model, use the text-to-speech model to perform speech synthesis on the target response text, and determine the target response speech that represents the personality style.

[0035] S105: Based on the current personality vector, determine the target value of at least one animation control parameter, determine the animation style vector based on the target value of at least one animation control parameter, inject the animation style vector into the animation generation model, and use the animation generation model to process the target response text and the target response speech to determine the target facial expression animation representing the personality style;

[0036] S106: Perform multimodal synchronization and consistency verification on the target reply text, the target reply voice, and the target facial expression animation, and output the verified target reply text, target reply voice, and target facial expression animation.

[0037] Digital clones refer to interactive entities presented as digital clones, capable of outputting multimodal content such as text, voice, and facial / motion animations.

[0038] Personality vectors are used to quantify the personality style and personality state of a digital avatar, denoted as The basic personality vector is a pre-constructed personality vector of the digital avatar, denoted as... It consists of parameter values ​​corresponding to n normalized personality dimensions. The current personality vector is the personality vector of the digital clone in the current scenario, denoted as... . and The value range for each personality dimension is [missing information]. .

[0039] Emotional labels are used to quantify a certain situation. Emotional labels include m emotion types and the corresponding emotional intensity for each emotion type. Emotion types must include at least two or more of the following: happiness, anger, sadness, calmness, surprise, and disgust; the emotional intensity ranges from [value missing]. The emotional intensity vector is composed of the emotional intensities corresponding to each emotional type. , where m is the number of emotion types.

[0040] The role setting field contains structured information that constrains role consistency, denoted as... The character setting fields should include at least the character's identity and background settings, speaking style, disabled or required expressions, interaction boundaries and behavior rules.

[0041] Large-Scale Character Model (LLM) refers to adding character-specific fields to a pre-trained large-scale language model. With the current personality vector Coupled and injected during inference (or solidified through instruction fine-tuning), the model maintains role and personality consistency across multiple rounds of dialogue, outputting target response text that aligns with the personality's style. .

[0042] As an example, in step S101, when the user interacts with the digital avatar via voice and / or text, the electronic device can perform emotion recognition on the user input information collected during the interaction process. For example, it can use natural language modeling for emotion classification or voice intonation analysis to determine the emotion label corresponding to the user input information. In this example, the emotion label includes m emotion types and the emotion intensity corresponding to each emotion type, with the emotion intensity ranging from [value missing]. .

[0043] As an example, in step S102, after determining the emotion tag corresponding to the user's input information, the electronic device needs to determine its corresponding emotion offset based on the emotion tag. This emotion offset can be determined by processing the emotion tag using a pre-set mapping function, or by other technical means. Next, the emotion offset is... Superimposed basic personality vector This yields the new current personality vector. And immediately update the current personality vector in the memory cache. For subsequent calls; also record the update log for this round (timestamp, sentiment tag, ... and In this example, if the current dialogue turn t reaches the preset dialogue turn number T of the optimization period (e.g., a cumulative 100 interactions), then the average offset of the sentiment shift within that optimization period is calculated. , and according to Fine-tuning the basic personality vector The version number is incremented, and then the cumulative statistics are cleared before entering the next cycle.

[0044] As an example, in step S103, the electronic device uses a pre-trained role big model to process the role setting field, dialogue context and current personality vector to determine the target response text that represents personality style. The role big model injects the role setting field and current personality vector during the inference period to constrain the role consistency and personality consistency of the target response text in multi-turn dialogues.

[0045] As an example, in step S104, the electronic device can determine target values ​​corresponding to multiple voice control parameters based on the current personality vector. These voice control parameters can be, but are not limited to, speech rate and fundamental frequency, and can also include timbre or other parameters that affect the characteristics of the speech signal. The target values ​​corresponding to the multiple voice control parameters are fused to form a voice style vector, which is then injected into a text-to-speech model used to achieve text-to-speech conversion. The text-to-speech model is used to synthesize the target response text. The speech rate, fundamental frequency, and timbre parameters in the speech synthesis process are adjusted using the voice style vector to make the generated target response speech reflect the personality style indicated by the current personality vector. Specifically, this is achieved by adjusting the speech rate, pitch, and emotional tone of the output speech to make the speech sound like the preset personality style of the digital avatar (such as enthusiastic or calm). Specifically, the electronic device can determine the target value of at least one voice control parameter based on the current personality vector, and combine the target values ​​of at least one voice control parameter to form a voice style vector. The generated speech style vector is passed as additional input along with the target response text to the text-to-speech model (e.g., a TTS model). In implementation, a neural network TTS model supporting style control is preferred (e.g., Tacotron series models or VITS models with global style embedding). The input interface for the speech style vector can be reserved in the TTS model architecture, or the model can be trained using multi-style speech data to learn to adjust its speech synthesis performance based on the input speech style vector. When calling the speech synthesis API, the generated speech style vector is passed to the model. The text-to-speech model is injected (e.g., by calling the interface synthesize_speech(text,style=S)) to generate the target response speech that matches the personality's tone. At the same time, the text-to-speech model outputs the target response speech, such as the phoneme sequence and corresponding duration, during the inference process for subsequent multimodal synchronous processing.

[0046] As an example, in step S105, the electronic device can determine target values ​​corresponding to multiple animation control parameters based on the current personality vector. These animation control parameters can be, but are not limited to, parameters such as facial expression intensity, movement amplitude, movement frequency, movement rhythm, and facial expression switching smoothness. The target values ​​of at least one animation control parameter are fused to form an animation style vector, which is then injected into the animation generation model used to generate the animation. The animation generation model is used to generate animations for the target reply text and target reply speech. The animation style vector is used to adjust the intensity of the facial expression (such as the amplitude of a smile), the amplitude and frequency of the movement (such as the speed of nodding and the frequency of gestures) in the target facial expression animation, so that it is synchronized with the target reply text and target reply speech, reflecting that the facial expressions and movements in the target facial expression animation have the preset personality style of the digital clone (such as frequent and large movements for extroverted personalities, and restrained and small movements for introverted personalities).

[0047] In this example, the electronic device drives the lip movements in the target facial expression animation of the digital avatar based on the phoneme timestamps output by the text-to-speech model. This involves playing pre-defined lip shapes (such as opening and closing mouth shapes) according to a timeline, ensuring that lip movements are synchronized with each word of speech. Next, based on the target response text, sentiment analysis is performed, or the aforementioned sentiment tags are directly used to select an expression / action script that matches the semantics (e.g., selecting a smiling expression or nodding action when responding to happy content). Then, the current personality vector is... Substitute into the predefined mapping function The system calculates the target value of at least one animation control parameter corresponding to the facial expression and animation, and scales the selected facial expression and action. For example, in software implementation, several parameters (such as smile_curve and gesture_amp) can be defined in the animation engine. The values ​​of these parameters are determined by the current personality vector, thereby amplifying or reducing the exaggeration of the preset action. If the current personality vector shows that the character is extroverted and enthusiastic, the smile curve value and gesture frequency are increased accordingly; if the personality is introverted and conservative, these parameters are decreased to make the animation more subtle. The animation generation model can be implemented based on a game engine (Unity3D, Unreal, etc.) or a graphics library. It generates a sequence of animation frames or animation control flow synchronized with the voice by calling the engine's animation control interface (such as setting bone joint rotation and BlendShape weights) through scripts, which can be displayed in real time on the user interface by the rendering engine.

[0048] As an example, in step S106, the electronic device needs to perform multimodal synchronization and consistency verification on the target response text, target response voice, and target facial expression animation, and obtain a multimodal output that passes the consistency verification. Specifically, this includes: First, multimodal synchronization processing is required for the target response voice and target facial expression animation, specifically aligning the phoneme timestamps corresponding to the target response voice with the timeline of the target facial expression animation to ensure strict synchronization between lip movements and voice, and matching the rhythm of movements with the pauses in voice, thereby maintaining temporal consistency in both visual and auditory senses. Next, the emotions of the target response text, target response voice, and target facial expression animation are analyzed to determine whether the emotions of the three are consistent or whether there is any deviation from the personality style corresponding to the current personality vector, in order to determine the consistency analysis result. If the consistency analysis result is a failure, the voice control parameters or animation control parameters need to be adjusted, or at least one of the target response text, target response voice, and target facial expression animation needs to be regenerated for multimodal synchronization and consistency verification to be performed again. If the consistency analysis result is a pass, the verified target response text, target response voice, and target facial expression animation are output.

[0049] To ensure seamless integration of speech, lip movements, and actions in the digital avatar output, enhancing the realism for the user, a sophisticated time alignment strategy was implemented. Specifically, when the text-to-speech model generates the target response speech, it simultaneously outputs corresponding phoneme timestamps, indicating the start and end times of each phoneme. The animation generation model aligns the phoneme timestamps with the timeline of the target facial expression animation, ensuring that the digital avatar's lip movements are precisely aligned with the speech in a precise time sequence: when the speech enters the pronunciation period of a certain phoneme, the digital avatar's mouth is driven to produce the corresponding opening and closing shape; when there are pauses or gaps in the speech, the lip movements are paused to match silence. Taking the pronunciation of "Hello" as an example, if the synthesized speech has phoneme timestamps of / ha / 0.00s–0.20s and / lou / 0.20s–0.50s, then during the period of 0.00s–0.20s, the digital clone's mouth shape gradually expands to form the "ha" sound, and at 0.20s it switches to the "lou" sound until it closes completely at 0.50s. This frame-by-frame alignment ensures that the lip movements seen by the user are completely synchronized with the pronunciation heard, avoiding common glitches such as "the sound stops but the mouth is still moving." Not only lip movements, but other facial expressions and body movements accompanying the speech also require timing coordination. In this example, the electronic device can adjust the timing of facial expressions and actions based on the rhythm and audio energy changes of the target's response speech: when it detects that the speech is at the end of a sentence or semantic segment (often accompanied by a drop in fundamental frequency and a brief silence), it can control the digital avatar to perform a corresponding facial expression ending action (such as a nodding action completing a cycle at the end of a sentence), or insert appropriate body gestures during long sentences to coordinate with semantic emphasis points. In summary, by fully utilizing the temporal characteristics of speech to guide the animation process, this invention achieves audiovisual synchronization, making the dynamic image of the digital avatar match the rhythm of the speech expression, thus enhancing the naturalness of the interaction.

[0050] In one embodiment, the emotion label includes m emotion types and an emotion intensity corresponding to each emotion type, wherein the value range of the emotion intensity is: ;

[0051] The emotional shift is determined according to the following formula: ;in, For emotional offset; This is a weight matrix for the emotion-personality dimension. n represents the number of personality dimensions; m represents the number of emotion types. Let the emotion intensity vector be... , Determined based on the intensity of the emotion corresponding to multiple emotion types; This is the emotional attenuation coefficient. The range of values ​​is ; Adapt weights to different scenarios. The range of values ​​is ;

[0052] The current personality vector is determined according to the following formula: ; Based on the basic personality vector, Let i be the current personality vector, and let i be the index of the personality dimension. min and max form the interval truncation to ensure that the current personality vector is within the specified range. Within the range.

[0053] As an example, the emotion labels determined by emotion recognition based on user input include m emotion types and the corresponding emotion intensity for each emotion type. The range of emotion intensity values ​​is... The emotion types here include two or more of the following: happiness, anger, sadness, calmness, surprise, and disgust. The emotion intensity vector is constructed by combining the emotion intensities corresponding to the m emotion types. .

[0054] The emotion-personality dimension weight matrix is ​​constructed by calling the pre-built n personality dimensions and m emotion types. , m represents the supported emotion type, and the emotion-personality dimension weight matrix. Each emotional personality weight vector in This matrix represents the degree of influence of the j-th emotion type on the i-th personality dimension. It is obtained through training on a large amount of labeled data (typical value range). A positive value indicates that the emotion will enhance a certain personality dimension, while a negative value indicates that it will weaken it. The personality dimensions here include five or more of the following: extraversion, agreeableness, conscientiousness, neuroticism, and openness.

[0055] Next, according to This formula determines the emotional shift. .in, This is the emotional attenuation coefficient, with a value range of... This is used to control the impact of emotions as they gradually weaken with each round of dialogue, preventing long-term emotional lingering. Let the current round of dialogue be t, and the emotional offset from the previous round be... Then the offset in this round should satisfy That is, when there is no new strong emotional input, the emotional offset gradually decreases to 0; Weights are assigned to adapt to different scenarios, and their value ranges are specified. It is used to adjust the intensity of emotional impact according to the application scenario, such as taking into account the emotional impact in a serious work scenario. Suppressing emotional fluctuations, leisure and entertainment scenarios Amplify the emotional effect. and Simultaneously acting on This yields the sentiment shift that takes into account time decay and contextual factor correction. (A vector of n dimensions), mood offset After calculation, the data is stored in a memory cache for use when updating the current personality vector. Once the update is complete, the data can be cleared or replaced with the next value.

[0056] In this example, when determining the emotional offset Then, it is superimposed on the basic personality vector. Above, calculate the current personality vector. The calculation formula is as follows: That is, for each personality dimension, the basic personality vector Corresponding emotional offset Add them together to determine the current personality vector. And limit the result to the current personality vector exist Within the specified range, values ​​outside the range are truncated to boundary values ​​of 0 or 1 to prevent emotions from excessively perturbing the personality. These constraints ensure that the current personality vector does not overflow its normal value range, thus preventing numerical anomalies caused by excessive emotional influence (e.g., an excessively high parameter in one dimension leading to an extremely unnatural speech rate). Current Personality Vector The system reflects the current personality state of the digital clone in real time, and serves as the input for each generation module, enabling the output to instantly reflect emotional changes.

[0057] For example, suppose the digital clone's underlying personality vector In five dimensions (extroversion) Affinity Conscientiousness neurotic Openness The values ​​for all values ​​are 0.5. The current user input, after emotion recognition, yields an "anger" emotion intensity of 0.9, which is the emotion intensity vector. The scenario is a customer service environment. Let's define the emotional personality weight vector corresponding to "anger". Emotional attenuation coefficient Scene adaptation weight Then the emotional offset according to Calculations can yield the following results. =[0.0, -0.2916, 0.1458, 0.2187, -0.0729]. The calculated current personality vector. The values ​​are [0.5, 0.2084, 0.6458, 0.7187, 0.4271]. In subsequent text, speech, and animation generation processes, this current personality vector is used as a basis. Adjustments were made to enable the digital avatar to output more serious, concise, and problem-solving oriented text and voice in angry situations, along with lower smile intensity and less affectionate facial animations, thereby reducing cross-modal style conflict and maintaining personality control.

[0058] In one embodiment, such as Figure 2 As shown, after step S106, the personality-driven multimodal digital avatar interaction method for large role models further includes:

[0059] S201: Adopted Update the current dialogue round number, where t is the current dialogue round number;

[0060] S202: When the current number of dialogue turns is greater than the preset number of dialogue turns corresponding to the optimization period, determine the average offset based on multiple emotion offsets corresponding to the optimization period;

[0061] S203: Based on the last emotional offset and the average offset in the optimization cycle, update the basic personality vector. The update formula for the basic personality vector is as follows: ,in, Let be the basic personality vector at time k+1. Let k be the basic personality vector at time k. For personality optimization coefficients, This represents the average offset.

[0062] As an example, a pre-defined list of emotion tags (type: list of structures) is used to store the emotion tags corresponding to the most recently detected user input. Each emotion tag includes an emotion type (an enumerated value, such as 6 categories like happy / angry / sad), emotion intensity (a floating-point number between 0 and 1), and a timestamp. The system retains only the 5 most recent emotion records; when a new record arrives, the oldest one is discarded. This list is cached in memory for later retrieval.

[0063] This example includes a long-term personality optimization mechanism, which supports periodic optimization and adjustment of the basic personality vector based on long-term interaction data to adapt to gradual changes in user preferences. The optimization cycle is set to T rounds of dialogue (e.g., T=100). At the end of each optimization cycle, the average offset of the accumulated emotional shift within that cycle is calculated. Then, it is superimposed on the basic personality vector according to a certain ratio, that is... ,in, Let represent the basic personality vector at time k, which is the last basic personality vector within the optimization cycle. The personality optimization coefficient (learning rate, typical value) ), , Let t be the sentiment shift at time t within the optimization period. This is the average offset. By selecting a smaller value... Ensure that the basic personality vector undergoes only subtle changes, without disrupting the original personality settings, while gradually converging towards the preferences reflected in actual user interactions. For example, if the average shift of the "extraversion" dimension over a long period of interaction... A value of +0.1 indicates that the user prefers an extroverted and lively demeanor. The extroversion dimension of the basic personality vector will gradually increase, making the digital avatar more extroverted in subsequent interactions. This long-term personality optimization mechanism allows the digital avatar to "learn" user preferences and continuously improve its personality settings, thus maintaining freshness and compatibility in long-term companionship. It should be noted that all basic personality optimization operations are recorded with version numbers. Users can view the personality evolution log, and if they find that the personality has deviated from its initial intention, they can roll back the personality configuration to a previous version, ensuring that personality evolution remains within the user's control.

[0064] In this embodiment, the basic personality vector is dynamically optimized and adjusted based on a preset optimization cycle. When the cumulative number of current dialogue rounds reaches the preset number of dialogue rounds corresponding to the optimization cycle, the average offset of the emotional offset within the optimization cycle is calculated and superimposed according to the personality optimization coefficient, thereby updating the basic personality vector so that the personality setting of the digital clone can be adaptively adjusted according to the user's long-term interaction preferences.

[0065] In one embodiment, the large character model includes a natural language model, a vocabulary layer, a sentence structure layer, and a catchphrase layer;

[0066] like Figure 3 As shown, step S103, namely, using the large character model to process the character setting fields, dialogue context, and the current personality vector to determine the target response text representing the personality style, includes:

[0067] S301: The natural language model is used to process the role setting fields and the dialogue context to determine the initial response text;

[0068] S302: Adjust the initial response text at the vocabulary level using the emotional vocabulary corresponding to the current personality vector; adjust the initial response text at the sentence structure level using the sentence style corresponding to the current personality vector; adjust the initial response text at the catchphrase level using the catchphrase corresponding to the current personality vector; fuse the outputs of the vocabulary level, the sentence structure level, and the catchphrase level to determine the target response text.

[0069] To ensure that the personality style of the digital avatar is fully reflected in the text, voice, and facial expression output modalities, a mapping algorithm from the parameter values ​​of the personality dimension to the output features was designed for each modal, enabling the precise control of personality style over language wording, tone of voice, and facial expressions.

[0070] As an example, in processing character setting fields, dialogue context, and the current personality vector using a large character model, a natural language model is first used to generate text for the character setting fields and dialogue context to determine the initial response text. This initial response text is then transmitted to the lexical layer, sentence structure layer, and catchphrase layer. At the lexical layer, the pre-injected current personality vector and the initial response text are processed to convert some words in the initial response text into emotional words that match the current personality vector. At the sentence structure layer, the pre-injected current personality vector and the initial response text are processed to convert the sentence structure in the initial response text into a sentence style that matches the current personality vector. At the catchphrase layer, the pre-injected current personality vector and the initial response text are processed to add catchphrases that match the current personality vector to the initial response text. Finally, the outputs from the lexical layer, sentence structure layer, and catchphrase layer are fused to output the target response text. In this example, after the natural language model outputs the initial response text, the vocabulary, sentence structure, and catchphrases of the initial response text need to be adjusted to ensure that the wording and tone of the target response text are consistent with the personality style corresponding to the digital avatar.

[0071] The lexical layer invokes a pre-stored personality-lexical mapping, which establishes a corresponding emotional lexical for each personality dimension. The emotional lexical is categorized by the intensity of emotion within that dimension. For example, a high "extraversion" dimension (value > 0.8) corresponds to a lexical containing many enthusiastic and high-energy phrases (e.g., "Great!", "So happy!"), a moderate "extraversion" dimension uses milder words (e.g., "Very good," "Quite happy"), and a low "extraversion" dimension favors brief and restrained expressions. Similarly, a high "conscientiousness" dimension can incorporate rigorous wording (e.g., "Please pay attention," "Make sure it's accurate"), while a high "affability" dimension uses more intimate interjections (e.g., "Oh~," "Nah!"). When generating text, the system selects words from the corresponding emotional lexical indices in a certain proportion based on the strength of each dimension in the current personality vector to replace or refine the words in the initial response text. For example, when extraversion = 0.9 and affinity = 0.8, more than 60% of the emotional words in the target response text will come from the extraversion and affinity lexicon, such as replacing "very good" with "fantastic" and adding interjections like "ya" to make the text full of enthusiasm and affinity.

[0072] The sentence structure layer invokes a pre-stored personality-sentence structure rule mapping, which pre-sets sentence style rules corresponding to each personality dimension. For example, highly extroverted personalities prefer to use emotionally strong sentences such as exclamations and rhetorical questions, with shorter sentence lengths (<15 characters) to maintain a brisk pace; highly conscientious personalities prefer declarative sentences and well-organized compound sentences, often using long sentences or bullet points to ensure information completeness (sentence length ≥20 characters). Furthermore, highly open personalities may allow unconventional and novel expressions, while highly conservative personalities follow more formal and standard syntax. A syntactic analyzer is integrated into the sentence structure layer to scan the initial response text generated by LLM, activating corresponding sentence structure rules based on the current personality vector, and adjusting and reconstructing sentences that do not conform to the personality style. For example, if the digital avatar's personality is extroverted and open, and the initial response text is too long and bland, it will be broken down into multiple short sentences and rhetorical questions or exclamations will be introduced to enhance the tone and tension. Conversely, for a meticulous and responsible personality, short sentences will be combined and numbered or linked words will be added to make the expression more logical and complete. Through these adjustments at the sentence level, the structure and rhythm of the final output response text are made to better match the digital avatar's personality profile.

[0073] The catchphrase layer calls a pre-stored personality-catchphrase mapping to accommodate different personality styles, each with their own commonly used modal particles or catchphrases. For example, the friendly and approachable personality often adds "ne," "ya," or "yo" to the end of sentences to express intimacy and a gentler tone; the lively and extroverted personality likes to use "haha" or "hey" to express laughter or greetings; and the open and curious personality may frequently say "wa," "I really didn't expect," or "interesting." In this example, default catchphrase lists are configured for different digital avatars across different personality dimensions, and frequency thresholds are set based on the strength of different personality dimensions in the current personality vector. For example, when affinity ≥ 0.8, each sentence contains at least one modal particle or ending particle on average; when the humor parameter is high, each dialogue contains laughter such as "haha." After determining the initial response text, the large character model needs to select whether to insert these catchphrases based on the current personality vector, making the output target response text more characteristic of the character. At the same time, the frequency of use of catchphrases is positively correlated with the corresponding personality dimension. The higher the parameter value of the personality dimension, the more frequent and obvious the catchphrases are, ensuring the consistency between language style and personality style.

[0074] Electronic devices also need to integrate the outputs from the lexical, sentence structure, and verbal expression layers to determine the target response text. This ensures that the final output of the target response text closely matches the preset personality in terms of word choice, sentence structure, and tone. For example, an "extroverted + friendly" digital avatar might answer a user's question with phrases like, "I'm so happy! This question is super interesting, let me tell you the answer~," conveying a warm and friendly personality to the user. Conversely, an "introverted + highly conscientious" avatar might respond with, "Okay, I've carefully checked the relevant information. Here are the detailed steps," using concise and polite language to fully demonstrate its composure and seriousness, effectively ensuring the personality consistency of the digital avatar's dialogue content.

[0075] In this embodiment, dialogue text generation is based on a pre-trained large-scale natural language model (LLM). In implementation, the current personality vector is injected during the language model decoding process: for example, the current personality vector or its corresponding personality label is fed as additional input into the LLM context, or a finely tuned, style-controlled large-scale personality model is selected. After the large-scale personality model generates the initial response text, the module's built-in personality post-processing rules further refine the text, ensuring that the word choice and tone match the preset personality (e.g., checking for the inclusion of catchphrases consistent with the personality), ultimately outputting the target response text with the personality style.

[0076] In one embodiment, the voice control parameters include at least one of speech rate and fundamental frequency;

[0077] Step S104, namely determining the target value of at least one voice control parameter based on the current personality vector, includes:

[0078] When the parameter value of the i-th personality dimension in the current personality vector is within a preset range, the target values ​​of speech rate and fundamental frequency are determined based on a linear mapping function, wherein the linear mapping function includes: , ;

[0079] When the parameter value of the i-th personality dimension in the current personality vector is not within a preset range, the target values ​​of speech rate and fundamental frequency are determined based on a nonlinear mapping function, wherein the nonlinear mapping function includes: , ;

[0080] in, The target value for speech rate, This serves as the baseline value for speech rate. Let be the linear coefficient of the i-th personality dimension with respect to speech rate. Let be the non-linear coefficient of the i-th personality dimension with respect to speech rate. The target value for the baseband. This is the reference value for the fundamental frequency. Let be the linear coefficient of the i-th personality dimension with respect to the fundamental frequency. Let be the nonlinear coefficient of the i-th personality dimension with respect to the fundamental frequency. represents the parameter value of the i-th personality dimension in the current personality vector.

[0081] Speech is an important medium for conveying emotions and personality style. This solution incorporates personality dimension parameter control into the traditional text-to-speech process. It uses a mapping function combining linear and nonlinear methods to adjust the speech control parameters, including speech rate and fundamental frequency, so that the output speech matches the speaking style of the target personality.

[0082] Linear mapping between speech rate and fundamental frequency: Parameter values ​​for personality dimensions are within a preset range (e.g., parameter values ​​for personality dimensions are within...). In the normal case (between), a linear mapping function is used for control. Assume the baseline value for speech rate is... Words per minute, the base frequency reference value is Hz (corresponding to a neutral voice setting of around 0.5 for a personality dimension). Definition Let i be the linear coefficient of personality dimension i with respect to speech rate. Let be the linear coefficients of personality dimension i with respect to fundamental frequency. These coefficients can be obtained by training on a multi-style speech dataset or preset by configuration parameters. Then, the target value of the speech rate of the target response speech output satisfies the target value of the fundamental frequency: , .

[0083] For example, training might yield the linear coefficient of extraversion on speech rate. (For every 1.0 increase in extraversion, speech rate increases by 20 words per minute), affinity And neurotic (People with high neuroticism speak slightly slower). Another example is the linear coefficient of extraversion on the fundamental frequency. Hz (more extroverted, higher-pitched voice), neurotic Hz (a high neurotic person's voice is lower in pitch), etc. Then, when the current personality vector is extraversion 0.8, affinity 0.6, and neuroticism 0.2, words per minute Hz. It can be seen that the digital clone's speech rate and fundamental frequency are slightly faster and higher than the baseline value, and it is more enthusiastic, which is consistent with the expectation of an extroverted and friendly personality.

[0084] Nonlinear mapping of extreme personality values: This refers to situations where the parameter values ​​of a personality dimension are outside the preset range, for example, at extremely low values. ) or extremely high value ( When using linear mapping alone, speech feature changes may become too drastic or distorted. In this case, a nonlinear mapping with a quadratic term is introduced to constrain the trends in speech rate and fundamental frequency. The adjusted formula is: , ,in , These are nonlinear coefficients, and their typical range of values ​​is... Used in Adjust the sum appropriately when it approaches 0 or 1. For example, for extroversion, when... At very high, ,like This is equivalent to reducing the linear part. This slows down the increase in speaking speed as extroversion increases, preventing the speaking speed from becoming unacceptably fast when extroversion is extremely high (e.g., limiting it to less than 200 words per minute). Similarly, for introversion (e.g. ), quadratic term 0.01 coordination This prevents the speech rate from decreasing excessively and becoming unbearably slow. By combining linear and nonlinear methods, the above mapping function can both ensure that the influence direction of the personality dimension parameter values ​​remains unchanged and smooth out extreme values, thus making the speech style adjustment both sensitive and robust.

[0085] It should be noted that the present invention adopts a segmentation strategy of "linear mapping within a preset range and nonlinear mapping outside the preset range" in the speech modality. The main reason is that the target values ​​of speech control parameters such as speech rate and fundamental frequency are highly sensitive to intelligibility and naturalness. If linear extrapolation is still performed when the parameter values ​​of personality dimension are at extreme values, it is easy to cause excessively fast speech rate, abnormal fundamental frequency or timbre distortion, and may cause the style control vector of the text-to-speech model to deviate from the training distribution, resulting in unstable phenomena such as prosodic jitter and breakage. Therefore, quadratic terms or saturation terms are introduced in the extreme range to softly limit the parameter changes, so as to improve the stability and usability of speech synthesis. Correspondingly, facial expressions and action modalities are typically generated based on preset animation templates and engine constraints (including the range of BlendShape values, limits on skeletal movement amplitude, and upper and lower limits of animation parameters). In this embodiment, linear mapping combined with parameter boundary constraints can ensure the stability and naturalness of the action performance. In other embodiments, facial expression and action mapping can also use piecewise linear, quadratic nonlinear, or S-shaped saturation functions to further suppress exaggerated actions or rhythmic abrupt changes caused by extreme personality parameters. All of the above alternative methods fall within the protection scope of this invention.

[0086] In one embodiment, the voice control parameters include timbre;

[0087] Step S104, which is to determine the target value of at least one voice control parameter based on the current personality vector, further includes:

[0088] The target value of timbre is determined based on a linear mapping function, which includes: ;

[0089] in, The target value for timbre, As the baseline value for timbre, Let be the linear coefficient of the i-th personality dimension with respect to timbre. This is the minimum value for timbre. This represents the maximum value of the timbre. Let be the parameter value of the i-th personality dimension in the current personality vector. This is a function that cuts off the interval.

[0090] As an example, in addition to speech rate and fundamental frequency, this invention also introduces other speech style dimensions such as timbre, allowing digital avatars with different personalities to differentiate themselves in terms of voice and tone. Assuming the baseline value for timbre is neutral... ,definition The linear coefficient of personality dimension i on the brightness and darkness of timbre can be obtained through... Determine the target value for timbre. In this example, the determination of the target value for timbre is not predicated on the determination of the target values ​​for speech rate and fundamental frequency. The speech style vector is composed of the target values ​​for speech rate, fundamental frequency, and / or timbre, and at least one of these is selected as needed and injected into the text-to-speech model.

[0091] For example, people with high extroversion have brighter and clearer voices, which can make... (Increasing the timbre value makes the sound clearer); the voice of a highly introverted person is softer and more ethereal, which can make... (Lowering the timbre value makes the voice softer and calmer.) Similarly, the tone intensity parameter can be set to control the intonation and rhythm of speech. Adjusting these additional voice control parameters ensures that the digital avatar not only matches the personality style in terms of speech speed and tone, but also in terms of timbre and overall atmosphere, thus presenting a specific personality's speaking style. For example, an extroverted and lively person has a bright timbre and a wide intonation, while a serious and steady person has a lower voice and a calm tone. In this example, by adjusting the timbre parameter to control the brightness or softness of the voice, the output voice is made more closely match the corresponding personality style.

[0092] In this example, the target values ​​of speech control parameters such as speech rate, fundamental frequency, and timbre are combined to form a speech style vector. The input is fed into a text-to-speech model (e.g., a TTS model), driving the model to synthesize a target response text with a personality style. Overall, by quantitatively adjusting multiple elements of speech, this invention achieves a mapping from personality parameters to speech features, so that the words spoken by the digital avatar not only conform to the personality in content, but also sound like the person speaking that personality.

[0093] In one embodiment, the animation control parameters include at least one of facial expression intensity, movement amplitude, movement frequency, movement rhythm, and facial expression switching smoothness;

[0094] The target value of at least one animation control parameter is determined by the following formula:

[0095] ;

[0096] ;

[0097] ;

[0098] ;

[0099] ;

[0100] in, The target value for facial expression intensity. This serves as a baseline value for facial expression intensity. is the linear coefficient of the i-th personality dimension with respect to facial expression intensity; The target value for the range of motion, This serves as a baseline value for the range of motion. Let be the linear coefficient of the i-th personality dimension with respect to the amplitude of movement. This represents the minimum range of motion. This represents the maximum range of motion. The target value for the action frequency, This serves as a baseline value for the action frequency. Let be the linear coefficient of the i-th personality dimension with respect to action frequency. This represents the minimum frequency of the action. This represents the maximum value of the action frequency; The target value for the rhythm of the movement. This serves as the baseline value for the rhythm of the movement. Let be the linear coefficient of the i-th personality dimension with respect to the rhythm of action. This represents the minimum value of the action rhythm. This represents the maximum value of the action rhythm; The target value for the smoothness of facial expression transitions. This serves as the baseline value for the smoothness of facial expression transitions. is the linear coefficient of the i-th personality dimension on the smoothness of expression switching; Let be the parameter value of the i-th personality dimension in the current personality vector. This is a function that cuts off the interval.

[0101] In this example, to avoid symbolic conflict, the emotion intensity vector is uniformly denoted as... The intensity of facial expressions is uniformly recorded as These are different concepts.

[0102] As an example, in terms of visual modality, in order to reflect personality differences, the system pre-sets baseline values ​​for animation control parameters such as facial expression intensity, movement amplitude, movement frequency, movement rhythm, and facial expression transition smoothness. Then, through a mapping function, the baseline values ​​of the above animation control parameters are dynamically mapped and adjusted according to the current personality vector to determine the target values ​​of animation control parameters such as facial expression intensity, movement amplitude, movement frequency, movement rhythm, and facial expression transition smoothness. Based on the target values ​​of all animation control parameters, the animation style vector is determined to enrich the non-verbal communication features of the digital avatar.

[0103] Facial expression intensity is used to characterize the prominence of facial muscle movements. A baseline value for facial expression intensity is assumed. (Medium smile intensity), the linear coefficient of personality dimension on facial expression intensity (Corresponding to extraversion, affinity, conscientiousness, neuroticism, and openness, respectively). Current personality vector. The target value for facial expression intensity is calculated as follows:

[0104]

[0105] The final target value for facial expression intensity is The digital clone displays the widest possible smile, which aligns with a highly extroverted and friendly personality.

[0106] Range of motion refers to the extent of limb movement. Let's define a baseline value for range of motion. (Medium range of motion), minimum range of motion maximum value The linear coefficient of personality dimensions on the amplitude of movement Current personality vector The target value of the motion amplitude is calculated as follows:

[0107]

[0108] The target value of the final movement range is Approaching the maximum threshold The digital clone's body movements are quite large, reflecting extroverted and open personality traits.

[0109] Action frequency is defined as the average number of times per minute that a digital clone's limb movements (such as nodding or gesturing) occur. Let a baseline value for action frequency be set. times / minute (approximately one physical movement every 30 seconds for a neutral personality), Let be the linear coefficient of personality dimension i with respect to action frequency, obtained through... Determine the target value of the action frequency For example, extroverted personalities typically prefer frequent gestures and movements, and therefore can be trained. For a relatively large positive value, such as +3, to make the extrapolation = 0.9 Movements per minute, or one movement every 12 seconds on average, significantly increases the frequency of movement; conversely, highly introverted people move less, which can make... If the value is negative, such as -1, then introversion = 0.8. Times per minute means a halved frequency of movement, making the digital avatar appear more reserved and less active. By adjusting the frequency of movement, the overall activity level of the digital avatar changes according to personality: extroverted individuals are more active, while introverted individuals are more sedentary.

[0110] Movement rhythm is used to control the duration or speed of a single movement. Let's define a baseline value for movement rhythm. seconds / time, define linear coefficients ,pass Determine the target value of the movement rhythm For example, extroverted people are decisive and efficient in their actions, which can be described as... When extroversion is 0.9 A second, or about 30% faster pace of movement; a highly conscientious person may move more calmly and steadily, which can be set... When the due diligence score is 0.8 Slowing down the pace slightly makes the movements appear more composed. By adjusting the rhythm of the movements, digital clones with different personalities can have different feel to their actions, appearing either "fierce and energetic" or "steady and slow."

[0111] Facial expression smoothness refers to the smoothness of the transition from one facial expression state to another, with a value ranging from 0 to 1, where 1 represents the smoothest and most seamless transition. Let's define a baseline value for facial expression smoothness. ,definition If the coefficient is linear, then the target value for the smoothness of facial expression transitions is... pass Certainly. For example, people with high neuroticism (low emotional stability) have more sudden and abrupt changes in expression, which can be attributed to... A negative value, such as -0.2, indicates a neuroticism score of 0.7. Abrupt transitions in facial expressions can convey tension and anxiety; conversely, a person with high affability is gentle and their facial expressions shift more smoothly and naturally, which can convey... When the affinity is 0.8 The transition from one emotion to another is smoother and more natural. This is achieved by controlling the target value for the smoothness of the expression transition. When users observe digital clones, they can feel the differences in micro-expressions between different personalities: the expressions of highly approachable characters change like a gentle breeze without leaving a trace, while the expressions of highly neurotic characters may suddenly turn sour or change from anger to joy in a rather abrupt way.

[0112] Similarly, the target values ​​for facial expression intensity and movement amplitude can be determined using the above formulas. In addition to the parameter mappings mentioned above, the influence of personality on the choice of specific facial expressions and movements can also be considered. For example, a digital avatar with a highly open personality might use exaggerated body language to express "surprise," while a highly introverted personality might only show slight facial changes. Through a series of mapping adjustments and rule controls, this invention ensures that the digital avatar's facial expressions and body language are highly consistent with its personality settings, reflecting its personality traits from the arc of a smile to the frequency of nodding.

[0113] In one embodiment, such as Figure 4 As shown, step S106, which involves performing multimodal synchronization and consistency verification on the target reply text, the target reply voice, and the target facial expression animation, and outputting the verified target reply text, target reply voice, and target facial expression animation, includes:

[0114] S401: Extract emotional features from the target reply text, the target reply voice, and the target facial animation respectively to determine the text emotion vector, the voice emotion vector, and the facial emotion vector;

[0115] S402: Calculate the intermodal similarity of the text sentiment vector, speech sentiment vector, and facial expression sentiment vector to determine the consistency score;

[0116] S403: When the consistency score is greater than the first score threshold, output the target response text, target response voice and target facial animation that have passed the verification;

[0117] S404: When the consistency score is not greater than the first score threshold and the consistency score is greater than the second score threshold, a slight adjustment is triggered to adjust at least one of the target values ​​of the voice control parameters and the target values ​​of the animation control parameters, and to repeatedly perform multimodal synchronization and consistency verification on the target response text, the target response voice and the target facial expression animation.

[0118] S405: When the consistency score is not greater than the second score threshold, trigger regeneration to regenerate at least one of the target response text, the target response voice, and the target facial animation, and repeatedly perform multimodal synchronization and consistency verification on the target response text, the target response voice, and the target facial animation.

[0119] The first scoring threshold is a pre-set threshold used to assess whether the consistency score reaches a higher standard. It can be implemented using... This indicates that the second scoring threshold is a pre-set threshold used to assess whether the consistency score has reached a lower standard. It can be implemented using... Indicates, that is .

[0120] As an example, in step S401, the electronic device can extract emotional features from the target reply text, the target reply voice, and the target facial expression animation respectively, and determine the text sentiment vector corresponding to the target reply text. The emotional vector of the target response speech The emotional vector corresponding to the target facial expression animation These sentiment vectors reflect emotional information from different modalities.

[0121] As an example, in step S402, the electronic device can evaluate the text sentiment vector based on a pre-set consistency assessment algorithm. Voice emotion vector and facial expression vectors These sentiment vectors are used to obtain a consistency score. Then, the consistency score value Compared with the first scoring threshold Second rating threshold The comparisons are made so that different control strategies can be implemented based on the different comparison results. The consistency evaluation algorithm here can be a pairwise calculation using the cosine similarity algorithm, followed by comprehensive analysis, or other algorithms can be used for evaluation.

[0122] As an example, in step S403, when the electronic device's consistency score is greater than the first score threshold, i.e. At that time, determine the text sentiment vector. Voice emotion vector and facial expression vectors These emotion vectors are quite similar, so their consistency check is deemed to have passed. The target response text, target response voice, and target facial animation that have passed the consistency check are then output.

[0123] As an example, in step S404, when the electronic device's consistency score is not greater than the first scoring threshold and the consistency score is greater than the second scoring threshold, i.e. At that time, determine the text sentiment vector. Voice emotion vector and facial expression vectors If the similarity between these emotion vectors reaches a low standard but not a high standard, then at least one of the target values ​​of the voice control parameters and the animation control parameters is adjusted. Specifically, at least one of the parameters of speech rate, fundamental frequency, and timbre can be adjusted, and / or at least one of the parameters of facial expression intensity, movement frequency, movement rhythm, and facial expression switching smoothness can be adjusted. After making slight adjustments to the relevant parameters, the multimodal synchronization and consistency check of the target response text, the target response voice, and the target facial expression animation is repeated until the readjusted consistency score is greater than the first score threshold. Then, the target response text, target response voice, and target facial expression animation that have passed the consistency check are output.

[0124] As an example, in step S405, when the electronic device's consistency score is not greater than the second score threshold, i.e. At that time, the text sentiment vector Voice emotion vector and facial expression vectors If the similarity between these emotion vectors does not reach the minimum standard, at this point, it is necessary to regenerate at least one of the target response text, target response voice, and target facial animation. That is, repeat steps S103-S105, and then repeat the multimodal synchronization and consistency check of the target response text, the target response voice, and the target facial animation until the readjusted consistency score is greater than the first score threshold. Then, output the target response text, target response voice, and target facial animation that have passed the consistency check.

[0125] In one embodiment, the consistency score is determined using the following formula:

[0126] ;

[0127] ;

[0128] ;

[0129] ;

[0130] in, This is the consistency score. For text sentiment vectors; For speech emotion vectors, For facial expression and emotion vectors; Cosine similarity between speech sentiment vector and text sentiment vector; The cosine similarity between the speech emotion vector and the facial expression emotion vector; The cosine similarity between the facial expression sentiment vector and the text sentiment vector; , and Three pre-set weighting coefficients.

[0131] As an example, electronic devices determine the sentiment vector of text. Voice emotion vector and facial expression vectors After obtaining these sentiment vectors, cosine similarity needs to be calculated pairwise to determine the cosine similarity between any two sentiment vectors. , , Then, adopt The three sets of cosine similarities are weighted and averaged to determine their corresponding consistency scores, so that the consistency scores can effectively reflect whether the emotions of the multimodal outputs are consistent.

[0132] In one embodiment, such as Figure 5 As shown, in step S403, when the consistency score is greater than the first score threshold, the target response text, target response voice, and target facial expression animation that have passed the verification are output, including:

[0133] S501: When the consistency score is greater than the first score threshold, a comprehensive emotion vector is determined based on the text emotion vector, the voice emotion vector, and the facial expression emotion vector;

[0134] S502: Based on the current personality vector, query the pre-set personality emotion mapping relationship to determine the expected emotion vector corresponding to the current personality vector;

[0135] S503: Calculate the similarity between the comprehensive emotion vector and the expected emotion vector to determine the personality compatibility.

[0136] S504: When the personality compatibility is less than the preset compatibility, adjust the overall style of the target reply text, the target reply voice, and the target facial expression animation until the personality compatibility is not less than the preset compatibility, and then output the verified target reply text, target reply voice, and target facial expression animation.

[0137] As an example, in step S501, the electronic device can also be based on text sentiment vectors. Voice emotion vector and facial expression vectors Determine the comprehensive sentiment vector Specifically, regarding the sentiment vector of the text Voice emotion vector and facial expression vectors Perform mean or weighted processing to determine the comprehensive sentiment vector. ,For example, .

[0138] As an example, in step S502, the electronic device can also be based on the current personality vector. Query the pre-set personality-emotion mapping relationship to dynamically determine the current personality vector. Corresponding expected sentiment vector This expected emotional vector Overall sentiment information used to reflect the multimodal information of the desired output.

[0139] As an example, in step S503, the electronic device can analyze the comprehensive emotion vector. and expected emotional vector Perform similarity calculations to determine personality compatibility. Specifically, but not limited to, cosine similarity algorithms can be used to analyze the comprehensive sentiment vector. and expected emotional vector The cosine similarity between the two is calculated and determined as the degree of personality compatibility. and the personality compatibility Compare with a preset fit and perform different operations based on the comparison results.

[0140] As an example, in step S504, the electronic device determines the compatibility of the individuals. Less than the preset fit At that time, determine based on the text sentiment vector Voice emotion vector and facial expression vectors Overall deviation from the current personality vector Corresponding expected sentiment vector Further overall style adjustments are still needed until the personality is a good fit. Not less than the preset fit Upon successful verification, the output includes the target reply text, target reply voice, and target facial animation. This preset fit... It is a pre-set threshold used to assess whether the personality compatibility meets a high standard, and can be set to 0.8 or other thresholds.

[0141] This invention provides a personality-driven, large-scale, multimodal digital avatar interaction system, which corresponds one-to-one with the personality-driven, large-scale, multimodal digital avatar interaction method described in the previous embodiments. For example... Figure 6 As shown, the personality-driven, large-scale, multimodal digital avatar interaction system includes:

[0142] The emotion recognition module 601 is used to perform emotion recognition on user input information collected during the interaction between the digital clone and the user, and to determine the emotion label.

[0143] The personality center module 602 is used to determine the emotion offset based on the emotion tag, and to determine the current personality vector corresponding to the digital clone based on the emotion offset and the basic personality vector corresponding to the digital clone.

[0144] The character big model text generation module 603 is used to process the character setting fields, dialogue context and the current personality vector using the character big model to determine the target response text that represents the personality style. The character big model injects the character setting fields and the current personality vector during the inference period to constrain the target response text to maintain character consistency and personality consistency in multiple rounds of dialogue.

[0145] The speech synthesis module 604 is used to determine the target value of at least one speech control parameter based on the current personality vector, determine the speech style vector based on the target value of at least one speech control parameter, inject the speech style vector into a text-to-speech model, use the text-to-speech model to perform speech synthesis on the target response text, and determine the target response speech that represents the personality style.

[0146] The facial expression animation module 605 is used to determine the target value of at least one animation control parameter based on the current personality vector, determine the animation style vector based on the target value of at least one animation control parameter, inject the animation style vector into the animation generation model, and use the animation generation model to process the target response text and the target response speech to determine the target facial expression animation representing the personality style.

[0147] The multimodal synchronization module 606 is used to perform multimodal synchronization and consistency verification on the target reply text, the target reply voice, and the target facial expression animation, and output the target reply text, target reply voice, and target facial expression animation that have passed the verification.

[0148] In one embodiment, the personality central module determines an emotion offset based on emotion tags and updates the current personality vector, wherein the emotion offset and the current personality vector at least satisfy:

[0149] , ;

[0150] in, For emotional offset; This is a weight matrix for the emotion-personality dimension. Let the emotion intensity vector be... , Determined based on the emotional intensity corresponding to multiple emotional types This is the emotional attenuation coefficient. The range of values ​​is ; Adapt weights to different scenarios. The range of values ​​is ; Based on the basic personality vector, This represents the current personality vector. This is a function that cuts off the interval.

[0151] In one embodiment, the personality central module 602 further includes a personality version number field and a status update log field, which are used to record the update history of the basic personality vector and the current personality vector and support the rollback of personality parameters.

[0152] In one embodiment, the speech synthesis module 604 maps the current personality vector to speech control parameters, the speech control parameters including at least one of speech rate, fundamental frequency and timbre;

[0153] The target values ​​of the speech rate and the base frequency are obtained by a segmentation strategy that uses linear mapping within a preset range and nonlinear mapping outside the preset range.

[0154] The target value of the timbre is obtained by linear mapping and interval truncation.

[0155] In one embodiment, the facial expression animation module 605 maps the current personality vector to animation control parameters, which include at least one of facial expression intensity, movement amplitude, movement frequency, movement rhythm, and facial expression switching smoothness. Boundary truncation is performed on each animation control parameter to suppress exaggerated movements or abrupt changes in rhythm caused by extreme personality values.

[0156] In one embodiment, the multimodal synchronization module 606 calculates a consistency score based on text emotion vectors, speech emotion vectors, and facial expression emotion vectors, and sets a first scoring threshold and a second scoring threshold. When the consistency score is greater than the first scoring threshold, it outputs the target response text, target response speech, and target facial expression animation that have passed the verification. When the consistency score is not greater than the first scoring threshold but greater than the second scoring threshold, it triggers a slight adjustment to adjust at least one of the target values ​​of the speech control parameters and the animation control parameters. When the consistency score is not greater than the second scoring threshold, it triggers regeneration to regenerate at least one of the target response text, the target response speech, and the target facial expression animation.

[0157] In one embodiment, the multimodal synchronization module 606 determines the desired emotion vector based on the current personality vector, and determines the personality fit based on the similarity calculation between the comprehensive emotion vector and the desired emotion vector; when the personality fit is less than a preset fit threshold, the overall style is adjusted.

[0158] In one embodiment, the personality central module 602 performs long-term optimization of the basic personality vector according to the optimization cycle. The update of the basic personality vector includes at least: calculating the average offset of the emotion offset within the optimization cycle and superimposing it onto the basic personality vector according to the personality optimization coefficient.

[0159] To avoid ambiguity in terminology, the terms "text generation module" and "character model text generation module" in this specification refer to the same functional module unless otherwise specified: a module used to generate target response text by injecting character setting fields and the current personality vector into the character model during the reasoning phase; if the term "character model text generation module" is used only in subsequent claims and embodiments, it should be understood as an overall module implementation including the character model reasoning engine and personality constraint injection logic.

[0160] The system of this invention uses the personality central module 602 as its core, and integrates the current personality vector. The unified control parameters are distributed to the character large model text generation module 603, speech synthesis module 604, and facial animation module 605, and the multimodal synchronization module 606 performs alignment and consistency checks on the output, thus forming a closed-loop control. In this example, the core generation engine of the character large model text generation module 603 is a pre-trained character large model (LLM), and character setting fields are injected during the inference period. With the current personality vector This is to constrain the consistency of generated responses in terms of role and personality across multiple rounds of dialogue. If simply described as a "pre-trained language model" or "natural language model," it should be understood as a text generation mechanism implemented using a large role model.

[0161] This invention adopts a layered and modular system architecture, comprising five core modules: a personality central module 602, a character large-scale model text generation module 603, a speech synthesis module 604, an facial expression animation module 605, and a multimodal synchronization module 606. The personality central module 602, acting as a unified personality control center, maintains the basic personality vector. And update the current personality vector. The character model text generation module 603 injects character setting fields. With the current personality vector Then generate the target response text The speech synthesis module 604 responds with text based on the target. With the current personality vector Generate target response voice and phoneme timestamps The facial animation module 605 is based on the current personality vector. With phoneme timestamps Generate target facial animation The multimodal synchronization module 606 responds to the target text. Target reply voice With target facial animation Alignment and consistency scores and personality compatibility scores are calculated, and fine-tuning or regeneration is triggered under a threshold grading strategy, thus forming a closed-loop control architecture of "generation-verification-adjustment". Through this architecture, the digital avatar's expression in different modalities always revolves around the same personality state, improving the coherence and realism of the interaction.

[0162] To support the aforementioned personality-driven and multimodal collaborative functions, this invention designs a comprehensive system data structure to store important parameters and status information of each module. The main data structures and fields are as follows:

[0163] (1) Data stored in the personality central module 602

[0164] In this example, the personality central module 602 has a personality parameter record table, which includes a basic personality vector field, a current personality vector field, an emotion label field, an emotion offset field, a personality version number field, and a status update log field. The fields form a traceable personality evolution link through update rules.

[0165] Basic Personality Vector Field (Type: Floating-point array, Length n): Stores the basic personality vector of the digital clone. The array elements correspond to parameter values ​​for each personality dimension (such as extraversion, agreeableness, conscientiousness, neuroticism, and openness), with values ​​ranging from [value range missing]. The vector is preserved to four decimal places. It is initially defined by the user or quantified through a questionnaire and can be stored in a local database (SQLite) and periodically backed up to the cloud. For example, each personality style can be set as a field (e.g., "Extraversion," type floating-point), and the entire personality vector... Stored as arrays or database records.

[0166] Current personality vector field (type: floating-point array, length n): Stores the currently updated current personality vector. This vector is available for real-time use by various generation modules. Its dimensions are consistent with the base personality vector, and it is updated frequently based on emotional input (e.g., updated after each dialogue round, or with a minimum update interval of <= 100ms). The current personality vector is stored in a high-performance memory cache (such as Redis) and can be synchronized to a local database to prevent loss.

[0167] The Personality Version Number field (type: string) identifies the version of the base personality vector setting, for example, using the format "Vx.y" (V1.0, V1.1, etc.). The version number increments whenever the base personality vector is manually modified or adjusted by the optimization mechanism. Several recent (e.g., 3) historical version numbers and their corresponding personality parameter configurations are also retained to support version rollback functionality. Version information is stored in a local database and uploaded to cloud logs.

[0168] Scene Adaptation Weight Field (Type: Floating-point array, Length n): Stores the scene adaptation weights of each personality dimension in different scenarios. By default, each dimension has a weight of 1.0, meaning the sentiment bias applies to all dimensions to its original degree; however, the scenario-appropriate weights for certain dimensions can be adjusted for specific scenarios. A value higher or lower than 1 indicates a change in personality performance in that dimension, either strengthening or weakening it. Weights are applied across all scenarios. The sum can be regularized to n for easier calculation. Scene adaptation weight configuration is stored in a local configuration file (JSON format) and can be synchronized to the cloud to maintain consistency across multiple devices.

[0169] Status Update Log Field (Type: List): Records important status changes and events of the Personality Central Module 602. Each log entry includes a timestamp, event description, and parameter change details (e.g., "Emotion Input: Anger 0.8, ..."). The log entries are recorded as follows: [-0.10,…], extraversion decreased from 0.5 to 0.4, etc., and the operator (marked as "system" if automatically adjusted by the system, and user ID if manually adjusted by the user). The log list is stored in a local database and a cloud database (such as MySQL), retaining the most recent 1000 entries, with subsequent entries archived every 30 days for future auditing. For example, when a user praises their digital avatar, triggering a "happy" emotion, the log records the event time, the "happy" tag, and the magnitude of the increase in extraversion or pleasantness parameters in the personality vector. Through state update records, the system can reconstruct the evolution of personality parameters over time, ensuring that each module obtains a consistent historical state reference at different interaction stages.

[0170] Personality optimization parameter field (type: floating point): Stores personality optimization coefficients. This is used to control the magnitude of long-term optimization of the base personality vector. The default value is, for example, 0.08, which can be adjusted by the user according to their expectations (it can be increased if a significant change in personality over time is desired). (Conversely, reduce). This parameter is stored in a local configuration file, and the user interface provides adjustment options.

[0171] This invention provides multiple methods for constructing the basic personality vector of a digital clone, supporting standardized personality settings and personalized customization, and storing the constructed basic personality vector of the digital clone in local memory, cloud server or other storage devices so that the personality central module 602 can directly call it.

[0172] Questionnaire Quantification Method: Authoritative personality psychology questionnaires (such as the Big Five Personality Inventory NEO-PI-3, MBTI index, Eysenck Personality Questionnaire EPQ, etc.) are used to obtain survey data. The scores of each dimension of the survey data are then normalized to... Intervals form the basic personality vector. For example, an extraversion dimension questionnaire score of 80 can be converted into an extraversion personality parameter of 0.80. The personality parameters obtained by this method have a high degree of standardization and are suitable for setting the basic personality vector for general digital avatars.

[0173] User-defined method: This method provides a visual interface that allows users to directly set the strength of each personality dimension to determine personality parameter data. Based on this data, the basic personality vector corresponding to the digital avatar is determined. Users can adjust parameter values ​​such as "Extraversion = 0.9" and "Affinity = 0.7" by scrolling, and can also add custom personality dimensions (such as additional dimensions like "Sense of Humor" and "Patience," with a value of 0.8 indicating a strong sense of humor). This method gives users complete control over the digital avatar's personality and is suitable for personalized digital avatar scenarios.

[0174] Behavioral data training method: This method trains and optimizes personality vectors based on historical interaction data between digital avatars and users over a long period. Specifically, it uses historical interaction data corresponding to the digital avatars to train the model and determine the basic personality vectors for each avatar. Specifically, it collects data such as text style features of past conversations of the digital avatars, user feedback (satisfaction ratings, emotional responses), and interaction durations, and constructs a loss function with user satisfaction and personality consistency as the objective. The personality parameters are updated through iterative optimization algorithms. For example, logistic regression or reinforcement learning models can be used, with improved user satisfaction and personality consistency as positive rewards, and the personality vector is adjusted through backpropagation to gradually approach the user's preferred personality settings.

[0175] The iterative formula for training and updating can be expressed as: ,in, Let be the value of the i-th dimension of the personality vector at the k-th iteration. For learning rate (e.g.) ), The target loss function is defined as follows (e.g., including user satisfaction score and personality consistency score). By updating personality parameters at regular intervals, the personality style of the digital avatar can be gradually fine-tuned towards the user's preferences, making it suitable for the personality evolution of long-term companion-type digital avatars.

[0176] Scene Adaptation Method: Several typical scene personality templates are preset. The personality configuration is automatically switched according to the current scene in which the digital avatar is located. Based on the current scene of the digital avatar, a pre-set scene personality mapping table is queried to determine the basic personality vector corresponding to the digital avatar. For example, in a customer service scene, a personality parameter combination of "high conscientiousness (0.9) + high affinity (0.8) + moderate extraversion (0.6)" is loaded, while in an entertainment chat scene, the personality configuration of "high extraversion (0.9) + high openness (0.8) + moderate conscientiousness (0.5)" is switched. The scene adaptation module supports hot switching, quickly loading the corresponding personality vector when switching scenes, improving the performance and professionalism of the digital avatar in multiple scenarios.

[0177] The basic personality vector obtained through the above methods is denoted as... It contains n elements ( For example, it typically includes five dimensions: extraversion, affinity, conscientiousness, neuroticism, and openness. Each element's value ranges from [value range missing]. A higher value indicates a stronger personality style. Basic Personality Vector As a stable personality setting for the digital clone, it is loaded during system initialization, stored in the personality central module 602, and can be adjusted by the user or updated through training as needed.

[0178] (2) Data stored in the character large model text generation module 603

[0179] Personality-Vocabulary Mapping (Type: Dictionary): The key is a "personality dimension + intensity range" identifier (e.g., "Extraversion_H" indicates a high extraversion range), and the value is a list of corresponding emotional words. This vocabulary collects commonly used words or phrases under various personality styles and supports regular updates (e.g., retrieving new vocabulary from the cloud every 7 days) and user-defined expansions. Dictionary data is stored locally in CSV or JSON format and converted into an in-memory Trie tree upon loading for fast retrieval and replacement.

[0180] Personality-Sentence Pattern Rule Mapping (Type: Dictionary): The key is a "personality dimension + intensity range" identifier, and the value is the corresponding set of sentence pattern rules. Each rule can be represented by a regular expression or a simple script, used to match and replace specific sentence structures. For example, the rules corresponding to "Extraversion_H" might include rules that match adding an exclamation mark to the end of a statement, and "Conscientiousness_H" might include rules that merge parallel short sentences, etc. Each personality dimension has at least 10 preset sentence pattern rules, and advanced users can add custom rules through an XML configuration file. The rule dictionary is stored in a local XML file and backed up in the cloud for sharing by multiple users.

[0181] Personality-Catchphrase Mapping (Type: Dictionary): The key is a "personality dimension + intensity range" identifier, and the value is the corresponding catchphrase set. This is used to create a unique catchphrase set for digital avatars with different personality styles. For example, the friendly and approachable personality often adds "ne," "ya," or "yo" to the end of sentences to express intimacy and a gentler tone; the lively and extroverted personality likes to use "haha" or "hey" to express laughter or greetings; and the open and curious personality may frequently say "wa," "I really didn't expect that," or "interesting."

[0182] (3) Data stored in speech synthesis module 604

[0183] Mapping coefficient library (type: dictionary): Stores the coefficients of the voice control parameters, including at least the linear coefficients of the voice control parameters. , , and the linear coefficients of the voice control parameters , The dictionary keys are the names of the control parameters (e.g., "alpha_vr", "beta_f0", "epsilon_timbre", "gamma_vr", "delta_f0", "zeta_freq", "eta_rhythm", "theta_smooth"), and the dictionary values ​​are floating-point arrays corresponding to the number of personality dimensions, n. Mapping coefficients can be obtained through training on labeled corpora or pre-configured from configuration files, supporting version updates and rollback. The mapping coefficient library is stored in a local database and synchronized to a cloud parameter file for backup after each parameter update.

[0184] (4) Data stored in the facial animation module 605

[0185] Animation Parameter Templates (Type: Dictionary / Resource File): Stores baseline values ​​for animation control parameters such as expressions and actions. The key is the name of the expression or action type (e.g., "smile," "nod," "wave"), and the value is a parameter object for that action, including baseline values ​​for expression intensity, action amplitude, action frequency, action rhythm, and expression transition smoothness. The Expression Animation Module 605 generates specific animations by overlaying personality offsets onto these templates. Animation parameter templates are typically bound to a specific animation engine, such as a Unity3D or Unreal Engine project file, and are stored in a local resource folder. Developers can adjust and optimize them using engine tools, and new action templates can be imported to expand the digital avatar's action library.

[0186] (5) Data stored in the multimodal synchronization module 606

[0187] Consistency Threshold (Type: Floating Point): The first score threshold for determining consistency in storage consistency scoring. Second rating threshold This can be set according to actual conditions. Preset fit. This is used to assess whether the personality compatibility meets the preset threshold.

[0188] Expected Emotional Vector (Type: Floating-point Array): Stores the expected emotional vector corresponding to the current personality vector. This expected affective vector is obtained through table lookup or calculation. For example, it can be mapped based on personality style to reveal the basic emotional composition typically conveyed by that personality (e.g., the expected affective vector of a highly extroverted and amicable personality). Primarily focused on "happiness and peace"). Whenever the current personality vector... During the update, the multimodal synchronization module 606 updates the expected sentiment vector in real time. (To ensure consistency verification can obtain the latest personality expectations). This expected emotional vector It resides in memory (e.g., Redis cache), is refreshed synchronously with the current personality vector, and is also recorded in the log for easy analysis.

[0189] The process of the personality-driven large-scale role model multimodal digital avatar interaction system provided in this embodiment of the invention can be summarized into three stages: initial loading, real-time interactive update, and multimodal generation output. Each stage is described below.

[0190] (1) Initialization loading logic

[0191] When the digital avatar interaction system starts, configuration parameters and model resources need to be loaded sequentially, and module connections need to be initialized to ensure smooth operation of subsequent interactions. The initialization process is as follows:

[0192] Loading basic configuration: The system first reads the local configuration file to obtain the number of personality dimensions and the first scoring threshold. Second rating threshold Preset compatibility Personality optimization coefficient Scene adaptation weight The system checks global parameters and performs validity checks. For example, it checks whether the personality dimensions include at least five basic dimensions. If an anomaly is found (such as missing configuration or incorrect format), the system will load the default configuration and log a warning.

[0193] Loading Personality Vectors: Retrieves the base personality vector of the digital clone from the local database. And its version number. Then check if there are any dynamic emotion tags left over from the previous session in the memory cache. If so, retrieve the most recent emotion tag and calculate the emotion offset. Otherwise, the emotional offset will be adjusted. Initialize as an all-zero vector. Calculate the current personality vector. (Usually during system startup) Therefore Based on this, the personality central module 602 prepares the initial current personality vector. It will be distributed to other modules later.

[0194] Loading mapping coefficients: The system reads the mapping coefficient library from the speech synthesis module 604 (e.g., coefficient dictionaries for alpha_vr, beta_f0, epsilon_timbre, gamma_vr, delta_f0, zeta_freq, eta_rhythm, theta_smooth, etc.) and the animation parameter template from the facial expression animation module 605, loading these parameters into memory for efficient calculation. Simultaneously, it loads the personality-vocabulary mapping, personality-sentence pattern rule mapping, and personality-catchphrase mapping required by the character large model text generation module 603. If some configuration files are missing or corrupted, an error log is recorded and default parameters are loaded to ensure system availability.

[0195] Loading Model Resources: Each deep learning model is initialized sequentially, including the pre-trained LLM (Large Character Model) weights for the LLM text generation module 603, the text-to-speech (TTS) weights for the speech synthesis module 604, the emotion recognition model, and necessary external dependencies (such as word segmentation / syntactic analysis models). Model files are preferably stored locally; however, when files are large or device performance is limited, on-demand loading, quantization, or distillation of the simplified model can be used. This invention also supports updating to the latest model version from the cloud and records the model version number after the update for comparison and rollback.

[0196] Module Startup and Connection: The processes or service threads of the following modules are started sequentially: Personality Central Module 602, Character Model Text Generation Module 603, Speech Synthesis Module 604, Facial Animation Module 605, and Multimodal Synchronization Module 606. The communication interfaces between them are then tested for smooth operation. For example, an empty personality parameter update broadcast is sent, and the responses of each module are checked: Character Model Text Generation Module 603 should return an acknowledgment within 500ms; Speech Synthesis Module 604 and Facial Animation Module 605 should establish an RPC connection and wait for the call, etc. If a module connection fails, the system automatically retryes up to 3 times. If the connection still fails, an error message is output to the user interface (e.g., "Speech Synthesis Module 604 initialization failed, please check the TTS model file"), and the system enters a safe degrade mode (which may only provide text output).

[0197] Initialization complete: Once all the above steps are successful, the system records an initialization log (including the loaded personality version, connection status of each module, main parameter configurations, etc.), and then enters standby mode to await user input. At this point, all modules of the digital avatar have been loaded and share the initial personality state parameters, ready for consistent personality output.

[0198] (2) Logic for dynamic updating of personality vectors

[0199] During system operation, whenever new user input (text or voice) is received, the system dynamically updates the personality state of the digital avatar and drives multimodal output according to the following process:

[0200] Emotion Input Detection: User input information is first analyzed by the emotion recognition module 601 to extract emotion tags. During text interaction, sentiment analysis is performed on the user input; during voice interaction, emotion recognition is performed on the user's voice tone. The system sets the time resolution for emotion detection, for example, updating the emotion results every 50 milliseconds. If no obvious emotion is present at a certain moment, a "calm" emotion intensity of 0 is output. This emotion tag is sent to the personality center module 602 in real time.

[0201] Calculating the emotion offset: After receiving the emotion label, the personality center module 602 can determine the corresponding emotion intensity vector based on the emotion label. Utilize the scene adaptation weights of the current scene and the preset emotional attenuation coefficient According to the formula Calculate the mood offset To prevent emotional shifts The changes are too drastic; first, adjust the emotion intensity vector. Perform a smoothing filter (e.g., take the average of the emotion intensities from the three most recent tests), and then calculate the resulting emotion shift. The parameter values ​​for each personality dimension exceeded If so, then truncate to the boundary value to protect the emotional offset. The parameter values ​​for each personality dimension in Inside.

[0202] Update current personality vector: Personality central module 602 based on emotion offset Superimposed basic personality vector This yields the new current personality vector. And immediately update the current personality vector in the memory cache. Provided for downstream modules to call; simultaneously records the update log for this round (timestamp, sentiment tag, ... and Furthermore, adopting Update the current dialogue round number t. If the current dialogue round number t is greater than the preset dialogue round number T corresponding to the optimization period (e.g., a cumulative 100 rounds of interaction), then calculate the average offset of the sentiment shift within that optimization period. , and according to Fine-tuning The version number is incremented, and then the accumulated statistics are cleared before starting the next cycle. This update is usually completed in milliseconds, allowing the new personality parameters to be quickly made available to downstream modules.

[0203] Real-time parameter push: The personality central module 602 pushes the updated current personality vector via RPC or local call interface. The message is pushed to the character model text generation module 603, the speech synthesis module 604, and the facial animation module 605. Each module receives the personality control parameter and updates its internal cache to ensure that subsequent generated content uses the latest personality state. For example, when generating a response, the character model text generation module 603 will call the latest current personality vector. Before synthesizing the next sentence, the speech synthesis module 604 updates speech control parameters such as speech rate. To ensure real-time performance, each module is required to complete the update within 10ms of receiving the parameters; otherwise, a warning is recorded for performance optimization.

[0204] Update Log: The personality central module 602 records key information about this personality state update in the log list, including the timestamp, the detected emotion type and intensity, and the emotion offset. Updated current personality vector The logs track various dimensions, including whether they trigger basic personality optimization. They are stored locally in a database for development, debugging, and tracking personality changes, and asynchronously uploaded to a cloud server for big data analysis (such as plotting user emotion and personality changes over time, and training and improving emotion mapping models).

[0205] Anomaly Handling: If a module fails to update parameters during the push process (e.g., due to network latency or module failure causing a timeout), the system will retry twice. If it still fails, the system will temporarily maintain the previous value for the module's personality parameters and warn the user interface that the output of a certain modality may be temporarily inaccurate. Simultaneously, the system background will attempt to automatically restart or repair the failed module. In this degraded mode, the digital clone can still interact; however, for security reasons, the multimodal consistency verification threshold will be temporarily lowered or the failed modality will be skipped to ensure the entire system remains uninterrupted.

[0206] (3) Multimodal content generation and output

[0207] After updating the current personality vector, the various generation modules of the system work together to produce multimodal personalized response content, which is then presented to the user after synchronous calibration:

[0208] Target response text generation: The character model text generation module 603 receives user input information and the latest current personality vector. First, an initial response text is generated using LLM (Local Language Modeling). Then, "personality-vocabulary mapping," "personality-sentence structure rule mapping," and "personality-catchphrase mapping" are applied to process the initial response text. The character model text generation module 603 extracts words from the personality vocabulary as needed to replace neutral words in the initial response text, matches corresponding rules from the sentence structure rule library to adjust sentence structure and tone, and inserts personalized catchphrases, ultimately forming a target response text with a personality style. For example, for a current personality that is extroverted and humorous, the initial response "Okay, I understand" might be adjusted to a more lively sentence like "Okay! I know!" The complete text is temporarily stored in the synchronization module before being sent to the user for consistency verification (extracting text sentiment).

[0209] Speech signal synthesis: The speech synthesis module 604 obtains the target response text and the current personality vector. Then, the target values ​​of the speech control parameters are calculated using the mapping function, specifically determining the target value of the speech rate. Target value of fundamental frequency Target value of timbre These parameters are then loaded as speech style vectors into the TTS model, and together with the text, speech waveform data and corresponding phoneme timestamps are generated. During the generation process, the speech synthesis module 604 automatically embeds speech style vectors (such as background mood sound effects or special intonation processing for interjections) into the speech signal to improve the accuracy of emotion recognition. The final output speech signal is in a standard audio format (such as a PCM stream or MP3 file) and includes a precise phoneme timestamp sequence, which is then processed by the synchronization module.

[0210] It is important to note that timbre parameters (such as brightness / fullness / softness) belong to a speech style dimension that is relatively independent of speech rate and fundamental frequency. The determination of their target values ​​is not predicated on the pre-determined target values ​​for speech rate and fundamental frequency. During speech synthesis, the speech style vector can be composed of any one or more parameters from speech rate, fundamental frequency, and timbre and injected into the text-to-speech model. Preferably, the target value of timbre employs a linear mapping and performs interval truncation.

[0211]

[0212] in, The target value for timbre, As the baseline value for timbre, Let be the linear mapping coefficient of the i-th personality dimension to timbre. Let be the parameter value of the i-th personality dimension in the current personality vector. As a preset boundary. Compared to speech rate and fundamental frequency, timbre parameters are less sensitive to intelligibility. Using linear mapping and truncation can achieve differentiated expression of personality style while ensuring stability.

[0213] To further improve feasibility and stability, the nonlinear mapping preferably adopts a quadratic soft-limiting form: when (For example When ), the target values ​​for speech rate and fundamental frequency are respectively:

[0214] ,

[0215] in The nonlinear adjustment coefficient is preferably within the following range: Used in When the amplitude of linear extrapolation approaches 0 or 1, soft limiting is applied to avoid decreased intelligibility or unstable prosodic speech caused by excessively fast speech rate or abnormal fundamental frequency; when... A linear mapping is employed to ensure that speech parameters change smoothly with personality dimensions. This segmentation strategy makes speech style adjustment both sensitive and robust.

[0216] Expression and motion animation generation: Expression animation module 605 obtains the target reply text, target reply voice, and phoneme timestamp. and current personality vector After setting the parameters, the digital avatar begins to operate. First, it infers the necessary facial expression and action sequences based on text semantics and speech stress; for example, it adds a head tilt when the text contains a question, and a nod when the speech ends with a falling intonation. Then, it uses the current personality vector... Adjust the target values ​​for each animation control parameter: calculate the target value for at least one of the animation control parameters, such as expression intensity, movement amplitude, movement frequency, movement rhythm, and expression transition smoothness. Then, input the animation instructions into the rendering engine (Unity3D, etc.): reset the expression to its initial state at 0 milliseconds, and then strictly follow the speech timestamp timeline to drive the face blendshape weights to change with the phoneme's mouth opening degree, driving the skeletal joints to perform nodding, gestures, and other movements according to the calculated rhythm. For example, during the 0.1s to 0.5s of the speech corresponding to the pronunciation of "I'm so happy," the character's mouth opens to its maximum, the facial expression is a high-intensity smile, the upper body performs a small nodding movement at 0.3s, and then the expression slightly softens and the movement pauses at the 0.5s pause in the speech. This continues until the end of the speech, generating a series of animation frames that match the length of the speech.

[0217] Motion amplitude mapping: Motion amplitude parameters control the spatial range of a single limb or head movement, reflecting personality differences such as "extroverted / restrained" and "enthusiastic / composed." Motion amplitude can be represented by joint rotation angles (e.g., the head pitch angle in a nodding motion), skeletal displacement (e.g., the distance of a hand gesture), or a BlendShape amplitude scaling factor. Let the baseline value for motion amplitude be... ,set up and These are the lower and upper limits of the range of motion (e.g., the head nodding and pitch angle are limited to a certain range). Or the scaling factor is limited to ),definition Let be the linear mapping coefficient of the i-th personality dimension to the amplitude of movement. Then the target value of the amplitude of movement can be determined as follows:

[0218]

[0219] in, This represents the parameter value of the i-th personality dimension in the current personality vector. By truncating the action amplitude, we can avoid excessive exaggeration or distortion of actions caused by extreme personality values. Together with parameters such as action frequency, action rhythm, and smoothness of expression switching, this forms a robust animation style control vector.

[0220] Multimodal Synchronization and Calibration: After text, speech, and animation content have been generated, the multimodal synchronization module 606 performs final integration and calibration. First, it uses phoneme timestamps... Adjust the timeline of the animation frame sequence: ensure that each lip movement strictly corresponds to the corresponding phoneme interval, and align the occurrence time of major actions (such as nodding, gestures) with the emphasized parts of the text content or pauses between sentences (e.g., a nod during a 0.5-0.7s silence period). If the generated animation does not match the audio duration (e.g., the animation is too long or too short), align the audio length by adjusting the action interval or slightly stretching / compressing the animation time. Next, the multimodal synchronization module 606 extracts the speech emotion vector. Text sentiment vector Facial expression and emotion vectors Calculate the consistency score. Compatibility with personality The system then determines whether content needs to be regenerated or adjusted based on the aforementioned rules. If adjustments are needed, possible operations include: increasing or decreasing the intonation intensity of the synthesized speech (e.g., making the speech gentler to match facial expressions), fine-tuning the intensity of facial animation (e.g., increasing the smile amplitude to match the pleasantness of the speech), or even having the character model text generation module 603 use different wording (e.g., if the original text content lacks clear emotional color, leading to modal inconsistency, try using sentences with exclamation marks to enhance the emotion). These adjustments are coordinated by the synchronization module to complete the work of each generation module until the new output meets the consistency conditions. Because the generation speed of each module in this invention is relatively fast (typically text < 200ms, speech < 500ms), even after one or two iterations of adjustments, users will not subjectively perceive any significant delay.

[0221] Output Presentation: The final, calibrated text, voice, and animation will be synchronously presented to the user through the user interface. Text content can be displayed as subtitles or dialog boxes, while the voice audio plays automatically, allowing the user to intuitively feel that the digital avatar is communicating with them: the words spoken, tone of voice, and facial expressions are harmonious and consistent with their personality setting. This multimodal fusion output provides users with a highly immersive interactive experience.

[0222] In a preferred embodiment, when calculating the consistency score, the multimodal synchronization module not only calculates the overall consistency score, but also calculates the similarity between each pair of modalities and locates the "conflict source modality." Specifically, it calculates... , , The system selects the modality pair with the minimum value as the conflict pair. When the consistency score is in the "mildly unacceptable" range, targeted fine-tuning is performed only on the modality with more adjustable parameters in the conflict pair (e.g., prioritizing fine-tuning of facial expression intensity / motion amplitude / smoothness, followed by fine-tuning of speech rate / fundamental frequency / timbre) to minimize regeneration costs and improve convergence speed. When the consistency score is in the "severely unacceptable" range, the conflict modality is regenerated, and the consistency score and personality fit are checked again after regeneration until the output meets the threshold. Through the closed-loop control of "conflict location - targeted fine-tuning / regeneration", cross-modal personality conflict can be significantly reduced and the efficiency and stability of consistency verification can be improved.

[0223] Compared to existing solutions that only introduce persona traits in a single module, this invention proposes a "unified personality vector bus" mechanism: the personality center uses... A single, shared personality state source is simultaneously injected into the text generation, speech synthesis, and facial animation modules. The multimodal synchronization module uses this unified personality state as a reference to score the consistency of the emotional vectors output from each of the three modalities, and further performs a secondary verification of personality fit using the expected personality emotional vector. This transforms "personality consistency" from a subjective description into a calculable, threshold-determinable, and closed-loop corrective control problem. This unified bus and dual-verification closed-loop control enable the system to maintain controllable and traceable role and personality consistency even with inherent coupling errors and model illusions in multimodal outputs.

[0224] The technical solution of this invention can be implemented by combining software architecture with a deep learning model. The system adopts a modular design, with each functional module calling and connecting according to a predetermined process, and parameters being passed between modules to form a complete data flow loop. The following describes the implementation path of the software and model and the parameter injection method in conjunction with each module:

[0225] Personality Central Module 602: The core service is built using the Python Flask framework, and Redis is used to store real-time personality vectors. Along with emotion tags, MySQL stores historical logs and basic personality vectors. Emotion recognition employs a pre-trained BERT emotion classification model (for text input) and a Wav2Vec2+CNN model (for speech input), outputting an emotion type and emotion intensity vector. .

[0226] Character Model Text Generation Module 603: Based on the LLaMA-7B or ChatGLM-6B model, fine-tunes instructions and injects "character setting fields". +Personality Vector The system provides prompt templates and integrates a regularized sentence structure rule engine to enable word replacement and sentence pattern adjustment. The retrieval memory unit uses the FAISS vector database to store dialogue history embeddings and supports Top-K similarity retrieval.

[0227] Speech synthesis module 604: It adopts the VITS model as the basic TTS architecture, adds a speech style vector interface to the model input layer, and trains the mapping coefficients through a multi-style speech dataset (including annotations of different personality dimensions). , etc. Phoneme timestamps The model's built-in alignment module outputs an accuracy of up to 10ms.

[0228] Facial Expression Animation Module 605: Developed using the Unity3D engine, it reads the current personality vector via C# script. With phoneme timestamps It drives the BlendShape (facial expressions) and skeletal animation of the digital avatar. Preset motion templates are stored in Animation Clip format, and personality parameters such as Speed ​​and Intensity of the animation are dynamically adjusted by scripts.

[0229] Multimodal synchronization module 606: Developed in C++, this module features a high-efficiency timing alignment engine based on phoneme timestamps. Achieve millisecond-level animation frame synchronization; consistency verification is performed using a sentiment vector extraction model deployed in PyTorch (CNN+LSTM for speech sentiment and ResNet for facial feature extraction for facial expressions), and cosine similarity scores are calculated in real time. .

[0230] Through the above technical solutions, the present invention has achieved many beneficial effects, specifically manifested in the following aspects:

[0231] (1) A highly scalable general framework

[0232] The personality-driven architecture of this invention adopts a modular design, with each personality dimension and modality control parameter existing in a configurable form, thus possessing high scalability. Developers can easily expand personality dimensions (e.g., adding traits such as "sense of humor" or "artistic taste" by simply adding corresponding parameters and mapping rules), add new output modalities (e.g., haptic feedback devices, olfactory simulation devices, etc. by simply adding corresponding personality mapping and synchronization control modules to the new modality), and expand application scenarios (from customer service and education to industrial training, psychological counseling, etc. by simply adjusting personality templates and scenario weights). The framework's versatility ensures that this invention can be adapted to the needs of different industries and has the ability to continuously evolve without having to overhaul the core architecture and redevelop it.

[0233] (2) Scene adaptation and user personality adaptation

[0234] This invention uses scene adaptation weights It achieves scene adaptation, automatically adjusting personality behavior according to the environment (e.g., suppressing liveliness in professional settings and enhancing friendliness in entertainment scenarios), allowing the digital avatar to "use multiple roles" and be competent in various situations. Simultaneously, combined with a long-term personality optimization mechanism, it achieves user personality adaptation; the system gradually learns the user's communication preferences and emotional needs, dynamically adjusting the basic personality vector. The dual adaptation of scenarios and personalities allows the digital avatar to provide an appropriate interactive experience in different situations. In a sample test, user satisfaction increased by more than 40% after applying this mechanism compared to the fixed personality model, and the sense of familiarity and stickiness in long-term use also significantly improved.

[0235] (3) The project is feasible and efficient.

[0236] This invention employs a mature AI model combined with lightweight rules, ensuring both innovation and engineering feasibility. Each module can be developed and tested independently, and clear interface integration reduces system complexity. It supports both local and cloud deployment modes: the optimal model can be run on a high-performance server for maximum quality, while a customized model can be run independently on mobile or embedded devices, ensuring low latency. Testing shows that the system's initialization time on a PC is less than 3 seconds, with a response latency of approximately 0.5 seconds per round for text mode and 1-2 seconds for voice mode, demonstrating excellent performance. In cloud-based collaborative mode, after optimized network transmission compression, the additional latency is no more than 0.3 seconds, fully meeting real-time interaction requirements. Furthermore, all core algorithms can be executed in parallel or asynchronously (e.g., parallel emotion detection and text generation), fully utilizing multi-core CPUs and GPUs for acceleration. Large-scale concurrent user scenarios can also be addressed through elastic scaling of cloud resources, demonstrating high engineering practical value.

[0237] (4) Personality traits are traceable and controllable

[0238] Thanks to the introduction of version management and logging mechanisms, the evolution of the digital avatar personality in this invention is transparent and traceable. When an anomaly occurs (e.g., an output style deviating from expectations), the cause can be quickly located through logs (e.g., misjudgment of emotion recognition, mapping rule conflicts, or overly strict threshold settings), and the version rollback function can be used to restore the personality configuration to a normal state, thereby reducing the risk of personality "going out of control." Furthermore, the system allows developers and users to fine-tune personality parameters: developers can adjust the mapping function coefficients (e.g., ... , , , , , , Consistency threshold The output effect can be fine-tuned by adjusting the rule base parameters; professional users can also directly modify the personality vector and its mapping parameters to shape their ideal digital avatar personality. This high controllability ensures that the personality-driven multimodal output conforms to the algorithm logic and respects the user's wishes, making the digital avatar a truly "shapeable and manageable virtual personality".

[0239] The core innovation of this invention lies in proposing a unified personality vector bus mechanism with a personality central module at its core, which uses basic personality vectors... Emotional skew With the current personality vector Personality parameters, such as those related to text, speech, facial expressions, and actions, are integrated throughout the entire generation process to avoid personality fragmentation caused by separate generation of multiple modalities. A text-based sentiment vector-based approach is proposed. Voice emotion vector and facial expression vectors Equal sentiment vectors are used to determine the consistency score. Compatibility with personality A dual-verification mechanism is implemented, with threshold-based fine-tuning or regeneration triggered in stages, forming an executable "generation-verification-adjustment" closed-loop control. This is further integrated with role-setting fields. Output the large character model and the current personality vector. The linkage enables digital avatars to maintain consistency in their roles and multimodal expression styles across multiple interactions; and through long-term personality optimization and version rollback mechanisms, the personality becomes evolvable, traceable, and controllable.

[0240] In summary, this invention significantly surpasses existing solutions in terms of logical structure, content completeness, and technological innovation. It not only addresses the pain point of multimodal incoordination in existing digital avatars, achieving unified personality control and consistent modal output, but also, through dynamic adaptation and refined algorithm design, endows digital avatars with more intelligent and human-like interactive capabilities, possessing broad application prospects and commercial value.

[0241] To more clearly illustrate the implementation of the present invention, this section describes the above technical solution through specific embodiments, including a numerical embodiment, several alternatives, and system deployment methods. It should be understood that the following embodiments are used to help understand the present invention, but should not be used to limit the scope of protection of the present invention.

[0242] Let's define an application scenario—a digital avatar acts as a user's virtual friend, with a personality set as extroverted and approachable (enthusiastic, cheerful, friendly, and talkative). This embodiment will demonstrate the specific working process of the system in this scenario, including how to calculate personality updates, regulate multimodal output, and perform consistency checks.

[0243] (1) Set basic parameters

[0244] Personality Dimensions Definition: Using the Big Five personality model, including extraversion ( ), affinity ( ), due diligence ( ), neuroticism ) and openness Five dimensions.

[0245] Basic personality vector The initial personality traits were set as follows: Extraversion 0.9, Agreeableness 0.8, Conscientiousness 0.6, Neuroticism 0.3, and Openness 0.7. This indicates that the initial personality of the digital clone is very extroverted and friendly, relatively responsible, with moderate to high emotional stability and an open mind.

[0246] Scene adaptation weight This is a casual chat scenario. (Settings) In contrast, it strengthens the influence of affinity (1.1) and slightly reduces the influence of conscientiousness (0.9) and neuroticism (0.8) to highlight a relaxed and friendly interactive atmosphere.

[0247] First rating threshold Second scoring threshold :set up A modal consistency score of ≥0.7 is required to be considered as style consistent.

[0248] Personality optimization coefficient: (Every 100 rounds of interaction) One optimization, with minor adjustments.

[0249] Baseline values ​​for voice control parameters: Baseline value for speech rate Words per minute, baseband reference value Hz, the baseline value for timbre (0 represents very soft, 1 represents very bright).

[0250] Baseline values ​​for animation control parameters: Baseline values ​​for facial expression intensity (Smile intensity 60%, considered a slight smile), baseline value for movement frequency. The baseline value for the movement rhythm is times per minute (approximately once every 30 seconds for a nod). Seconds per nod (a single nod lasts 1.5 seconds).

[0251] (2) Interaction process

[0252] (2A) Obtain user input information and perform emotion recognition on the user input information to determine the emotion label; determine the corresponding emotion intensity vector based on the emotion label. For example, a user sends a message to a digital clone via text: "It was so nice chatting with you today!" The emotion recognition module 601 analyzes this sentence and identifies the user's emotion as "happy", with an intensity of 0.8 and a timestamp of 2025-05-20 10:00:00.123 (year-month-day hour:minute:second.millisecond). This emotion label is sent to the personality center module 602 so that the personality center module 602 can determine the emotion intensity vector based on the emotion label. .

[0253] (2B) Calculate the emotional offset : Personality Central Module 602 searches for the emotion-personality dimension weight matrix corresponding to the emotion of "happiness". The column vector (assuming the "happiness" column is one of the six basic emotions): This indicates the weighting of positive emotions on personality dimensions: Extraversion +0.2, Agreeableness +0.3, Conscientiousness -0.1, Neuroticism -0.2, Openness +0.1. Combined with the emotion intensity vector... Emotional attenuation coefficient Current scene weight Dimensional calculation :

[0254] Extraversion dimension:

[0255] Affinity dimension:

[0256] Due diligence dimension:

[0257] Neuroticism dimension:

[0258] Openness dimension:

[0259] In summary, the calculated sentiment shift... All values ​​are within the range of [-0.3, 0.3], and no truncation is needed. It can be seen that the user's display of happiness subtly adjusts the digital clone's personality state towards a more extroverted and friendly direction, while slightly reducing seriousness, conscientiousness, and tension. This is logical: when the other party is happy, the digital clone becomes more enthusiastic and easygoing.

[0260] (2C) Update the current personality vector : Adjust the above emotional shift Add to the basic personality vector The current personality vector is obtained by limiting the range. :

[0261] Extraversion: (Since 0.9 + 0.144 exceeds 1, it is truncated to 1.0)

[0262] Affinity: (Truncation to 1.0)

[0263] Due diligence:

[0264] Neuroticism:

[0265] Openness:

[0266] In summary, the calculated current personality vector Compared to the baseline personality vector, extraversion and agreeableness both rise to the upper limit of 1.0 (extremely extroverted and enthusiastic, very friendly), conscientiousness slightly decreases to 0.5352 (slightly more easygoing and agreeable), neuroticism decreases to 0.1848 (more peaceful and relaxed), and openness slightly increases to 0.7720 (slightly more active and curious). The personality central module 602 caches this updated vector and pushes it to the text, voice, and emoji modules.

[0267] (2D) Target Response Text Generation: Character Large Model Text Generation Module 603 receives the character setting field at this time. Dialogue context and current personality vector The system begins generating the target response text for the digital clone. First, it is based on user input and the dialogue context. The pre-trained LLM is invoked to generate the initial response text: "It's great to hear you're happy, I also enjoy chatting with you." Then, the character large model text generation module 603 generates the text based on the current response text. The initial response text is then refined: Extraversion = 1.0, Affinity = 1.0, falling into the high extraversion and high affinity ranges, respectively. Therefore, the corresponding high extraversion and high affinity vocabulary libraries are loaded. The high extraversion library contains passionate words such as "super happy" and "fantastic," while the high affinity library contains affectionate terms and interjections such as "ne," "ya," and "~." The character model text generation module 603 modifies the wording of the initial response text, replacing "very happy" with "so happy," breaking down straightforward sentences and adding interjections, and applying high extraversion sentence structure rules (preferring the use of exclamation marks and elongated sounds). The final target response text is: "So happy! Chatting with you is so much fun, I'll share again next time~." This response uses exclamation marks and onomatopoeic words, contains positive words such as "so happy" and "super happy," and the ending "ya~," fully reflecting the digital avatar's current extroverted, enthusiastic, and friendly personality. The character model text generation module 60 then... The data is sent to the multimodal synchronization module 606 for temporary storage, in preparation for subsequent consistency checks.

[0268] (2E) Speech signal synthesis: The speech synthesis module 604 receives the target response text. "I'm so happy! Chatting with you is so much fun, I want to share again next time~" and current personality vector The speech synthesis module 604 first determines the speech based on the current personality vector. Calculate the target values ​​for the voice control parameters.

[0269] Assume the following mapping coefficients are obtained in advance through training:

[0270] Linear coefficient of speech rate (Extroversion +20, Affinity +10, Conscientiousness +5, Neuroticism -5, Openness +8)

[0271] Linearity coefficient of fundamental frequency (Extroversion +30 Hz, Affinity +20 Hz, Conscientiousness +10 Hz, Neuroticism -15 Hz, Openness +15 Hz)

[0272] Nonlinear coefficient of speech rate (This scenario assumes that no special handling is done when there is an extreme value, so we temporarily take 0.)

[0273] Nonlinear coefficient of fundamental frequency (Same as above)

[0274] The linear coefficient of timbre (These correspond to the effects of extraversion, affinity, conscientiousness, neuroticism, and openness on timbre brightness, respectively), and the calculation process is as follows:

[0275]

[0276] It has increased by 25% compared to the baseline of 150 words per minute, indicating that the digital avatar's speaking speed has significantly increased at this moment, which is in line with its happy and excited state.

[0277]

[0278] It has increased significantly (by 32%) compared to the baseline of 200Hz, meaning that the fundamental frequency is significantly higher and more lively and shrill.

[0279]

[0280] The target value of the timbre exceeds the upper limit of the defined range by 1.0, so it is truncated to 1.0 (indicating the brightest and clearest timbre).

[0281] After obtaining the above speech style parameters, the speech synthesis module 604 constructs a speech style vector , and inputs it together with the target reply text into a neural network TTS model such as VITS. The TTS model generates a target reply voice with specified speaking speed, fundamental frequency, and timbre characteristics , and exports the target reply voice The corresponding phoneme timestamps . For example, the pronunciation of "So happy!" at the beginning of the text lasts for about 0.5 seconds. Among them, "So" corresponds to the phoneme / tai / from 0.10s to 0.25s, "happy" / kai / from 0.25s to 0.40s, "heart" / xin / from 0.40s to 0.50s, and "la" as a modal particle is连读extended with the previous phoneme to 0.50s. The whole sentence "Chatting with you is extremely happy" lasts for 1.5 seconds, and the final modal particle "ya~" occupies 0.2 seconds and fades out gradually, making the total duration of the whole paragraph of speech about 2.2 seconds. Finally, the speech synthesis module 604 outputs the target reply voice (PCM data stream with a sampling rate of 16kHz, length about 2.2 seconds) and the corresponding phoneme timestamps , the target reply voice is sent to the audio playback queue, and the phoneme timestamps are handed over to the multimodal synchronization module 606 for animation alignment.

[0282] (2F) Generation of expression and action animations: The expression animation module 605 also starts to work at this time, using the target reply text , the target reply voice and the current personality vector to determine the target expression animation , and the calculation process is as follows:

[0283] Basic motion planning: Since the reply expresses a happy emotion, the facial animation module 605 decides to use a "smiling" emoticon throughout the sentence, and add a cheerful nod at the end of the sentence to show affirmation or excitement. The text detects two sentences with exclamation marks, indicating emotional excitement; therefore, a quick nod will be added at the end of each of the two sentences corresponding to the voice to reinforce the expression.

[0284] Animation parameter adjustment: Target value for smile expression intensity The baseline value for the intensity of a smiling expression.

[0285] Use coefficient ,

[0286] calculate ;

[0287] If the value exceeds the upper limit of 1, it is truncated to 1.0 (maximum smile).

[0288] Target value of action frequency : Baseline value of action frequency times / minute;

[0289] Use coefficient ,

[0290] calculate times / minute;

[0291] Approximately 9 times per minute (an average of one nod every 6.4 seconds). Considering that the voice message is only 2.2 seconds long, only 2 nods were actually performed (once at the end of each of the two sentences).

[0292] Target value of action rhythm The baseline value of movement rhythm seconds / time;

[0293] Use coefficient ,calculate seconds / time;

[0294] coefficient Each nod takes approximately 0.32 seconds, and the rhythm is rapid.

[0295] Target value for smoothness of facial expression transitions The baseline value for the smoothness of facial expression transitions. ;

[0296] Use coefficient ;

[0297] calculate ;

[0298] Compared to the benchmark, the facial expressions transition more naturally.

[0299] Animation generation: The rendering engine of the Facial Animation Module 605 (such as Unity3D) generates animations based on phoneme timestamps. The animation is driven by the following mouth shapes: at 0.10s, the mouth opens wide to produce the "tai" sound; at 0.50s, the mouth closes slightly to reveal a smile as the "la" sound ends; during the pause between 0.50s and 0.70s, the smile is maintained and a nod is performed (the head lowers at 0.55s and returns to its original position at 0.87s); a second nod is inserted at the end of the second sentence (approximately 1.8s) (completed within 0.18 seconds); during the drawn-out "ya~" sound, the digital clone blinks slightly and tilts its head. The entire animation process is closely synchronized with the speech, ensuring that the mouth shape matches the sound, and that the facial expressions and movements are semantically consistent, ultimately generating the target facial expression animation. .

[0300] (2G) Multimodal Synchronous Output and Verification:

[0301] Timing Alignment: Multimodal Synchronization Module 606 receives target response voice. Corresponding phoneme timestamp And target facial animation Afterwards, it is necessary to control the target to reply with voice. Corresponding phoneme timestamp And target facial animation The timeline is synchronized, with only minor adjustments made (such as ensuring the animation remains still during voice pauses).

[0302] Consistency score calculation:

[0303] Emotional Vector Extraction: Speech Emotional Vectors (90% happiness, 10% calmness), text sentiment vector (95% happiness, 5% calmness), emotional expression vector (98% happiness, 2% peace).

[0304] Pairwise cosine similarity calculation: , , .

[0305] Calculate the consistency score It is far above the first scoring threshold. The decision was approved.

[0306] Personality compatibility verification: Expected emotional vector (90% happiness + 10% peace);

[0307] Comprehensive sentiment vector ;

[0308] Personality compatibility The personality compatibility is greater than the preset compatibility. It's a perfect match.

[0309] (2H) Output Presentation: The multimodal synchronization module 606 will display the verified target response text. Target reply voice Target facial animation Combined output. Users see the reply text "So happy! Chatting with you is so much fun, I'll share again next time~", hear a bright and cheerful voice, and at the same time, the virtual friend on the screen has a bright smile on his face, nodding excitedly as he speaks. The actions and voices are perfectly matched, providing a highly immersive interactive experience.

[0310] The core idea of ​​this invention is to achieve consistent multimodal output under unified personality control. Around this core, each functional module can have different implementation variations to adapt to different application requirements and technical conditions. The following lists several preferred and alternative solutions for different types of modules:

[0311] Personality Central Module 602

[0312] Preferred solution: A solution combining emotional shift overlay and long-term optimization to achieve automatic learning in order to adjust the basic personality vector.

[0313] Alternative Solution 1: Reinforcement Learning Update. This involves modeling personality state updates as a reinforcement learning task, using long-term user satisfaction as the reward, and dynamically adjusting the base personality vector using DQN or policy gradient methods. While it can uncover personality regulation strategies, it is complex to implement and requires extensive training with interactive data.

[0314] Alternative Option 2: Manual User Update. Provide users with an open personality control interface, offering sliders or options for direct adjustment of the base personality vector (e.g., "reduce extraversion"). After the user adjusts, the system immediately regenerates a response based on the new personality. Users have complete control over the personality, but this requires some operational effort and is suitable for personalized virtual partners.

[0315] Character Large Model Text Generation Module 603

[0316] The preferred approach is a pre-trained large-scale language model (LLM) combined with personality parameter injection and rule correction. This approach achieves the best results by integrating the fluency of language generated through deep learning with the consistency of personality controlled by rules.

[0317] Alternative Solution 1: Prompt Engineering. This involves converting personality parameters into dialogue prompts, supplemented with prefixes or few-sample examples, to guide a general LLM in generating responses that match the personality style. For example, constructing a prompt like, "You are an extroverted and friendly virtual friend who speaks warmly and affectionately:", requires no modification to the model's internals, but demands high precision in prompt design.

[0318] Alternative Option 2: Personality-Specific Model Fine-Tuning. Collect a large amount of dialogue data (such as lines from film and television characters) tailored to a specific personality style, and fine-tune it on a medium-sized model (such as LLaMA-7B) to teach the model the personality's speaking style. This reduces rule intervention but requires training different models for each personality, making it more costly and suitable for dialogues with specific IP characters.

[0319] Speech Synthesis Module 604

[0320] Preferred solution: Neural network TTS + personality style vector injection, directly generating styled speech through an end-to-end model. .

[0321] Alternative Solution 1: Speech Style Transfer. First, neutral speech is synthesized, then speech conversion technologies (such as CycleGAN and AutoVC) are used to transfer the neutral speech to the speech style of the target personality. This is suitable for scenarios with limited device computing power, reducing the real-time computational burden through a two-stage processing approach.

[0322] Alternative Solution 2: Multi-style speech library concatenation. Pre-record or synthesize a series of short speech segments (such as greetings and common responses) for various personality styles, and output the target response text when needed. It performs template matching and audio splicing with extremely low latency, making it suitable for short, fixed sentence scenarios.

[0323] Facial Animation Module 605

[0324] Preferred solution: A traditional 3D animation approach combining skeletal animation and BlendShape, where personalized movements are driven by adjusting parameters. .

[0325] Alternative Solution 1: End-to-end dynamic image generation. Utilizing a diffusion model or GAN, the input target text is the response text. Target reply voice and current personality vector It can generate corresponding 2D animation sequences or 3D motion sequences in one go. It can represent complex facial expressions and subtle movements, but it requires a large amount of computation and is suitable for pre-rendering high-quality animations.

[0326] Alternative Solution 2: Motion Capture Data Adaptation. Utilize footage of live actors performing different personality styles (obtaining skeletal animation data through motion capture), and adapt it to the current personality vector. The amplitude and frequency of these motion data are adjusted. Based on real data, the movements are highly realistic, making them suitable for high-fidelity scenarios such as virtual anchors and digital avatars.

[0327] Multimodal synchronization module 606

[0328] Preferred solution: Phoneme timestamp Alignment and algorithm verification enable automatic synchronization.

[0329] Alternative Solution 1: Audio-driven video frame synchronization. Extract the time-frequency energy value of the speech and align it with the frame rate of change of facial expression animation (e.g., speed up the action when the audio energy is high, and pause the action when there is silence). No text information is required, and it is suitable for speech or real-time audio input without time stamps.

[0330] Alternative Solution 2: Manual Rule Validation. A pre-defined list of cross-modal consistency rules is used (e.g., "When expressing 'happiness,' the voice fundamental frequency rises, the facial expression smiles, and exclamation appears in the text"). If these rules are not met, corrections are executed. This approach is easy to implement but not comprehensive enough, and is suitable for virtual customer service representatives with fixed emotional categories.

[0331] All of the above alternative solutions are feasible under certain conditions and fall within the protection scope of this invention. This invention is not limited to the preferred embodiments described above; any of the above alternative solutions or combinations thereof can be used to achieve the same core objective, depending on the requirements.

[0332] The system of this invention can be flexibly deployed in different environments to meet the requirements of different application scenarios for response latency and computing resources. It mainly has the following two typical architectures:

[0333] (1) Local deployment mode

[0334] Deployment method: All modules are deployed on the user's local device and run, including the personality central module 602, the character large model text generation module 603, the speech synthesis module 604, the facial expression animation module 605, and the multimodal synchronization module 606.

[0335] Applicable scenarios: Scenarios with extremely high requirements for real-time interaction or unstable network, such as offline voice assistants on mobile phones, local virtual passenger assistants in unmanned vehicles, and exhibition hall guide robots that need to operate offline.

[0336] Architectural features: Personality parameters ( , , and threshold All data, including model resources, is stored locally to avoid network transmission latency. The entire process of user input processing, multimodal generation, and consistency verification is completed locally. Typically, system initialization time is approximately 2 seconds, single-round text dialogue latency can be controlled to around 0.5 seconds, and single-round voice dialogue latency can be controlled to within 1 second. All data remains locally, improving privacy, security, and usability.

[0337] Resource Requirements: A certain level of computing power is required for the terminal device. PCs need at least a 4-core CPU, 8GB of RAM, and an entry-level GPU; mobile devices can reduce model size through model distillation and quantization to adapt to the AI ​​acceleration module of the mobile SoC. Within the limits of computing power, local deployment provides the best real-time interactive experience.

[0338] (2) Cloud-Terminal Collaboration Mode

[0339] Deployment method: Modules with high computational demands are deployed on cloud servers, while the terminal only runs lightweight modules and is responsible for the user interface. The two work together through network communication. Typically, the personality central module 602, the character large model text generation module 603, and the speech synthesis module 604 are located in the cloud, while the facial expression animation module 605, the multimodal synchronization module 606, and the user UI run locally on the terminal.

[0340] Applicable scenarios: Situations requiring high-performance computing or serving a large number of users, such as virtual customer service systems with tens of thousands of concurrent users, cloud gaming virtual anchors with complex facial expression modeling, and hardware-limited AR / VR terminals.

[0341] Architectural features:

[0342] Communication and latency: The terminal uploads user input (text or voice, pre-processed and compressed) to the cloud, and the cloud calculates the target response text. Target reply voice and key animation parameters (phoneme timestamps) The system sends the facial expression sequence back to the terminal. It uses HTTP / 2 or WebSocket persistent connections, compresses the audio into MP3 / OGG ​​format, and keeps the round-trip latency within 300 milliseconds (server response 100ms + network transmission 200ms). After receiving the data, the terminal locally renders the target facial expression animation. It also plays audio, providing a real-time experience similar to local deployment.

[0343] Fault tolerance mechanism: When the terminal detects a timeout in the cloud response or a connection interruption, it automatically switches to local degradation mode and calls the basic personality vector stored on the terminal. The lightweight response model continues to provide core functionalities (such as text mode), and the interface displays network status information to the user. Once cloud services are restored, it will switch back to cloud mode to ensure service continuity and reliability.

[0344] In summary, whether for small-scale applications on a single device or large-scale online services, the system architecture of this invention can be flexibly deployed, leveraging the powerful computing capabilities of the cloud while utilizing the terminal's rapid rendering and caching capabilities to provide a stable and efficient user experience.

[0345] This invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the personality-driven role-model multimodal digital avatar interaction method described in the above embodiments. To avoid repetition, it will not be described again here.

[0346] This invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the personality-driven role-model multimodal digital avatar interaction method described in the above embodiments. To avoid repetition, it will not be described again here.

[0347] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A personality-driven, large-scale role model, multimodal digital avatar interaction method, characterized in that, include: Emotion recognition is performed on user input information collected during the interaction between the digital avatar and the user to determine emotion tags. Each emotion tag includes m emotion types and the corresponding emotion intensity for each emotion type. The range of the emotion intensity is... ; The emotion offset is determined based on the emotion label, and the emotion offset is determined according to the following formula: ;in, For emotional offset; This is a weight matrix for the emotion-personality dimension. n represents the number of personality dimensions; m represents the number of emotion types. Let the emotion intensity vector be... , Determined based on the intensity of the emotion corresponding to multiple emotion types; This is the emotional attenuation coefficient; Assign scene-adaptive weights; determine the current personality vector corresponding to the digital clone based on the emotion offset and the base personality vector corresponding to the digital clone; A large role model is used to process the role setting fields, dialogue context, and the current personality vector to determine the target response text that represents personality style. The large role model is injected with the role setting fields and the current personality vector during the inference period to constrain the target response text to maintain role consistency and personality consistency in multiple rounds of dialogue. Based on the current personality vector, a target value for at least one voice control parameter is determined. Based on the target value of at least one voice control parameter, a voice style vector is determined. The voice style vector is injected into a text-to-speech model. The text-to-speech model is used to synthesize the target response text into speech, and a target response speech representing the personality style is determined. Based on the current personality vector, a target value for at least one animation control parameter is determined. Based on the target value of at least one animation control parameter, an animation style vector is determined. The animation style vector is injected into an animation generation model. The animation generation model is used to process the target response text and the target response speech to determine the target facial expression animation that represents the personality style. Perform multimodal synchronization and consistency checks on the target response text, the target response voice, and the target facial expression animation, and output the target response text, target response voice, and target facial expression animation that pass the checks.

2. The method according to claim 1, characterized in that, The current personality vector is determined according to the following formula: ; Based on the basic personality vector, Let i be the current personality vector, and let i be the index of the personality dimension. min and max form the interval truncation to ensure that the current personality vector is within the specified range. Within the range.

3. The method according to claim 1, characterized in that, The personality-driven, large-scale role model multimodal digital avatar interaction method also includes: use Update the current dialogue round number, where t is the current dialogue round number; If the current number of dialogue rounds is greater than the preset number of dialogue rounds corresponding to the optimization period, the average offset is determined based on multiple emotion offsets corresponding to the optimization period. Based on the last emotional offset and the average offset in the optimization cycle, the basic personality vector is updated. The update formula for the basic personality vector is as follows: ,in, Let be the basic personality vector at time k+1. Let k be the basic personality vector at time k. For personality optimization coefficients, This represents the average offset.

4. The method according to claim 1, characterized in that, The character model includes a natural language model, a vocabulary layer, a sentence structure layer, and a catchphrase layer; The process of using a large character model to process character setting fields, dialogue context, and the current personality vector to determine the target response text representing personality style includes: The natural language model is used to process the role setting fields and the dialogue context to determine the initial response text; The initial response text is adjusted using the emotional vocabulary corresponding to the current personality vector at the vocabulary layer; the initial response text is adjusted using the sentence style corresponding to the current personality vector at the sentence structure layer; the initial response text is adjusted using the catchphrases corresponding to the current personality vector at the catchphrase layer; the outputs of the vocabulary layer, the sentence structure layer, and the catchphrase layer are fused to determine the target response text.

5. The method according to claim 1, characterized in that, The voice control parameters include at least one of speech rate and fundamental frequency; Determining the target value of at least one voice control parameter based on the current personality vector includes: When the parameter value of the i-th personality dimension in the current personality vector is within a preset range, the target values ​​of speech rate and fundamental frequency are determined based on a linear mapping function, wherein the linear mapping function includes: , ; When the parameter value of the i-th personality dimension in the current personality vector is not within a preset range, the target values ​​of speech rate and fundamental frequency are determined based on a nonlinear mapping function, wherein the nonlinear mapping function includes: , ; in, The target value for speech rate, This serves as the baseline value for speech rate. Let be the linear coefficient of the i-th personality dimension with respect to speech rate. Let be the non-linear coefficient of the i-th personality dimension with respect to speech rate. The target value for the baseband. This is the reference value for the fundamental frequency. Let be the linear coefficient of the i-th personality dimension with respect to the fundamental frequency. Let be the nonlinear coefficient of the i-th personality dimension with respect to the fundamental frequency. represents the parameter value of the i-th personality dimension in the current personality vector.

6. The method according to claim 1 or 5, characterized in that, The voice control parameters also include timbre; The step of determining the target value of at least one voice control parameter based on the current personality vector further includes: The target value of timbre is determined based on a linear mapping function, which includes: ; in, The target value for timbre, As the baseline value for timbre, Let be the linear coefficient of the i-th personality dimension with respect to timbre. This is the minimum value for timbre. This represents the maximum value of the timbre. Let be the parameter value of the i-th personality dimension in the current personality vector. This is a function that cuts off the interval.

7. The method according to claim 1, characterized in that, The animation control parameters include at least one of the following: facial expression intensity, movement amplitude, movement frequency, movement rhythm, and facial expression transition smoothness. The target value of at least one animation control parameter is determined by the following formula: ; ; ; ; ; in, The target value for facial expression intensity. This serves as a baseline value for facial expression intensity. is the linear coefficient of the i-th personality dimension with respect to facial expression intensity; The target value for the range of motion, This serves as a baseline value for the range of motion. Let be the linear coefficient of the i-th personality dimension with respect to the amplitude of movement. This represents the minimum range of motion. This represents the maximum range of motion. The target value for the action frequency, This serves as a baseline value for the action frequency. Let be the linear coefficient of the i-th personality dimension with respect to action frequency. This represents the minimum frequency of the action. This represents the maximum value of the action frequency; The target value for the rhythm of the movement. This serves as the baseline value for the rhythm of the movement. Let be the linear coefficient of the i-th personality dimension with respect to the rhythm of action. This represents the minimum value of the action rhythm. This represents the maximum value of the action rhythm; The target value for the smoothness of facial expression transitions. This serves as the baseline value for the smoothness of facial expression transitions. is the linear coefficient of the i-th personality dimension on the smoothness of expression switching; Let be the parameter value of the i-th personality dimension in the current personality vector. This is a function that cuts off the interval.

8. The method according to claim 1, characterized in that, The step of performing multimodal synchronization and consistency verification on the target reply text, the target reply voice, and the target facial expression animation, and outputting the verified target reply text, target reply voice, and target facial expression animation, includes: Emotional features are extracted from the target response text, the target response voice, and the target facial animation to determine the text emotion vector, the voice emotion vector, and the facial emotion vector; Intermodal similarity is calculated for text sentiment vectors, speech sentiment vectors, and facial expression sentiment vectors to determine consistency scores; When the consistency score is greater than the first score threshold, the target response text, target response voice, and target facial animation that have passed the verification are output. When the consistency score is not greater than the first score threshold and the consistency score is greater than the second score threshold, a slight adjustment is triggered to adjust at least one of the target values ​​of the voice control parameters and the target values ​​of the animation control parameters, and to repeatedly perform multimodal synchronization and consistency verification on the target response text, the target response voice and the target facial expression animation. When the consistency score is not greater than the second score threshold, regeneration is triggered to regenerate at least one of the target response text, the target response voice, and the target facial expression animation, and the multimodal synchronization and consistency verification of the target response text, the target response voice, and the target facial expression animation are repeatedly performed.

9. The method according to claim 8, characterized in that, The consistency score is determined using the following formula: ; ; ; ; in, This is the consistency score. For text sentiment vectors; For speech emotion vectors, For facial expression and emotion vectors; Cosine similarity between speech sentiment vector and text sentiment vector; The cosine similarity between the speech emotion vector and the facial expression emotion vector; The cosine similarity between the facial expression sentiment vector and the text sentiment vector; , and Three pre-set weighting coefficients.

10. The method according to claim 8, characterized in that, When the consistency score is greater than the first score threshold, the target response text, target response voice, and target facial animation that have passed the verification are output, including: When the consistency score is greater than the first score threshold, a comprehensive emotion vector is determined based on the text emotion vector, the voice emotion vector, and the facial expression emotion vector. Based on the current personality vector, a pre-set personality emotion mapping relationship is queried to determine the expected emotion vector corresponding to the current personality vector; The similarity between the comprehensive emotion vector and the expected emotion vector is calculated to determine the personality compatibility. When the personality compatibility is less than a preset compatibility, the overall style of the target reply text, the target reply voice, and the target facial expression animation is adjusted until the personality compatibility is not less than the preset compatibility, and then the verified target reply text, target reply voice, and target facial expression animation are output.

11. A personality-driven, large-scale, multimodal digital avatar interaction system, characterized in that, include: The emotion recognition module is used to identify emotions in user input information collected during the interaction between the digital avatar and the user, and to determine emotion tags. The emotion tags include m emotion types and the corresponding emotion intensity for each emotion type. The value range of the emotion intensity is... ; The personality center module is used to determine an emotion offset based on the emotion label, and the emotion offset is determined according to the following formula: ;in, For emotional offset; This is a weight matrix for the emotion-personality dimension. n represents the number of personality dimensions; m represents the number of emotion types. Let the emotion intensity vector be... , Determined based on the intensity of the emotion corresponding to multiple emotion types; This is the emotional attenuation coefficient; Assign scene-adaptive weights; determine the current personality vector corresponding to the digital clone based on the emotion offset and the base personality vector corresponding to the digital clone; The character big model text generation module is used to process the character setting fields, dialogue context and the current personality vector using the character big model to determine the target response text that represents the personality style. The character big model is injected with the character setting fields and the current personality vector during the inference period to constrain the target response text to maintain character consistency and personality consistency in multiple rounds of dialogue. The speech synthesis module is used to determine the target value of at least one speech control parameter based on the current personality vector, determine the speech style vector based on the target value of at least one speech control parameter, inject the speech style vector into the text-to-speech model, use the text-to-speech model to perform speech synthesis on the target response text, and determine the target response speech that represents the personality style. The facial expression animation module is used to determine the target value of at least one animation control parameter based on the current personality vector, determine the animation style vector based on the target value of at least one animation control parameter, inject the animation style vector into the animation generation model, and use the animation generation model to process the target response text and the target response speech to determine the target facial expression animation representing the personality style. The multimodal synchronization module is used to perform multimodal synchronization and consistency verification on the target reply text, the target reply voice, and the target facial expression animation, and output the target reply text, target reply voice, and target facial expression animation that pass the verification.

12. The system according to claim 11, characterized in that, The personality central module determines the emotion offset based on emotion tags and updates the current personality vector; The emotional offset and the current personality vector at least satisfy the following: , ; in, For emotional offset; This is a weight matrix for the emotion-personality dimension. Let the emotion intensity vector be... , Determined based on the emotional intensity corresponding to multiple emotional types, where m is the number of emotional types. This is the emotional attenuation coefficient. The range of values ​​is ; Adapt weights to different scenarios. The range of values ​​is ; Based on the basic personality vector, This represents the current personality vector. This is a function that cuts off the interval.

13. The system according to claim 11, characterized in that, The personality central module also includes a personality version number field and a status update log field, which are used to record the update history of the basic personality vector and the current personality vector and support the rollback of personality parameters.

14. The system according to claim 11, characterized in that, The speech synthesis module maps the current personality vector to speech control parameters, which include at least one of speech rate, fundamental frequency and timbre. The target values ​​of the speech rate and the base frequency are obtained by a segmentation strategy that uses linear mapping within a preset range and nonlinear mapping outside the preset range. The target value of the timbre is obtained by linear mapping and interval truncation.

15. The system according to claim 11, characterized in that, The facial expression animation module maps the current personality vector to animation control parameters, which include at least one of facial expression intensity, movement amplitude, movement frequency, movement rhythm, and facial expression switching smoothness. Boundary truncation is performed on each animation control parameter to suppress exaggerated movements or abrupt changes in rhythm caused by extreme personality values.

16. The system according to claim 11, characterized in that, The multimodal synchronization module calculates a consistency score based on text emotion vectors, speech emotion vectors, and facial expression emotion vectors, and sets a first scoring threshold and a second scoring threshold. When the consistency score is greater than the first scoring threshold, it outputs the verified target response text, target response speech, and target facial expression animation. When the consistency score is not greater than the first scoring threshold but greater than the second scoring threshold, it triggers a slight adjustment, adjusting at least one of the target values ​​of the speech control parameters and the animation control parameters. When the consistency score is not greater than the second scoring threshold, it triggers regeneration, regenerating at least one of the target response text, the target response speech, and the target facial expression animation.

17. The system according to claim 11, characterized in that, The multimodal synchronization module determines the expected emotion vector based on the current personality vector, and calculates the personality fit based on the similarity between the comprehensive emotion vector and the expected emotion vector; when the personality fit is less than the preset fit threshold, the overall style is adjusted.

18. The system according to claim 11, characterized in that, The personality central module performs long-term optimization of the basic personality vector according to the optimization cycle. The update of the basic personality vector includes at least: calculating the average offset of the emotion offset within the optimization cycle and superimposing it on the basic personality vector according to the personality optimization coefficient.