An intelligent interaction and decision-making method and system integrated with an AI virtual image

By using multi-dimensional interactive data encryption transmission and emotion-intention joint reasoning, combined with a multi-level Bayesian decision framework, the problems of insufficient data transmission security, disconnect between emotion and intention understanding, poor decision reliability, and insufficient personalized interaction in AI virtual avatar systems are solved, achieving a highly secure, accurate, and personalized interactive experience.

CN122472219APending Publication Date: 2026-07-28LOOTOM TELCOVIDEO NETWORK WUXI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LOOTOM TELCOVIDEO NETWORK WUXI
Filing Date
2026-06-29
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Existing AI virtual avatar systems suffer from insufficient data transmission security, a disconnect between emotion and intent understanding, poor decision-making reliability and lack of interpretability, and insufficient personalized interaction capabilities.

Method used

It employs multi-dimensional interactive data encryption transmission, sentiment-intent joint reasoning, and a multi-level Bayesian decision framework, combined with adaptive hybrid encryption, a dual-stream Transformer architecture, and a cross-attention mechanism, to achieve dynamic security level adjustment and personalized interaction.

Benefits of technology

It improves data transmission security, the accuracy of emotion-intent understanding, the reliability of decision-making, and the personalized interactive experience, meeting the requirements of real-time interactive performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122472219A_ABST
    Figure CN122472219A_ABST
Patent Text Reader

Abstract

The application relates to the field of artificial intelligence and man-machine interaction technology, and particularly discloses an intelligent interaction and decision method and system integrated with an AI virtual image, which comprises the following steps: acquiring multi-dimensional interaction data of a user through a client; transmitting the multi-dimensional interaction data to a server end through the client; performing emotion-intention joint reasoning according to the multi-dimensional interaction data through the server end to obtain an emotion-intention joint vector of the user; executing intelligent decision based on a multi-level Bayesian decision framework according to the emotion-intention joint vector of the user through the server end to obtain an optimal response strategy; outputting multi-modal response information on the client according to the optimal response strategy through the server end, and then receiving user feedback information on the client. The application can improve the security of adaptive hybrid encryption transmission, the accuracy of emotion-intention joint reasoning, the reliability and explainability of multi-level Bayesian decision, and the personalized interaction experience based on voiceprint recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and human-computer interaction technology, and in particular to an intelligent interaction and decision-making method and a system integrating AI virtual avatars. Background Technology

[0002] With the rapid development of artificial intelligence technology, AI virtual avatars have been widely used in customer service, education, healthcare, entertainment, and other fields. However, existing technologies have shortcomings in the following aspects: First, regarding secure data transmission, existing AI virtual avatar systems typically use the traditional TLS / SSL protocol for communication encryption. While this method provides transport layer security, it has the following problems: (1) Fixed encryption mode cannot dynamically adjust security policies according to the threat level of the network environment; (2) The encrypted data packets have obvious structural features and are easily identified and located by traffic analysis tools; (3) The lack of a fine-grained data sensitivity classification mechanism leads to all data being encrypted at the same level, which wastes computing resources and fails to provide enhanced protection for highly sensitive data.

[0003] Second, regarding interaction logic, existing emotion recognition and intent understanding are usually performed independently, neglecting the deep connection between emotional state and user intent. For example, the same phrase "I don't need it anymore" might express disappointment in a negative emotional state, while in a calm emotional state it might simply mean rejection. The lack of joint modeling of emotion and intent leads to insufficient accuracy and naturalness in interaction responses. Furthermore, state tracking in multi-turn dialogues often uses fixed-length context windows, which cannot effectively handle long-range dependencies and topic switching.

[0004] Third, in terms of decision-making, existing AI virtual avatar decision-making systems mostly employ rule-based decision trees or simple reinforcement learning methods. Rule-based systems lack flexibility and struggle to cope with complex and ever-changing interaction scenarios; while traditional reinforcement learning methods require large amounts of training data, and the decision-making process lacks interpretability. Furthermore, existing systems typically do not consider the safety boundary constraints of the policy when making decisions, which may lead to the generation of inappropriate or harmful responses.

[0005] Fourth, in terms of user identification and personalized interaction, existing AI virtual avatar interaction systems lack automatic user identification capabilities based on biometrics. The system cannot automatically identify the identity of the current user through voice interaction, resulting in all users receiving the same response strategy and interaction style, failing to provide differentiated services based on individual user preferences, knowledge levels, emotional inclinations, and historical interaction memories. Traditional methods require users to manually log in or enter identity identifiers, which feels disjointed and unnatural in virtual avatar scenarios where voice is the primary means of interaction. Furthermore, even if the system can identify the user, it lacks a mechanism to dynamically associate and dynamically recall voiceprint identity with personalized response styles and cross-conversation historical memories, resulting in a lack of continuity and personalization in the interaction, making it difficult to establish long-term user trust relationships. Summary of the Invention

[0006] In view of the defects and shortcomings of the existing technology, the present invention provides an intelligent interaction and decision-making method integrating AI virtual avatars to solve the technical problems of insufficient data transmission security, separation of emotion and intent understanding, poor decision reliability and lack of interpretability, and insufficient personalized interaction capabilities in the existing technology.

[0007] As a first aspect of the present invention, a method for intelligent interaction and decision-making integrating an AI virtual avatar is provided, the method comprising: Step S1: Obtain multi-dimensional interaction data of the user through the client; wherein, the multi-dimensional interaction data includes voice signal data, facial image data, text input data, and public behavior data; Step S2: The client encrypts and transmits the user's multi-dimensional interaction data to the server. Step S3: The server performs sentiment-intent joint reasoning based on the user's multi-dimensional interaction data to obtain the user's sentiment-intent joint vector; Step S4: The server executes intelligent decision-making based on a multi-level Bayesian decision-making framework according to the user's sentiment-intent joint vector to obtain the optimal response strategy; Step S5: The server outputs multimodal response information to the client according to the optimal response strategy, and then receives user feedback information on the client.

[0008] Further, step S1 includes: A multi-stage echo cancellation scheme based on frequency domain block adaptive filtering is adopted to perform adaptive echo cancellation processing on the speech signal data in the multi-dimensional interactive data to obtain clean speech signal data after echo cancellation; specifically, it includes the following three stages: Phase 1: Linear echo cancellation based on frequency domain block adaptive filtering Let d(n) be the mixed speech signal data collected by the microphone, which includes the user's voice, speaker echo, and ambient noise; and let d(n) be the echo signal data played by the speaker. Define the time-domain error signal. : ; in, This refers to the number of frames acquired for the mixed speech signal data. This is the time-domain error signal, i.e., the clean speech signal data after linear echo cancellation; Let L be the weight vector of a frequency domain filter of length L; The weight vector is generated using the frequency domain block normalized least mean square algorithm. The update uses normalized step size control: ; in, Step size factor; For regularization parameters; for Conjugate operation; For the first The weight vector of the frequency domain filter of the block. The first one calculated according to the update formula The weight vector of the frequency domain filter with +1 block, For the first The frequency domain vector of the echo signal data of the block. For the first The frequency domain vector of the time-domain error signal of the block; Phase Two: Dual-talk detection and adaptive step size factor control When the user and the AI ​​virtual avatar speak simultaneously, i.e., in a dual-talk scenario, the weight vector of the frequency domain filter... An offset may occur, leading to echo signal data leakage; this stage dynamically controls the weight vector of the first stage based on the acquisition results of this frame. Update: ; in, This is the echo estimation signal output by the frequency domain filter; Let i be the i-th dimension weight vector of the frequency domain filter in the n-th frame; For the first Echo signal data of the frame; Calculate the energy ratio statistic DTD: ; in, α is the smoothing constant; α is the echo attenuation factor. Calculate the normalized cross-correlation statistic. : ; The rules for determining whether a conversation takes place are as follows: when and When the scenario is determined to be a dual-talk scenario, the step size factor is controlled. To freeze the weight vector Update; otherwise, determine it as a single-talk or silent state, and control the step size factor. To recover the weight vector Update; in, The energy threshold for dual-talk detection; The cross-correlation threshold for dual-talk detection; Phase 3: Residual Echo Suppression and Post-processing The residual echo suppression method based on frequency domain Wiener filtering is used to process the clean speech signal data after linear echo cancellation. Nonlinear residual echo cancellation is performed to obtain clean speech signal data after nonlinear echo cancellation. ; ; ; in, For residual echo power spectrum estimation; For near-end speech power spectrum estimation; To suppress the depth factor; This is the lower limit of the power spectrum of near-end speech. The Wiener filter gain function has a range of values ​​[ [,1]; k is the block index, For frequency index, For the k-th block, the Frequency domain value of the time-domain error signal at each frequency point.

[0009] Further, step S2 includes: Step S21: ECDH Key Negotiation Phase The client generates the first temporary elliptic curve key pair. , For the client's private key, The client uses the client public key. Send to the server; the server also generates a second temporary elliptic curve key pair. , This is the server-side private key. The server-side public key, and the server-side public key Return to the client; wherein, , = • G, where G is the base point of the temporary elliptic curve; both parties compute the shared key. : ; From the shared key on the server side Derived session key and fragmented misdirected seeds : ; Where HKDF is the HMAC-based key derivation function; salt is the salt value of HKDF; info is the context information field of HKDF; and 64 is the output length of HKDF. and Take 32 bytes each; For string / byte concatenation operators; Step S22: Perform sensitivity-level encryption on the multi-dimensional interaction data to obtain encrypted multi-dimensional interaction data. First, the sensitivity levels of the multi-dimensional interactive data are defined; wherein, the public behavior data is defined as low-sensitivity data, the text input data and speech-to-text data are defined as medium-sensitivity data, and the facial image data and voice signal data are defined as high-sensitivity data. Secondly, data of different sensitivity levels are encrypted using the following encryption formula: Low-sensitivity data: ; Medium / high sensitivity data: ; in, For low-sensitivity data to be encrypted; For medium / highly sensitive data to be encrypted; This is encrypted, low-sensitivity data; For encrypted medium / high sensitivity data; For additional authentication data; This is a combined encryption algorithm using the ChaCha20 stream encryption algorithm and the Poly1305 authentication tag algorithm. It is an Advanced Encryption Standard (AES) algorithm; For combined encryption algorithms A one-time random number; For Advanced Encryption Standard Algorithm A one-time random number; Step S23: Dynamic Fragmentation Obfuscation Algorithm (1) Divide the encrypted multi-dimensional interactive data into N actual fragments, where N is determined by the following formula: ; in, The length of the encrypted multi-dimensional interactive data. Where N is the standard fragment size, and N is the total number of actual fragments into which the encrypted multi-dimensional interactive data is divided. The current security level; This is a function to find the maximum value. It is a rounding function; (2) Generated using Reed-Solomon encoding A redundant check chip is used for integrity verification and loss recovery of the encrypted multi-dimensional interactive data; (3) Determine the obfuscated order of the N actual fragments to be sent based on the pseudo-random sequence generator: ; PRNG stands for Pseudo-random Number Generator, which uses a sharded, scrambled seed. For key, current timestamp For seed; " represents output; For actual sharding An arrangement used to define the obfuscated arrangement order of the actual fragments sent. ; (4) Embed a 64-bit timestamp watermark in the header of each actual slice. : ; in, It is a hash message authentication code algorithm based on SHA-256; This is a unique identifier for the current actual shard. To extract the first 8 bytes of the HMAC-SHA256 output; Step S24: According to the obfuscated arrangement order of the N actual fragments sent, the encrypted multi-dimensional interactive data is transmitted to the server via network encryption; Among them, the current comprehensive security index value of the network is defined. To adaptively adjust the network's current security level. ; When SI < This is extremely dangerous; adjust the network's current security level. = 5; when ≤ SI< At that time, standard security is maintained, and the current security level of the network is adjusted. = 3; When SI ≥ At that time, it is very safe; adjust the current security level of the network. = 1; ; in: This represents the current overall security index value of the network; the smaller the value, the more dangerous the network environment. These are weighting coefficients, corresponding to network latency respectively. Packet loss rate Abnormal traffic detection score The weights satisfy =1; The normalization function maps indices of different dimensions to the interval [0, 1]. This is the lower bound of the safety threshold. This represents the upper bound of the safety threshold; the current safety level. Used to control the actual number of shards and the number of obfuscation rounds.

[0010] Further, step S3 includes: Step S31: Extract the speech prosody temporal feature vector from the speech signal data. Extract the facial AU temporal feature vector from the facial image data. Extract contextual sentiment word vectors from the text input data. The speech prosodic temporal feature vector is obtained through a feature fusion gating mechanism. The facial AU temporal feature vector and the context sentiment word vectors Integrate into a fusion emotional vector : ; ; in, This is the gate vector; Use the Sigmoid activation function; This is the gate weight matrix; This is the gated bias vector; This is a vector concatenation operation; This is element-wise multiplication; The fused emotion vector is processed by the emotion flow encoder. Encoding as emotional flow features And determine the probability distribution of emotion categories. ; ; ; in, For emotion stream encoder; The linear projection weight matrix for sentiment classification; This is the bias vector for sentiment classification; Features of emotional flow The vector of [CLS] marker bits; It is a normalized exponential function; Given emotional flow features The probability distribution of emotion category e; emotion category e ∈ {positive, negative, neutral, anxious, excited, calm}; Step S32: Extract the intent vector from the text input data. Simultaneously, the user's historical dialogues are encoded into contextual historical dialogue vectors through a hierarchical attention network. ; The intent vector is processed by the intent flow encoder. and contextual history dialogue vectors Encoding as intent stream features And determine the probability distribution of intent categories. ; ; ; in, For intent stream encoder; Linear projection weight matrix for intention classification; The bias vector for intention classification; Intent Flow Features The vector of [CLS] marker bits; It is a normalized exponential function; For a given intent flow feature When, the probability distribution of intent category i; intent category i ∈ {query, instruction, casual conversation, complaint, request for help, feedback}; Intent vector With contextual history dialogue vectors splicing; Step S33: Apply the emotion stream features using a cross-attention function. and intent flow features To conduct two-way information exchange: ; ; ; in, For cross-attention function, For querying the matrix, The key matrix, It is a value matrix; Based on the characteristics of emotional flow For query, intent stream features The output is the cross-attention for key-value pairs; In order to be based on the characteristics of the intent flow For query, sentiment stream features The output is the cross-attention for key-value pairs; This is the normalized sentiment-intention joint vector; For layer normalization operation; Determine the joint probability distribution of emotion and intention : ; in, The linear projection weight matrix for joint sentiment-intention classification; The bias vector for joint sentiment-intention classification; For sentiment-intention joint vector The vector of [CLS] marker bits; Given an emotion-intention joint vector The probability distribution of the joint emotion-intention category (e, i) at that time; Step S34: Multi-turn dialogue state tracking (1) Under normal circumstances, maintain a variable-length dialogue state vector. Updated cyclically via gating: ; ; in, Let be the gating vector for the t-th round of dialogue; Use the Sigmoid activation function; These are the weight matrix and bias vector of the update gate, respectively; Let be the joint emotion-intent vector of the t-th round of dialogue; Let be the state vector of the (t-1)th round of dialogue; Let t be the state vector of the t-th round of dialogue; These are the weight matrix and bias vector for the state transition, respectively; (2) When a sudden change in emotion or a shift in intent is detected, the state reset gate is triggered: ; ; in, Reset the gate scalar for the state of the t-th round of dialogue; These are the weight matrix and bias vector of the state reset gate, respectively; Let be the probability distribution of the sentiment category in the t-th round of dialogue; Let be the probability distribution of the sentiment category in the (t-1)th round of dialogue; for and L1 norm difference; For the Kronecker delta function; To retrieve the sentiment category index corresponding to the highest probability; Number the round of the dialogue.

[0011] Furthermore, in step S4, the multi-level Bayesian decision framework includes a situational awareness layer, a response strategy generation layer, and an execution planning layer, specifically including: Step S41: Situational Awareness Layer Construct the user state space S = { , , ..., } contains M user states; where each user state It is a vector containing emotion-intention joint vector A composite vector of dialogue turns and user profile knowledge graph; Approximate posterior probability distribution using particle filters : (1) Initialization = 200 particles { , Each particle represents a user state; in, The number of particles; Let be the state value of the j-th particle; Let be the weight of the j-th particle; (2) Prediction steps: ; Where A is the state transition matrix, which describes the evolution of the user's state between adjacent rounds of dialogue; B is the control input matrix, which maps externally observed actions to the state space. This is the currently observed user action vector; It follows a multivariate Gaussian distribution; This represents the state of a user in the t-th round of conversation. Given all observations from the 1st round of dialogue to the tth round of dialogue, let be the approximate posterior probability distribution of the user state; The Gaussian noise covariance matrix for user state transitions; Let j be the user state value of the j-th particle in the (t-1)th round of dialogue. Let j be the user state value of the j-th particle in the t-th round of dialogue; (3) Update steps:

[0012] in, Let be the observation value of the t-th round of dialogue, i.e., the joint probability distribution of sentiment and intention in the t-th round of dialogue; Proportional to the sign; Refers to the assumed user state value Under these conditions, the current joint probability distribution of emotion and intention is observed. The possibility; This refers to the reliability of the state prior to the j-th particle itself; (4) Resampling: When the effective number of particles At that time, perform system resampling; in, ; Step S42: Response Strategy Generation Layer First, define the response strategy space. For all candidate response strategies The set of candidate response strategies is defined. Expected return function : ; ; in, For a candidate response strategy, i.e., a mapping function from state to action; Let be the discount factor for the t-th round of dialogue; In the state Next action Instant reward value; For mathematical expectation operators; The total number of rounds of dialogue for decision-making; In the state Next action Task completion reward value, In the state Next action Emotional fit reward value, In the state Next action The security constraint reward value; This is the weighted coefficient of the three components in the reward function: task completion reward value, emotional fit reward value, and safety constraint reward value. Secondly, a candidate response search strategy combining Bayesian optimization and Monte Carlo tree search is employed: (1) Choice: ; in, This is the formula for the upper confidence bound; Candidate response strategy Average return estimate; For the exploration coefficient of UCB1; This represents the number of times a parent node has been visited in a Monte Carlo tree search tree. Candidate response strategy The number of times the corresponding node was accessed; (2) Extension: Model the response strategy-reward mapping through Gaussian process and use the expected improvement acquisition function to select the next candidate response strategy; (3) Simulation: Conduct 50 Monte Carlo simulations to evaluate the expected returns; (4) Backhaul: Update the Q value and access count of all nodes on the path; Search target: ; in, The optimal response strategy obtained from the search; For reference response strategies; To find candidate response strategies that maximize the overall value within the parentheses ; To improve the acquisition function; This is the KL divergence adjustment coefficient; Candidate response strategy Compared to the reference response strategy KL divergence; Step S43: Execute the planning layer According to the optimal response strategy Generating the multimodal response information to control the AI ​​virtual avatar to execute the multimodal response information on the client; specifically including: (1) Voice content generation: The AI ​​virtual character's emotional voice content is generated through a large language model; wherein, the emotion-intention joint vector is used to generate the emotional voice content. Conditional control over the wording and tone of the emotional speech content; (2) Facial expression parameter mapping: The emotion flow features are mapped using a differentiable renderer. The facial expression parameters are mapped to the AI ​​virtual avatar. (3) Motion parameter generation: The limb motion parameters of the AI ​​virtual image are generated based on a hybrid method of motion graph and reinforcement learning; wherein, the limb motion parameters are generated by the emotion flow features. Decide; (4) Timeline orchestration: The timeline orchestration engine synchronizes and schedules the emotional voice content, facial expression parameters and body movement parameters of the AI ​​virtual image.

[0013] Further, step S5 includes: After the AI ​​virtual avatar executes the multimodal response information, it receives user feedback information; and updates the Bayesian posterior based on the user feedback information. ; in, This refers to user feedback information from the t-th round of dialogue. To receive user feedback information After that, user status The posterior probability; Let be the likelihood function, in user state The following user feedback information was observed. The probability of; User status The prior probability is provided by the particle filter; It is proportional to the sign.

[0014] Furthermore, step S3 also includes: (1) Voiceprint feature extraction and embedding Valid speech segments are extracted from the user's speech signal data, and current voiceprint features are extracted from the valid speech segments. Then from the current voiceprint features Extract the current voiceprint embedding vector Finally, the current voiceprint embedding vector Intra-class compactness enhancement is performed to obtain the enhanced current speaker embedding vector. : ; Where α is the enhancement coefficient; MLP is a multilayer perceptron used for nonlinear transformation of the embedding space; (2) Voiceprint comparison and identity recognition Voiceprint registration phase: When a user interacts with the system for the first time, the system guides the user to register their voiceprint, collects at least three speech signal samples from different contexts, and extracts the sample voiceprint embedding vector for each speech signal sample. Then, the voiceprint registration template of a certain user was calculated. And obtain the user voiceprint registration template library This includes voiceprint registration templates for N users; ; in, This represents the number of segments of a user's voice signal sample collected. Let d be the sample voiceprint embedding vector of a user's d-th segment of speech signal; Voiceprint comparison stage: embedding a user's current voiceprint into a vector. The fusion score was calculated by comparing the user's voiceprint registration template with the template library U and using cosine similarity and probabilistic linear discriminant analysis. : ; ; ; Among them, T n Create a voiceprint registration template for the nth user in the user voiceprint registration template library U; For T n and Cosine similarity score between them; For T n and Probability linear discriminant analysis scoring between them; For T n and The probability that they come from the same user; For T n and The probability of coming from different users; For T n and The fusion score between them; These are the corresponding weight coefficients; Identity determination stage: based on fusion score User identity verification is performed, and the rules for identity verification are as follows: like Then, a user's identity is directly determined to be the user with the highest fusion score. ; like Then, a second determination is made on the identity of a user; like If so, a user is determined to be an unregistered user, and the new user registration process is triggered; in, The high confidence threshold; The low confidence threshold; (3) Personalized response style adaptation After identifying a user, the user's personalized response style configuration is loaded from the user's profile knowledge graph to form the user's personalized response style vector. ; ; in, Tone preference dimension, controlling the language style and facial expression tone of the AI ​​virtual avatar; To control the formality level, the wording and complexity of the AI ​​virtual avatar are controlled; To control the level of detail, the information density and extent of expansion of the AI ​​virtual avatar are controlled; To control the speech rate and response timing of the AI ​​virtual avatar in terms of interaction rhythm; The weight of the emotional response strategy of the AI ​​virtual avatar is controlled to determine the degree of empathy. (4) Injection of sentiment-intention joint reasoning engine and injection of Bayesian decision framework This user's personalized response style vector As an additional conditional input, to modulate the query matrix of the cross-attention function in step S33. This leads to the query matrix modulated by style vectors. : ; in, , These are the weight matrix and bias vector for style vector modulation, respectively; This is element-wise multiplication; Use the Sigmoid activation function; Based on the style vector modulated query matrix Determine the joint sentiment-intention vector modulated by style vector. ; ; (5) Bayesian decision framework injection Based on the style vector modulated sentiment-intention joint vector Update the instant reward value in step S42. : ; in, In state Next action Style matching bonus value; Weighting coefficients for style matching reward values; Among them, the optimal response strategy is determined based on the updated instant reward value. The generated multimodal response information is updated to match the user's preferences.

[0015] Furthermore, it also includes: Each registered user maintains an independent cross-dialogue interaction memory, employing a hierarchical memory architecture: The first layer is working memory. Store the state vector of the current dialogue. Engage with the complete historical dialogue, and compress the summary at the end of the current dialogue: ; ; in, For the first The state vector of the turn-based dialogue, i.e., the state vector of the current dialogue; For the first Multi-dimensional interaction data of a turn-based dialogue, i.e., multi-dimensional interaction data of the current dialogue; For the first Multimodal response information of a turn-based dialogue, i.e., the multimodal response information of the current dialogue; For the first The sentiment category of the turn-based dialogue, i.e., the sentiment category of the current dialogue; For the first The intent category of the turn-based dialogue, i.e., the intent category of the current dialogue; For the first A structured summary of the turn-based dialogue, i.e., a structured summary of the current dialogue; A function for generating summaries; The second layer is contextual memory. Store structured summaries of historical conversations: ; in, This is the total number of rounds of historical dialogue; Utilizing vector-based similarity from the contextual memory Retrieve the current conversation Structured summary: ; in, (·) is a text embedding function. For the first A structured summary of the historical dialogue; For the first Structured summary of historical dialogue The vector; cos(·,·) is the cosine similarity function; Input text data for the user in the current conversation. Input text data for the user in the current conversation Vector representation of; For the first Structured summary of historical dialogue Text data input by the user in the current conversation The correlation score between them; This includes providing a structured summary of the current conversation after it ends. Write the aforementioned scene memory .

[0016] Furthermore, it also includes: (1) The aforementioned contextual memory The search weight decays exponentially over time: ; in, The attenuation coefficient is... This is the current timestamp, used for comparison with timestamps from historical sessions; For the first Timestamps of historical dialogues; An index for historical sessions; For the first The search weight of each round of historical dialogue; (2) When new contextual memories conflict with old contextual memories, conflict resolution is carried out based on the proximate nature of timestamps and confidence levels: ; in, For memorizing new scenarios; Memories of old scenes; Refers to the conflict resolution function; The output of the conflict resolution function; (·) represents the confidence function; (·) represents the recency function; The confidence level of new context memory; This is due to the proximate nature of memories of new situations; The confidence level of memories of past scenes; The "→" indicates the proximate nature of old scene memories; "→" indicates mapping to, used to indicate the output of the conflict resolution function.

[0017] As a second aspect of the present invention, an intelligent interaction and decision-making system integrating AI virtual avatars is provided to implement the intelligent interaction and decision-making method integrating AI virtual avatars described above. The intelligent interaction and decision-making system integrating AI virtual avatars includes: The client is used to acquire multi-dimensional interaction data of users and encrypt and transmit the multi-dimensional interaction data of users to the server; wherein, the multi-dimensional interaction data includes voice signal data, facial image data, text input data and public behavior data; On the server side, the system performs joint sentiment-intent reasoning based on the user's multi-dimensional interaction data to obtain the user's sentiment-intent joint vector; executes intelligent decision-making based on a multi-level Bayesian decision framework based on the user's sentiment-intent joint vector to obtain the optimal response strategy; and outputs multimodal response information to the client based on the optimal response strategy, and then receives user feedback information on the client.

[0018] The intelligent interaction and decision-making method integrating AI virtual avatars provided by this invention has the following advantages: (1) Enhanced Security of Adaptive Hybrid Encrypted Transmission: The dynamic fragmentation obfuscation algorithm proposed in this invention breaks the structural regularity of traditional encrypted data packets, making it difficult for attackers to infer communication content through traffic analysis. Combined with a data sensitivity grading mechanism and adaptive security level adjustment, the overall encryption overhead is effectively reduced while ensuring the security of highly sensitive data. Experiments show that, compared with traditional TLS encryption schemes, the scheme of this invention reduces encryption overhead by about 35% and improves resistance to traffic analysis by about 60% at the same security strength.

[0019] (2) Improved accuracy of joint emotion-intention reasoning: This invention achieves deep joint modeling of emotion and intention through a dual-stream Transformer architecture and a cross-attention mechanism, effectively capturing the influence of emotional state on intention understanding. On the standard test set, the joint reasoning model achieves an intention recognition accuracy of 94.7%, which is 6.3 percentage points higher than the independent reasoning model; the F1 score for emotion recognition reaches 91.2%, which is 4.8 percentage points higher than the independent reasoning model.

[0020] (3) Reliability and interpretability of multi-level Bayesian decision-making: The hierarchical Bayesian decision-making framework proposed in this invention achieves interpretability through probabilistic reasoning, and the decision-making process can be clearly traced back to each step of probabilistic reasoning. The KL divergence constraint ensures the safety boundary of the generated strategy, avoiding extreme or inappropriate responses. The online learning and update mechanism enables the system to continuously adapt to the behavior patterns of individual users, and user satisfaction in A / B testing is 23.5% higher than that of traditional methods.

[0021] (4) Overall system performance: Through modular design and asynchronous message bus architecture, this invention achieves decoupling and horizontal scaling of each functional module. In concurrent stress testing, the system can stably support processing more than 500 interactive requests per second, with a median end-to-end response latency of less than 800 milliseconds, meeting the performance requirements of real-time interaction.

[0022] (5) Enhanced Personalized Interaction Experience Based on Voiceprint Recognition: This invention is the first to achieve personalized interaction based on voiceprint recognition in an AI virtual avatar interaction system. Through a TDNN voiceprint encoder and a cosine similarity + PLDA fusion scoring mechanism, the voiceprint recognition accuracy reaches 98.2% (EER=1.8%), with a recognition latency of less than 200ms, meeting the requirements for real-time interaction. Personalized Style Vector By influencing the reasoning, decision-making, and response generation stages simultaneously through conditional injection, the system ensures consistency in interaction style across the entire process. A cross-session hierarchical memory architecture (working memory → episodic memory → semantic memory) enables the system to "remember" users' long-term preferences and historical interactions. In A / B testing, enabling voiceprint recognition and personalization increased user retention by 31.2%, interaction satisfaction by 27.8%, and average interaction duration by 42%, significantly enhancing user stickiness and interaction depth. Attached Figure Description

[0023] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the following detailed description to explain the invention, but do not constitute a limitation thereof.

[0024] Figure 1 A flowchart of the intelligent interaction and decision-making method for integrating AI virtual avatars provided by the present invention.

[0025] Figure 2 A flowchart illustrating the specific implementation of the intelligent interaction and decision-making method for integrating AI virtual avatars provided by this invention.

[0026] Figure 3 This is a schematic diagram of the adaptive hybrid encryption transmission process provided by the present invention.

[0027] Figure 4 This is a schematic diagram of the data packet structure of the dynamic fragmentation and obfuscation algorithm provided by the present invention.

[0028] Figure 5 A schematic diagram of the structure of the emotion-intention joint reasoning engine provided by the present invention.

[0029] Figure 6 This is a schematic diagram of the structure of the multi-level Bayesian decision-making framework provided by the present invention.

[0030] Figure 7 This is a schematic diagram of the personalized interaction process based on voiceprint recognition provided by the present invention.

[0031] Figure 8 This is an architecture diagram of the intelligent interaction and decision-making system integrating AI virtual avatars provided by the present invention. Detailed Implementation

[0032] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0033] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of the invention described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0035] This embodiment provides an intelligent interaction and decision-making method integrating AI virtual avatars. Figure 1 The flowchart of the intelligent interaction and decision-making method for integrating AI virtual avatars provided by this invention is as follows: Figure 1 As shown, the intelligent interaction and decision-making method integrating AI virtual avatars includes: Step S1: Obtain multi-dimensional interaction data of the user through the client; wherein, the multi-dimensional interaction data includes voice signal data, facial image data, text input data, and public behavior data; Preferably, step S1 includes: In AI virtual avatar interaction scenarios, while the system speaker plays the virtual avatar's voice, the microphone captures a mixed signal of the user's voice and the speaker echo. To ensure the accuracy of speech recognition and sentiment analysis, adaptive echo cancellation (AEC) must be performed before feature extraction. This invention employs a multi-stage echo cancellation scheme based on frequency domain block adaptive filtering to perform adaptive echo cancellation (AEC) processing on the speech signal data in the multi-dimensional interaction data to obtain clean speech signal data after echo cancellation; specifically, it includes the following three stages: Phase 1: Linear echo cancellation based on frequency domain block adaptive filtering Let d(n) be the mixed speech signal data collected by the microphone, which includes the user's voice, speaker echo, and ambient noise; and let d(n) be the echo signal data played by the speaker. Using the echo path filter as the reference input for echo cancellation, the core problem of echo cancellation is to estimate the echo path filter so that the filter output is approximately equal to the echo component; define the time-domain error signal. : ; in, This refers to the number of frames acquired for the mixed speech signal data. This is the time-domain error signal, i.e., the clean speech signal data after linear echo cancellation; The weight vector (time domain) of the frequency domain filter of length L simulates the acoustic echo path between the speaker and the microphone; The weight vector is generated using the frequency domain block normalized least mean square algorithm (FDAF-NLMS). The update uses normalized step size control: ; in, This is the step size factor (default 0.5), which controls the trade-off between convergence speed and steady-state error; Regularization parameter (default 10) -6 This prevents the denominator from being zero and improves numerical stability. for Conjugate operation; For the first The weight vector of the frequency domain filter of the block. The first one calculated according to the update formula The weight vector of the frequency domain filter with +1 block, For the first The frequency domain vector of the echo signal data of the block. For the first The frequency domain vector of the time-domain error signal of the block; Phase Two: Dual-talk detection and adaptive step size factor control When the user and the AI ​​virtual avatar speak simultaneously, i.e., in a dual-talk scenario, the weight vector of the frequency domain filter... An offset may occur, leading to echo signal data leakage; this stage dynamically controls the weight vector of the first stage based on the acquisition results of this frame. The updates (μ=0 or μ>0) together constitute a stable and robust echo cancellation system. This invention employs a dual-talk detector based on signal-to-energy ratio and cross-correlation.

[0036] ; in, This is the echo estimation signal output by the frequency domain filter; Let i be the i-th dimension weight vector of the frequency domain filter in the n-th frame; For the first Echo signal data of the frame; Calculate the energy ratio statistic DTD (which measures the proportion of non-echo components in the near-end signal): ; in, Smoothing constant (default 10) -8 ), to prevent the denominator from being zero; α is the echo attenuation factor (default 0.8), which compensates for the difference between the actual echo path and the estimated path; Calculate the normalized cross-correlation statistic. (Measuring the correlation between the far-end reference signal and the microphone signal): ; The rules for determining whether a conversation takes place are as follows: when and When the scenario is determined to be a dual-talk scenario, the step size factor is controlled. To freeze the weight vector Update; otherwise, determine it as a single-talk or silent state, and control the step size factor. To recover the weight vector Update; in, The energy threshold for dual-talk detection (default 0.3); The cross-correlation threshold for dual-talk detection (default 0.6). Phase 3: Residual Echo Suppression and Post-processing Even after frequency domain filters eliminate linear echoes, nonlinear residual echoes may still exist. This invention employs a residual echo suppression method based on frequency domain Wiener filtering to suppress residual echoes in the clean speech signal data after linear echo elimination. Nonlinear residual echo cancellation is performed to obtain clean speech signal data after nonlinear echo cancellation. ; ; ; in, The residual echo power spectrum is estimated by smoothing the cross spectrum of the error signal and the echo estimate; The power spectrum of near-end speech is estimated by smoothing the self-power spectrum of the error signal; To suppress depth factor (default 2.0), the larger the value, the stronger the residual echo suppression; Set as the lower limit of the power spectrum of near-end speech (default 0.05, i.e. -26dB) to prevent excessive suppression from causing speech distortion; The Wiener filter gain function has a range of values ​​[ [,1], used for frequency domain masking of residual echoes; k is the block index, For frequency index, For the k-th block, the Frequency domain value of the time-domain error signal at each frequency point.

[0037] After the above three-stage echo cancellation processing, the echo component in the microphone signal is effectively suppressed, and the output is the clean speech signal data after nonlinear echo cancellation. The data is then transferred to the server for subsequent MFCC feature extraction, speech recognition, and emotional prosody analysis.

[0038] Step S2: The client encrypts and transmits the user's multi-dimensional interaction data to the server. Preferred, such as Figure 3-4 As shown, step S2 includes: Step S21: ECDH (Elliptic Curve Diffie-Hellman) Key Negotiation Phase The client generates the first temporary elliptic curve key pair. , This is the client's private key (scalar). The client uses the client public key. Send to the server; the server also generates a second temporary elliptic curve key pair. , This is the server-side private key (scalar). The server-side public key, and the server-side public key Return to the client; wherein, , = • G, where G is the base point (generator) of the temporary elliptic curve secp256r1; both parties compute the shared key. : ; in, Let be the shared key obtained through ECDH negotiation, and be a point on the temporary elliptic curve. From the shared key on the server side Derived session key and fragmented misdirected seeds : ; Where HKDF is an HMAC-based Key Derivation Function; salt is the salt value of HKDF, which increases the randomness of key derivation; info is the context information field of HKDF, used to bind the purpose of the key; and 64 is the output length of HKDF (64 bytes). and Take 32 bytes each; For string / byte concatenation operators; The dialogue key, derived from HKDF, is used for subsequent data encryption; This is a fragmented confusion seed derived from HKDF, used for pseudo-random sequence generation; Step S22: Perform sensitivity-level encryption on the multi-dimensional interaction data to obtain encrypted multi-dimensional interaction data. First, the sensitivity levels of the multi-dimensional interactive data are defined; wherein, the public behavior data (page dwell time, scroll distance, etc.) are defined as low-sensitivity data Level 0, the text input data and speech-to-text data are defined as medium-sensitivity data Level 1, and the facial image data and voice signal data are defined as high-sensitivity data Level 2. Secondly, data of different sensitivity levels are encrypted using the following encryption formula: Low-sensitivity data: ; Medium / high sensitivity data: ; in, For low-sensitivity data to be encrypted; For medium / highly sensitive data to be encrypted; This is encrypted, low-sensitivity data; For encrypted medium / high sensitivity data; Additional authentication data, including timestamps and packet sequence numbers, is used to prevent replay attacks but is not encrypted; It is a combined encryption algorithm of the ChaCha20 stream encryption algorithm and the Poly1305 authentication tag algorithm, which has higher computational efficiency than AES-GCM; It uses Advanced Encryption Standard (AES) algorithms, 256-bit keys, and GCM (Galois / Counter Mode) certified encryption mode to provide confidentiality and integrity. For combined encryption algorithms A one-time random number is used to ensure that the encryption result is different each time under the same key; For Advanced Encryption Standard Algorithm A one-time random number is used to ensure that the encryption result is different each time under the same key; Step S23: Dynamic Fragmentation Obfuscation Algorithm (1) Divide the encrypted multi-dimensional interactive data into N actual fragments (the last fragment may be less than N), where N is determined by the following formula: ; in, The length of the encrypted multi-dimensional interactive data. N is the standard fragment size (default 512 bytes), and N is the total number of actual fragments (integer) into which the encrypted multi-dimensional interactive data is divided, which is at least 4. The current security level (1-5) is dynamically determined by the adaptive security level adjustment algorithm; This is a function to find the maximum value. It is a rounding function; (2) Generated using Reed-Solomon encoding A redundant checksum is provided for integrity verification and loss recovery of the encrypted multi-dimensional interactive data; wherein... For Reed-Solomon error correction coding.

[0039] (3) Determine the obfuscated order of the N actual fragments to be sent based on the pseudo-random sequence generator: ; PRNG stands for Pseudo-random Number Generator, which uses a sharded, scrambled seed. For key, current timestamp For seed; " represents output; For actual sharding An arrangement used to define the obfuscated arrangement order of the actual fragments sent. ; The current timestamp is used as the PRNG seed to ensure the reproducibility of the permutation; (4) Embed a 64-bit (8-byte) timestamp watermark in the header of each actual fragment. Used for replay protection and fragmentation verification: ; in, It is a hash message authentication code algorithm based on SHA-256; This is a unique identifier for the current actual shard. To extract the first 8 bytes of the HMAC-SHA256 output; Step S24: According to the obfuscated arrangement order of the N actual fragments sent, the encrypted multi-dimensional interactive data is transmitted to the server via network encryption; Among them, the current comprehensive security index value of the network is defined. (Refers to network communication security level, used to combat network threats such as traffic analysis, packet loss attacks, malicious probing, etc.) to adaptively adjust the current security level of the network. The safety level adjustment cycle is 500 milliseconds, and the PID controller smoothly transitions to avoid oscillation.

[0040] When SI < This is extremely dangerous; adjust the network's current security level. = 5 (Highest security protection, N times the protection, 3 rounds of obfuscation); when ≤ SI< At that time, standard security is maintained, and the current security level of the network is adjusted. = 3 (Standard safety protection); When SI ≥ At that time, it is very safe; adjust the current security level of the network. = 1 (Minimum security protection, reducing redundant check chips); ; in: This is the current overall security index value of the network; the smaller the value, the more dangerous the network environment. These are weighting coefficients, corresponding to network latency respectively. Packet loss rate Abnormal traffic detection score The weights satisfy =1; The normalization function maps indices of different dimensions to the interval [0, 1]. This is the lower bound of the safety threshold; below this value, the highest safety level is triggered. This is the upper limit of the safety threshold; anything above this value is downgraded to the lowest safety level. Current network latency (milliseconds); This represents the current network packet loss rate (percentage). This is the abnormal traffic detection score, ranging from 0 to 1, with higher values ​​indicating more suspicious traffic detected; current security level. Used to control the actual number of shards and the number of obfuscation rounds.

[0041] Step S3: The server performs sentiment-intent joint reasoning based on the user's multi-dimensional interaction data to obtain the user's sentiment-intent joint vector; Preferred, such as Figure 5 As shown, step S3 includes: Step S31: Extract 39-dimensional MFCC features, fundamental frequency F0, and speech energy from the speech signal data, and extract the speech prosodic temporal feature vector using 1D-CNN. Facial AU temporal feature vectors are extracted from the facial image data using bidirectional LSTM. The contextual sentiment word vectors are extracted from the text input data using a pre-trained RoBERTa model. The speech prosodic temporal feature vector is obtained through a feature fusion gating mechanism. The facial AU temporal feature vector and the context sentiment word vectors Integrate into a fusion emotional vector : ; ; in, This is the gate vector (values ​​are between 0 and 1). Let σ(x) be the sigmoid activation function. -x The output range is (0, 1). This is the gating weight matrix, which controls the fusion ratio of the three modal features; This is the gated bias vector; This is a vector concatenation operation. To implement gating filtering for element-wise multiplication (Hadamard product); The fused emotion vector is processed by the emotion flow encoder. Encoding as emotional flow features And determine the probability distribution of emotion categories. ; ; ; in, It is an emotion flow encoder (6 layers, 512 hidden dimensions, and 8 heads for multi-head attention). The linear projection weight matrix for sentiment classification; This is the bias vector for sentiment classification; Features of emotional flow The vector of [CLS] marker bits is used as the aggregate representation of the entire input; The normalized exponential function maps real number vectors to probability distributions. Given emotional flow features The probability distribution of emotion category e; emotion category e ∈ {positive, negative, neutral, anxious, excited, calm}; Step S32: Extract the intent vector from the text input data using the BERT-wwm model. Simultaneously, the user's historical dialogues are encoded into contextual historical dialogue vectors through a hierarchical attention network. ; The intent vector is processed by the intent flow encoder. and contextual history dialogue vectors Encoding as intent stream features And determine the probability distribution of intent categories. ; ; ; in, It is an intent-flow encoder (6 layers, 512 hidden dimensions, and 8 heads for multi-head attention). Linear projection weight matrix for intention classification; The bias vector for intention classification; Intent Flow Features The vector of [CLS] marker bits; It is a normalized exponential function; For a given intent flow feature When, the probability distribution of intent category i; intent category i ∈ {query, instruction, chat, complaint, request for help, feedback}; Intent vector With contextual history dialogue vectors splicing; Step S33: Apply the emotion stream features using a cross-attention function. and intent flow features To conduct two-way information exchange: ; ; ; in, For cross-attention function, To query the matrix, determine "what to focus on"; The key matrix provides "the matching information that is being followed"; For value matrices, provide "the actual content that is of interest"; Based on the characteristics of emotional flow For query, intent stream features The output is the cross-attention for key-value pairs; In order to be based on the characteristics of the intent flow For query, sentiment stream features The output is the cross-attention for key-value pairs; This is the normalized sentiment-intention joint vector; For layer normalization operation; Determine the joint probability distribution of emotion and intention : ; in, The linear projection weight matrix for joint sentiment-intention classification; The bias vector for joint sentiment-intention classification; For sentiment-intention joint vector The vector of [CLS] marker bits; Given an emotion-intention joint vector The probability distribution of the joint emotion-intention category (e, i) at that time; the probability distribution of this joint emotion-intention category (e, i) It captures the conditional dependence between emotions and intentions; for example, the probability of an intention to "seek help" increases significantly during an "anxious" emotional state.

[0042] Step S34: Multi-turn dialogue state tracking (1) Under normal circumstances, maintain a variable-length dialogue state vector. (Variable length), updated via gated loop: ; ; in, Let t be the gating vector for the t-th round of dialogue, which controls the proportion of new information incorporated into the old state. The closer the value is to 1, the more new information there is. Use the Sigmoid activation function; These are the weight matrix and bias vector of the update gate, respectively; Let be the joint emotion-intent vector of the t-th round of dialogue; Let be the state vector of the (t-1)th round of dialogue; Let t be the state vector of the t-th round of dialogue; These are the weight matrix and bias vector for the state transition, respectively; Hyperbolic tangent activation function, output range (-1, 1); (2) When a sudden change in emotion or a shift in intent is detected, the state reset gate is triggered: ; ; in, Reset the state scalar for the t-th round of dialogue; the closer the value is to 1, the more historical states are discarded. These are the weight matrix and bias vector of the state reset gate, respectively; Let be the probability distribution of the sentiment category in the t-th round of dialogue; Let be the probability distribution of the sentiment category in the (t-1)th round of dialogue; for and The L1 norm difference measures the magnitude of emotional shifts; This is the Kronecker delta function, which returns 0 if the two argmax values ​​are the same, and 1 if they are different. To retrieve the sentiment category index corresponding to the highest probability; Number the round of the dialogue.

[0043] Step S4: The server executes intelligent decision-making based on a multi-level Bayesian decision-making framework according to the user's sentiment-intent joint vector to obtain the optimal response strategy; Preferred, such as Figure 6 As shown, in step S4, the multi-level Bayesian decision framework includes a situational awareness layer, a response strategy generation layer, and an execution planning layer, specifically including: Step S41: Situational Awareness Layer Construct the user state space S = { , , ..., } contains M user states; where each user state It is a vector containing emotion-intention joint vector A composite vector of dialogue turns and user profile knowledge graph; Approximate posterior probability distribution using particle filters : (1) Initialization = 200 particles { , Each particle represents a user state; in, The number of particles (default 200); Let be the state value of the j-th particle; The weight of the j-th particle reflects the credibility of the particle's (state hypothesis); (2) Prediction steps: ; Where A is the state transition matrix, which describes the evolution of the user's state between adjacent rounds of dialogue; B is the control input matrix, which maps externally observed actions to the state space. This is the currently observed user action vector; It follows a multivariate Gaussian (normal) distribution; This represents the state of a user in the t-th round of conversation. Given all observations from the 1st round of dialogue to the tth round of dialogue, let be the approximate posterior probability distribution of the user state; The Gaussian noise covariance matrix for user state transitions acts like a "noise knob," controlling the degree of state diffusion during prediction. Let j be the user state value of the j-th particle in the (t-1)th round (previous round). The user state value of the j-th particle in the t-th round (current round); (3) Update steps: ; in, Let be the observation value of the t-th round of dialogue, i.e., the joint probability distribution of sentiment and intention in the t-th round of dialogue; Proportional to the sign; Refers to the assumed user state value Under these conditions, the current joint probability distribution of emotion and intention is observed. The possibility; This refers to the reliability of the state prior (or the state probability after the prediction step) as the j-th particle itself. (4) Resampling: When the effective number of particles At that time, perform system resampling; in, ; Step S42: Response Strategy Generation Layer First, define the response strategy space. For all candidate response strategies The set of candidate response strategies is defined. Expected return function : ; ; in, For a candidate response strategy, i.e., a mapping function from state to action; Let be the discount factor for the t-th round of dialogue (decaying exponentially with time t). in, It is a fixed discount factor. for The power of t; The default value is 0.95; the closer it is to 1, the more emphasis is placed on future returns. In the state Next action Instant reward value; For mathematical expectation operators; The total number of rounds of dialogue for decision-making; In the state Next action The task completion reward value measures whether the response effectively solves the user's problem; In the state Next action The emotional fit reward value measures whether the emotional response matches the user's emotional state; In the state Next action The safety constraint reward value, and the penalty for unsafe or inappropriate responses; This is the weighted coefficient of the three components in the reward function: task completion reward value, emotional fit reward value, and safety constraint reward value. Let be a candidate response strategy, which is a mapping function from state to action. Candidate actions for the strategy output; Secondly, a candidate response search strategy combining Bayesian optimization and Monte Carlo Tree Search (MCTS) is employed: (1) Choice: ; in, The upper confidence bound formula balances the known rewards and exploration potential of a strategy. Candidate response strategy Average return estimate; The exploration coefficient for UCB1 controls the trade-off between exploration and utilization; This represents the number of times a parent node has been visited in a Monte Carlo tree search tree. Candidate response strategy The number of times the corresponding node was accessed; (2) Extension: Model the response strategy-reward mapping through Gaussian process and select the next candidate response strategy using the Expected Improvement acquisition function; (3) Simulation: Conduct 50 Monte Carlo simulations to evaluate the expected returns; (4) Backhaul: Update the Q value and access count of all nodes on the path; Search target: ; in, The optimal response strategy obtained from the search; The reference response strategy (pre-trained policy model) serves as a benchmark for security constraints. To find candidate response strategies that maximize the overall value within the parentheses ; To improve the expected improvement function, Bayesian optimization is used to select the next evaluation strategy. The KL divergence adjustment coefficient (initially 1.0) represents the tolerance of the control strategy to deviate from the reference strategy. Candidate response strategy Compared to the reference response strategy The KL divergence measures the degree of difference between two distributions; The KL divergence constraint ensures that the optimal response strategy does not deviate too far from the reference response strategy, thus avoiding the generation of extreme responses. Automatic adjustment is achieved through Bayesian optimization.

[0044] Step S43: Execute the planning layer The optimal response strategy The process involves decomposing the data into executable action sequence instructions to generate the multimodal response information, which is then used by the client to control the AI ​​virtual avatar to execute the multimodal response information; specifically, this includes: (1) Voice content generation: The AI ​​virtual character's emotional voice content is generated through a large language model; wherein, the emotion-intention joint vector is used to generate the emotional voice content. Conditional control over the wording and tone of the emotional speech content; (2) Facial expression parameter mapping: The emotion flow features are mapped using a differentiable renderer. The facial expression parameters (including 52 blendshape weights) are mapped to the AI ​​virtual avatar. (3) Motion parameter generation: The limb motion parameters of the AI ​​virtual image are generated based on a hybrid method of motion graph and reinforcement learning; wherein, the limb motion parameters are generated by the emotion flow features. Decide; (4) Time axis orchestration: The emotional voice content, facial expression parameters and body movement parameters of the AI ​​virtual image are synchronously scheduled through the time axis orchestration engine to ensure that the multimodal output is accurately aligned in the time dimension; the time accuracy is at the frame level (33ms@30fps).

[0045] Specifically, such as Figure 7 As shown, step S3 further includes: (1) Voiceprint feature extraction and embedding: After acquiring user speech signals in the multimodal perception layer (after echo cancellation preprocessing), this invention uses a voiceprint embedding model based on a time-delay neural network (TDNN) to perform front-end processing and feature extraction on the speech signals: Voice Activity Detection (VAD): Endpoint detection is performed on the preprocessed speech signal to remove silence segments and non-speech segments, in order to extract effective speech segments from the user's preprocessed speech signal data; a VAD model based on a deep neural network (DNN) is adopted, with a detection accuracy of F1 > 98%; Voiceprint feature extraction: Extract 80-dimensional log Mel-filterbank features from valid speech segments, i.e., the current voiceprint features. The frame length is 25ms, the frame shift is 10ms, and the mean-variance normalization is applied (CMVN). Voiceprint embedding extraction: This involves extracting the normalized current voiceprint features. Input the TDNN-based x-vector encoder to obtain Voiceprint embedding vector : ; ; Among them, the latency context window of the TDNN layer The frames are [5, 3, 3, 1, 1], and the statistical pooling layer aggregates frame-level features. For segment-level representation, fully connected layer Output 512-dimensional speaker embedding vector TDNN stands for Time Delay Neural Network, which captures long-term dependent voiceprint features. The statistical measures calculated for the statistical pooling layer used after TDNN, where mean is the mean of each frame and stddev is the standard deviation of each frame. It is a classic architecture for voiceprint embedding.

[0046] Voiceprint embedding enhancement: Adaptive affinity propagation (AAP) is used for the... Voiceprint embedding vector Intra-class compactness enhancement is performed to obtain the enhanced current speaker embedding vector. : ; Where α is the enhancement coefficient (default 0.3), which controls the strength of the residual connection; MLP is a multilayer perceptron used for nonlinear transformation of the embedding space; (2) Voiceprint comparison and identity recognition Voiceprint registration phase: When a user interacts with the system for the first time, the system guides the user to register their voiceprint, collecting at least three speech signal samples from different contexts (each 5-10 seconds long; the more samples, the better the cross-context stability of the voiceprint). The system then extracts the sample voiceprint embedding vector for each speech signal sample. Then, the voiceprint registration template of a certain user was calculated. And obtain the user voiceprint registration template library This includes voiceprint registration templates for N users; ; in, This represents the number of segments of a user's voice signal sample collected. Let d be the sample voiceprint embedding vector of a user's d-th segment of speech signal; The voiceprint registration template (512-dimensional) for a user is obtained by averaging the voiceprint embedding vectors of multiple speech signal samples; Voiceprint comparison stage: embedding a user's current voiceprint into a vector. The fusion score was calculated by comparing the user's voiceprint registration template with the template library U and using cosine similarity and probabilistic linear discriminant analysis (PLDA). : ; ; ; Among them, T n Create a voiceprint registration template for the nth user in the user voiceprint registration template library U; For T n and The cosine similarity score between the two vectors measures how close their directions are, with a value of [-1, 1]. For T n and The probability linear discriminant analysis score (log-likelihood ratio) between the two audio segments measures the probability that they come from the same speaker. For T n and The probability that they come from the same user; For T n and The probability of coming from different users; For T n and The fusion score between them; These are the corresponding weight coefficients (default) =0.4, =0.6), PLDA has a higher weight because it has stronger discriminative power; Identity determination stage: based on fusion score User identity verification is performed, and the rules for identity verification are as follows: like Then the identity of a user can be directly determined. The user with the highest integrated rating ; like Then, the identity of a user is further determined by combining contextual auxiliary information (such as the current device, time period, and historical users in conversations); like If so, a user is determined to be an unregistered user, and the new user registration process is triggered; in, The index n is the user with the highest fusion score; The high confidence threshold (default 0.85); Set as the low confidence threshold (default 0.6); (3) Personalized response style adaptation Determine a user's identity Then, the user's personalized response style configuration is loaded from the user's profile knowledge graph to form the user's personalized response style vector. ; ; in, The tone preference dimension, with values ​​of {friendly, formal, lively, calm}, controls the language style and facial expression tone of the AI ​​virtual avatar; Formalness dimension, continuous value [0, 1], 0 is extremely colloquial, 1 is extremely formal, controls the word choice and word complexity of the AI ​​virtual image; The level of detail is a continuous value [0, 1], where 0 represents a minimalist response and 1 represents a detailed explanation, controlling the information density and extent of expansion of the AI ​​virtual avatar. The interaction rhythm dimension has continuous values ​​[0, 1], where 0 represents fast and concise, and 1 represents slow and meticulous, controlling the speech rate and response timing of the AI ​​virtual character. The empathy level dimension has continuous values ​​[0, 1], where 0 represents task-oriented and 1 represents emotion-oriented, controlling the weight of the AI ​​virtual character's emotional response strategy; Additionally, based on a user's identity Retrieve the most similar data from episodic memory (Default 3) summaries, serving as personalized context for the current conversation, are used in the sentiment-intent joint reasoning engine.

[0047] (4) Analyze the user's personalized response style vector. Injecting an emotion-intention joint inference engine and a Bayesian decision framework through a style conditionation mechanism. This user's personalized response style vector As an additional conditional input, to modulate the query matrix of the cross-attention function in step S33. This leads to the query matrix modulated by style vectors. : ; in, , These are the weight matrix and bias vector for style vector modulation, respectively; This is element-wise multiplication; Use the Sigmoid activation function; Based on the style vector modulated query matrix Determine the joint sentiment-intention vector modulated by style vector. ; ; This allows the results of the joint emotion-intention reasoning to shift towards the style preferred by the user while maintaining accuracy.

[0048] (5) Bayesian decision framework injection Will As an additional factor to the immediate reward value of Bayesian decision-making, and based on the sentiment-intention joint vector modulated by the style vector. Update the instant reward value in step S42. : ; in, In the state Next action The style matching reward value measures the degree of match between the candidate response strategy and the user's style preferences. It is calculated by matching the style vector of the candidate response strategy with the user's personalized response style vector. The scalar value obtained from the similarity between them (e.g., cosine similarity or negative Euclidean distance); Weighting factor for style matching reward value (default 0.15); Among them, the optimal response strategy is determined based on the updated instant reward value. The generated multimodal response information is updated to match the user's preferences. (6) Injection response generation: based on Conditional control of the prosodic parameters of speech synthesis and the facial expression and action parameters of the virtual avatar ensures that the output speech rate, pitch, volume, and the virtual avatar's smile and gesture amplitude are consistent with the user's preferences. This process enables AI virtual avatars to recognize users the moment they speak and automatically switch to the user's unique response style and historical memory context, achieving a personalized intelligent interactive experience that is "one face for every user".

[0049] Step S5: The server outputs multimodal response information to the client according to the optimal response strategy, and then receives user feedback information on the client.

[0050] Preferred, such as Figure 6 As shown, step S5 includes: After the AI ​​virtual avatar executes the multimodal response information, it receives user feedback information; and updates the Bayesian posterior based on the user feedback information. ; Simultaneously update the policy prior distribution and the state transition model parameters of the particle filter, enabling the system to continuously adapt to the behavior patterns of individual users; in, For user feedback information in the t-th round of dialogue, including explicit ratings (such as satisfaction scores) or implicit behavioral signals (such as dwell time, click behavior); To receive user feedback information After that, user status The posterior probability; Let be the likelihood function, in user state The following user feedback information was observed. The probability of; User status The prior probability is provided by the particle filter; It is proportional to the sign.

[0051] Preferred options also include: Each registered user maintains an independent cross-dialogue interaction memory, employing a hierarchical memory architecture: The first layer is working memory. Store the state vector of the current dialogue. Engage with the complete historical dialogue, and compress the summary at the end of the current dialogue: ; ; in, For the first The state vector of the turn-based dialogue, i.e., the state vector of the current dialogue; For the first Multi-dimensional interaction data of a turn-based dialogue, i.e., multi-dimensional interaction data of the current dialogue; For the first Multimodal response information of a turn-based dialogue, i.e., the multimodal response information of the current dialogue; For the first The sentiment category of the turn-based dialogue, i.e., the sentiment category of the current dialogue; For the first The intent category of the turn-based dialogue, i.e., the intent category of the current dialogue; For the first A structured summary (256 dimensions) of the turn-based dialogue, i.e., a structured summary of the current dialogue; For the summary generation function, represent any feasible method of "compressing working memory into a fixed-length summary"; The second layer is contextual memory. (Episodic Memory), storing structured summaries of historical conversations: ; in, This is the total number of rounds of historical dialogue; Utilizing vector-based similarity from the contextual memory Retrieve the current conversation (Default 3) Structured summaries to serve as context for the current conversation: ; in, (·) is a text embedding function that maps text to a vector representation; For the first A structured summary of the historical dialogue; For the first Structured summary of historical dialogue The vector; cos(·,·) is the cosine similarity function; Input text data for the user in the current conversation (i.e., the user's question or utterance in the current round of conversation). Input text data for the user in the current conversation Vector representation of; For the first Structured summary of historical dialogue User input text data in the current conversation The correlation score between them; This includes providing a structured summary of the current conversation after it ends. Write the aforementioned scene memory .

[0052] Specifically, it also includes: (1) The aforementioned contextual memory Search weight Decays exponentially over time: ; in, This is the decay coefficient (default 0.05), which controls the rate at which memories are forgotten. This is the current timestamp, used for comparison with timestamps from historical sessions; For the first Timestamps of historical dialogues; An index for historical sessions; For the first The search weight of each round of historical dialogue; Among them, decayed episodic memories enter an archived state when they are not retrieved for a long time; (2) When new contextual memories conflict with old contextual memories, conflict resolution is carried out based on the proximate nature of timestamps and confidence levels: ; in, For memorizing new scenarios; Memories of old scenes; Refers to the conflict resolution function; The output of the conflict resolution function; (·) represents the confidence function; (·) is the recency function, where the value is larger for more recent events; The confidence level of new context memory; This is due to the proximate nature of memories of new situations; The confidence level of memories of past scenes; The "→" indicates the proximate nature of old scene memories; "→" indicates mapping to, used to indicate the output of the conflict resolution function.

[0053] like Figure 2 As shown, the specific process of the intelligent interaction and decision-making method integrating AI virtual avatars provided by this invention is as follows: (1) Establish a multimodal perception layer: By integrating echo cancellation speech acquisition, facial expression capture, text semantic parsing and user public behavior data tracking modules, multi-dimensional user interaction data is collected; at the same time, speech signals are collected for voiceprint feature extraction to realize user identity recognition; Among them, voiceprint recognition and personalized context loading: voiceprint features are extracted and identity is matched on the collected voice signals to identify the identity of the current interactive user; based on the recognition results, the user's personalized response style configuration (including tone preference, knowledge level adaptation, and interaction rhythm) and cross-historical dialogue interaction memory are loaded to provide user identity context for subsequent reasoning and decision-making. (2) Adaptive hybrid encryption transmission: Secure transmission of multi-dimensional interactive data is achieved by using a key negotiation protocol based on elliptic curve cryptography to generate a dialogue key, and combining a dynamic fragmentation and obfuscation algorithm to encrypt and reassemble the transmitted data packets; (3) Joint reasoning of sentiment and intent: Based on the dual-stream Transformer network with fusion attention mechanism, the sentiment state and user intent are reasoned in parallel, and the joint modeling of sentiment and intent is achieved through cross attention layer; (4) Multi-level Bayesian decision-making: Execute intelligent decisions based on a multi-level Bayesian decision-making framework, including a situational awareness layer, a policy generation layer, and an execution planning layer; (5) Multimodal response generation: drive the AI ​​virtual image to perform multimodal response output.

[0054] The intelligent interaction and decision-making method for integrated AI virtual avatars provided by this invention specifically involves technical solutions for secure data transmission, joint emotion-intent reasoning, and multi-level Bayesian intelligent decision-making in multimodal interaction scenarios. (1) Adaptive hybrid encryption transmission mechanism: For the first time, a combination scheme of data sensitivity hierarchical encryption + dynamic fragmentation obfuscation + adaptive security level adjustment is introduced into the AI ​​virtual avatar interaction system. Through ECDH key negotiation, AES-256-GCM / ChaCha20-Poly1305 hybrid encryption, Reed-Solomon redundancy check and pseudo-random sequence fragmentation obfuscation, encryption overhead is reduced by 35% and anti-traffic analysis capability is improved by 60% while ensuring data security. (2) Joint emotion-intent reasoning engine: The first dual-stream Transformer + cross-attention joint modeling architecture is created to realize deep bidirectional interaction between emotion and intent. Combined with gated loop dialogue state tracking and attention decay mechanism, the long-range dependency and topic switching problems in multi-turn dialogue are effectively solved. The intent recognition accuracy is 94.7% and the emotion recognition F1 score is 91.2%. (3) Multi-level Bayesian decision framework: For the first time, a hierarchical decision framework of particle filter situational awareness + Bayesian optimization / MCTS policy search + KL divergence safety constraint is applied to AI virtual character interaction. The decision process is interpretable, the safety boundary is controllable, and online learning and updating are supported. User satisfaction is improved by 23.5%. (4) End-to-end multimodal synchronous drive: Based on the complete multimodal output pipeline of NeRF expression rendering + VITS emotional speech synthesis + motion graph / RL body movement + time axis orchestration engine, multimodal synchronous output with frame-level accuracy (33ms) is achieved. (5) Integrated adaptive echo cancellation speech acquisition: For the first time in AI virtual character interaction scenario, a three-level AEC scheme of frequency domain block adaptive filtering (FDAF-NLMS) + dual-talk detection + residual echo Wiener filtering suppression is introduced. The computational complexity is reduced from O(L^2) to O(L*logL) by frequency domain overlap preservation method, and adaptive step size freezing is achieved by combining dual-talk detector based on energy ratio and cross-correlation to avoid weight shift in dual-talk scenario. The residual echo Wiener filter further eliminates nonlinear echo components, with an ERLE of 25-35dB. This scheme effectively solves the echo interference problem between virtual avatar voice playback and user voice acquisition, ensuring the accuracy of speech recognition and emotional prosody analysis. (6) Personalized interaction mechanism based on voiceprint recognition: For the first time, voiceprint recognition-driven personalized interaction is realized in the AI ​​virtual avatar interaction system. The TDNN voiceprint encoder is used to extract 512-dimensional voiceprint embedding vectors, and high-precision user recognition is achieved through cosine similarity + PLDA fusion scoring (accuracy 98.2%, EER=1.8%, delay <200ms). The personalized style vector is innovatively integrated into the system. By simultaneously influencing three stages—emotion-intent joint reasoning (style modulation cross-attention query projection), Bayesian decision-making (style matching reward items), and response generation (prosody / expression parameter control)—through a conditional injection mechanism, end-to-end style consistency is achieved. A three-layer cross-session memory architecture—working memory, episodic memory, and semantic memory—is proposed, supporting vector-based historical memory retrieval and time-decay forgetting mechanisms. Enabling voiceprint personalization improves user retention by 31.2% and interaction satisfaction by 27.8%.

[0055] The following example application scenario (intelligent customer service scenario) will be used to illustrate in detail the intelligent interaction and decision-making method of the integrated AI virtual avatar provided by the present invention.

[0056] Scenario Description: Users use the AI ​​virtual customer service avatar on the company's official website to inquire about products and provide feedback on after-sales issues.

[0057] The interaction process is as follows: (1) The user opens a webpage and begins a conversation with the AI ​​virtual customer service avatar. The multimodal perception module collects the user's voice input "The product I bought has quality problems, and it has happened three times already". At the same time, the camera captures the user's facial expression as frowning (AU4 activated), and the voice prosody features show that the speech speed is fast and the pitch is rising.

[0058] The voiceprint recognition module simultaneously extracts voiceprint embedding vectors from the user's speech, compares them with the registered user template library, and identifies the current user as "Mr. Zhang". =0.93, higher than the high confidence threshold. The system loads Mr. Zhang's personalized style vector. Tone preference =Friendliness, formality =0.4 (speaking-oriented), level of detail =0.7 (too detailed), empathy level =0.8 (emotionally oriented). Simultaneously, a search of episodic memory revealed that Mr. Zhang had inquired about the logistics of the same order three days prior. =0.89). This personalized style vector This will be incorporated into subsequent reasoning and decision-making processes.

[0059] (2) The adaptive encrypted communication module encrypts multi-dimensional interactive data according to sensitivity levels before transmission. Facial AU features and voice signals are marked as Level 2 (high sensitivity), encrypted with AES-256-GCM, and subjected to 4 rounds of fragmented obfuscation. Text content and behavioral data are marked as Level 1 and Level 0, respectively.

[0060] (3) Results of the emotion-intention joint reasoning engine: Emotional output: P (negative) ) = 0.72, P(anxiety| = 0.18, P(neutral| = 0.06; Intent stream output: P(complaint| ) = 0.81, P(Help | ) = 0.12, P(query| = 0.04; Joint reasoning output: P(negative, complaint|) ) = 0.65, P(anxiety, seeking help| = 0.15; Dialogue state tracking: Dialogue state vector The update reads: "Users are dissatisfied with the product quality, their emotions are negative, and they need a solution." (Note: Mr. Zhang's personalized style vector has been loaded into the inference engine.) Query matrix of cross-attention layer The projection is style-modulated, shifting towards a more "friendly and detailed" approach while maintaining the accuracy of the reasoning.

[0061] (4) Multi-level Bayesian decision-making: Situational awareness layer: The particle filter estimates the user's state as "dissatisfied customer, needs immediate solution, and tends to escalate complaint"; Response strategy generation layer: MCTS search generates a set of candidate response strategies, which are then selected using Bayesian optimization to find the optimal response strategy. First, express your apology and empathy → provide specific solutions (return / exchange / compensation) → inquire about satisfaction → provide follow-up commitments; Execution planning layer: Generates specific response text, empathic facial parameters, reassuring body language, and timeline arrangement; (Note: The decision reward function includes a style matching term.) The optimal response strategy must simultaneously satisfy task completion, emotional fit, safety, and Mr. Zhang's preference for a "friendly and detailed" style.

[0062] (5) Virtual Avatar Response: The AI ​​virtual customer service replies in a friendly tone and detailed style preferred by Mr. Zhang: "Hello Mr. Zhang, I am very sorry for the bad experience you had again. I completely understand how you feel—the logistics problem last time was already very inconvenient for you, and now there is a quality problem. I really feel bad. Let me handle it for you immediately. I have the following solutions for you to choose from..." At the same time, the system makes a slight head tilt to express apology and makes concerned eye contact. The system also mentions the "logistics problem last time" in the historical dialogue, which reflects the retrieval of Mr. Zhang's cross-dialogue memory.

[0063] User feedback processing: The user responded, "Okay, then please exchange the item for me," shifting their emotional state from negative to neutral-to-positive. The online learning module updates the user state estimate and policy priors based on user feedback, providing a more accurate decision-making basis for subsequent interactions.

[0064] As another embodiment of the present invention, such as Figure 8 As shown, an intelligent interaction and decision-making system integrating AI virtual avatars is provided, wherein the intelligent interaction and decision-making system integrating AI virtual avatars includes a client and a server: The client is used to acquire multi-dimensional interaction data of the user and encrypt and transmit the multi-dimensional interaction data of the user to the server; wherein, the multi-dimensional interaction data includes voice signal data, facial image data, text input data and public behavior data; The server is configured to perform joint sentiment-intent reasoning based on the user's multi-dimensional interaction data to obtain the user's sentiment-intent joint vector; and to execute intelligent decision-making based on a multi-level Bayesian decision framework based on the user's sentiment-intent joint vector to obtain the optimal response strategy; and to output multimodal response information to the client based on the optimal response strategy, and then receive user feedback information on the client.

[0065] Specifically, the intelligent interaction and decision-making system integrating AI virtual avatars also includes the following modules: (1) Multimodal perception module: Deployed on the client side, it includes a microphone array, camera, text input interface, behavior tracker, and integrated echo cancellation (AEC) submodule. It is responsible for collecting multi-dimensional interaction data such as user voice, facial expressions, text input, and user public behavior data. The AEC submodule performs multi-level echo cancellation based on frequency domain block adaptive filtering, including three stages: frequency domain filter echo cancellation, dual-talk detection and step size control, and residual echo Wiener filtering suppression. This ensures that user voice is accurately collected while the virtual avatar is playing voice, and eliminates the interference of speaker echo on speech recognition and emotion analysis.

[0066] (2) Adaptive Encrypted Communication Module: Deployed simultaneously on both the client and server sides (the client is responsible for encrypting and sending data, and the server is responsible for decrypting and receiving data; key negotiation is jointly completed by both parties), including a key negotiation submodule, a sensitivity grading submodule, a fragmentation and obfuscation submodule, and a security level adjustment submodule. It is responsible for executing the adaptive hybrid encryption transmission method to ensure the security and anti-analysis capability of interactive data during transmission.

[0067] (3) Sentiment-Intent Joint Reasoning Module: Deployed on the server side, it includes a sentiment feature encoder, an intent feature encoder, a two-stream Transformer network, a cross-attention layer, and a dialogue state tracker. It is responsible for executing the sentiment-intent joint reasoning method described above.

[0068] (4) Multi-level Bayesian decision module: Deployed on the server side, it includes a particle filter situational awareness unit, a Bayesian optimization policy generator, and an action sequence planner. It is responsible for executing the multi-level Bayesian decision method described above and generating the optimal policy. It is then broken down into executable action sequences (voice, facial expressions, and actions) instructions.

[0069] (5) Virtual Avatar Driving Module: The server generates instructions, and the client renders them (speech synthesis can be completed on the server and the audio stream can be transmitted; facial expressions and body movements can be sent by parameters from the server, and the client renders them in real time through NeRF, motion graphs, etc.). It includes a speech synthesis submodule (emotional TTS based on the VITS model), an facial expression animation submodule (differentiable rendering based on NeRF), a body movement submodule (a hybrid method based on motion graphs + RL), and a timeline orchestration engine. It is responsible for converting executable action sequence instructions into multimodal (speech, facial expression, and action synchronization) response information that the client can render.

[0070] (6) Knowledge graph management module: Deployed on the server side, it maintains user profile knowledge graph (including user preferences, historical interaction records and personalized parameters), domain knowledge graph (including business domain knowledge base and FAQ library) and interaction history graph (structured historical dialogue records), providing knowledge support for sentiment-intent reasoning and Bayesian decision making.

[0071] (7) Voiceprint Recognition and Personalized Interaction Module: Deployed on the server side, it includes a voiceprint feature extraction submodule (TDNN-based x-vector voiceprint encoder), a voiceprint comparison and identity recognition submodule (cosine similarity + PLDA fusion scoring), a personalized style adaptation submodule (style vector generation and conditional injection), and a hierarchical memory management submodule (a two-layer architecture of working memory and contextual memory). It is responsible for extracting voiceprint features from speech signals and identifying user identities, loading corresponding personalized response style configurations and cross-historical dialogue memories, and providing user identity context and personalized parameters for sentiment-intent joint reasoning and Bayesian decision-making.

[0072] (8) System Communication Architecture: The modules on the server side communicate through an asynchronous message bus based on gRPC (in addition, the server sends response information to the client through the asynchronous message bus), and the message format is defined using Protocol Buffers. The system supports horizontal scaling, and each module can be deployed and expanded independently. Communication between modules is also protected by an encrypted channel.

[0073] (9) System Deployment Architecture: The client deploys a multimodal perception module and a lightweight encrypted client, which communicate with the server via a WebSocket long connection. The server adopts a microservice architecture, with each inference and decision-making module deployed in a containerized manner in a Kubernetes cluster, and communication and load balancing managed through a service mesh.

[0074] It is understood that the above embodiments are merely exemplary implementations used to illustrate the principles of the present invention, and the present invention is not limited thereto. For those skilled in the art, various modifications and improvements can be made without departing from the spirit and essence of the present invention, and these modifications and improvements are also considered to be within the scope of protection of the present invention.

Claims

1. A method for intelligent interaction and decision-making integrating AI virtual avatars, characterized in that, The intelligent interaction and decision-making method integrating AI virtual avatars includes: Step S1: Obtain multi-dimensional interaction data of the user through the client; wherein, the multi-dimensional interaction data includes voice signal data, facial image data, text input data, and public behavior data; Step S2: The client encrypts and transmits the user's multi-dimensional interaction data to the server. Step S3: The server performs sentiment-intent joint reasoning based on the user's multi-dimensional interaction data to obtain the user's sentiment-intent joint vector; Step S4: The server executes intelligent decision-making based on a multi-level Bayesian decision-making framework according to the user's sentiment-intent joint vector to obtain the optimal response strategy; Step S5: The server outputs multimodal response information to the client according to the optimal response strategy, and then receives user feedback information on the client.

2. The intelligent interaction and decision-making method integrating AI virtual avatars according to claim 1, characterized in that, Step S1 includes: A multi-stage echo cancellation scheme based on frequency domain block adaptive filtering is adopted to perform adaptive echo cancellation processing on the speech signal data in the multi-dimensional interactive data to obtain clean speech signal data after echo cancellation; specifically, it includes the following three stages: Phase 1: Linear echo cancellation based on frequency domain block adaptive filtering Let d(n) be the mixed speech signal data collected by the microphone, which includes the user's voice, speaker echo, and ambient noise; and let d(n) be the echo signal data played by the speaker. Define the time-domain error signal. : ; in, This refers to the number of frames acquired for the mixed speech signal data. This is the time-domain error signal, i.e., the clean speech signal data after linear echo cancellation; Let L be the weight vector of a frequency domain filter of length L; The weight vector is generated using the frequency domain block normalized least mean square algorithm. The update uses normalized step size control: ; in, Step size factor; For regularization parameters; for Conjugate operation; For the first The weight vector of the frequency domain filter of the block. The first one calculated according to the update formula The weight vector of the frequency domain filter with +1 block, For the first The frequency domain vector of the echo signal data of the block. For the first The frequency domain vector of the time-domain error signal of the block; Phase Two: Dual-talk detection and adaptive step-size factor control When the user and the AI ​​virtual avatar speak simultaneously, i.e., in a dual-talk scenario, the weight vector of the frequency domain filter... An offset may occur, leading to echo signal data leakage; this stage dynamically controls the weight vector of the first stage based on the acquisition results of this frame. Update: ; in, This is the echo estimation signal output by the frequency domain filter; Let i be the i-th dimension weight vector of the frequency domain filter in the n-th frame; For the first Echo signal data of the frame; Calculate the energy ratio statistic DTD: ; in, α is the smoothing constant; α is the echo attenuation factor. Calculate the normalized cross-correlation statistic. : ; The rules for determining whether a conversation takes place are as follows: when and Determined to be a dual-talk scenario, control the step size factor. To freeze the weight vector Update; otherwise, determine it as a single-talk or silent state, and control the step size factor. To recover the weight vector Update; in, The energy threshold for dual-talk detection; The cross-correlation threshold for dual-talk detection; Phase 3: Residual Echo Suppression and Post-processing The residual echo suppression method based on frequency domain Wiener filtering is used to process the clean speech signal data after linear echo cancellation. Nonlinear residual echo cancellation is performed to obtain clean speech signal data after nonlinear echo cancellation. ; ; ; in, For residual echo power spectrum estimation; For near-end speech power spectrum estimation; To suppress the depth factor; This is the lower limit of the power spectrum of near-end speech. The Wiener filter gain function has a range of values ​​[ [,1]; k is the block index, For frequency index, For the k-th block, the Frequency domain value of the time-domain error signal at each frequency point.

3. The intelligent interaction and decision-making method integrating AI virtual avatars according to claim 1, characterized in that, Step S2 includes: Step S21: ECDH Key Negotiation Phase The client generates the first temporary elliptic curve key pair. , For the client's private key, The client uses the client public key. Send to the server; the server also generates a second temporary elliptic curve key pair. , This is the server-side private key. The server-side public key, and the server-side public key Return to the client; wherein, , = • G, where G is the base point of the temporary elliptic curve; both parties compute the shared key. : ; From the shared key on the server side Derived session key and fragmented misdirected seeds : ; Where HKDF is the HMAC-based key derivation function; salt is the salt value of HKDF; info is the context information field of HKDF; and 64 is the output length of HKDF. and Take 32 bytes each; For string / byte concatenation operators; Step S22: Perform sensitivity-level encryption on the multi-dimensional interaction data to obtain encrypted multi-dimensional interaction data. First, the sensitivity levels of the multi-dimensional interactive data are defined; wherein, the public behavior data is defined as low-sensitivity data, the text input data and speech-to-text data are defined as medium-sensitivity data, and the facial image data and voice signal data are defined as high-sensitivity data. Secondly, data of different sensitivity levels are encrypted using the following encryption formula: Low-sensitivity data: ; Medium / high sensitivity data: ; in, For low-sensitivity data to be encrypted; For medium / highly sensitive data to be encrypted; This is encrypted, low-sensitivity data; For encrypted medium / high sensitivity data; For additional authentication data; This is a combined encryption algorithm using the ChaCha20 stream encryption algorithm and the Poly1305 authentication tag algorithm. It is an Advanced Encryption Standard (AES) algorithm; For combined encryption algorithms A one-time random number; For Advanced Encryption Standard Algorithm A one-time random number; Step S23: Dynamic Fragmentation Obfuscation Algorithm (1) Divide the encrypted multi-dimensional interactive data into N actual fragments, where N is determined by the following formula: ; in, The length of the encrypted multi-dimensional interactive data. Where N is the standard fragment size, and N is the total number of actual fragments into which the encrypted multi-dimensional interactive data is divided. The current security level; This is a function to find the maximum value. It is a rounding function; (2) Generated using Reed-Solomon encoding A redundant check chip is used for integrity verification and loss recovery of the encrypted multi-dimensional interactive data; (3) Determine the obfuscated order of the N actual fragments to be sent based on the pseudo-random sequence generator: ; PRNG stands for Pseudo-random Number Generator, which uses a sharded, scrambled seed. For key, current timestamp For seed; " represents output; For actual sharding An arrangement used to define the obfuscated arrangement order of the actual fragments sent. ; (4) Embed a 64-bit timestamp watermark in the header of each actual slice. : ; in, It is a hash message authentication code algorithm based on SHA-256; This is a unique identifier for the current actual shard. To extract the first 8 bytes of the HMAC-SHA256 output; Step S24: According to the obfuscated arrangement order of the N actual fragments sent, the encrypted multi-dimensional interactive data is transmitted to the server via network encryption; Among them, the current comprehensive security index value of the network is defined. To adaptively adjust the network's current security level. ; When SI < This is extremely dangerous; adjust the network's current security level. = 5; when ≤ SI< At that time, standard security is maintained, and the current security level of the network is adjusted. = 3; When SI ≥ At that time, it is very safe; adjust the current security level of the network. = 1; ; in: This represents the current overall security index value of the network; the smaller the value, the more dangerous the network environment. These are weighting coefficients, corresponding to network latency respectively. Packet loss rate Abnormal traffic detection score The weights satisfy =1; The normalization function maps indices of different dimensions to the interval [0, 1]. This is the lower bound of the safety threshold. This represents the upper bound of the safety threshold; the current safety level. Used to control the actual number of shards and the number of obfuscation rounds.

4. The intelligent interaction and decision-making method integrating AI virtual avatars according to claim 1, characterized in that, Step S3 includes: Step S31: Extract the speech prosody temporal feature vector from the speech signal data. Extract the facial AU temporal feature vector from the facial image data. Extract contextual sentiment word vectors from the text input data. The speech prosodic temporal feature vector is obtained through a feature fusion gating mechanism. The facial AU temporal feature vector and the context sentiment word vectors Integrate into a fusion emotional vector : ; ; in, This is the gate vector; Use the Sigmoid activation function; This is the gate weight matrix; This is the gated bias vector; This is a vector concatenation operation; This is element-wise multiplication; The fused emotion vector is processed by the emotion flow encoder. Encoding as emotional flow features And determine the probability distribution of emotion categories. ; ; ; in, For emotion stream encoder; The linear projection weight matrix for sentiment classification; This is the bias vector for sentiment classification; Features of emotional flow The vector of [CLS] marker bits; It is a normalized exponential function; Given emotional flow features The probability distribution of emotion category e; emotion category e ∈ {positive, negative, neutral, anxious, excited, calm}; Step S32: Extract the intent vector from the text input data. Simultaneously, the user's historical dialogues are encoded into contextual historical dialogue vectors through a hierarchical attention network. ; The intent vector is processed by the intent flow encoder. and contextual history dialogue vectors Encoding as intent stream features And determine the probability distribution of intent categories. ; ; ; in, For intent stream encoder; Linear projection weight matrix for intention classification; The bias vector for intention classification; Intent Flow Features The vector of [CLS] marker bits; It is a normalized exponential function; For a given intent flow feature When, the probability distribution of intent category i; intent category i ∈ {query, instruction, chat, complaint, request for help, feedback}; Intent vector With contextual history dialogue vectors splicing; Step S33: Apply the emotion stream features using a cross-attention function. and intent flow features To conduct two-way information exchange: ; ; ; in, For cross-attention function, For querying the matrix, The key matrix, It is a value matrix; Based on the characteristics of emotional flow For query, intent stream features The output is the cross-attention for key-value pairs; In order to be based on the characteristics of the intent flow For query, sentiment stream features The output is the cross-attention for key-value pairs; This is the normalized sentiment-intention joint vector; For layer normalization operation; Determine the joint probability distribution of emotion and intention : ; in, The linear projection weight matrix for joint sentiment-intention classification; The bias vector for joint sentiment-intention classification; For sentiment-intention joint vector The vector of [CLS] marker bits; Given an emotion-intention joint vector The probability distribution of the joint emotion-intention category (e, i) at that time; Step S34: Multi-turn dialogue state tracking (1) Under normal circumstances, maintain a variable-length dialogue state vector. Updated cyclically via gating: ; ; in, Let be the gating vector for the t-th round of dialogue; Use the Sigmoid activation function; These are the weight matrix and bias vector of the update gate, respectively; Let be the joint emotion-intent vector of the t-th round of dialogue; Let be the state vector of the (t-1)th round of dialogue; Let t be the state vector of the t-th round of dialogue; These are the weight matrix and bias vector for the state transition, respectively; (2) When a sudden change in emotion or a shift in intent is detected, the state reset gate is triggered: ; ; in, Reset the gate scalar for the state of the t-th round of dialogue; These are the weight matrix and bias vector of the state reset gate, respectively; Let be the probability distribution of the sentiment category in the t-th round of dialogue; Let be the probability distribution of the sentiment category in the (t-1)th round of dialogue; for and L1 norm difference; For the Kronecker delta function; To retrieve the sentiment category index corresponding to the highest probability; Number the round of the dialogue.

5. The intelligent interaction and decision-making method integrating AI virtual avatars according to claim 4, characterized in that, In step S4, the multi-level Bayesian decision framework includes a situational awareness layer, a response strategy generation layer, and an execution planning layer, specifically including: Step S41: Situational Awareness Layer Construct the user state space S = { , , ..., } contains M user states; where each user state It is a vector containing emotion-intention joint vector A composite vector of dialogue turns and user profile knowledge graph; Approximate posterior probability distribution using particle filters : (1) Initialization = 200 particles { , Each particle represents a user state; in, The number of particles; Let be the state value of the j-th particle; Let be the weight of the j-th particle; (2) Prediction steps: ; Where A is the state transition matrix, which describes the evolution of the user's state between adjacent rounds of dialogue; B is the control input matrix, which maps externally observed actions to the state space. This is the currently observed user action vector; It follows a multivariate Gaussian distribution; This represents the state of a user in the t-th round of conversation. Given all observations from the 1st round of dialogue to the tth round of dialogue, let be the approximate posterior probability distribution of the user state; The Gaussian noise covariance matrix for user state transitions; Let j be the user state value of the j-th particle in the (t-1)th round of dialogue. Let j be the user state value of the j-th particle in the t-th round of dialogue; (3) Update steps: ; in, Let be the observation value of the t-th round of dialogue, i.e., the joint probability distribution of sentiment and intention in the t-th round of dialogue; Proportional to the sign; Refers to the assumed user state value Under these conditions, the current joint probability distribution of emotion and intention is observed. The possibility; This refers to the reliability of the state prior to the j-th particle itself; (4) Resampling: When the effective number of particles At that time, perform system resampling; in, ; Step S42: Response Strategy Generation Layer First, define the response strategy space. For all candidate response strategies The set of candidate response strategies is defined. Expected return function : ; ; in, For a candidate response strategy, i.e., a mapping function from state to action; Let be the discount factor for the t-th round of dialogue; In the state Next action Instant reward value; For mathematical expectation operators; The total number of rounds of dialogue for decision-making; In the state Next action Task completion reward value, In the state Next action Emotional fit reward value, In the state Next action The security constraint reward value; This is the weighted coefficient of the three components in the reward function: task completion reward value, emotional fit reward value, and safety constraint reward value. Secondly, a candidate response search strategy combining Bayesian optimization and Monte Carlo tree search is employed: (1) Choice: ; in, This is the formula for the upper confidence bound; Candidate response strategy Average return estimate; For the exploration coefficient of UCB1; This represents the number of times a parent node has been visited in a Monte Carlo tree search tree. Candidate response strategy The number of times the corresponding node was accessed; (2) Extension: Model the response strategy-reward mapping through Gaussian process and use the expected improvement acquisition function to select the next candidate response strategy; (3) Simulation: Conduct 50 Monte Carlo simulations to evaluate the expected returns; (4) Backhaul: Update the Q value and access count of all nodes on the path; Search target: ; in, The optimal response strategy obtained from the search; For reference response strategies; To find candidate response strategies that maximize the overall value within the parentheses ; To improve the acquisition function; This is the KL divergence adjustment coefficient; Candidate response strategy Compared to the reference response strategy KL divergence; Step S43: Execute the planning layer According to the optimal response strategy Generating the multimodal response information to control the AI ​​virtual avatar to execute the multimodal response information on the client; specifically including: (1) Voice content generation: The AI ​​virtual character's emotional voice content is generated through a large language model; wherein, the emotion-intention joint vector is used to generate the emotional voice content. Conditional control over the wording and tone of the emotional speech content; (2) Facial expression parameter mapping: The emotion flow features are mapped using a differentiable renderer. The facial expression parameters are mapped to the AI ​​virtual avatar. (3) Motion parameter generation: The limb motion parameters of the AI ​​virtual image are generated based on a hybrid method of motion graph and reinforcement learning; wherein, the limb motion parameters are generated by the emotion flow features. Decide; (4) Timeline orchestration: The timeline orchestration engine synchronizes and schedules the emotional voice content, facial expression parameters and body movement parameters of the AI ​​virtual image.

6. The intelligent interaction and decision-making method integrating AI virtual avatars according to claim 5, characterized in that, Step S5 includes: After the AI ​​virtual avatar executes the multimodal response information, it receives user feedback information; and updates the Bayesian posterior based on the user feedback information. ; in, This refers to user feedback information from the t-th round of dialogue. To receive user feedback information After that, user status The posterior probability; Let be the likelihood function, in user state The following user feedback information was observed. The probability of; User status The prior probability is provided by the particle filter; It is proportional to the sign.

7. The intelligent interaction and decision-making method integrating AI virtual avatars according to claim 5, characterized in that, Step S3 further includes: (1) Voiceprint feature extraction and embedding Valid speech segments are extracted from the user's speech signal data, and current voiceprint features are extracted from the valid speech segments. Then from the current voiceprint features Extract the current voiceprint embedding vector Finally, the current voiceprint embedding vector is... Intra-class compactness enhancement is performed to obtain the enhanced current speaker embedding vector. : ; Where α is the enhancement coefficient; MLP is a multilayer perceptron used for nonlinear transformation of the embedding space; (2) Voiceprint comparison and identity recognition Voiceprint registration phase: When a user interacts with the system for the first time, the system guides the user to register their voiceprint, collects at least three speech signal samples from different contexts, and extracts the sample voiceprint embedding vector for each speech signal sample. Then, the voiceprint registration template of a certain user was calculated. And obtain the user voiceprint registration template library This includes voiceprint registration templates for N users; ; in, This represents the number of segments of a user's voice signal sample collected. Let d be the sample voiceprint embedding vector of a user's d-th segment of speech signal; Voiceprint comparison stage: embedding a user's current voiceprint into a vector. The fusion score was calculated by comparing the user's voiceprint registration template with the template library U and using cosine similarity and probabilistic linear discriminant analysis. : ; ; ; Among them, T n Create a voiceprint registration template for the nth user in the user voiceprint registration template library U; For T n and Cosine similarity score between them; For T n and Probability linear discriminant analysis scoring between them; For T n and The probability that they come from the same user; For T n and The probability of coming from different users; For T n and The fusion score between them; These are the corresponding weight coefficients; Identity determination stage: based on fusion score User identity verification is performed, and the rules for identity verification are as follows: like Then, a user's identity is directly determined to be the user with the highest fusion score. ; like Then, a second determination is made on the identity of a user; like If so, a user is determined to be an unregistered user, and the new user registration process is triggered; in, The high confidence threshold; The low confidence threshold; (3) Personalized response style adaptation After identifying a user, the user's personalized response style configuration is loaded from the user's profile knowledge graph to form the user's personalized response style vector. ; ; in, Tone preference dimension, controlling the language style and facial expression tone of the AI ​​virtual avatar; To control the formality level, the wording and complexity of the AI ​​virtual avatar are controlled; To control the level of detail, the information density and extent of expansion of the AI ​​virtual avatar are controlled; To control the speech rate and response timing of the AI ​​virtual avatar in terms of interaction rhythm; The weight of the emotional response strategy of the AI ​​virtual avatar is controlled to determine the degree of empathy. (4) Injection of sentiment-intention joint reasoning engine and injection of Bayesian decision framework This user's personalized response style vector As an additional conditional input, to modulate the query matrix of the cross-attention function in step S33. This leads to the query matrix modulated by style vectors. : ; in, , These are the weight matrix and bias vector for style vector modulation, respectively; This is element-wise multiplication; Use the Sigmoid activation function; Based on the style vector modulated query matrix Determine the joint sentiment-intention vector modulated by style vector. ; ; (5) Bayesian decision framework injection Based on the style vector modulated sentiment-intention joint vector Update the instant reward value in step S42. : ; in, In state Next action Style matching bonus value; Weighting coefficients for style matching reward values; Among them, the optimal response strategy is determined based on the updated instant reward value. The generated multimodal response information is updated to match the user's preferences.

8. The intelligent interaction and decision-making method integrating AI virtual avatars according to claim 7, characterized in that, Also includes: Each registered user maintains an independent cross-dialogue interaction memory, employing a hierarchical memory architecture: The first layer is working memory. Store the state vector of the current dialogue. Engage with the complete historical dialogue, and compress the summary at the end of the current dialogue: ; ; in, For the first The state vector of the turn-based dialogue, i.e., the state vector of the current dialogue; For the first Multi-dimensional interaction data of a turn-based dialogue, i.e., multi-dimensional interaction data of the current dialogue; For the first Multimodal response information of a turn-based dialogue, i.e., the multimodal response information of the current dialogue; For the first The sentiment category of the turn-based dialogue, i.e., the sentiment category of the current dialogue; For the first The intent category of the turn-based dialogue, i.e., the intent category of the current dialogue; For the first A structured summary of the turn-based dialogue, i.e., a structured summary of the current dialogue; A function for generating summaries; The second layer is contextual memory. Store structured summaries of historical conversations: ; in, This is the total number of rounds of historical dialogue; Utilizing vector-based similarity from the contextual memory Retrieve the current conversation Structured summary: ; in, (·) is a text embedding function. For the first A structured summary of the historical dialogue; For the first Structured summary of historical dialogue The vector; cos(·,·) is the cosine similarity function; Input text data for the user in the current conversation. Input text data for the user in the current conversation Vector representation of; For the first Structured summary of historical dialogue Text data input by the user in the current conversation The correlation score between them; This includes providing a structured summary of the current conversation after it ends. Write the aforementioned scene memory .

9. The intelligent interaction and decision-making method integrating AI virtual avatars according to claim 8, characterized in that, Also includes: (1) The aforementioned contextual memory The search weight decays exponentially over time: ; in, The attenuation coefficient is... It is the current timestamp; For the first Timestamps of historical dialogues; An index for historical sessions; For the first The search weight of each round of historical dialogue; (2) When new contextual memories conflict with old contextual memories, conflict resolution is carried out based on the proximate nature of timestamps and confidence levels: ; in, For memorizing new scenarios; Memories of old scenes; Refers to the conflict resolution function; The output of the conflict resolution function; (·) represents the confidence function; (·) represents the recency function; The confidence level of new context memory; This is due to the proximate nature of memories of new situations; The confidence level of memories of past scenes; "→" indicates the proximate nature of old scene memories and is used to indicate the output of the conflict resolution function.

10. An intelligent interaction and decision-making system integrating AI virtual avatars, used to implement the intelligent interaction and decision-making method integrating AI virtual avatars as described in any one of claims 1-9, characterized in that, The intelligent interaction and decision-making system integrating AI virtual avatars includes a client and a server: The client is used to acquire multi-dimensional interaction data of the user and encrypt and transmit the multi-dimensional interaction data of the user to the server; wherein, the multi-dimensional interaction data includes voice signal data, facial image data, text input data and public behavior data; The server is configured to perform joint sentiment-intent reasoning based on the user's multi-dimensional interaction data to obtain the user's sentiment-intent joint vector; and to execute intelligent decision-making based on a multi-level Bayesian decision framework based on the user's sentiment-intent joint vector to obtain the optimal response strategy; and to output multimodal response information to the client based on the optimal response strategy, and then receive user feedback information on the client.