AI multi-mode emotion interaction memory terminal

The multi-threaded acquisition engine and multi-head cross-attention mechanism realize microsecond synchronization of multi-modal data. Combined with the graph convolution network and differential privacy mechanism, misjudgment and personalization problems in multi-modal emotion recognition are solved, and the accuracy and continuity of emotional interaction are improved.

CN120372536APending Publication Date: 2025-07-25SHENZHEN XINZHI FUTURE TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510447104.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing technology has problems of misjudgment and lack of personalization in multimodal emotion recognition, especially in the lack of cross-modal timing correlation and context semantic modeling, which makes it difficult to accurately capture the emotional progressive relationship and the coordinated changes of micro-expressions and speech tone in dialogue scenarios. Moreover, the cloud multimodal analysis engine fails to effectively solve the problem of dynamic weight allocation, resulting in a lack of personalization and continuity of interaction strategies.

Method used

A multi-threaded acquisition engine is used to realize microsecond synchronization of voice, facial expressions and text data, dynamically allocate modal weights through a multi-head cross-attention mechanism, and activate the gated LSTM conflict dissolution module when cross-modal confidence differences are different; emotional memory modeling constructs emotional state transfer probability through graph convolution networks, and protects user privacy with differential privacy mechanisms; the terminal deploys a pruning quantitative decision model, and implements sparse federated learning optimization decisions in the cloud.

Benefits of technology

It realizes accurate synchronization and dynamic weight allocation of multimodal data, builds a user-specific emotional state map, supports millisecond back search and hard real-time emotional feedback within 200ms, and improves the accuracy and personalization of emotional interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372536A_ABST
    Figure CN120372536A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of AI interaction, and discloses an AI multi-modal emotion interaction memory terminal, which realizes microsecond-level synchronization of voice, facial expression and text data through a multi-thread acquisition engine, dynamically allocates each modal weight by adopting a multi-head cross attention mechanism, and adaptively adjusts modal importance based on a conversation context hidden state; when the cross-modal confidence difference exceeds a threshold value, a gating LSTM conflict resolution module is activated, and the multi-source data collaboration problem is solved; the emotional memory modeling constructs an emotional state transition topology based on a graph convolutional network, protects user privacy in combination with a differential privacy mechanism, and realizes associated event storage of millisecond backtracking of short-term memory and long-term memory. The technology integrates multi-modal dynamic perception, privacy security calculation and adaptive learning ability, significantly improves the real-time performance and personification degree of emotion interaction, and can be applied to the fields of intelligent customer service, emotion accompanying, health monitoring and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of AI interaction technologies, and more specifically, to an AI multi-modal emotional interaction memory terminal. Background Art

[0002] With the rapid development of artificial intelligence technologies, emotional interaction has become the core direction for enhancing the human-computer interaction experience. Currently, emotional recognition technologies have gradually developed from single-modal to multi-modal fusion, improving the comprehensiveness of emotional analysis by combining multiple information sources such as speech, text, and vision. For example, multi-modal models based on deep learning can process emotional features in context conversations through recurrent neural networks, attention mechanisms, etc. At the same time, the improvement in the computing power of edge intelligent terminals and the popularization of dedicated AI chips provide hardware support for real-time emotional interaction.

[0003] After retrieval, the patent with Chinese patent number CN119476488A discloses an AI interaction method and system based on emotional recognition. The method includes the following steps: in a cloud data center, build an emotion relief database and construct an emotional AI interaction engine; according to real-time multi-modal monitoring data, use the emotional AI interaction engine to generate real-time emotional recognition results and real-time AI interaction data; according to the real-time emotional recognition results, generate a real-time AI response strategy and extract real-time emotion relief data; this invention solves the problems of low accuracy of emotional recognition, low intelligence level of AI, and lack of evolution function existing in the prior art.

[0004] After retrieval, the patent with Chinese patent number CN119476488A discloses an AI interaction method and system based on emotional recognition. The method includes building an emotion relief database and constructing an emotional AI interaction engine; according to real-time multi-modal monitoring data, using the emotional AI interaction engine to generate real-time emotional recognition results and real-time AI interaction data; according to the real-time emotional recognition results, generating a real-time AI response strategy and extracting real-time emotion relief data; this invention solves the problems of low accuracy of emotional recognition, low intelligence level of AI, and lack of evolution function existing in the prior art.

[0005] Although the above technologies adopt multi-modal data, the modeling of cross-modal temporal correlation and context semantics still remains at the shallow fusion stage. The model relies on independently extracting single-modal features and then performing weighted fusion, resulting in the difficulty of accurately capturing the emotional progression relationship in the dialogue scenario and the coordinated changes between micro-expressions and speech intonations. In addition, the cloud multi-modal analysis engine proposed in the above patent does not clearly solve the problem of cross-modal dynamic weight allocation, which may lead to misjudgment due to modal conflicts. Moreover, the above system mostly uses a static emotion alleviation database, which cannot construct a time series portrait of the user's emotional state and does not consider long-term features such as the user's emotional preferences and stress patterns in historical interactions, resulting in the lack of personalization and continuity in the interaction strategy. Based on this, the present invention designs an AI multi-modal emotional interaction memory terminal to solve the above problems. Summary of the Invention

[0006] The purpose of the present invention is to provide an AI multi-modal emotional interaction memory terminal, which solves the problems of misjudgment and lack of personalization in the background technology.

[0007] To solve the above technical problems, the present invention provides the following technical solutions:

[0008] An AI multi-modal emotional interaction memory terminal, comprising:

[0009] Step S1, deploy a multi-threaded acquisition engine to synchronously obtain voice, facial expression, and text data, and achieve microsecond-level synchronization through hardware timestamp alignment; adopt a multi-head cross-attention mechanism to dynamically allocate modal weights, and the weight calculation satisfies: αv = ∑m∈{v,t,i}exp(sv)exp(sm), sm = Wm·hctx, where v, t, i respectively represent the voice, text, and visual modalities, hctx is the current dialogue context hidden state. When the cross-modal confidence difference is detected to exceed the threshold (Δ>0.4), activate the conflict resolution module based on gated LSTM;

[0010] Step S2, emotional memory modeling, model the emotional state transition probability through a graph convolutional network: P(et+1∣et) = Softmax(A·Wg et), where A is the user's emotional association matrix, et is the emotional vector at time t, and adopt a differential privacy mechanism to add noise, satisfying the privacy budget constraint of ≤∈1.5;

[0011] Step S3, end-cloud collaborative decision-making, the terminal deploys a pruning and quantization decision model, which is implemented on the NPU and supports policy optimization driven by the Bellman equation: Q(s,a)←r+γmaxa′Q(s′,a′), and the cloud aggregates feature gradients through sparse federated learning, and the update formula is: where η≤0.-2 is the compression learning rate;

[0012] Step S4, Incremental Evolution Mechanism. When the frequency of the detected new emotion pattern exceeds 5 times per hour, trigger an online update: L = λ∥fteacher(x) - fstudent(x)∥22 + (1 - λ)Ltask, where the knowledge transfer weight λ ≥ 0.6.

[0013] Preferably, in the step S1, construct a voice fundamental frequency - micro - expression dynamic association map, and analyze the real - time matching degree between the change rate of the voice signal fundamental frequency and the intensity of the facial action unit through the Pearson correlation coefficient; when the cross - modal confidence difference exceeds the threshold, trigger the visual modality weight attenuation mechanism and synchronously improve the semantic parsing priority of the text modality; use the sliding window mechanism to trace back the context of the previous three rounds of conversations to generate a modality calibration compensation vector.

[0014] Preferably, the implementation method of the emotion memory modeling in the step S2 includes: the short - term memory layer uses a circular buffer to store the original multi - modal data stream, supporting millisecond - level backtracking retrieval based on timestamps; the long - term memory layer constructs an emotion state transition topology through a graph convolutional network, and the node representation uses a two - channel encoding structure to store the emotion feature vector and the associated event metadata respectively; the implementation process of the differential privacy mechanism includes two - stage protection of trusted execution environment verification and gradient update noise injection.

[0015] Preferably, in the step S3, the edge - cloud collaborative decision - making includes: the terminal decision - making model uses channel pruning and layer fusion technologies to compress the scale of the computational graph and is deployed on a dedicated emotion - computing NPU core; during the aggregation process of cloud - side federated learning, feature gradient sparsification processing is implemented, and only the high - dimensional feature projections of cross - user emotion patterns are synchronized; the policy optimization module driven by the Bellman equation integrates a two - factor evaluation mechanism of an emotion consistency reward function and a user feedback score.

[0016] Preferably, the triggering conditions of the incremental evolution mechanism in the step S4 also include: the new emotion pattern reaches a preset consistency threshold in cross - modal verification; the terminal computing resource load rate is lower than the safe operation critical value.

[0017] Preferably, the calculation formula for the voice fundamental frequency change rate is:

[0018] ΔF0 / Δt = (F0(t) - F0(t - 1)) / Δt, where F0(t) is the voice fundamental frequency at the current moment, F0(t - 1) is the voice fundamental frequency at the previous moment, and Δt is the time interval;

[0019] The generation formula for the modality calibration compensation vector is:

[0020] C = α*V + β*T + γ*I, where α, β, and γ are weight coefficients, and V, T, and I are the feature vectors of the voice, text, and visual modalities respectively.

[0021] Preferably, the size of the circular buffer in the short-term memory layer is N, and the storage formula is: M = [m1, m2,..., mN], where mi is the multimodal data of the i-th round of conversation;

[0022] The formula for constructing the emotional state transition topology of the long-term memory layer is:

[0023] G = (V, E), where V is the set of nodes, E is the set of edges, the nodes represent emotional states, and the edges represent state transition probabilities.

[0024] Preferably, the trusted execution environment verification is achieved through digital signature and certificate authentication, and the formula is: Verify(σ, M) = true / false, where σ is the digital signature and M is the message;

[0025] The formula for gradient update noise injection is: g' = g + Lap(λ), where g is the original gradient, Lap(λ) is the Laplace noise, and λ is the privacy parameter.

[0026] Preferably, the implementation formula of the channel pruning and layer fusion technology is: M' = Prune(M, θ), where M is the original model, θ is the pruning threshold, and M' is the pruned model;

[0027] The formula for feature gradient sparsification processing is:

[0028] G' = Sparse(G, k), where G is the original feature gradient, k is the sparsification ratio, and G' is the sparsified feature gradient.

[0029] Compared with the prior art, the beneficial effects achieved by the present invention are:

[0030] 1. In the present invention, through the multi-threaded acquisition engine and the hardware timestamp alignment technology, the microsecond-level synchronization of multimodal data is achieved, ensuring the precise matching of voice intonation changes and facial expressions and movements during the conversation; at the same time, the multi-head cross-attention mechanism is adopted to dynamically allocate modal weights, which can automatically adjust the importance of each modality according to the current conversation context, thereby more accurately capturing the emotional progression relationship and the coordinated changes between micro-expressions and voice intonations.

[0031] 2. In the present invention, the emotional state transition probability is modeled through a graph convolutional network to construct a user-specific emotional state map, which can analyze the transition rules between different emotional nodes. The short-term memory layer uses a circular buffer to dynamically store multimodal raw data for nearly 30 minutes, supporting millisecond-level backtracking retrieval of the user's historical emotional fluctuations; the long-term memory layer constructs an emotional state transition topology through a graph convolutional network, and the node representation adopts a dual-channel encoding structure to ensure the secure isolation of the user's emotional data during storage and transmission.

[0032] 3. In the present invention, the terminal runs a lightweight decision-making model with a dedicated emotion computing NPU chip. Through channel pruning technology, the model's computational workload is compressed to 40% of the original scale, achieving a hard real-time guarantee for emotional feedback response within 200 ms. The cloud collaborative platform adopts a federated learning framework, only aggregating cross-user emotional feature projection data, and dynamically optimizing the global emotional understanding model, further improving the accuracy and efficiency of decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is the multi-modal data acquisition and fusion flow chart of the present invention;

[0034] Figure 2 is the emotional memory modeling flow chart of the present invention;

[0035] Figure 3 is the end-cloud collaborative decision-making flow chart of the present invention;

[0036] Figure 4 is the emotional interaction flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0038] Embodiment 1;

[0039] Please refer to Figures 1 - 4 , in the embodiment of the present invention, an AI multi-modal emotional interaction memory terminal is characterized by including:

[0040] Step S1, deploying a multi-threaded acquisition engine to synchronously obtain voice, facial expression, and text data, and achieving microsecond-level synchronization through hardware timestamp alignment;

[0041] Adopting a multi-head cross-attention mechanism to dynamically allocate modal weights, and the weight calculation satisfies:

[0042] αv = ∑m∈{v,t,i}exp(sv)exp(sm), sm = Wm·hctx,

[0043] where v, t, i respectively represent the voice, text, and visual modalities, hctx is the current dialogue context hidden state. When it is detected that the cross-modal confidence difference exceeds the threshold (Δ>0.4), activate the conflict resolution module based on gated LSTM;

[0044] Step S2, Emotional Memory Modeling, modeling the emotional state transition probability through a graph convolutional network:

[0045] P(et+1∣et) = Softmax(A·Wg et), where A is the user emotion association matrix, et is the emotion vector at time t, and noise is added using the differential privacy mechanism, satisfying the privacy budget constraint of ≤∈1.5;

[0046] Step S3, Edge-Cloud Collaborative Decision Making, the terminal deploys a pruning and quantization decision model, which is implemented on the NPU and supports policy optimization driven by the Bellman equation:

[0047] Q(s,a) ← r + γmaxa′Q(s′,a′),

[0048] The cloud aggregates feature gradients through sparse federated learning, and the update formula is: where η ≤ 0.-2 is the compression learning rate;

[0049] Step S4, Incremental Evolution Mechanism, when it is detected that the frequency of the new emotion pattern exceeds 5 times per hour, trigger online update:

[0050] L = λ∥fteacher(x) - fstudent(x)∥22 + (1 - λ)Ltask, where the knowledge transfer weight λ ≥ 0.6.

[0051] In Step S1, construct a voice fundamental frequency - micro-expression dynamic association map, and analyze the real-time matching degree between the change rate of the voice signal fundamental frequency and the intensity of facial action units through the Pearson correlation coefficient; when it is detected that the cross-modal confidence difference exceeds the threshold, trigger the visual modality weight attenuation mechanism and synchronously improve the semantic parsing priority of the text modality; use the sliding window mechanism to trace back the context of the previous three rounds of conversations and generate a modality calibration compensation vector.

[0052] The implementation method of emotional memory modeling in Step S2 includes: the short-term memory layer uses a circular buffer to store the original multi-modal data stream, supporting millisecond-level backtracking retrieval based on timestamps; the long-term memory layer constructs an emotional state transition topology through a graph convolutional network, and the node representation adopts a dual-channel coding structure, storing the emotional feature vector and associated event metadata respectively. The implementation process of the differential privacy mechanism includes two-stage protection of trusted execution environment verification and gradient update noise injection.

[0053] In Step S3, edge-cloud collaborative decision making includes: the terminal decision model uses channel pruning and layer fusion technologies to compress the scale of the computational graph and is deployed on a dedicated emotion computing NPU core; during the cloud federated learning aggregation process, feature gradient sparsification processing is implemented, and only the high-dimensional feature projections of cross-user emotion patterns are synchronized; the policy optimization module driven by the Bellman equation integrates a dual-factor evaluation mechanism of an emotion consistency reward function and user feedback scoring.

[0054] The triggering conditions for the incremental evolution mechanism in step S4 also include: the new emotion pattern reaches a preset consistency threshold in cross-modal verification; the load rate of the terminal computing resources is lower than the critical value for safe operation.

[0055] The working principle of the embodiments of the present invention is as follows: In the multi-modal data collection and fusion stage, the terminal synchronously obtains voice, facial expression, and text data by deploying a multi-threaded collection engine, and realizes microsecond-level synchronization by means of the hardware timestamp alignment technology, ensuring the precise alignment of multi-source data on the time axis. Subsequently, the multi-head cross-attention mechanism is used to dynamically allocate modal weights, and the weight calculation formula is αv = ∑m∈{v,t,i}exp(sv)exp(sm), sm = Wm·hctx, where v, t, i respectively represent the voice, text, and visual modalities, and hctx is the hidden state of the current dialogue context. When it is detected that the cross-modal confidence difference exceeds the threshold (Δ>0.4), the conflict resolution module based on gated LSTM is activated to resolve the conflicts between modalities and ensure the accuracy of data fusion.

[0056] In terms of emotion memory modeling, the terminal uses a graph convolutional network to model the emotion state transition probability, and the formula is P(et+1∣et) = Softmax(A·Wg et), where A is the user emotion association matrix and et is the emotion vector at time t. At the same time, the differential privacy mechanism is used to add noise to meet the privacy budget constraint of ≤∈1.5 to ensure the privacy security of user emotion data. The short-term memory layer uses a circular buffer to store the original multi-modal data stream, supporting millisecond-level retrospective retrieval based on timestamps; the long-term memory layer constructs an emotion state transition topology through a graph convolutional network, and the node representation adopts a dual-channel coding structure, storing the emotion feature vector and the associated event metadata respectively.

[0057] In the process of terminal-cloud collaborative decision-making, the terminal deploys a pruned quantization decision model, which is implemented on the NPU and supports policy optimization driven by the Bellman equation. The formula is Q(s,a)←r+γmaxa′Q(s′,a′). The cloud aggregates the feature gradients through sparse federated learning, and the update formula is where η≤0.-2 is the compression learning rate. The terminal decision model compresses the scale of the computational graph by using channel pruning and layer fusion technologies and is deployed on a dedicated emotion computing NPU core; during the cloud federated learning aggregation process, feature gradient sparsification processing is implemented, and only the high-dimensional feature projections of cross-user emotion patterns are synchronized; the policy optimization module driven by the Bellman equation integrates a dual-factor evaluation mechanism of an emotion consistency reward function and a user feedback score.

[0058] The incremental evolution mechanism triggers online updates by detecting the frequency of new emotion patterns. When it exceeds 5 times per hour, the formula is L = λ∥fteacher(x)-fstudent(x)∥22+(1 - λ)Ltask, where the knowledge transfer weight λ ≥ 0.6. In addition, the new emotion pattern reaches the preset consistency threshold in cross-modal verification, and the terminal computing resource load rate is lower than the safe operation critical value, enabling the terminal to continuously learn and adapt to new emotion patterns, and continuously improving the intelligent level of emotional interaction and user experience.

[0059] Example 2;

[0060] Please refer to Figures 1 - 4 , in the embodiment of the present invention, the calculation formula of the voice fundamental frequency change rate is:

[0061] ΔF0 / Δt = (F0(t)-F0(t - 1)) / Δt, where F0(t) is the voice fundamental frequency at the current moment, F0(t - 1) is the voice fundamental frequency at the previous moment, and Δt is the time interval;

[0062] The generation formula of the modal calibration compensation vector is:

[0063] C = α*V + β*T + γ*I, where α, β, and γ are weight coefficients, and V, T, and I are the feature vectors of the voice, text, and visual modalities respectively.

[0064] The size of the circular buffer in the short-term memory layer is N, and the storage formula is: M = [m1,m2,...,mN], where mi is the multi-modal data of the i-th round of conversation;

[0065] The formula for constructing the emotional state transition topology of the long-term memory layer is:

[0066] G = (V,E), where V is the set of nodes, E is the set of edges, the nodes represent emotional states, and the edges represent state transition probabilities.

[0067] The trusted execution environment verification is achieved through digital signature and certificate authentication. The formula is: Verify(σ,M) = true / false, where σ is the digital signature and M is the message;

[0068] The formula for gradient update noise injection is: g' = g + Lap(λ), where g is the original gradient, Lap(λ) is the Laplace noise, and λ is the privacy parameter.

[0069] The implementation formula of the channel pruning and layer fusion technology is: M' = Prune(M,θ), where M is the original model, θ is the pruning threshold, and M' is the pruned model;

[0070] The formula for feature gradient sparsification processing is:

[0071] G' = Sparse(G, k), where G is the original feature gradient, k is the sparsification ratio, and G' is the sparsified feature gradient.

[0072] The working principle of the embodiments of the present invention is as follows: In the multi-modal data acquisition and processing stage, the terminal synchronously acquires voice, facial expression, and text data through a multi-threaded acquisition engine, and uses a hardware timestamp alignment to achieve microsecond-level synchronization. The voice fundamental frequency change rate is calculated by the formula ΔF0 / Δt = (F0(t) - F0(t - 1)) / Δt, where F0(t) is the voice fundamental frequency at the current moment, F0(t - 1) is the voice fundamental frequency at the previous moment, and Δt is the time interval. A multi-head cross-attention mechanism is used to dynamically allocate modal weights, and the weight calculation formula is αv = exp(sv) / ∑m∈{v,t,i}exp(sm), where v, t, i represent the voice, text, and visual modalities, and sm = Wm·hctx, and hctx is the hidden state of the current dialogue context.

[0073] When the cross-modal confidence difference exceeds the threshold (Δ > 0.4), the conflict resolution module based on gated LSTM is activated; at the same time, a voice fundamental frequency - micro-expression dynamic association graph is constructed, and the real-time matching degree between the voice signal fundamental frequency change rate and the facial action unit intensity is analyzed through the Pearson correlation coefficient. When the cross-modal confidence difference exceeds the threshold, a visual modality weight attenuation mechanism is triggered, the semantic parsing priority of the text modality is enhanced, and a sliding window mechanism is used to trace back the previous three rounds of dialogue context to generate a modality calibration compensation vector C = α·V + β·T + γ·I, where α, β, γ are weight coefficients, and V, T, I are the feature vectors of the voice, text, and visual modalities respectively.

[0074] In terms of emotional memory modeling, the short-term memory layer uses a circular buffer to store the original multi-modal data stream, and the storage formula is M = [m1, m2,..., mN], where mi is the multi-modal data of the i-th round of dialogue, supporting millisecond-level backtracking retrieval based on timestamps. The long-term memory layer constructs an emotional state transition topology through a graph convolutional network, and the formula is G = (V, E), where V is the set of nodes representing emotional states, and E is the set of edges representing state transition probabilities. The node representation adopts a dual-channel encoding structure, storing the emotional feature vector and the associated event metadata respectively.

[0075] The differential privacy mechanism is implemented through two-stage protection of trusted execution environment verification and gradient update noise injection. The trusted execution environment verification formula is Verify(σ, M) = true / false, where σ is the digital signature and M is the message; the gradient update noise injection formula is g′ = g + Lap(λ), where g is the original gradient, Lap(λ) is the Laplace noise, and λ is the privacy parameter.

[0076] During the end-cloud collaborative decision-making process, the terminal deploys a pruning and quantization decision-making model, which is implemented on the NPU and supports policy optimization driven by the Bellman equation. The formula is Q(s,a)←r+γmaxa′Q(s′,a′). The terminal decision-making model compresses the scale of the computational graph through channel pruning and layer fusion techniques. The formula is M′=Prune(M,θ), where M is the original model, θ is the pruning threshold, and M′ is the pruned model, and it is deployed on a dedicated emotion computing NPU core.

[0077] The cloud aggregates feature gradients through sparse federated learning, and the update formula is η≤0.2 is the compression learning rate, and feature gradient sparsification processing is implemented. The formula is G′=Sparse(G,k), where G is the original feature gradient, k is the sparsification ratio, and G′ is the sparsified feature gradient. Only the high-dimensional feature projections of cross-user emotion patterns are synchronized. The policy optimization module driven by the Bellman equation integrates a dual-factor evaluation mechanism of an emotion consistency reward function and user feedback scores.

[0078] The incremental evolution mechanism detects the frequency of new emotion patterns. When it exceeds 5 times / hour or meets other trigger conditions (such as the new emotion pattern reaches the preset consistency threshold in cross-modal verification, the terminal computing resource load rate is lower than the safe operation critical value, and the model version iteration exceeds three generations), it triggers an online update. The formula is L=λ∥fteacher(x)-fstudent(x)∥22+(1-λ)Ltask, where the knowledge transfer weight λ≥0.6.

[0079] Embodiment 3;

[0080] Please refer to Figures 1 - 4 , in the embodiment of the present invention, an AI multi-modal emotion interaction memory terminal is provided, and its implementation process is as follows: First, the terminal captures the user voice signal, facial micro-expression, and interactive text data stream in real time through a multi-thread parallel acquisition engine, and uses hardware-level timestamp alignment technology to achieve microsecond-level synchronization of multi-source data, ensuring the precise matching of voice intonation changes and expression actions during the conversation. In the data fusion stage, the system is built with a dynamic weight allocation module that automatically calculates the confidence weights of the voice, text, and visual modalities based on the current conversation context. When the detected cross-modal confidence difference exceeds 0.4, it preferentially improves the text semantic parsing level and activates the conflict resolution algorithm, generating a compensation vector by backtracking the previous three rounds of conversation history, effectively improving the accuracy of emotion interpretation.

[0081] Emotional memory modeling adopts a two-layer storage architecture. The short-term memory layer dynamically stores multimodal raw data for nearly 30 minutes through a circular buffer, supporting millisecond-level retrospective retrieval of the user's historical emotional fluctuations. The long-term memory layer constructs a user-specific emotional state map, analyzes the transfer rules between different emotional nodes using a graph neural network, and injects random noise that meets the 1.5 privacy budget to ensure the secure isolation of the user's emotional data during storage and transmission. When it is recognized that the user repeatedly shows an emotional migration path of joy-surprise-contemplation, the system automatically strengthens the weight of this path to improve the prediction ability of subsequent interactions.

[0082] In the decision response link, the terminal runs a lightweight decision-making model on a dedicated emotion computing NPU chip. Through channel pruning technology, the model's computational workload is compressed to 40% of the original scale, achieving a hard real-time guarantee for emotional feedback response within 200 ms. The cloud collaborative platform adopts a federated learning framework to aggregate cross-user emotional feature projection data and dynamically optimize the global emotion understanding model. When it is detected that the user continuously triggers a new composite emotion pattern 5 times per hour, the terminal automatically starts an online learning mechanism to transfer the knowledge of the cloud model to the local model, and the transfer weight is set to 0.65 to ensure that the system continues to evolve while maintaining stability.

[0083] The working principle of the embodiments of the present invention is as follows: Through a multimodal time series alignment and compensation mechanism, the emotional recognition response speed of the AI is increased to the natural rhythm of human conversation; an emotional state map with causal reasoning ability is constructed to enable the AI to have the long-term memory characteristics of growing with the user; a flexible learning architecture of end-cloud collaboration is adopted to dynamically optimize the emotional interaction model on the premise of ensuring user privacy, and finally achieve a highly anthropomorphic empathy interaction experience.

[0084] Working principle: In the multimodal data acquisition and fusion stage, the terminal synchronously acquires voice, facial expression, and text data through a multi-threaded acquisition engine, and uses hardware timestamp alignment technology to achieve microsecond-level synchronization; subsequently, a multi-head cross-attention mechanism is used to dynamically allocate modal weights. When it is detected that the cross-modal confidence difference exceeds the threshold, a conflict resolution module based on gated LSTM is activated to solve the conflicts between modalities and ensure the accuracy of data fusion.

[0085] In terms of emotional memory modeling, the terminal uses a graph convolutional network to model the emotional state transition probability. At the same time, a differential privacy mechanism is used to add noise to ensure the privacy and security of the user's emotional data. The short-term memory layer uses a circular buffer to store the original multimodal data stream, supporting millisecond-level retrospective retrieval based on timestamps; the long-term memory layer constructs an emotional state transition topology through a graph convolutional network, and the node representation adopts a dual-channel encoding structure to store emotional feature vectors and associated event metadata respectively.

[0086] During the end-cloud collaborative decision-making process, the terminal deploys a pruning and quantization decision model, which is implemented on the NPU and supports policy optimization driven by the Bellman equation. The cloud aggregates feature gradients through sparse federated learning, performs feature gradient sparsification processing, and only synchronizes the high-dimensional feature projections of cross-user sentiment patterns. The policy optimization module driven by the Bellman equation integrates a dual-factor evaluation mechanism of an emotion consistency reward function and a user feedback score. The incremental evolution mechanism triggers an online update by detecting the frequency of new sentiment patterns. When the frequency exceeds 5 times per hour or other triggering conditions are met, an online update is triggered. In addition, conditions such as the new sentiment pattern reaching the preset consistency threshold in cross-modal verification, the terminal computing resource load rate being lower than the safe operation critical value, and the cloud knowledge model version iterating more than three generations will also trigger the incremental evolution mechanism, enabling the terminal to continuously learn and adapt to new sentiment patterns, and continuously improving the intelligent level of emotional interaction and user experience.

[0087] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An AI multi-modal emotional interaction memory terminal, characterized in that, Including: Step S1, deploy a multi-threaded acquisition engine to synchronously obtain voice, facial expression, and text data, and achieve microsecond-level synchronization through hardware timestamp alignment; Adopt a multi-head cross-attention mechanism to dynamically allocate modal weights, and the weight calculation satisfies: αv = ∑m∈{v,t,i}exp(sv)exp(sm), sm = Wm·hctx, where v, t, i respectively represent the voice, text, and visual modalities, hctx is the current dialogue context hidden state. When the cross-modal confidence difference is detected to exceed the threshold (Δ>0.4), activate the conflict resolution module based on gated LSTM; Step S2, emotional memory modeling, model the emotional state transition probability through a graph convolutional network: P(et+1∣et) = Softmax(A·Wg et), where A is the user emotional association matrix, et is the emotional vector at time t, and add noise using the differential privacy mechanism to satisfy the privacy budget constraint of ≤∈1.5; Step S3, edge-cloud collaborative decision-making, deploy a pruning and quantization decision model at the terminal, implement it on the NPU, and support policy optimization driven by the Bellman equation: Q(s,a)←r+γmaxa′Q(s′,a′), The cloud aggregates feature gradients through sparse federated learning, and the update formula is: where η ≤ 0.-2 is the compression learning rate; Step S4, incremental evolution mechanism, when the frequency of the detected new emotional pattern exceeds 5 times per hour, trigger an online update: L = λ∥fteacher(x)-fstudent(x)∥22+(1-λ)Ltask, where the knowledge transfer weight λ≥0.

6.

2. The AI multimodal emotion interaction memory terminal according to claim 1, wherein: In the said step S1, construct a voice fundamental frequency - micro-expression dynamic association map, and analyze the real-time matching degree between the fundamental frequency change rate of the voice signal and the intensity of the facial action unit through the Pearson correlation coefficient; when the cross-modal confidence difference is detected to exceed the threshold, trigger the visual modality weight attenuation mechanism, and synchronously improve the semantic parsing priority of the text modality; adopt a sliding window mechanism to trace back the previous three rounds of dialogue context to generate a modal calibration compensation vector.

3. An AI multi-modal emotional interaction memory terminal according to claim 1, characterized in that: The implementation method of emotional memory modeling in the said step S2 includes: the short-term memory layer uses a circular buffer to store the original multi-modal data stream, supporting millisecond-level backtracking retrieval based on timestamps; the long-term memory layer constructs an emotional state transition topology through a graph convolutional network, and the node representation adopts a dual-channel coding structure to store the emotional feature vector and the associated event metadata respectively.

4. An AI multimodal emotional interaction memory terminal according to claim 1, characterized in that: In the said step S2, the implementation process of the differential privacy mechanism includes two-stage protection of trusted execution environment verification and gradient update noise injection.

5. An AI multi-modal emotional interaction memory terminal according to claim 1, characterized in that: In the said step S3, edge-cloud collaborative decision-making includes: the terminal decision model uses channel pruning and layer fusion technologies to compress the scale of the computational graph and is deployed on a dedicated emotional computing NPU core; During the cloud federated learning aggregation process, implement feature gradient sparsification processing, and only synchronize the high-dimensional feature projections of cross-user emotional patterns; the policy optimization module driven by the Bellman equation integrates a dual-factor evaluation mechanism of an emotional consistency reward function and a user feedback score.

6. An AI multi-modal emotional interaction memory terminal according to claim 1, characterized in that: The triggering conditions of the incremental evolution mechanism in the said step S4 also include: the new emotional pattern reaches a preset consistency threshold in cross-modal verification; the terminal computing resource load rate is lower than the safe operation critical value.

7. An AI multimodal emotion interaction memory terminal according to claim 2, characterized in that: The calculation formula for the voice fundamental frequency change rate is: ΔF0 / Δt = (F0(t) - F0(t - 1)) / Δt, where F0(t) is the fundamental frequency of speech at the current moment, F0(t - 1) is the fundamental frequency of speech at the previous moment, and Δt is the time interval; The generation formula for the modal calibration compensation vector is as follows: C = α*V + β*T + γ*I, where α, β, and γ are weight coefficients, and V, T, and I are the feature vectors of the speech, text, and visual modalities respectively.

8. An AI multi-modal emotional interaction memory terminal according to claim 3, characterized in that: The size of the circular buffer in the short-term memory layer is N, and the storage formula is: M = [m1, m2,..., mN], where mi is the multi-modal data of the i-th round of conversation; The formula for constructing the emotional state transition topology of the long-term memory layer is as follows: G = (V, E), where V is the set of nodes, E is the set of edges, the nodes represent emotional states, and the edges represent state transition probabilities.

9. An AI multi-modal emotional interaction memory terminal according to claim 4, characterized in that: The verification of the trusted execution environment is achieved through digital signature and certificate authentication, and the formula is: Verify(σ, M) = true / false, where σ is the digital signature and M is the message; The formula for gradient update noise injection is: g' = g + Lap(λ), where g is the original gradient, Lap(λ) is the Laplace noise, and λ is the privacy parameter.

10. An AI multi-modal emotional interaction memory terminal according to claim 5, characterized in that: The implementation formula for the channel pruning and layer fusion technology is: M' = Prune(M, θ), where M is the original model, θ is the pruning threshold, and M' is the pruned model; The formula for feature gradient sparsification processing is: G' = Sparse(G, k), where G is the original feature gradient, k is the sparsification ratio, and G' is the sparsified feature gradient.

Citation Information

Patent Citations

  • AI interaction method and system based on emotion recognition

    CN119476488A

Cited By

  • AI psychological tutoring system with emotion intelligence and privacy calculation double engines

    CN120585332A

  • Artificial intelligence psychological assessment method and device based on multiple modes

    CN120600318A

  • Operation intention recognition method, system and equipment based on multi-modal fusion and medium

    CN120873982A

  • Digital human interaction system and method based on multi-modal emotion recognition

    CN121116129A

  • Digital human interaction system and method based on multi-modal emotion recognition

    CN121116129B