Multi-modal fusion anti-immersion psychological intervention closed-loop method and system thereof

By employing a multimodal fusion-based anti-immersion psychological intervention closed-loop method, which combines EEG and speech signals to dynamically adjust virtual interaction and real-world behavioral tasks, the problem of insufficient recovery of real-world social functions caused by excessive virtual participation in existing systems has been solved, achieving individualized and effective recovery of real-world social functions.

CN122624801APending Publication Date: 2026-08-25ZHENGZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610802393.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-04
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing digital psychological intervention systems are designed to maximize virtual participation, which hinders the recovery of real-world social function in socially withdrawn individuals with mild depression and makes it impossible to achieve low-cost, large-scale implementation of behavioral activation therapy.

Method used

By employing a multimodal fusion-based anti-immersive psychological intervention closed-loop method, and utilizing the fusion of EEG emotional truth scores and voice emotional features, the frequency and intensity of embodied physical interaction, conversational AI emotional support, and structured behavioral activation tasks are dynamically adjusted to construct a longitudinal closed-loop analysis mechanism, thereby achieving a decreasing proportion of virtual interaction and an increasing proportion of real-world behavioral activation tasks.

Benefits of technology

It improved the robustness and ecological validity of emotion assessment, optimized individualized intervention strategies, promoted the gradual transition of users from virtual interaction to real-world social interaction, increased the frequency of real-world social behavior, and reduced digital dependence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122624801A_ABST
    Figure CN122624801A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of artificial intelligence and digital mental health, and specifically relates to a multi-modal fusion anti-immersion psychological intervention closed-loop method and system, which comprises: collecting multi-channel electroencephalogram signals and voice signals and fusing and outputting real-time emotion true value scores; calling three types of intervention means of somatic physical interaction device accompaniment, conversational artificial intelligence emotional support and structured behavior activation task according to the grading of the scores; performing dynamic reduction on virtual interaction means and dynamic increase on behavior activation task according to the anti-immersion principle; performing longitudinal closed-loop analysis on three types of core indicators according to the preset evaluation period; and reversely optimizing intervention strategy parameters according to the analysis results, so that the present application takes the recovery of real social functions as a goal orientation and avoids that the intervention system itself becomes a new source of digital dependence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of interdisciplinary technology of artificial intelligence and digital mental health, specifically involving a closed-loop method and system for anti-immersion psychological intervention that integrates EEG emotion perception, embodied physical interaction and behavior activation therapy for digital implementation. Background Technology

[0002] With the rapid development of digital mental health services, AI-based psychological intervention systems are gradually becoming an important supplement to traditional clinical psychotherapy. While existing digital psychological intervention products have made some progress in emotion recognition and immediate emotional support, they still face deep-seated structural challenges in intervention goal design, objective verification of intervention effects, and the digital transformation of clinical evidence-based therapies.

[0003] Chinese patent application CN115670463A discloses a depression detection system based on EEG emotional neural feedback signals. This system employs a technical solution combining multi-channel EEG signal acquisition, EEG emotional feature extraction, real-time emotion classification and prediction, and neural feedback presentation, forming a closed-loop link from EEG signal acquisition to emotion feedback presentation, achieving objective assessment of depression levels based on EEG signals. However, the core objective of this solution stops at the level of detecting depressive states, failing to further construct an execution link from detection results to proactive intervention behavior. It lacks a strategy scheduling mechanism for dynamically linking objective EEG indicators with specific intervention methods, and it does not address a systematic solution for guiding users from virtual interactive environments to real-world social scenarios.

[0004] Chinese patent application CN116661607A discloses an emotion regulation method and system based on multimodal emotional interaction. This scheme obtains the user's mental health status grading results through a mental health assessment report and uses a virtual avatar to match corresponding emotion regulation schemes according to three levels: normal, stressed, and crisis. However, this scheme's intervention relies entirely on purely digital interaction guided by the virtual avatar, lacking the tactile feedback provided by embodied physical interaction devices and the companionship experience in a real space. Furthermore, its grading mechanism is a static state matching rather than a strategy optimization that dynamically adjusts with the intervention process, failing to construct a feedback loop driving the intervention strategy based on the intervention effect.

[0005] The aforementioned existing technologies reveal a deep-seated, yet unrecognized, technological bottleneck in the field of digital psychological intervention: current digital intervention systems are generally designed with maximizing user virtual engagement in mind, using immersive interaction to increase user retention and usage time. For individuals with mild depression exhibiting social withdrawal, this design logic structurally conflicts with the treatment goal of promoting the recovery of real-world social function. The immediate emotional satisfaction users gain in the virtual companionship environment replaces their intrinsic motivation to actively seek real-world interpersonal contact. Increased virtual interaction time is accompanied by a simultaneous decrease in the frequency of real-world social behavior, making the intervention system itself a new source of digital dependence rather than a transitional bridge to real-world social interaction. Simultaneously, behavioral activation therapy, the most evidence-based approach in evidence-based psychology, requires therapists to guide patients back to real-world activities and social scenarios. This core requirement is fundamentally incompatible with the design logic of highly immersive digital products, preventing behavioral activation therapy from achieving faithful, low-cost, and scalable deployment on digital platforms to date. Summary of the Invention

[0006] To address the technical bottleneck of existing digital psychological intervention systems that prioritize maximizing virtual participation, which often exacerbates digital dependence rather than promoting real-world social function recovery in socially withdrawn individuals with mild depression, this invention provides a multimodal fusion anti-immersion psychological intervention closed-loop method and system. This system constructs a multi-layered, interconnected intervention framework based on EEG emotional truth scores as an objective criterion, the principle of diminishing anti-immersion as a core strategy, and a three-indicator vertical closed-loop feedback mechanism. While ensuring the user's emotional safety, it achieves individualized closed-loop psychological intervention with a goal of restoring real-world social function by adaptively decreasing the proportion of virtual interaction based on objective neurophysiological indicators while simultaneously increasing the intensity of real-world behavioral activation task guidance.

[0007] The technical solution of this invention is as follows:

[0008] A multimodal fusion anti-immersion psychological intervention closed-loop method, applied to the psychological intervention of mild depression and social withdrawal groups, includes the following steps: collecting multi-channel EEG and speech signals from users; extracting EEG emotional features from the multi-channel EEG signals and speech emotional features from the speech signals; fusing the EEG emotional features and speech emotional features to output a real-time emotional truth score; based on the score range of the real-time emotional truth score, invoking at least three different types of virtual and real gradient intervention methods, including embodied physical interaction device companionship, conversational AI emotional support, and structured behavioral activation tasks; and, according to the anti-immersion principle, based on the real-time emotional truth score... The difference between the current preset assessment period mean of the truth score and the initial intervention period mean represents the improvement in the real-time emotional truth score. The frequency and intensity of use of embodied physical interaction devices and conversational AI emotional support are dynamically reduced, while the guidance intensity of structured behavior activation tasks is dynamically increased. According to the preset assessment period, a longitudinal closed-loop analysis is performed on three core indicators: the change in the user's emotional truth score baseline, the increase in the frequency of real-world social behavior, and the decrease in the digital interaction dependence index. Based on the results of the longitudinal closed-loop analysis, the scoring interval threshold and the dynamically decreasing deceleration rate parameter of the tiered call are adjusted in reverse to drive the individualized rolling optimization of the intervention strategy.

[0009] This invention also provides a multimodal fusion anti-immersion psychological intervention closed-loop system for psychological intervention in groups with mild depression and social withdrawal. It includes: an emotion perception layer comprising a multi-channel EEG acquisition module and a voice acquisition module. The multi-channel EEG acquisition module acquires multi-channel EEG signals from the user and extracts EEG emotional features, while the voice acquisition module acquires the user's voice signals and extracts voice emotional features. The emotion perception layer fuses the EEG emotional features and voice emotional features to output a real-time emotion truth score. The intervention execution layer includes an embodied physical interaction device, a conversational AI emotional support module, and a behavior activation task generation module. The intervention execution layer calls three types of interventions based on the score range of the real-time emotion truth score. The intervention employs various methods, including a dynamic reduction in the frequency and intensity of use of the embodied physical interaction device companionship and conversational AI emotional support modules based on the improvement in real-time emotional truth scores, and a dynamic increase in the structured behavioral activation tasks output by the behavioral activation task generation module, in accordance with the anti-immersion principle. The effect evaluation layer includes a longitudinal closed-loop analysis module and a strategy optimization module. The longitudinal closed-loop analysis module performs longitudinal analysis on three core indicators—the change in the emotional truth score baseline, the increase in the frequency of real-world social behavior, and the decrease in the digital interaction dependence index—according to a preset evaluation cycle. The strategy optimization module adjusts the scoring interval threshold and deceleration rate parameters of the intervention execution layer based on the longitudinal analysis results, driving individualized rolling optimization of the intervention strategy.

[0010] The beneficial effects of this invention are as follows:

[0011] First, this invention provides a real-time emotion truth score by fusing multi-channel EEG signals and speech signals, offering objective neurophysiological evidence that surpasses self-report scales for intervention strategy scheduling. The mechanism lies in the fact that EEG signals directly reflect the electrical activity state of the cerebral cortex, independent of the user's subjective will, while speech emotion characteristics reflect the immediate state of the autonomic nervous system. The fusion of these two signals allows for cross-validation to eliminate the inherent noise bias of a single modality, producing a more reliable emotion assessment result than using either EEG or speech alone. Compared to using only a single EEG modality for emotion classification, this invention's dual-modal fusion scheme improves both the robustness and ecological validity of emotion assessment.

[0012] Second, this invention is the first to propose the anti-immersion principle and systematically embed it into the intervention strategy scheduling mechanism, achieving a gradual transition where the proportion of virtual interaction dynamically decreases as the user's real-time emotional truth score improves, while the intensity of guidance for real-world behavioral activation tasks increases synchronously. The mechanism lies in the fact that the dorsolateral prefrontal cortex of socially withdrawn individuals with mild depression exhibits low activation in real social situations, while the social cognitive processing requirements of virtual interaction scenarios are far lower than in real-world scenarios. Continuously providing virtual support will deprive the prefrontal social cognitive circuits of training opportunities. The anti-immersion reduction strategy, by gradually reducing low-cognitive-load virtual support and increasing high-cognitive-load real-world social tasks, forces the prefrontal social cognitive circuits to repeatedly activate under controllable pressure and complete adaptive reshaping. Compared to static matching strategies using fixed gradations, the dynamic reduction mechanism of this invention can achieve refined adaptive adjustment of the intervention intensity based on individual recovery progress.

[0013] Third, this invention constructs a complete feedback loop from intervention effect evaluation to reverse adjustment of strategy parameters through longitudinal closed-loop analysis of three core indicators: changes in the baseline of emotion truth score, the increase in the frequency of real-world social behavior, and the decrease in the digital interaction dependence index. The mechanism lies in the fact that the three types of indicators capture different aspects of the intervention effect from the neurophysiological, behavioral, and dependence dimensions, respectively. A lag in improvement in any dimension can be promptly identified and trigger strategy adjustments. The synergistic effect of the three-dimensional joint analysis makes the sensitivity and accuracy of strategy optimization far exceed that of single-indicator driven schemes. Attached Figure Description

[0014] Figure 1 This is a flowchart of the multimodal fusion anti-immersion psychological intervention closed-loop method provided in the embodiments of the present invention.

[0015] Figure 2 This is an architecture diagram of the multimodal fusion anti-immersion psychological intervention closed-loop system provided in the embodiments of the present invention. Detailed Implementation

[0016] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings. The following embodiments are used to illustrate the present invention but are not intended to limit the scope of the present invention.

[0017] See Figure 1 This invention provides a multimodal fusion anti-immersion psychological intervention closed-loop method, applicable to psychological intervention for individuals with mild depression and social withdrawal. The method comprises five core steps, each forming a deeply coupled closed-loop relationship.

[0018] Step S1: Multimodal emotion perception and real-time emotion truth score output. Step S1 collects the user's multi-channel EEG signals and speech signals, extracts EEG emotion features from the multi-channel EEG signals, extracts speech emotion features from the speech signals, and fuses the EEG emotion features and speech emotion features to output a real-time emotion truth score.

[0019] In terms of EEG emotion feature extraction, a consumer-grade multi-channel EEG acquisition headband (4 to 8 channels, sampling rate 256Hz) was worn to continuously acquire multi-channel EEG signals from the user. The acquired multi-channel EEG signals were bandpass filtered across a frequency range of 0.5Hz to 45Hz to remove power line interference and EMG artifacts. The filtered signals were then segmented with a window length of 1s and a sliding step size of 0.5s. Differential entropy features were extracted from the EEG signals within each time window, and differential entropy values ​​were calculated in the θ band (4 to 8Hz), α band (8 to 13Hz), β band (13 to 30Hz), and γ band (30 to 45Hz) to form a multi-band EEG emotion feature vector.

[0020] In terms of speech emotion feature extraction, the user's speech signal is continuously collected through the microphone of the terminal device at a sampling rate of 16kHz. Mel-frequency cepstral coefficients, fundamental frequency change rate, speech rate, energy envelope, and formant parameters are extracted from the speech signal to form a speech emotion feature vector.

[0021] In terms of dual-modal fusion output of real-time emotion ground truth scores, EEG emotion feature vectors and phonological emotion feature vectors are fused into a unified emotion state representation through an attention-weighted fusion mechanism and mapped to a continuous real-time emotion ground truth score. The fusion calculation formula is as follows:

[0022] ,

[0023] in: The real-time emotion truth score is a scalar with a value range of 0 to 100. It is dimensionless and is calculated by this formula. It represents the user's current comprehensive emotional state. The higher the value, the more positive the emotion. It is the sigmoid activation function. Map any real number to the interval (0, 1) to ensure the boundedness of the score value; The weight vector is a fusion vector with the same dimension as the concatenated feature vector. Its value range is (-1, 1), and it is dimensionless. It is obtained through an offline trained emotion classification model. Its function is to map the fused feature vector to a scalar emotion score. The EEG emotion feature vector has a dimension of [missing information]. (In this embodiment) (corresponding to 5 differential entropy features for each of the 4 frequency bands), dimensionless, obtained from the EEG feature extraction step S1, carrying emotion-related information about the electrical activity of the cerebral cortex. The speech emotion feature vector has a dimension of . (In this embodiment) (), dimensionless, obtained from the speech feature extraction step S1, carrying real-time state information of the autonomic nervous system; The attention weights for the EEG modal are scalars, ranging from (0, 1), dimensionless, and dynamically calculated by the attention mechanism. They reflect the contribution of the current EEG modality to emotion assessment. When the EEG signal quality is high... Tend to larger values; Let be the speech modal attention weight, which is a scalar with a value range of (0, 1), dimensionless, and satisfies the constraints. It is dynamically calculated by the attention mechanism and reflects the contribution of the current speech modality to the emotion assessment. The fusion bias term is a scalar with a value range of (-5, 5), is dimensionless, and is obtained through offline training. It is used to correct the offset of the fusion linear transformation. Superscript indicates vector transpose operation Dimensions , with weighted fusion vector dimension Multiply to get a scalar, then add the scalar. It is still a scalar. After being mapped by sigmoid, it is multiplied by 100 to obtain a dimensionless scalar with consistent dimensions.

[0024] Step S2: Initiate three categories of virtual-real gradient intervention methods in a tiered manner. Based on the real-time emotion truth score's current range, Step S2 initiates at least three different categories of virtual-real gradient intervention methods in a tiered manner. These at least three categories include embodied physical interaction device companionship, conversational AI emotional support, and structured behavioral activation tasks, arranged in a virtual-reality gradient.

[0025] In this embodiment, the scoring range is set to three levels and a left-closed, right-open rule is used to ensure unique boundary assignment: When the real-time emotional truth score is in the range [0, 35), it is determined to be a depressed state. The embodied physical interaction device is used first, and conversational AI emotional support is activated for immediate emotional stabilization intervention, while the push of structured behavior activation tasks is temporarily suspended. When the real-time emotional truth score is in the range [35, 65), it is determined to be a transitional emotional state. The embodied physical interaction device is used at a low frequency, and cognitive reconstruction guidance is mainly carried out through conversational AI emotional support, gradually introducing low-difficulty structured behavior activation tasks. When the real-time emotional truth score is in the range [65, 100], it is determined to be a stable emotional state. The frequency of embodied physical interaction device companionship and conversational AI emotional support is reduced to the lowest level, and structured behavior activation tasks are used as the main intervention means to guide users back to real-world social scenarios.

[0026] In this embodiment, the embodied physical interaction device adopts the form of a quadrupedal bionic robot dog, possessing a haptic feedback mode (delivering tactile stimulation such as hugs and pats through flexible surface materials), an accompanying behavior mode (autonomously moving within the user's living space while maintaining an appropriate distance from the user), and an emotion-synchronized expression mode (synchronously reflecting the currently detected user's emotional state through ear posture, tail wagging frequency, and vocalization frequency). The intensity of the interactive behavior is adjusted based on a real-time emotion truth score: the lower the score, the more proactive the interactive behavior and the higher the expression intensity; the higher the score, the more restrained the interactive behavior and the gradually decreasing presence.

[0027] Conversational AI-powered emotional support provides users with emotional support based on cognitive behavioral therapy principles through a natural language dialogue interface, including a three-tiered dialogue strategy: emotion naming guidance, cognitive distortion identification, and positive behavior suggestions.

[0028] Structured behavioral activation tasks are automatically generated by the behavioral activation task generation module based on social withdrawal behavior profiles. Each task includes a task name, task scenario (independent activity, two-person interaction, or small group participation), estimated duration, social difficulty level (level 1 to 5), and completion criteria. The task content gradually transitions from low-social-requirement independent activities (such as a 15-minute outdoor walk or a single supermarket shopping trip) to two-person interactions (such as a 10-minute phone call with an acquaintance or greeting a neighbor) and small group participation (such as attending a community event or having dinner with three or more friends).

[0029] Step S3: Execution of the anti-immersion reduction strategy. Step S3, in accordance with the anti-immersion principle, dynamically reduces the frequency and intensity of the use of embodied physical interaction devices and conversational AI emotional support based on the difference between the current preset assessment cycle mean of the real-time emotion truth score and the initial cycle mean of the intervention, i.e., the improvement of the real-time emotion truth score, while dynamically increasing the guidance intensity of structured behavior activation tasks.

[0030] The core logic of the anti-immersion principle is that the ultimate goal of intervention systems is not to retain users, but to make users no longer need the system. Virtual interactive intervention methods (embody physical interactive device companionship and conversational AI emotional support) establish user trust and emotional baselines with high usage frequency and intensity in the early stages of intervention. Subsequently, their usage frequency and intensity are systematically reduced according to a decreasing function subject to multiple constraints, while the frequency and difficulty level of structured behavior activation tasks are increased according to an increasing function.

[0031] The dynamic decreasing function of the proportion of virtual interaction is defined as:

[0032] ,

[0033] in: For the first The proportion of virtual interaction at the end of each preset evaluation period is a scalar, with a value range of [value missing]. , dimensionless, is calculated by this formula, and represents the proportion of virtual interactive intervention methods in the total intervention plan; The decayable portion of the initial virtual interaction proportion is a scalar with a value of 0.70 (meaning the decayable range of the virtual interaction proportion at the start of intervention is 70%). It is dimensionless and is set by the system initialization parameters. This value reflects the dominant role of virtual support in the early stages of intervention. The deceleration rate parameter is a scalar quantity with an initial value range of 0.05 to 0.20, and its unit is dimensionless (because...). Also dimensionless, the individualized rolling optimization in step S5 dynamically adjusts the weight of virtual interaction to control the rate at which the weight of virtual interaction diminishes as the emotional improvement progresses. The larger the value, the faster the decay. Too large a value may cause user emotional instability, while too small a value may prolong the unnecessary virtual dependency time.

[0034] As of the date The cumulative emotional improvement over a preset assessment period is a scalar quantity, ranging from 0 to positive infinity, dimensionless, and calculated as follows: ,in For the first The mean of the real-time emotional truth score within a preset assessment period is normalized by dividing by 100. This normalization quantifies the cumulative positive improvement of the user's emotional truth score from the start of the intervention to the present. Only the cumulative positive improvement is considered, while fluctuations and declines are ignored, to avoid unreasonable rebounds in the decrement function when emotions fluctuate. The minimum safe value for the proportion of virtual interaction is a scalar value of 0.10 (meaning that even if the user recovers well, the proportion of virtual interaction will not be less than 10%). It is dimensionless and is set by the system initialization parameters. This value ensures that the user retains at least a minimum virtual emotional support channel throughout the entire intervention period to prevent a sharp rebound in the user's loneliness after the complete withdrawal of virtual support. It is a natural constant, approximately equal to 2.71828. Dimensionless Dimensionless, and their product is dimensionless. The dimensionless exponent is still a dimensionless scalar, multiplied by a dimensionless power. Add dimensionless The result is dimensionless, while the dimension is consistent.

[0035] The proportion of guidance from real-world activities will be synchronized. Calculate to ensure that the sum of the virtual and real weights is always 1.

[0036] Step S4: Vertical Closed-Loop Analysis of Three Indicators. Step S4 performs a vertical closed-loop analysis on three core indicators—the change in the user's emotional truth score baseline, the increase in the frequency of real-world social behavior, and the decrease in the digital interaction dependence index—according to a preset evaluation period. In this embodiment, the preset evaluation period is 14 days.

[0037] The change in the baseline of the emotional truth score is calculated as follows: the difference between the mean of all real-time emotional truth scores within the current preset assessment period and the mean of the previous preset assessment period. A positive value indicates emotional improvement, and a negative value indicates emotional deterioration.

[0038] The incremental frequency of real-world social behavior is obtained by: using calendar event markers on the user's mobile terminal, completion records of structured behavior activation tasks, and user-reported social activity logs to count the number of times the user participates in real-world social activities within the current preset evaluation period, and calculating the difference with the previous preset evaluation period.

[0039] The formula for calculating the digital interaction dependency index is:

[0040] ,

[0041] in: For the first The digital interaction dependence index for a preset evaluation period is a scalar with a value range of 0 to positive infinity (the actual range is approximately 0 to 3). It is dimensionless and is calculated by this formula. It represents the degree of user dependence on virtual interaction intervention methods. The higher the value, the stronger the dependence. The intervention goal is to make the index continuously decrease. The duration dimension weight is a scalar with a value of 0.40. It is dimensionless and is determined by the contribution of usage duration to behavioral dependence in behavioral psychology literature, reflecting the influence weight of usage duration on the degree of dependence.

[0042] For the first The cumulative interaction time between the user and the virtual interactive intervention method within a preset evaluation period is a scalar, with a value range from 0 to positive infinity, and the unit is min, which is automatically recorded by the system log. The reference duration baseline is a scalar value, which is the cumulative interaction duration within the first preset evaluation period of the intervention, in minutes. It is determined and locked once when the intervention is initiated, so that the duration values ​​of subsequent periods are normalized based on the first period. The frequency dimension weight is a scalar with a value of 0.35. It is dimensionless and reflects the influence of the frequency of actively initiated interactions on the degree of dependence.

[0043] For the first The number of times a user initiates virtual interactions within a preset evaluation period is a scalar, ranging from 0 to positive infinity, and is dimensionless (number of times). It is obtained by statistically analyzing records of user-initiated sessions or active wake-up of physical interaction devices in the system log. Its purpose is to distinguish between interactions actively pushed by the system and interactions actively sought by the user, the latter being a more direct indicator of dependent behavior. The reference frequency baseline is a scalar value, which is the number of active initiations in the first preset assessment cycle of the intervention. It is dimensionless and is locked when the intervention is initiated. The passive acceptance ratio is the weight of the dimension. It is a scalar with a value of 0.25. It is dimensionless and reflects the weight of the passive acceptance ratio on the degree of dependence.

[0044] For the first The proportion of user interactions passively received by the system within a preset evaluation period to the total number of interactions is a scalar quantity, ranging from 0 to 1, dimensionless, and derived from... The calculation yielded, where The higher the percentage of passively received interactions, the more the user's interactions are driven by passive system pushes rather than by the user's intrinsic needs, reflecting a stronger user-driven dependency (this percentage should gradually decrease as intervention progresses, consistent with the other two dimensions). The unit is min / min, which is dimensionless. Dimensionless Dimensionless, weighted sum of three terms It is a dimensionless scalar with consistent dimensions.

[0045] The calculation of the decline in the digital interaction dependence index is based on the current preset evaluation period. Compared with the previous preset evaluation cycle The negative value of the difference, i.e. Positive values ​​indicate a decrease (improvement) in dependence, while negative values ​​indicate an increase (worsening) in dependence.

[0046] The longitudinal closed-loop analysis performs trend analysis on the time series of the above three types of indicators over multiple preset evaluation periods. When at least two of the three types of indicators show a continuous improvement trend (positive change for three consecutive preset evaluation periods), the intervention is deemed to be effective overall, and the current deceleration rate parameter is maintained; when only one type or no indicator shows a continuous improvement, the intervention is deemed to be progressing poorly, triggering the strategy adjustment in step S5.

[0047] Step S5: Vertical closed-loop analysis drives individualized rolling optimization of the intervention strategy. Based on the results of the vertical closed-loop analysis, Step S5 adjusts the scoring interval threshold and the dynamically decreasing deceleration rate parameter of the tiered call in reverse, driving the individualized rolling optimization of the intervention strategy.

[0048] The specific adjustment rule is as follows: when the longitudinal closed-loop analysis determines that the intervention progress is not ideal, the system will automatically reduce the deceleration rate parameter. The step size is adjusted downwards by one step (0.02) to slow the decline in the proportion of virtual interaction, and the scoring interval threshold is simultaneously shifted downwards by 5 points (i.e., initiating lower-difficulty interventions more leniently) to avoid deteriorating user emotions due to excessively rapid decline. When the longitudinal closed-loop analysis determines that the intervention is effective overall and all three indicators show continuous improvement, the system automatically adjusts the deceleration rate parameter. Adjusting the step size upwards accelerates the decline in the proportion of virtual interaction and speeds up the transition of users to real-world social scenarios.

[0049] This closed-loop adjustment mechanism ensures that the intervention strategy can adapt to the different recovery speeds of different users. Users who recover faster will reduce their reliance on virtual support more quickly, while users who recover slower will have a longer period of virtual emotional buffer. The output of step S5 inversely affects the scoring interval threshold of step S2 and the deceleration rate parameter of step S3, forming a closed-loop driving loop from the effect evaluation layer to the intervention execution layer.

[0050] The following provides a further explanation of the specific implementation methods involved in the rights.

[0051] In constructing a social withdrawal behavior profile, a user's social withdrawal behavior profile is built before initiating at least three different types of virtual and real-world interventions. The profile records the user's avoidance intensity and duration for different social scenarios (five scenarios in total: alone time, phone conversations, one-on-one meetings, group interactions, and public social interactions). Avoidance intensity is assessed using a self-rating scale on a scale of 1 to 10, combined with behavioral observation. Avoidance duration records the cumulative number of days the user avoided that scenario type over the past 30 days. The social withdrawal behavior profile serves as a personalized benchmark for structured behavioral activation tasks. The order in which behavioral activation tasks are pushed is determined by ranking avoidance intensity from low to high, with the scenario with the lowest avoidance intensity corresponding to the first task pushed.

[0052] In terms of calibrating the asymmetry of alpha wave power in the prefrontal cortex, the changes in asymmetry of alpha wave power in the prefrontal cortex were extracted when constructing a social withdrawal behavior profile, in response to different descriptions of social scenarios. Specifically, five text descriptions or images of social scenarios were presented to the user, and multi-channel EEG signals were collected simultaneously. Alpha wave power was extracted from the left prefrontal cortex (corresponding to the F3 electrode position) and the right prefrontal cortex (corresponding to the F4 electrode position) in the alpha band (8 to 13 Hz), and the asymmetry index was calculated. ,in and The values ​​represent the alpha wave power values ​​for the right and left prefrontal cortex, respectively. Since alpha wave power is negatively correlated with cortical activation, a negative value for this asymmetry index indicates stronger relative activation in the right prefrontal cortex (approach-avoidance motivation), while a positive value indicates stronger relative activation in the left prefrontal cortex (approach-approach motivation). The asymmetry variation between different social scenario descriptions reflects the user's true approach-avoidance motivation intensity for that scenario. This is used to calibrate the user's self-reported avoidance intensity values, and updating the avoidance intensity of the social withdrawal behavior profile yields the calibrated social withdrawal behavior profile.

[0053] In terms of joint time-frequency-space sparse decomposition, due to the limited number of channels in consumer-grade EEG acquisition headbands (typically 4 to 8 channels), the standard calculation of prefrontal alpha wave power asymmetry has insufficient spatial resolution under limited channel conditions, and the difference signals between the left and right prefrontal lobes may be submerged by spatial aliasing noise between channels. To address this bottleneck, joint time-frequency-space sparse decomposition is performed on multi-channel EEG signals when extracting the changes in prefrontal alpha wave power asymmetry.

[0054] The optimization objective of the joint time-frequency-space sparse decomposition is:

[0055] ,

[0056] in: The estimated high spatial resolution source signal matrix has dimensions of . ( The number of virtual source points is taken in this embodiment. ; ), which is the number of sampling points within the time window, dimensionless (after normalization), is obtained by solving the optimization problem using this formula, representing the equivalent high spatial resolution brain power signal distribution recovered from the few-channel observation signal, from which the α wave power of the left and right prefrontal regions can be extracted to calculate the high-precision asymmetry estimate; The observed multi-channel EEG signal matrix has a dimension of [missing information]. ( The actual number of channels is shown in this embodiment. (), the unit is μV, and it is obtained directly from the EEG acquisition device; The lead field matrix (also known as the forward model matrix) has dimensions of . Dimensionless (after normalization), it is calculated by combining the standard head model with the electrode position and describes the conduction relationship from each virtual source point to each channel electrode; The square of the Frobenius norm is used to calculate the sum of squares of all elements in the matrix, which represents the fitting error between the observed data and the model-reconstructed data. is the time-frequency sparsity regularization coefficient, a scalar with a value ranging from 0.01 to 0.10, dimensionless, determined through cross-validation, and controls the sparsity of the source signal in the time-frequency domain. If it is too large, it will be excessively sparse and cause signal loss; if it is too small, it will be impossible to effectively utilize the time-frequency sparse prior of the α band. For matrix The L1 norm (the sum of the absolute values ​​of all elements) is used to leverage the time-frequency sparsity of EEG signals in the alpha band (alpha waves exhibit intermittent bursts rather than continuous oscillations within a specific time window) to promote the sparse representation of the source signal matrix. is the regularization coefficient for spatial topology constraints, which is a scalar with a value ranging from 0.05 to 0.50. It is dimensionless, determined through cross-validation, and controls the strength of spatial topology constraints. Let be the spatial Laplace constraint matrix, with dimension . It is dimensionless and constructed from the spatial topological prior knowledge of the left and right hemispheres in emotional scenarios (adjacent source points should have similar activities, and left and right symmetrical source points should exhibit specific asymmetric patterns in emotional tasks). The terms constrain the spatial smoothness of the source signal distribution and the topological consistency between the left and right hemispheres. The unit is μV² (dimensionless after normalization). and All are dimensionless, the three terms can be added together, and their dimensions are consistent.

[0057] After solving the above optimization problem, from the estimated source signal matrix The alpha wave power of the left prefrontal region (F3 corresponding to the source point set) and the right prefrontal region (F4 corresponding to the source point set) is extracted, and the asymmetric change of the EEG prefrontal alpha wave power is calculated with high spatial resolution.

[0058] Regarding the temporal profile of alpha wave power suppression depth and task engagement scoring, multi-channel EEG signals were continuously acquired during the user's performance of a structured behavioral activation task. After time-frequency-space joint sparse decomposition processing, the temporal profile curve of alpha wave power suppression depth during task execution was extracted. Alpha wave power suppression depth was defined as:

[0059] ,

[0060] in: For the first The alpha wave power suppression depth at time t is a scalar with a value range of (-∞, 1) (positive values ​​indicate suppression, negative values ​​indicate enhancement). It is dimensionless and is calculated by this formula. It characterizes the change in the degree of activation of the cerebral cortex relative to the baseline during task execution. For the first The α-wave power value of the prefrontal cortex at time t is a scalar, taking the range of positive real numbers, with the unit μV². It is obtained by extracting the α-band power of the source signal in the prefrontal region from the output of the time-frequency-space joint sparse decomposition. The baseline value of the prefrontal alpha wave power in the resting state before the start of the mission is a scalar value, with a range of positive real numbers and a unit of μV². The average alpha wave power is taken as the value 60 seconds before the start of the mission. The unit is μV² / μV², which is dimensionless. It is still dimensionless, but has the same dimensions.

[0061] During the entire task execution period Arranging the sequence chronologically creates a time profile curve. The formula for calculating task engagement score is:

[0062] ,

[0063] in: The task engagement score is a scalar with a value range of 0 to 1 (after normalization). It is dimensionless and is calculated by this formula. It represents the user's true cognitive engagement during the execution of structured behavioral activation tasks. The higher the score, the deeper the engagement. Tasks that are mechanically completed tend to have lower scores, while tasks that involve genuine social interaction tend to have higher scores. The α value is a scalar with a value of 0.60, which is dimensionless and reflects the contribution weight of the α value to the determination of task engagement. Let be the time-averaged alpha wave power suppression depth during mission execution. It is a scalar quantity, ranging from approximately -0.5 to 0.8, dimensionless, and derived from all values ​​on the time profile curve. The arithmetic mean is obtained; the larger the mean, the more deeply the cerebral cortex remains activated during the task, and the higher the level of engagement. The normalized weight for recovery time is a scalar with a value of 0.40. It is dimensionless and reflects the contribution weight of α power recovery time to the determination of task engagement. The time required for the alpha wave power to recover from the suppressed state to the baseline level after the mission is completed is a scalar value ranging from 0 to positive infinity, with the unit being seconds (s). It is defined as follows: The longer the recovery time from the end of the task to below 0.05, the deeper and more lasting the cortical activation caused by the task is, reflecting true social cognitive processing rather than a superficial response. The total duration of the task is denoted as , which is a scalar with a range of positive real numbers and a unit of seconds. It is determined by the time span from the start to the end of the task. Dimensionless multiplied by dimensionless It is dimensionless; The unit is s / s, which is dimensionless multiplied by dimensionless. It is dimensionless; the sum of two items is dimensionless, and they have the same dimension.

[0064] The task engagement score and user self-report (where users rate their task experience on a scale of 1 to 5 after completion) are combined as the task completion confirmation result. When the task engagement score is greater than 0.50 and the user self-report score is greater than 3, the task is considered to be effectively completed; otherwise, it is considered to be ineffective and an alternative task of equal difficulty must be arranged in the next preset evaluation cycle.

[0065] Regarding the Hidden Markov Model (HMM) social recovery phase model, after accumulating task completion confirmation results over multiple preset evaluation periods, the task engagement score sequence, the longitudinal change sequence of real-time emotional ground truth score, and the longitudinal change sequence of digital interaction dependence index are jointly input into the HMM. This model models the user's social function recovery process as a five-stage hidden state sequence. The five hidden states are: withdrawal steady state (users avoid most social scenarios, emotional scores are low, and the digital dependence index is high), virtual dependence (users begin to rely on virtual interactions for emotional support but have not yet transitioned to real-world scenarios), transitional adaptation (users begin to attempt low-difficulty real-world social tasks, virtual dependence begins to decrease, but emotional scores may fluctuate), reality regression (users are able to actively participate in moderately difficult real-world social activities, and the proportion of virtual interaction is significantly reduced), and self-maintenance (users' real-world social functions are basically restored, the need for virtual interaction is reduced to a minimum, and emotional scores remain stable at a high level).

[0066] The core parameters of a Hidden Markov Model for Social Recovery include the state transition probability matrix and the observation probability distribution. The state transition probability matrix is ​​defined as:

[0067] ,

[0068] in: To emerge from the hidden state Transition to hidden state The transition probability is a scalar, taking values ​​in the range [0, 1], dimensionless, and satisfies the following condition. It is learned from model parameters and represents the probability that a user will jump from one recovery phase to another. For the first Hidden state variables corresponding to each preset evaluation period; and These are the five hidden states, each representing the first of the five hidden states. The and the first indivual, Based on clinical priors regarding social recovery, the state transition probability matrix was initialized as a semi-constrained structure with relatively large diagonal elements (self-transition probabilities) (0.70 to 0.85), moderate positive transition probabilities between adjacent states (0.10 to 0.20), extremely low jump transition probabilities (0.01 to 0.05), and limited reverse transition probabilities (0.03 to 0.10).

[0069] Each latent state corresponds to a set of optimal virtual-real intervention weight parameters. The system infers the user's most likely current recovery stage latent state using the Viterbi algorithm and matches the corresponding optimal virtual-real intervention weight parameters based on the inferred recovery stage latent state. The virtual weight is 0.85 for retreat steady state, 0.70 for virtual dependency, 0.50 for transitional adaptation, 0.30 for realistic regression, and 0.15 for autonomous maintenance. This model achieves an upgrade from the rule-based linear decrease in step S3 to a nonlinear adaptive adjustment based on latent state inference.

[0070] Regarding the safety baseline constraints of the intervention strategy, when the real-time emotion truth score drops more than a preset safety threshold within a consecutive preset assessment period (in this embodiment, the preset safety threshold is 15 points, that is, the baseline change of the emotion score is negative for two consecutive preset assessment periods and the cumulative decrease exceeds 15 points), the proportion of virtual interaction is triggered to urgently rebound to the initial proportion of the intervention. The system will send an alert to the designated guardian's terminal. The emergency recovery state will be maintained for at least one designated assessment cycle. Once the baseline emotional score has recovered to more than 90% of the pre-recovery level, the normal decline process will resume.

[0071] Regarding the generation of intervention process visualization reports, a visualization report is automatically generated according to a preset assessment cycle. The intervention process visualization report includes a trend chart of emotion truth score changes (a line graph with the preset assessment cycle as the horizontal axis and the average real-time emotion truth score as the vertical axis), a chart of changes in the frequency of real-world social behaviors (a bar chart with the preset assessment cycle as the horizontal axis and the frequency of social behaviors as the vertical axis), and a chart of changes in the digital interaction dependence index (a line graph with the preset assessment cycle as the horizontal axis and the digital interaction dependence index as the vertical axis). The visualization report is pushed to the user's terminal and the preset guardian's terminal in electronic document format, helping users and guardians to intuitively understand the intervention progress.

[0072] See Figure 2 This invention also provides a multimodal fusion anti-immersion psychological intervention closed-loop system, which is used to implement the multimodal fusion anti-immersion psychological intervention closed-loop method described in the above-described method embodiments. The system includes three levels: an emotion perception layer, an intervention execution layer, and an effect evaluation layer, with a closed-loop signal circuit formed between the three levels.

[0073] The emotion perception layer comprises a multi-channel EEG acquisition module and a speech acquisition module. The multi-channel EEG acquisition module collects multi-channel EEG signals from the user and extracts EEG emotional features. At the hardware level, the multi-channel EEG acquisition module uses a consumer-grade portable EEG acquisition headband with 4 to 8 dry electrode channels, supporting Bluetooth wireless transmission to a mobile terminal. The speech acquisition module collects the user's speech signals and extracts speech emotional features, achieved through the mobile terminal's built-in microphone or an external noise-canceling microphone. The emotion perception layer includes a feature fusion submodule, which fuses EEG emotional features and speech emotional features according to the attention-weighted fusion mechanism in step S1, outputting a real-time emotion truth score, which is then transmitted to the intervention execution layer.

[0074] The intervention execution layer includes an embodied physical interaction device, a conversational AI emotional support module, and a behavior activation task generation module. In this embodiment, the embodied physical interaction device is a quadrupedal bionic robot dog equipped with a haptic feedback sensor array, an autonomous navigation module, an emotion synchronization expression controller, and a wireless communication module. It can move autonomously within the user's living space and adjust the intensity of its interactive behavior based on the real-time emotion truth score output by the emotion perception layer. The conversational AI emotional support module runs on a mobile terminal or cloud server, providing emotional support based on cognitive behavioral therapy principles to the user through a natural language dialogue interface. The behavior activation task generation module automatically generates a list of structured behavior activation tasks that meet the requirements of the current intervention stage based on the social withdrawal behavior profile and the current virtual-real intervention ratio parameter, and pushes it to the user's terminal. The intervention execution layer calls upon the above three types of intervention methods in a graded manner according to the score range of the real-time emotion truth score, and dynamically reduces the frequency and intensity of the use of the embodied physical interaction device and the conversational AI emotional support module according to the progress of emotion improvement, while dynamically increasing the structured behavior activation tasks output by the behavior activation task generation module, based on the anti-immersion principle.

[0075] The effectiveness evaluation layer comprises a longitudinal closed-loop analysis module and a strategy optimization module. The longitudinal closed-loop analysis module extracts three indicators from the system logs according to a preset evaluation cycle: changes in the baseline of the emotional truth score, the increase in the frequency of real-world social behaviors, and the decrease in the digital interaction dependence index. It performs longitudinal trend analysis and outputs evaluation conclusions. The strategy optimization module, based on the longitudinal analysis results, adjusts the scoring interval thresholds and deceleration rate parameters of the intervention execution layer, driving individualized rolling optimization of the intervention strategy. The output of the effectiveness evaluation layer, through the system's internal parameter update interface, acts back on the intervention execution layer, forming a complete closed-loop architecture of perception → execution → evaluation → optimization → execution.

[0076] In terms of system deployment, the multi-channel EEG acquisition module and voice acquisition module of the emotion perception layer are deployed on the user-end device. The computational load of EEG emotion feature extraction and voice emotion feature extraction is relatively small and can be completed locally on the mobile terminal. The feature fusion submodule, the conversational AI emotion support module and behavior activation task generation module of the intervention execution layer are deployed on cloud servers or edge computing nodes to ensure the computational resource requirements for dialogue generation and task matching. The embodied physical interaction device, as an independent physical terminal, maintains real-time communication with the mobile terminal and cloud server through a wireless communication module. The longitudinal closed-loop analysis module and strategy optimization module of the effect evaluation layer are deployed on the cloud server. At the end of each preset evaluation cycle, the three-indicator longitudinal analysis and strategy parameter update are automatically performed. The updated scoring interval threshold and deceleration rate parameters are distributed to each submodule of the intervention execution layer through the parameter synchronization interface.

[0077] The embodiments of the present invention are not limited to the specific embodiments described above. Those skilled in the art can make various equivalent changes or substitutions based on the technical solutions of the present invention, and all such changes or substitutions should be included within the protection scope of the present invention.

Claims

1. A multimodal fusion anti-immersion psychological intervention closed-loop method, applied to the psychological intervention of mild depression and social withdrawal groups, characterized by: Includes the following steps: Collect multi-channel EEG signals and speech signals from users, extract EEG emotion features from the multi-channel EEG signals, extract speech emotion features from the speech signals, and fuse the EEG emotion features and the speech emotion features to output a real-time emotion truth score. Based on the scoring range of the real-time emotion truth score, at least three different types of virtual and real gradient intervention methods are invoked in a graded manner. The at least three different types of virtual and real gradient intervention methods include embodied physical interaction device companionship, conversational artificial intelligence emotional support, and structured behavioral activation tasks. According to the anti-immersion principle, the difference between the current preset evaluation cycle mean of the real-time emotion truth score and the initial cycle mean of the intervention is used as the improvement range of the real-time emotion truth score. The frequency and intensity of the use of the embodied physical interaction device companionship and the conversational artificial intelligence emotional support are dynamically reduced, while the guidance intensity of the structured behavior activation task is dynamically increased. According to a preset evaluation period, a longitudinal closed-loop analysis is performed on three core indicators: the change in the baseline of the user's emotional truth score, the increase in the frequency of real-world social behavior, and the decrease in the digital interaction dependence index. Among them, the change in the baseline of the emotional truth score, the increase in the frequency of real-world social behavior, and the decrease in the digital interaction dependence index are determined by the negative values ​​of the difference between the mean of the real-time emotional truth score across the preset evaluation period, the difference between the number of effective completions of the structured behavior activation task across the preset evaluation period, and the difference between the digital interaction dependence index across the preset evaluation period, respectively. Based on the results of the longitudinal closed-loop analysis, the scoring interval threshold of the hierarchical call and the deceleration rate parameter of the dynamic decrease are adjusted in reverse to drive the individualized rolling optimization of the intervention strategy.

2. The method according to claim 1, characterized in that, Before invoking at least three different types of virtual and real gradient intervention methods in a tiered manner, the method further includes: constructing a user social withdrawal behavior profile, wherein the social withdrawal behavior profile records the user's avoidance intensity and duration for different types of social scenarios, using the social withdrawal behavior profile as a personalized benchmark for the structured behavior activation task, and determining the push order of the behavior activation task based on the avoidance intensity sorted from low to high.

3. The method according to claim 2, characterized in that, The process of constructing the social withdrawal behavior profile also includes: extracting the asymmetric change in the power of the prefrontal alpha waves of the EEG when the user faces different social scenario descriptions, calibrating the user's self-reported avoidance intensity value with the asymmetric change in the power of the prefrontal alpha waves of the EEG, and updating the avoidance intensity of the social withdrawal behavior profile to obtain the calibrated social withdrawal behavior profile.

4. The method according to claim 3, characterized in that, When extracting the asymmetric change in the power of the prefrontal alpha wave in the EEG, a time-frequency-space joint sparse decomposition is performed on the multi-channel EEG signal. By utilizing the time-frequency sparsity of the EEG signal in the alpha band and the spatial topological constraints of the left and right hemispheres, the asymmetric change in the power of the prefrontal alpha wave in the EEG with high spatial resolution is recovered from the limited-channel EEG signal.

5. The method according to claim 4, characterized in that, During the user's execution of the structured behavior activation task, the time profile curve of the alpha wave power suppression depth is extracted during the task execution process. The task engagement score is determined based on the suppression depth value and recovery time of the time profile curve. The task engagement score is combined with the user's self-report as the task completion confirmation result.

6. The method according to claim 5, characterized in that, After accumulating the task completion confirmation results of multiple preset evaluation cycles, the task engagement score sequence, the longitudinal change sequence of real-time emotion truth score, and the longitudinal change sequence of digital interaction dependence index are jointly input into the Hidden Markov Social Recovery Stage Model to infer the hidden state of the user's current recovery stage. Based on the inferred hidden state of the recovery stage, the corresponding optimal virtual-real intervention ratio parameter is matched, and the optimal virtual-real intervention ratio parameter is used as the target ratio to cover the current iteration value of the deceleration rate parameter.

7. The method according to claim 1, characterized in that, The interactive behavior modes of the embodied physical interaction device include tactile feedback mode, accompanying behavior mode and emotion synchronous expression mode, and the intensity of the interactive behavior is adjusted according to the real-time emotion truth score.

8. The method according to claim 1, characterized in that, When the real-time emotion truth score drops below a preset safety threshold within a consecutive preset evaluation period, the proportion of virtual interaction is triggered to urgently rebound to the initial proportion of intervention, and a warning notification is sent to the preset guardian's terminal.

9. The method according to claim 1, characterized in that, An intervention process visualization report is automatically generated according to the preset assessment cycle. The intervention process visualization report includes a trend chart of changes in the emotional truth score, a chart of changes in the frequency of real-world social behaviors, and a chart of changes in the digital interaction dependence index.

10. A multimodal fusion anti-immersion psychological intervention closed-loop system, applied to psychological intervention for individuals with mild depression and social withdrawal, characterized by: To implement the method according to any one of claims 1-9, comprising: The emotion perception layer includes a multi-channel EEG acquisition module and a voice acquisition module. The multi-channel EEG acquisition module is used to acquire multi-channel EEG signals from the user and extract EEG emotion features. The voice acquisition module is used to acquire the user's voice signals and extract voice emotion features. The emotion perception layer fuses the EEG emotion features and the voice emotion features to output a real-time emotion truth score. The intervention execution layer includes an embodied physical interaction device, a conversational AI emotional support module, and a behavior activation task generation module. The intervention execution layer calls three types of intervention methods according to the scoring range of the real-time emotion truth score. According to the anti-immersion principle, the frequency and intensity of the use of the embodied physical interaction device and the conversational AI emotional support module are dynamically reduced based on the improvement of the real-time emotion truth score, while the structured behavior activation tasks output by the behavior activation task generation module are dynamically increased. The effect evaluation layer includes a longitudinal closed-loop analysis module and a strategy optimization module. The longitudinal closed-loop analysis module performs longitudinal analysis on three core indicators—the change in the baseline of the emotional truth score, the increase in the frequency of real-world social behavior, and the decrease in the digital interaction dependence index—according to a preset evaluation cycle. The strategy optimization module adjusts the scoring interval threshold and deceleration rate parameters of the intervention execution layer in reverse based on the longitudinal analysis results, driving the individualized rolling optimization of the intervention strategy.

Citation Information

Patent Citations

  • Depression detection system based on electroencephalogram emotion nerve feedback signal

    CN115670463A

  • Emotion regulation method and system based on multi-modal emotion interaction

    CN116661607A