A psychological healing robot based on simulation analysis of high human emotional affinity

Through the system design of multimodal perception, dynamic decision-making, and brain science verification layers, the geographical limitations of traditional psychotherapy and the shortcomings of AI psychological products have been overcome. Personalized psychological intervention and effect quantification have been achieved, which is applicable to educational and clinical treatment scenarios and improves the efficiency and accuracy of emotion recognition and intervention.

CN120656649BActive Publication Date: 2025-12-02BEIJING PUJU HEALTH TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510811337.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-12-02
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

Traditional psychotherapy suffers from geographical limitations and a shortage of human resources. Existing AI-powered psychological products have shortcomings in emotion recognition, intervention strategies, and effect quantification, and cannot meet the needs of mental health services.

Method used

The system adopts an architecture design consisting of a multimodal perception layer, a dynamic decision-making layer, a generative interaction layer, and a neuroscience validation layer. It integrates affective computing, dynamic intervention, and neuroscience feedback technologies to form a closed-loop psychological healing system. The multimodal perception layer collects data, the dynamic decision-making layer generates personalized intervention strategies, the generative interaction layer provides empathic responses based on ethical norms, and the neuroscience validation layer regulates intervention strategies through EEG monitoring and neurofeedback.

Benefits of technology

It achieves accurate identification of complex emotional states, provides personalized psychological intervention, and quantifies the intervention effect through EEG monitoring. It is applicable to multiple scenarios such as education and clinical treatment, improves the efficiency and accuracy of emotion recognition, and achieves dual compliance with cross-cultural emotional metaphor mapping and technical ethics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656649B_ABST
    Figure CN120656649B_ABST
Patent Text Reader

Abstract

This invention discloses a psychological healing robot based on high human emotional affinity simulation analysis, comprising: a multimodal perception layer for collecting physiological, motor, and environmental interaction data; the multimodal perception layer includes a heterogeneous data acquisition module, a spatiotemporal feature extraction network, and an attention fusion mechanism module; a dynamic decision-making layer for generating intervention strategies based on interaction data; the dynamic decision-making layer includes a reinforcement learning strategy engine and a hierarchical intervention selection tree; a generative interaction layer for generating ethically compliant empathic responses based on intervention strategies; the generative interaction layer includes an ethical constraint system and an empathic response generator; and a neuroscience validation layer for real-time monitoring of neural feedback and adjustment of intervention strategies via EEG; the neuroscience validation layer includes a neural feedback regulation module and a multimodal feedback design module. This provides a precise and personalized psychological intervention decision-making closed loop, overcoming the shortcomings of existing AI psychological products in emotion recognition, intervention strategies, and effect quantification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary field of artificial intelligence and clinical psychology, specifically to an intelligent psychological healing system that integrates multimodal emotion computing, dynamic intervention engine and brain science feedback technology, thereby forming a psychological healing robot based on simulation analysis with high human emotional fit. Background Technology

[0002] Current mental health services face two major pain points:

[0003] 1. Traditional psychotherapy suffers from geographical limitations and a shortage of human resources;

[0004] 2. Existing AI-powered psychological products have three major flaws:

[0005] (1) Emotion recognition relies on a single text modality (such as using only the PHQ-9 scale), and the false negative rate for complex states such as "smiling depression" reaches 37%;

[0006] (2) Mechanized intervention strategies (such as fixed scripts) cannot achieve the dynamic cognitive reconstruction required for CBT treatment;

[0007] (3) Lack of brain function verification methods makes it difficult to quantify and evaluate the intervention effect.

[0008] Therefore, existing mental health service models and technologies urgently need to be innovated to meet the growing mental health needs. Summary of the Invention

[0009] The purpose of this invention is to provide a psychological healing robot based on simulation analysis with high human emotional affinity. The system, through a four-layer architecture, addresses the shortcomings of existing technologies, providing a precise and personalized closed-loop decision-making process for psychological intervention. It also overcomes the geographical limitations and human resource shortages inherent in traditional psychotherapy, and addresses the deficiencies of existing AI-powered psychological products in emotion recognition, intervention strategies, and effect quantification. The system integrates affective computing, dynamic intervention, and neuroscience feedback technologies through its architecture, which comprises a multimodal perception layer, a dynamic decision-making layer, a generative interaction layer, and a neuroscience verification layer, forming a closed-loop psychological healing system. This system can accurately identify complex emotional states, provide personalized psychological intervention, and quantify intervention effects through EEG monitoring, making it suitable for various scenarios such as education and clinical treatment.

[0010] The first aspect of this invention is to provide a psychological healing robot based on high human emotional fit simulation analysis, comprising:

[0011] A multimodal perception layer is used to collect physiological, motion, and environmental interaction data. This multimodal perception layer includes: a heterogeneous data acquisition module for integrating multi-type sensor data from the intelligent terminal's physiological, motion, and environmental interactions; a spatiotemporal feature extraction network for processing synchronous speech and video data streams based on a 3D-CNN spatiotemporal feature extraction network to obtain emotion calculation values ​​corresponding to spatiotemporal features; and an attention fusion mechanism module for dynamically adjusting the weights of each modality.

[0012] A dynamic decision layer is used to generate intervention strategies based on the interaction data through a reinforcement learning policy engine and a hierarchical intervention selection tree; wherein, the dynamic decision layer includes: a reinforcement learning policy engine, used to establish a state space and an action space and design a reward function; and a hierarchical intervention selection tree, used to trigger corresponding intervention protocols according to the crisis level and generate intervention strategies based on the reward function;

[0013] A generative interaction layer is used to generate an empathetic response that conforms to ethical norms based on the intervention strategy; wherein, the generative interaction layer includes: an ethical constraint system, used to impose ethical constraints based on an ethical rule base composed of a taboo content library, knowledge boundary constraints, privacy protection mechanisms, and an ethical knowledge graph; and an empathetic response generator, used to generate an empathetic response based on constraint decoding technology, and apply the ethical constraints to the empathetic response to generate an empathetic response that conforms to ethical norms;

[0014] A brain science validation layer is used to monitor neural feedback in real time via EEG and adjust the intervention strategy. The brain science validation layer includes: a neural feedback regulation module for monitoring changes in prefrontal alpha wave power in real time via EEG and using these changes as neural feedback; and a multimodal feedback design module for mapping neural feedback to virtual scene parameters based on the neural feedback regulation mechanism and dynamically adjusting the intervention strategy based on these virtual scene parameters.

[0015] Preferably, the heterogeneous data acquisition module employs four types of physiological monitoring sensors, three types of motion sensing sensors, and two types of environmental interaction sensors. The four types of physiological monitoring sensors include heart rate variability (PPG), skin conductance response (GSR), blood oxygen saturation (SpO2), and body temperature sensors. The three types of motion sensing sensors include a triaxial accelerometer for gait analysis, a gyroscope for posture recognition, and a barometer for spatial positioning. The two types of environmental interaction sensors include a microphone for voice emotion analysis and an ambient light sensor for circadian rhythm monitoring. The heterogeneous data acquisition module adopts a layered acquisition architecture, sequentially transmitting data from the sensor hardware layer, Bluetooth 5.3 or BLE protocol stack, edge computing nodes, and data preprocessing unit to the cloud with encryption. Noise reduction of raw data is achieved through edge nodes, and a unified data description framework is established to ensure compatibility with Android Health Connect and Apple. The HealthKit protocol enables spatiotemporal alignment of data from multiple devices, employing a dynamic interpolation algorithm to balance differences in sampling rates among different sensors. The heterogeneous data acquisition module is also used for multimodal transmission and synchronization, as well as preliminary fusion of multimodal data. The multimodal transmission and synchronization are implemented based on a low-power transmission protocol and a spatiotemporal synchronization mechanism. The preliminary fusion of multimodal data is implemented based on a feature-level fusion strategy. By inputting the original signal into a 3D-CNN network suitable for simultaneous voice and video analysis, spatiotemporal features are extracted. An energy value is calculated for each modality using a learnable feedforward neural network, and then the importance score is converted into a weight using a softmax function.

[0016] Preferably, the process of processing synchronized speech and video data streams based on the spatiotemporal feature extraction network 3D-CNN includes:

[0017] A 3D-CNN architecture for spatiotemporal feature extraction is established, comprising: a data input and preprocessing layer, wherein the data input to the data input and preprocessing layer is a three-dimensional input tensor constructed from synchronously acquired time-series speech waveforms and spatial RGB data video frames, with dimensions [T×H×W×C] (time×height×width×channel); a 3D convolutional kernel, which uses a three-dimensional convolutional kernel to slide synchronously on the time axis of the speech frame sequence and the spatial axis of the video frame to capture the spatiotemporal correlation between speech fundamental frequency jitter and facial micro-expressions; a spatiotemporal feature output layer, which extracts features stepwise through a 3D convolution + pooling structure based on hierarchical stacking and spatiotemporal correlation, and outputs a spatiotemporal feature map in the last layer of the spatiotemporal feature output layer; and a target optimization layer, which optimizes features based on a loss function, wherein the loss function is determined by fusing action localization error and emotion classification cross-entropy loss.

[0018] Initial sentiment values ​​are obtained by processing synchronous speech and video data streams using a 3D-CNN network based on spatiotemporal feature extraction.

[0019] Based on the initial sentiment value, sentiment calculation is performed using a contradiction index analysis algorithm to obtain the sentiment calculation value, wherein the contradiction index analysis algorithm includes:

[0020] (1) Standardize the optimized features, including: Z-Score standardization of text sentiment polarity and speech sentiment intensity to eliminate dimensional differences;

[0021] (2) Construct a contradiction index model and determine and arbitrate abnormal and contradictory data:

[0022] The formula for calculating the contradiction index C in the contradiction index model is shown in equation (1) below:

[0023] (1);

[0024] (2)

[0025] Where T text For the standardized text sentiment polarity, BERT_output represents the text sentiment polarity output by BERT, μ text For standard text sentiment polarity; T voice The speech emotion intensity is represented by the Pitch_Jitter fundamental frequency jitter rate, μ, after standardization. text Standard speech emotion intensity; σ text σ is the environment vector modulation factor; voice This is the sensor state attenuation factor; when the inconsistency index C > 2.5, the difference between the two modes exceeds 2.5 times the standard deviation, triggering a manual review process, including:

[0026] (A) Prioritize the use of physiological sensor data;

[0027] (B) If physiological data is missing, a manual review process will be initiated;

[0028] (3) Determining physiological arousal, including: calculating and classifying the physiological stress index through heart rate variability and skin conductance, wherein the formula for calculating the physiological stress index A is as follows:

[0029]

[0030] The classification is divided into three levels: low A < 0.5, medium 0.5 ≤ A ≤ 1.5, and high A > 1.5; among which, HRV LF / HF GSR indicates the lower or higher value of heart rate variability. Δ This indicates skin electrical conductance.

[0031] Preferably, the modal weights are dynamically adjusted based on an attention mechanism or a cross-modal attention mechanism based on Transformer; wherein, the formula for dynamically adjusting the modal weights based on the attention mechanism is shown in equation (3) below:

[0032] (3);

[0033] Among them, W i This represents the modal weight values. There are N modes in total, e f(xj) Attention scores are assigned to each modality; N modalities correspond to N sensors or feature extractors, and the feature vector extracted by each modality is f(x). j ), j=1,2,...,N; the goal is to obtain the weights of each modality through an attention mechanism, and then calculate the weighted feature vector as the final representation;

[0034] The formula for dynamically adjusting the modal weights of the Transformer-based cross-modal attention mechanism is shown in equation (4) below:

[0035] (4)

[0036] Where Q is the query matrix for the current sentiment state, and K... i The key matrix (Key) for each modality's eigenvectors; d k It is the dimension of the key vector, that is, the length of the key vector in each head;

[0037] Dynamically adjusting the modal weights includes:

[0038] (1) Feature projection: Mapping multimodal features to the same latent space;

[0039] (2) Attention aggregation: Calculate the weighted fusion features for downstream sentiment classification;

[0040] (3) Introduce a spatiotemporal attention gating mechanism to suppress the weight of features in time periods with low signal-to-noise ratio.

[0041] Preferably, establishing the state space includes:

[0042] (1) The intensity of emotion is calculated based on the numerical value of emotion computing. The intensity of emotion is represented by a 0-3 level quantitative index, which includes 0 representing calm, 1 representing mild anxiety, 2 representing moderate anxiety and 3 representing severe anxiety.

[0043] (2) Identify cognitive distortion types, including: a natural language processing-based thinking trap classifier identifies thinking traps in user input, identifies 12 types of cognitive distortion in user input and maps them to 12 predefined labels;

[0044] The types of cognitive distortions identified include:

[0045] (A) Establish a classification system for thinking traps to identify 12 predefined labels, which include: all or nothing, overgeneralization, psychological filtering, negation of positivity, jumping to conclusions, exaggeration or underestimation, emotional reasoning, should statements, labeling, personalization, catastrophizing, and mind reading;

[0046] (B) Key psycholinguistic features are obtained based on feature engineering, including: extreme lexical analysis, catastrophic lexical patterns, cognitive distortion keywords, emotional polarity intensity, and absolutist statement detection;

[0047] (C) Establish a deep learning model BiLSTM+Attention, including: text input layer, embedding layer, bidirectional LSTM, attention mechanism, cognitive feature input and multi-label classification output;

[0048] (D) A rule enhancement engine is established based on the post-processing of prediction rules in cognitive psychology, including: Rule 1: Increase confidence when "never / always" is included and the prediction is all or nothing; Rule 2: Strengthen the catastrophizing label with the "if...it's all over" pattern; Rule 3: Strengthen the "should" statement label with the second person "you should"; and Rule 4: Strengthen the mind-reading label with psychological verbs + thought assertions.

[0049] (E) Identify cognitive distortion types, including: text preprocessing including word segmentation and lemmatization, feature extraction, text serialization, model prediction, rule enhancement, result parsing based on confidence thresholds, interpretability analysis, creation of visualization models, acquisition of attention weights, and generation of heatmaps;

[0050] The establishment of the action space includes the establishment of an intervention strategy library corresponding to the action space. The intervention strategy library contains 6 core actions: CBT mind recording, mindfulness breathing training, virtual exposure therapy, medication dosage adjustment, crisis referral, and no intervention. Each core action corresponds to different resource consumption and expected efficacy.

[0051] The design reward function includes:

[0052] Design a short-term reward function The short-term reward function is characterized by the user's real-time mood decline ΔE and the weekly change rate ΔS of the PHQ-9 / GAD-7 scale scores, where the user's real-time mood decline... ΔEt Let E be the user's real-time emotion assessment value at time t. t Compared with the user's real-time emotion assessment value E at time t-1 t-1 The difference, i.e., the degree of emotional decline. Negative values ​​indicate improved mood, as stated in the short-term reward function. The reward formula is expressed as: ;

[0053] Wherein, α and β represent the weighting coefficients of the user's real-time emotional decline ΔE and the weekly change rate of the PHQ-9 / GAD-7 scale scores, respectively, which are determined through expert experience or optimization; γ represents the user's compliance reward, such as a positive compliance score and a bonus score if the intervention action is completed, and a negative compliance score and a deduction score if the intervention action is not completed.

[0054] Design a long-term reward function, which is a discounted cumulative reward based on the Q-learning objective, expressed as:

[0055] ;

[0056] Where k represents the daily intervention duration of the reward function;

[0057] The design constraint is that the daily intervention duration is ≤45 minutes, in order to prevent cognitive overload.

[0058] Preferably, the intervention protocol for triggering the corresponding intervention protocol according to the crisis level is triggered based on the crisis level determination rules, wherein the crisis level corresponds to high risk, medium risk and low risk, and the determination rules are PHQ-9 ≥ 20 or mention of suicidal ideation, emotional intensity lasting ≥ 2 for 72 hours and single emotional peak ≥ 1.5, and the corresponding intervention protocol triggered is immediate referral to offline diagnosis and treatment + 24-hour AI monitoring, CBT twice a day + mindfulness reinforcement and push relaxation audio + breathing guidance;

[0059] The intervention strategy based on the reward function includes using the Q-learning algorithm and state transition probability to calculate the real-time user sentiment decline ΔE obtained from the short-term reward function. t The weekly change rate of PHQ-9 and GAD-7 scale scores was assessed separately, and intervention strategies were generated based on the assessment results.

[0060] Preferably, the process of evaluating the real-time user sentiment decline ΔE and the weekly change rate of PHQ-9 / GAD-7 scale scores obtained from the short-term reward function based on the Q-learning algorithm and state transition probability, and generating an intervention strategy based on the evaluation results, includes a data modeling stage, a Q-learning algorithm implementation stage, and a strategy generation stage; wherein:

[0061] The data modeling phase includes defining the state space, defining the action space, and constructing a state transition probability model; defining the state space includes defining state variables and defining state representations, and defining state variables includes: the user's real-time emotion value E. t Historical PHQ-9 / GAD-7 ratings and User demographic characteristics and environmental context; the state is defined as follows: The defined action space includes determining intervention actions as follows: a1: pushing mindfulness meditation audio; a2: cognitive behavioral therapy (CBT) practice; a3: emergency human consultation; and a0: no intervention, silent observation; the goal of constructing the state transition probability model is to predict the action a. t Post-state s t →s t+1 The probability of state transition probability models includes the following methods: (1) training the probability model using historical data: ; and (2) the model is selected as a Hidden Markov Model (HMM) or a Bayesian Network, and the probability distribution is updated dynamically;

[0062] The Q-learning algorithm implementation stage includes determining the Q-table update rule and determining the state transition probability. The state transition probability is obtained by handling uncertainty through probabilistic Q-learning update.

[0063] The strategy generation phase includes establishing a dual-indicator evaluation system and generating dynamic strategies. The dual-indicator evaluation system includes measures for the magnitude of the sentiment decline ΔE. t Using statistical daily ΔE t The mean / variance is used to verify the effectiveness of the action. For the weekly change rate of the indicator scale, a linear regression analysis method is used to assess the improvement trend based on whether the slope is significantly negative. The dynamic generation strategy has safety constraints and fatigue control. The safety constraint is when ΔE... t When the threshold is reached, manual intervention is forcibly triggered; the fatigue control means that the same action will not be pushed repeatedly within 24 hours.

[0064] Preferably, the ethical constraint system includes:

[0065] (1) Taboo content library: Establish a dynamically updated blacklist, which includes sensitive words such as suicide, violence and / or discrimination, and covers ICD-11 psychological crisis entries;

[0066] (2) Knowledge boundary constraints: Prohibit the generation of medical diagnostic suggestions, and only allow the provision of general psychological support strategies;

[0067] (3) Privacy protection mechanism: The PII information mentioned by the user is automatically obfuscated by using a combination of regular expressions and NER recognition;

[0068] (4) Ethical knowledge graph, based on mermaid code and graph LR, is implemented with the following architecture:

[0069] A [User Statement] -> B {Ethics Review Node}

[0070] B->|Safety| C[Generate Empathic Response]

[0071] B->|Risk| D[Triggering Standard Script]

[0072] D->E ["Responding according to WHO Mental Health Guidelines, Section 2.3"];

[0073] The generation of empathic responses based on constraint decoding technology includes:

[0074] (1) The model architecture corresponding to the constraint decoding technology is determined to be the Llama-3-8B base model, and the model corresponding to the constraint decoding technology is trained based on the psychological counseling dialogue dataset;

[0075] (2) Dynamic constraints are injected into the model corresponding to the constraint decoding technique to generate an empathic response, wherein the dynamic constraints include:

[0076] A. Emotional state adaptation constraint, wherein the emotional state adaptation constraint is used to adjust the generated temperature parameter according to the user's current emotional intensity;

[0077] B. Multi-expert voting mechanism constraints, wherein the expert models corresponding to the multi-expert voting mechanism constraints include: ethical review model, crisis identification model, and emotional support assessment model;

[0078] C. Constraints on empathy enhancement strategies, including: embedding of psycholinguistic features and fusion of non-linguistic symbols, wherein:

[0079] The psycholinguistic feature embedding includes two levels of typical psycholinguistic embedding and cross-cultural emotional metaphor mapping library embedding; the two levels of typical psycholinguistic embedding include lexical and syntactic levels, the lexical level includes increasing the probability of words in the empathy dictionary during the decoding stage, and the syntactic level includes using interrogative sentences with a proportion >30% to promote user self-disclosure; the cross-cultural emotional metaphor mapping library embedding includes the embedding of a database formed by 327 localized expressions and expressions of Japanese onomatopoeia;

[0080] The non-linguistic symbol fusion includes: emotion embedding combined with a speech synthesis engine; virtual digital human facial expression synchronization; and keyboard keystroke interval analysis;

[0081] The method for verifying the ethical compliance of the ethically binding empathic response generated by applying the ethical constraints to the empathic response is as follows:

[0082] (1) Ethical compliance verification using an automated testing framework, including:

[0083] Form an adversarial test set: containing 500 high-risk, leading inputs to verify whether the system triggers standard crisis protocols;

[0084] Determine the ethical deviation index: D ethics =Number of non-compliant responses / Total number of test samples × 100%;

[0085] (2) Conducting ethical compliance verification based on human supervision mechanisms, including:

[0086] A double-blind review process was implemented: every 1,000 generated responses were independently reviewed by three licensed psychological counselors.

[0087] Dynamic update mechanism: The constraint rule library is updated monthly based on newly promulgated ethical guidelines.

[0088] Preferably, the neurofeedback modulation module includes:

[0089] (1) A signal acquisition system, comprising: a high-density EEG electrode and a real-time signal acquisition optimization unit; wherein, the high-density EEG electrode comprises an international 10-20 standard lead system and a flexible electrode array; the real-time signal acquisition optimization unit performs real-time signal acquisition optimization based on a combination of differential amplification technology, 120dB common-mode rejection ratio and bandpass filter;

[0090] (2) The alpha wave power real-time analysis system includes: a preprocessing and feature extraction unit and a reference power calibration unit; the preprocessing and feature extraction unit is used to remove physiological artifacts through independent component analysis and retain pure alpha wave signals; and to calculate the alpha wave power spectral density within a 0.5-second time window through short-time Fourier transform and output the dynamic change curve; the reference power calibration unit is used to establish a personalized baseline, and the process of establishing the personalized baseline includes: continuously monitoring for 3 minutes in a relaxed state with the user's eyes closed, calculating the average alpha wave power as the value of the personalized baseline, so as to quantify the change in neural activity.

[0091] Preferably, the step of mapping neural feedback to virtual scene parameters based on the neural feedback modulation mechanism includes:

[0092] The alpha wave power is mapped to virtual scene parameters through visual feedback; sound waves synchronized with the alpha wave are generated through auditory feedback and transmitted through bone conduction headphones.

[0093] The intervention strategy based on the dynamic adjustment of virtual scene parameters includes:

[0094] Based on the virtual scene parameters, a tiered difficulty control is implemented to dynamically adjust the intervention strategy; or

[0095] Closed-loop stimulation is performed based on the virtual scene parameters to dynamically adjust the intervention strategy.

[0096] Advantages of the system of the present invention:

[0097] 1. Based on the keyboard tapping interval analysis algorithm (>1.2 seconds to trigger the cognitive fatigue warning), the recognition efficiency and recognition accuracy are improved.

[0098] 2. Construct a cross-cultural emotional metaphor mapping library (a database formed by 327 localized expressions such as "feeling a heavy weight on the heart" in Chinese and expressions such as Japanese onomatopoeias).

[0099] 3. Propose a multi-modal arbitration mechanism driven by the contradiction index (to solve the scenario judgment of "crying while smiling").

[0100] 4. Deeply integrate constrained decoding with psychological theories to achieve double compliance with technology and ethics.

[0101] 5. Through precise α-wave dynamic monitoring and closed-loop regulation, directional intervention of neural plasticity is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0102] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the related art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the related art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0103] Figure 1 FIG. is a system architecture diagram provided according to an embodiment of the present invention;

[0104] Figure 2 FIG. is a schematic diagram of a hierarchical intervention decision tree provided according to an embodiment of the present invention;

[0105] Figure 3 FIG. is a multi-modal data fusion flowchart provided according to an embodiment of the present invention;

[0106] Figure 4 FIG. is a schematic diagram of EEG feedback regulation provided according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0107] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0108] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0109] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0110] See Figure 1 This embodiment provides a psychological healing robot based on high human emotional fit simulation analysis, including:

[0111] A multimodal sensing layer is used to collect physiological, motion, and environmental interaction data;

[0112] A dynamic decision-making layer is used to generate intervention strategies based on the interactive data through a reinforcement learning policy engine and a hierarchical intervention selection tree.

[0113] A generative interaction layer is used to generate ethically compliant empathetic responses based on the intervention strategy; and

[0114] A neuroscience validation layer is used to monitor neural feedback in real time via EEG and adjust the intervention strategy.

[0115] In a preferred embodiment, the multimodal sensing layer includes:

[0116] A heterogeneous data acquisition module is used to integrate multiple types of sensor data from intelligent terminals, including physiological, motion, and environmental interaction data.

[0117] In this embodiment, the working principle of the heterogeneous data acquisition module includes:

[0118] 1. Sensor classification and data acquisition architecture:

[0119] (1) Data source identification

[0120] Physiological monitoring (4 categories): heart rate variability (PPG), skin conductance (GSR), blood oxygen saturation (SpO2), and body temperature sensor;

[0121] Motion sensing devices (3 categories): triaxial accelerometer (gait analysis), gyroscope (body posture recognition), and barometer (spatial positioning);

[0122] Environmental interaction (2 categories): microphone (voice emotion analysis) and ambient light sensor (circadian rhythm monitoring).

[0123] (2) Layered acquisition architecture

[0124] A [Sensor Hardware Layer] -> B (Bluetooth 5.3 / BLE Protocol Stack); B -> C {Edge Computing Node}; C -> D [Data Preprocessing: Noise Reduction / Standardization]; D -> E [Encrypted Transmission to Cloud].

[0125] Denoising of raw data can be achieved through edge nodes (e.g., Kalman filtering to eliminate motion artifacts).

[0126] In this embodiment, the multiple sensor data refers to nine types of sensors, including:

[0127] (A) Heart Rate Variability (HRV) Sensor:

[0128] Type: Photoplethysmography (PPG) sensor

[0129] Function: Monitors the state of the autonomic nervous system (stress / relaxation level).

[0130] Metrics: RMSSD (Heart Rate Variability), HR (Real-time Heart Rate)

[0131] (B) Gestational Response (GSR) Sensor

[0132] Type: Dry electrode bioelectric sensor

[0133] Function: Measures skin conductance levels (emotional arousal).

[0134] Application: Quantitative assessment of anxiety levels

[0135] (C) Facial expression recognition camera

[0136] Type: Near-infrared 3D depth camera

[0137] Algorithm: Real-time running of the FER (Facial Expression Recognition) model

[0138] Identify: Six basic emotions (happiness / sadness / anger / surprise / fear / disgust)

[0139] (D) Voice emotion analysis microphone

[0140] Type: Directional MEMS microphone array

[0141] Analysis dimensions: intonation, speech rate, pause frequency, and spectral features

[0142] Output: Emotional tendency value (-1 anger to +1 pleasure)

[0143] (E) Eye-tracking sensor

[0144] Type: Infrared corneal reflective eye tracker

[0145] Parameters: Pupil diameter (cognitive load), fixation duration (attention), blink rate (stress level)

[0146] (F) Respiratory Pattern Monitor

[0147] Type: Piezoelectric breathing band / millimeter-wave radar

[0148] Characteristics: respiratory depth, respiratory rate, degree of respiratory rhythm disorder

[0149] Related: Early warning signs of anxiety attacks (>20 times / minute)

[0150] (G) Body motion attitude sensor

[0151] Type: 9-axis IMU (accelerometer + gyroscope + magnetometer)

[0152] Behavior recognition:

[0153] Curled-up posture (depressive tendency)

[0154] Restlessness (a sign of anxiety)

[0155] Trembling amplitude (stress response)

[0156] (H) Contact pressure sensor

[0157] Type: Capacitive touch sensor array (integrated on the robot surface)

[0158] Function:

[0159] A. Hug strength assessment (intensity of emotional needs)

[0160] B. Touch duration analysis (social avoidance tendency)

[0161] (I) Environmental context sensor

[0162] Composite module: Illuminance + Ambient noise + Temperature and humidity

[0163] Function: To correct environmental interference with physiological signals

[0164] For example, pupil constriction caused by strong light is not misinterpreted as a stress response.

[0165] 2. Standardization processing of heterogeneous data

[0166] Protocol Conversion: Establish a unified sensor description framework (USDF) that is compatible with Android Health Connect and Apple HealthKit protocols, enabling spatiotemporal alignment of data from multiple devices;

[0167] Sampling rate adaptation: A dynamic interpolation algorithm is used to balance the differences in sampling rates between different sensors (e.g., synchronizing 1Hz body temperature data with 25Hz acceleration data).

[0168] 3. Multimodal transmission and synchronization

[0169] (1) Low power transmission protocol: Multi-device concurrent transmission is achieved through the improved BLE protocol, and the bandwidth utilization rate is increased to 92% (37% higher than the traditional Bluetooth 4.2); QoS priority is set, with physiological data > motion data > environmental data.

[0170] (2) Time and space synchronization mechanism: The hardware level uses the PPS signal of the GPS module to achieve microsecond-level time synchronization; the software level uses the NTP protocol to compensate for network delay, and the error is controlled within ±15ms.

[0171] like Figure 3 As shown in Figure 4, preliminary fusion of multimodal data

[0172] Feature-level fusion strategy: Early fusion, inputting the original signal into a 3D-CNN network to extract spatiotemporal features (suitable for simultaneous speech + video analysis); The basic idea of ​​the attention mechanism is: calculate the importance score (energy value) of each modality through a learnable network (usually a feedforward neural network), and then convert the importance score into weights through the softmax function.

[0173] 5. Optimization of steps 1-4

[0174] By defining a universal metadata template, it supports the parsing of 327 device data formats (covering 95% of existing terminals); it adopts dual-channel redundant transmission: main channel (Wi-Fi) + backup channel (LoRaWAN); and it uses a dynamic sensor scheduling algorithm to turn off unnecessary sensors based on user status (such as disabling GPS during sleep).

[0175] A spatiotemporal feature extraction network is used to obtain sentiment calculation values ​​corresponding to spatiotemporal features after processing synchronous speech and video data streams based on a 3D-CNN spatiotemporal feature extraction network.

[0176] In this embodiment, the processing of synchronized speech and video data streams based on the spatiotemporal feature extraction network 3D-CNN includes:

[0177] 1. Establish a 3D-CNN architecture for spatiotemporal feature extraction, wherein the 3D-CNN architecture for spatiotemporal feature extraction includes:

[0178] The data input and preprocessing layer is constructed by using synchronously acquired speech waveforms (time series) and video frames (spatial RGB data) as a three-dimensional input tensor with dimensions [T×H×W×C] (time×height×width×channel).

[0179] 3D convolution kernels, including those employing three-dimensional convolution kernels, slide synchronously on the time axis (speech frame sequence) and the spatial axis (video frame) to capture the spatiotemporal correlation between speech fundamental frequency jitter and facial micro-expressions (such as the AU4 frowning unit).

[0180] The spatiotemporal feature output layer is used to extract features step by step through a 3D convolution + pooling structure based on hierarchical stacking and spatiotemporal correlation, and then outputs the spatiotemporal feature map in the last layer of the spatiotemporal feature output layer (example dimension: 8×16×16×512).

[0181] Target optimization layer: Feature optimization is performed based on a loss function, which is determined by fusing action localization error (such as voice-expression asynchrony) and emotion classification cross-entropy loss.

[0182] 2. Using a 3D-CNN network based on spatiotemporal feature extraction to process synchronized speech and video data streams to obtain initial sentiment values;

[0183] 3. Based on the initial sentiment value, sentiment calculation is performed using the contradiction index analysis algorithm to obtain the sentiment calculation value, wherein the contradiction index analysis algorithm includes:

[0184] (1) Standardize the optimized features, including Z-Score standardization of text sentiment polarity and speech sentiment intensity to eliminate dimensional differences.

[0185] (2) Construct a contradiction index model and determine and arbitrate abnormal and contradictory data:

[0186] Taking audio data as an example, the formula for calculating the contradiction index C in the contradiction index model is shown in equation (1) below:

[0187] (1);

[0188] (2)

[0189] Where T textFor the standardized text sentiment polarity, BERT_output represents the text sentiment polarity output by BERT, μ text For standard text sentiment polarity; T voice The speech emotion intensity is represented by the Pitch_Jitter fundamental frequency jitter rate, μ, after standardization. text Standard speech emotion intensity; σ text σ is the environment vector modulation factor; voice This is the sensor state attenuation factor;

[0190] An example of coding for arbitrating abnormal data is as follows:

[0191] def dynamic_weight_calculation(h_modalities, env_vector, sensor_status, prev_context):

[0192] """

[0193] h_modalities: A list of feature vectors for the 9 modalities

[0194] env_vector: Environmental sensor data [light intensity, noise level, temperature, humidity]

[0195] sensor_status: A dictionary of sensor statuses {'SNR', 'last_valid_time'}

[0196] prev_context: The LSTM state at the previous time step.

[0197] """

[0198] # 1. Calculate the basic attention score

[0199] e_scores = []

[0200] for h_i in h_modalities:

[0201] score = vT @ np.tanh(Wh @ h_i + Wc @ prev_context + b)

[0202] e_scores.append(score)

[0203] # 2. Calculate the environmental modulation factor

[0204] gamma_factors = []

[0205] for h_i in h_modalities:

[0206] combined = np.concatenate([h_i, env_vector])

[0207] gamma = sigmoid(U.T @ combined)

[0208] gamma_factors.append(gamma)

[0209] # 3. Calculate the time decay factor

[0210] tau_factors = []

[0211] current_time = time.time()

[0212] for status in sensor_status:

[0213] delta_t = current_time - status['last_valid_time']

[0214] tau = np.exp(-0.05 * delta_t * status['SNR'])

[0215] tau_factors.append(tau)

[0216] # 4. Calculate the final weights

[0217] numerator = []

[0218] for i in range(9):[[ID=;40]]

[0219] num = gamma_factors[i] * tau_factors[i] * np.exp(e_scores[i])

[0220] numerator.append(num)

[0221] denominator = sum(numerator)

[0222] [[ID=5;1]]alpha_weights = [num / denominator for num in numerator]

[0223] return alpha_weights

[0224] A manual review process is triggered when the inconsistency index C > 2.5. When the inconsistency index C > 2.5 (i.e., the difference between the two modes exceeds 2.5 times the standard deviation), the following strategy is triggered as a manual review process:

[0225] (A) Prioritize the use of physiological sensor data (e.g., skin conductance response (GSR) > 5 μS is considered as genuine anxiety);

[0226] (B) If physiological data is missing, a manual review process will be initiated.

[0227] The attention fusion mechanism module is used to dynamically adjust the weights of each modality.

[0228] In this embodiment, during the late fusion stage, the weights of each modality are dynamically adjusted based on an attention mechanism or a cross-modal attention mechanism based on Transformer.

[0229] 1. The formula for dynamically adjusting the weights of each modality based on the attention mechanism is shown in equation (3) below:

[0230] (3);

[0231] Among them, W i This represents the modal weight values. There are N modes in total, e f(xj) Attention scores are assigned to each modality.

[0232] Experiments show that this method improves the accuracy of identifying contradictory scenarios such as "voice trembling but facial expression calm" by 23%.

[0233] There are N modalities (corresponding to N sensors or feature extractors), and the feature vector extracted for each modality is f(x). j (j=1,2,...,N). The goal is to obtain the weights of each modality through an attention mechanism, and then calculate the weighted feature vector as the final representation.

[0234] 2. Based on the Transformer cross-modal attention mechanism, the weights are dynamically adjusted using formula (4):

[0235] (4)

[0236] Where Q is the query matrix for the current sentiment state, and K... i The key matrix (Key) for each modality's eigenvectors; d k It is the dimension of the key vector, that is, the length of the key vector in each head.

[0237] In a preferred embodiment, the dynamic adjustment of modal weights includes:

[0238] (1) Feature projection: Mapping multimodal features (text, speech, video) to the same latent space (dimension: 256);

[0239] (2) Attention aggregation: Calculate the weighted fusion features for downstream sentiment classification.

[0240] (3) Introduce a spatiotemporal attention gating mechanism to suppress the weight of features in low signal-to-noise ratio time periods (such as speech pauses).

[0241] The above approach achieves accurate analysis of multimodal emotional conflict scenarios by combining 3D-CNN spatiotemporal feature extraction with dynamic attention fusion.

[0242] In a preferred embodiment, the dynamic decision-making layer includes:

[0243] A reinforcement learning policy engine is used to establish the state space and action space and design the reward function;

[0244] In this embodiment, establishing the state space includes:

[0245] (1) The intensity of emotion is calculated based on the numerical value of emotion computing. The intensity of emotion is represented by a 0-3 level quantitative index, which includes 0 representing calm, 1 representing mild anxiety, 2 representing moderate anxiety and 3 representing severe anxiety.

[0246] (2) Identify cognitive distortion types, including: a natural language processing (NLP) based thinking trap classifier to identify thinking traps in user input (such as "all or nothing" and "catastrophic thinking"), identify 12 types of cognitive distortions in user input and map them to 12 predefined labels;

[0247] The Natural Language Processing (NLP)-based mind trap classifier combines deep learning, rule engines, and psycholinguistic features to form the following system architecture:

[0248] A [User-input text] --> B (Text preprocessing)

[0249] B --> C {Multi-level feature extraction}

[0250] C --> D [Word-level features]

[0251] C --> E [Syntactic Features]

[0252] C --> F [Semantic Features]

[0253] D --> G [Feature Fusion]

[0254] E --> G

[0255] F --> G

[0256] G --> H [Multi-label classifier]

[0257] H --> I [Output 12 types of probabilities]

[0258] I --> J [Rule Post-processing]

[0259] In a preferred embodiment, the identification of cognitive distortion types includes the following steps:

[0260] 1. Establish a classification system for mental traps (12 categories)

[0261] THINKING_TRAPS = {

[0262] 0: "all_or_nothing", # All or nothing

[0263] 1: "overgeneralization", # excessive generalization

[0264] 2: "mental_filter", # Mental filtering

[0265] 3: "disqualifying_positive", # Negate the positive

[0266] 4: "jumping to conclusions", # drawing conclusions prematurely

[0267] 5: "magnification_minimization", # Exaggerate / minimize

[0268] 6: "emotional_reasoning", # Emotional Reasoning

[0269] 7: "should_statements", # Statements should be made.

[0270] 8: "labeling", # Labeling

[0271] 9: "personalization", #personalization

[0272] 10: "catastrophizing", # catastrophizing

[0273] 11: "mind_reading" # Mind reading

[0274] }

[0275] 2. Obtaining key psycholinguistic features based on feature engineering

[0276] Python

[0277] def extract_cognitive_features(text):

[0278] features = {}

[0279] # Analysis of Extreme Vocabulary

[0280] extreme_words = ["always", "never", "every", "complete", "total","perfect"]

[0281] features["extreme_count"] = sum(text.lower().count(word) for wordin extreme_words)

[0282] # Disaster-related vocabulary pattern

[0283] catastrophe_patterns = [

[0284] r"disaster", r"worst ever", r"unbearable", r"ruin", r"destroy" ]

[0286] features["catastrophe_score"] = sum(len(re.findall(p, text)) forp in catastrophe_patterns)

[0287] # Cognitive Distortion Keywords

[0288] distortion_triggers = {

[0289] "should": r"\b(should|must|ought to)\b",

[0290] "labeling": r"\b(stupid|failure|loser|idiot)\b",

[0291] "mind_reading": r"\b(knows|thinks|believes)\b that I'm? "

[0292] }

[0293] for key, pattern in distortion_triggers.items():

[0294] features[key] = len(re.findall(pattern, text))

[0295] # Emotional Polarity Intensity

[0296] blob = TextBlob(text)

[0297] features["polarity_intensity"] = abs(blob.sentiment.polarity)

[0298] # Absolute statement detection

[0299] features["absolute_statements"] = len(re.findall(r"\b(all|none|noone|everyone)\b", text))

[0300] return features

[0301] 3. Build a deep learning model (BiLSTM + Attention)

[0302] Python

[0303] import tensorflow as tf

[0304] from tensorflow.keras.layers import Input, Embedding, Bidirectional,LSTM, Dense, Attention

[0305] def create_trap_detection_model(vocab_size=20000, max_len=100):

[0306] # Text Input Layer

[0307] text_input = Input(shape=(max_len,))

[0308] # Embedded layer

[0309] embedding = Embedding(vocab_size, 128)(text_input)

[0310] # Bidirectional LSTM

[0311] bilstm = Bidirectional(LSTM(64, return_sequences=True))(embedding)

[0312] # Attention Mechanism

[0313] attention = Attention()([bilstm, bilstm])

[0314] # Cognitive Feature Input

[0315] feature_input = Input(shape=(12,)) # 12 handcrafted features

[0316] concat = tf.keras.layers.concatenate([attention, feature_input])

[0317] # Multi-label classification output

[0318] output = Dense(64, activation='relu')(concat)

[0319] output = Dense(12, activation='sigmoid')(output) # 12 categories of mental traps

[0320] model = tf.keras.Model(inputs=[text_input, feature_input],outputs=output)

[0321] model.compile(loss='binary_crossentropy', optimizer='adam',metrics=['accuracy'])

[0322] return model

[0323] 4. Establish a rule enhancement engine

[0324] Python

[0325] def apply_cognitive_rules(text, predictions):

[0326] "Post-processing based on cognitive psychology rules"

[0327] # Rule 1: Increase confidence when "never / always" is included and the prediction is either all or none.

[0328] if ('never' in text or 'always' in text) and predictions[0] >0.3:

[0329] predictions[0] = min(1.0, predictions[0] + 0.2)

[0330] # Rule 2: The "If...then it's all over" pattern reinforces the catastrophic label.

[0331] if re.search(r"if .* (ruin|end|disaster)", text) and predictions

[10] > 0.4:

[0332] predictions

[10] = min(1.0, predictions

[10] + 0.3)

[0333] # Rule 3: Use the second person "you should" to strengthen the "should" statement label

[0334] if re.search(r"you (should|must|ought to)", text) and predictions[7] > 0.3:

[0335] predictions[7] = min(1.0, predictions[7] + 0.25)

[0336] # Rule 4: Mental verbs + thought assertions reinforce mind-reading tags

[0337] mind_read_verbs = ["know", "think", "believe", "feel"]

[0338] if any(verb in text for verb in mind_read_verbs) and "that" intext and predictions

[11] > 0.4:

[0339] predictions

[11] = min(1.0, predictions

[11] + 0.15)

[0340] return predictions

[0341] 5. Identify types of cognitive distortion

[0342] Python

[0343] def detect_thinking_traps(user_input):

[0344] # Text Preprocessing

[0345] cleaned_text = preprocess_text(user_input) # Includes word segmentation, lexical reconstruction, etc.

[0346] # Feature Extraction

[0347] linguistic_features = extract_linguistic_features(cleaned_text)

[0348] cognitive_features = extract_cognitive_features(cleaned_text)

[0349] # Text Serialization

[0350] sequence = tokenizer.texts_to_sequences([cleaned_text])

[0351] padded_seq = pad_sequences(sequence, maxlen=100)

[0352] # Model Prediction

[0353] raw_predictions = model.predict([padded_seq, np.array([cognitive_features])])[0]

[0354] # Rule Enhancement

[0355] final_predictions = apply_cognitive_rules(cleaned_text, raw_predictions)

[0356] # Result Analysis

[0357] detected_traps = []

[0358] for i, prob in enumerate(final_predictions):

[0359] if prob > 0.65: # Confidence threshold

[0360] detected_traps.append((THINKING_TRAPS[i], round(float(prob), 2)))

[0361] # Interpretability Analysis

[0362] explanation = generate_explanation(cleaned_text, detected_traps)

[0363] return {

[0364] "traps": detected_traps,

[0365] "explanation": explanation,

[0366] "original_text": user_input

[0367] }

[0368] def visualize_attention(text, model):

[0369] # Creating a Visual Model

[0370] attention_model = tf.keras.Model(

[0371] inputs = model.input,

[0372] outputs=model.get_layer("attention").output )

[0374] # Obtain attention weights

[0375] attn_weights = attention_model.predict(prepare_input(text))

[0376] # Generate heatmap

[0377] plt.figure(figsize=(15, 2))

[0378] sns.heatmap(attn_weights[0], annot=True, xticklabels=text.split())

[0379] plt.title("Cognitive Distortion Attention Weights")

[0380] plt.show()

[0381] This embodiment employs multimodal feature fusion and attention visualization interpretation, particularly utilizing word embedding features for semantic understanding, syntactic dependency features for relational analysis, cognitive linguistic features to expand domain knowledge, and incremental learning mechanisms to achieve good recognition results. These mechanisms include using user feedback loops and weekly incremental training: when a user marks a sample as "inaccurate," the sample enters a correction queue, while contrastive learning is used to enhance the recognition ability of difficult samples.

[0382] (3) Determining physiological arousal, including: calculating and classifying the physiological stress index through heart rate variability and skin conductance, wherein the formula for calculating the physiological stress index A is as follows:

[0383]

[0384] The classification is divided into three levels: low (A<0.5), medium (0.5≤A≤1.5), and high (A>1.5); among which, HRV LF / HF GSR indicates the lower or higher value of heart rate variability. Δ This indicates skin electrical conductance.

[0385] In this embodiment, establishing the action space includes establishing an intervention strategy library corresponding to the action space. The intervention strategy library contains 6 core actions: CBT mind recording, mindfulness breathing training, virtual exposure therapy, drug dosage adjustment, crisis referral and no intervention. Each core action corresponds to different resource consumption and expected efficacy.

[0386] In this embodiment, the design reward function includes:

[0387] Design a short-term reward function The short-term reward function is characterized by the user's real-time mood decline ΔE and the weekly change rate ΔS of the PHQ-9 / GAD-7 scale scores, where the user's real-time mood decline... ΔEt Let E be the user's real-time emotion assessment value at time t. t Compared with the user's real-time emotion assessment value E at time t-1 t-1 The difference, i.e., the degree of emotional decline. Negative values ​​indicate improved mood, as stated in the short-term reward function. The reward formula is expressed as: ;

[0388] Wherein, α and β represent the weighting coefficients of the user's real-time emotional decline ΔE and the weekly change rate of the PHQ-9 / GAD-7 scale scores, respectively, which are determined through expert experience or optimization; γ represents the user's compliance reward, such as a positive compliance score and a bonus score if the intervention action is completed, and a negative compliance score and a deduction score if the intervention action is not completed.

[0389] Design a long-term reward function, which is a discounted cumulative reward based on the Q-learning objective, expressed as:

[0390] ;

[0391] Where k represents the daily intervention duration of the reward function;

[0392] The design constraint is that the daily intervention duration is ≤45 minutes, in order to prevent cognitive overload.

[0393] like Figure 2 As shown, a tiered intervention selection tree is used to trigger corresponding intervention protocols according to the crisis level and generate intervention strategies based on the reward function.

[0394] In this embodiment, the intervention protocol for triggering the corresponding crisis level is triggered based on the crisis level determination rules shown in Table 1.

[0395] Table 1

[0396] Risk level Decision condition (logical OR) Triggering Protocol High risk PHQ-9 ≥ 20 or mention of suicidal ideation Immediate referral to in-person treatment + 24-hour AI monitoring Medium risk Emotional intensity lasting ≥ level 2 for 72 hours CBT twice daily + mindfulness reinforcement Low risk Single emotional peak ≥ 1.5 level Push relaxation audio + breathing guidance

[0397] In this embodiment, the intervention strategy generated based on the reward function includes the real-time user sentiment reduction ΔE obtained from the short-term reward function based on the Q-learning algorithm and state transition probability. t The weekly change rates of PHQ-9 and GAD-7 scale scores were evaluated separately, and intervention strategies were generated based on the evaluation results. The Q-learning algorithm was used to update the state-action value matrix based on historical intervention effects. Experiments showed that the strategy convergence speed was improved by 41% compared with the random strategy. The state transition probability was determined based on a Markov decision model constructed from 100,000+ clinical cases.

[0398] The process of evaluating the real-time user sentiment decline ΔE and the weekly change rate of PHQ-9 / GAD-7 scale scores obtained from the short-term reward function based on the Q-learning algorithm and state transition probability is divided into three stages: data modeling, Q-learning algorithm implementation, and strategy generation, combining reinforcement learning and probabilistic models.

[0399] I. Data Modeling Stage

[0400] 1. Define the state space.

[0401] Defined state variables include: the user's real-time sentiment value E t (Data collected in real time via sensors / questionnaires); Historical PHQ-9 / GAD-7 scores and (Updated weekly); User demographics (static features such as gender and age); and environmental context (time, location, activity type).

[0402] Define state representation:

[0403]

[0404] 2. Define the Action Space

[0405] The intervention actions include: a1: sending mindfulness meditation audio; a2: cognitive behavioral therapy (CBT) practice; a3: emergency human counseling; and a0: no intervention (silent observation).

[0406] 3. Construct a state transition probability model, the goal of which is to predict the execution of action a. t Post-state s t →s t+1 The probability of state transition probability models; methods for constructing state transition probability models include:

[0407] (1) Train the probabilistic model using historical data:

[0408] ;

[0409] (2) The model is selected as a Hidden Markov Model (HMM) or a Bayesian Network, and the probability distribution is updated dynamically.

[0410] II. Implementation Stage of Q-learning Algorithm

[0411] 1. Determine the Q-table update rules as follows:

[0412] ;

[0413] Where η represents the learning rate and γ represents the discount factor;

[0414] 2. Determine the state transition probability. Uncertainty is handled through probabilistic Q-learning updates to obtain the state transition probability as follows:

[0415] ;

[0416] 3. Algorithm Flow

[0417] For each user session:

[0418] Observe the current state s_t

[0419] For each timestep:

[0420] Choose action a_t using an ε-greedy strategy (exploration vs. exploitation)

[0421] Execute a_t, observe the reward r_t = R_short(s_t, a_t) and the transition state s_{t+1}.

[0422] Update the state transition probability model P(s_{t+1}|s_t, a_t).

[0423] Update Q-table:

[0424] Q(s_t, a_t) += η * [r_t + γ * max_a Q(s_{t+1}, a) - Q(s_t,a_t)]

[0425] s_t = s_{t+1}

[0426] III. Strategy Generation Phase

[0427] 1. Establish a dual-indicator evaluation system

[0428] index Evaluation methods <![CDATA[Emotional decline ΔE t > <![CDATA[Statistical daily ΔE t mean / variance, and verify the effectiveness of actions]]> Scale weekly change rate Does the slope of the linear regression analysis show a significantly negative trend (indicating improvement)?

[0429] 2. Generate dynamic strategies

[0430] Input: Current state ;

[0431] Output: Optimal action ;

[0432] • Constraints:

[0433] o Safety Limit: When ΔE t When the threshold is reached, manual intervention is forcibly triggered.

[0434] Fatigue control: The same action will not be pushed repeatedly within 24 hours;

[0435] IV. Verification and Iteration

[0436] 1. Offline validation: Backtesting using historical data to compare the performance improvement rates of Q-learning strategy versus rule-based strategy.

[0437] 2. Online A / B testing: Group users and verify the effectiveness of the new strategy on ΔE. t And the significance of the scale improvement rate (t-test).

[0438] 3. Dynamic parameter adjustment: Adjust the reward weights α, β based on user feedback (e.g., PHQ-9 has a higher weight than GAD-7).

[0439] Through a real-time decision-making mechanism, edge computing nodes perform a strategy evaluation every 5 minutes (latency <200ms). This solution achieves a precise and personalized psychological intervention decision-making closed loop by quantifying the state space and optimizing dynamic strategies. Of course, the foundation for this is the realization of "high human emotional fit simulation analysis".

[0440] The technical effects of hierarchical intervention selection trees:

[0441] 1. Probabilistic Q-learning: Introducing state transition probabilities to address the uncertainty of mental states;

[0442] 2. Dual-objective reward: Simultaneously optimize short-term mood decline (ΔE) and long-term clinical indicators (PHQ-9 / GAD-7).

[0443] 3. Safety mechanism: Set an emotional breakdown threshold to force a switch to manual intervention.

[0444] In a preferred embodiment, the generative interaction layer includes:

[0445] The ethical constraint system is used to impose ethical constraints based on an ethical rule base consisting of a taboo content library, knowledge boundary constraints, privacy protection mechanisms, and an ethical knowledge graph.

[0446] In this embodiment:

[0447] (1) Taboo content library: Establish a dynamically updated blacklist (including sensitive words such as suicide / violence / discrimination, covering ICD-11 psychological crisis entries).

[0448] (2) Knowledge boundary constraints: Prohibit the generation of medical diagnostic suggestions (such as "You should take XX drug"), and only allow the provision of general psychological support strategies;

[0449] (3) Privacy protection mechanism: Automatically obfuscates PII information such as names / addresses mentioned by users (using regular expressions + NER for joint recognition).

[0450] (4) Ethical knowledge graph, based on mermaid code and graph LR, is implemented with the following architecture:

[0451] A [User Statement] -> B {Ethics Review Node}

[0452] B->|Safety| C[Generate Empathic Response]

[0453] B->|Risk| D[Triggering Standard Script]

[0454] D-> E ["Responding according to WHO Mental Health Guidelines, Section 2.3"]

[0455] This map contains 2,000+ standard psychological support phrases (certified by the APA Ethics Committee).

[0456] An empathic response generator is used to generate empathic responses based on constraint decoding technology, and to apply the ethical constraints to the empathic responses to generate empathic responses that conform to ethical norms.

[0457] In this embodiment, the generation of empathic responses based on constraint decoding technology includes:

[0458] (1)Determine that the model architecture corresponding to the constrained decoding technology is the Llama-3-8B base model (Normalization layer: RMSNorm (Root Mean Square Normalization) is adopted, with a 30% higher computational efficiency than LayerNorm; Position encoding: upgraded to RoPE (Rotary Position Embedding), supporting dynamic context expansion to 128K tokens to solve the long-range dependence attenuation problem; The attention mechanism adopts grouped query attention (GQA) and sliding window attention (SWA); With MoE (Mixture of Experts) support, it can achieve dynamic sparse activation and video memory optimization, thereby stabilizing gated training through Z-loss regularization and avoiding routing oscillation), and train the model corresponding to the constrained decoding technology based on the psychological counseling dialogue dataset (including 100,000 labeled samples);

[0459] (2)Inject dynamic constraints into the model corresponding to the constrained decoding technology to generate empathic responses. The dynamic constraints include:

[0460] A. Emotional state adaptation constraint, which is used to adjust the generation temperature parameter according to the user's current emotional intensity; In this embodiment, when the user's current emotional intensity is in an anxious state, the temperature parameter temperature = 0.3, so as to correspond to the psychological healing robot to make a conservative response; When the user's current emotional intensity is in a calm state, the temperature parameter temperature = 0.7, so as to correspond to the psychological healing robot to make a creative response;

[0461] B. Multi-expert voting mechanism constraint, where the expert models corresponding to the multi-expert voting mechanism constraint include: ethical review model, crisis identification model and emotional support degree evaluation model.

[0462] C. Empathy enhancement strategy constraint, including: psycholinguistic feature embedding and non-verbal symbol fusion, where:

[0463] The psycholinguistic feature embedding includes two-level typical psycholinguistic embedding and cross-cultural emotional metaphor mapping library embedding; The two-level typical psycholinguistic embedding includes lexical level and syntactic level. The lexical level includes increasing the probability of words in the "empathy dictionary" (such as "feel the same" and "it's really hard for you") during the decoding stage. The syntactic level includes using a question sentence proportion > 30% (such as "Does this thing make you feel a lot of pressure?") to promote user self-disclosure; The cross-cultural emotional metaphor mapping library embedding includes the embedding of a database formed by 327 local expressions such as "having a heavy heart" in Chinese and Japanese onomatopoeias; <000103The non-verbal symbol fusion includes: emotion embedding combined with a speech synthesis engine (e.g., reducing speech rate by 20% and increasing fundamental frequency by 5% to express concern); virtual digital human facial expression synchronization (driving the 3D facial motion coding system FACS based on the response content); and keyboard tapping interval analysis (>1.2 seconds triggers a cognitive fatigue warning).

[0465] In this embodiment, the method for verifying the ethical compliance of the ethically binding empathic response generated by applying the ethical constraints is as follows:

[0466] (1) Ethical compliance verification using an automated testing framework, including:

[0467] Create an adversarial test suite: containing 500 high-risk, leading inputs (such as "I want to end everything") to verify whether the system triggers standard crisis protocols;

[0468] Determine the ethical deviation index: D ethics =Number of non-compliant responses / Total number of test samples × 100%;

[0469] After verification, the system D_{ethics} of this embodiment is 0.17% (better than the 2.3% error rate of human consultants).

[0470] (2) Conducting ethical compliance verification based on human supervision mechanisms, including

[0471] A double-blind review process is implemented: every 1,000 generated responses must be independently reviewed by 3 licensed psychological counselors.

[0472] Dynamic update mechanism: The constraint rule library is updated monthly based on newly promulgated ethical guidelines.

[0473] like Figure 4 As shown, in a preferred embodiment, the brain science verification layer includes:

[0474] The neural feedback modulation module is used to monitor the changes in the power of the alpha wave in the prefrontal cortex in real time via EEG, and to use the changes in the power of the alpha wave in the prefrontal cortex as neural feedback.

[0475] In a preferred embodiment, the neurofeedback modulation module includes:

[0476] (1) Signal acquisition system, including: high-density EEG electrodes and real-time signal acquisition optimization unit; wherein, the high-density EEG electrodes adopt the international 10-20 standard lead system, deploying silver / silver chloride dry electrodes in the prefrontal cortex region (Fp1 / Fp2 / Fz), with contact impedance controlled below 5kΩ, supporting a sampling rate of 256-512Hz per second, and adopting a new flexible electrode array to penetrate the stratum corneum and directly contact the dermis, improving the signal-to-noise ratio (SNR>10dB) of the α wave signal; the real-time signal acquisition optimization unit adopts differential amplification technology combined with a common-mode rejection ratio of 120dB to eliminate environmental electromagnetic interference (such as 50Hz power frequency noise), and also includes a bandpass filter (7-13Hz) to directly extract the α band signal and reduce the calculation delay to within 5ms.

[0477] (2) The alpha wave power real-time analysis system includes: a preprocessing and feature extraction unit and a reference power calibration unit; the preprocessing and feature extraction unit is used to remove physiological artifacts such as blinking and electromyography through independent component analysis (ICA) and retain pure alpha wave signals; and to calculate the alpha wave power spectral density (PSD) within a 0.5-second time window through short-time Fourier transform (STFT) and output the dynamic change curve; the reference power calibration unit is used to establish a personalized baseline, the process of establishing the personalized baseline includes: continuously monitoring for 3 minutes in a relaxed state with the user's eyes closed, calculating the average alpha wave power as the value of the personalized baseline, so as to quantify the change in neural activity.

[0478] The multimodal feedback design module is used to map neural feedback to virtual scene parameters based on the neural feedback modulation mechanism, and to dynamically adjust the intervention strategy based on the virtual scene parameters.

[0479] In a preferred embodiment, the step of mapping neural feedback to virtual scene parameters based on the neural feedback modulation mechanism includes:

[0480] The power of alpha waves is mapped to virtual scene parameters through visual feedback (e.g., in the "Mood Fruit Tree" game scene under the forest scene, the tranquility of the forest increases with the enhancement of alpha waves); sound waves (frequency 8-13Hz) synchronized with alpha waves are generated through auditory feedback and transmitted through bone conduction headphones.

[0481] As a preferred embodiment, the intervention strategy based on the dynamic adjustment of the virtual scene parameters includes:

[0482] The intervention strategy is dynamically adjusted by using tiered difficulty control based on the virtual scene parameters; for example, in this embodiment, when the alpha wave power exceeds the baseline of 1.5σ for 5 consecutive minutes, the complexity of the cognitive training task is automatically increased (e.g., memory matrix dimension +1); or

[0483] Closed-loop stimulation is performed based on the virtual scene parameters to dynamically adjust the intervention strategy; for example, in this embodiment, transcranial magnetic stimulation (TMS) is triggered to synchronize with alpha wave oscillations to enhance the functional connectivity of the prefrontal-limbic system (phase-locking accuracy ±15ms).

[0484] In this embodiment, the method for verifying and optimizing the metrics includes:

[0485] (1) Validation:

[0486] Alpha wave power enhancement rate: Clinical data shows that continuous training for 4 weeks can increase the power of prefrontal alpha waves by 22%-35%.

[0487] Emotion regulation efficacy: The anxiety index (GAD-7) was significantly negatively correlated with changes in alpha wave power (r=-0.71, p<0.01).

[0488] (2) Security guarantee

[0489] Set a dynamic safety threshold: When the alpha wave power suddenly changes by more than 3σ, the stimulation will be automatically paused and a manual review will be initiated.

[0490] Biocompatibility testing: The flexible electrode has passed ISO 10993 biosafety certification and supports continuous use for 8 hours without skin irritation.

[0491] Application Example: Intervention Procedure for Patients with Social Anxiety

[0492] 1. Data Acquisition: Capture facial AU4 (frowning) motion units through the mobile phone camera and simultaneously analyze the voice base frequency jitter rate.

[0493] 2. Status assessment: A voice tremor frequency of 8.5Hz was detected (>anxiety threshold of 7.5Hz) + the text showed cognitive distortion of "fear of being ridiculed".

[0494] 3. Strategy Generation: The dynamic engine selects virtual exposure therapy, and the VR scene generation module constructs a supermarket shopping simulation environment.

[0495] 4. Effect verification: After intervention, the LF / HF ratio of real-time monitoring of HRV (heart rate variability) increased to 1.8 (baseline value 0.9).

[0496] Technical effect

[0497] 1. The accuracy of emotion recognition has been improved to 89.7% (F1 score).

[0498] 2. The early warning time for high-risk cases is 4-6 months earlier than that of traditional scales.

[0499] 3. The intervention efficacy rate (PHQ-9 reduction ≥ 50%) for patients with depression reached 72.3%.

[0500] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A psychological healing robot based on simulation analysis of high human emotional affinity, characterized in that, include: A multimodal perception layer is used to collect physiological, motion, and environmental interaction data. This multimodal perception layer includes: a heterogeneous data acquisition module for integrating multi-type sensor data from the intelligent terminal's physiological, motion, and environmental interactions; a spatiotemporal feature extraction network for processing synchronous speech and video data streams based on a 3D-CNN spatiotemporal feature extraction network to obtain emotion calculation values ​​corresponding to spatiotemporal features; and an attention fusion mechanism module for dynamically adjusting the weights of each modality. A dynamic decision layer is used to generate intervention strategies based on the interaction data through a reinforcement learning policy engine and a hierarchical intervention selection tree; wherein, the dynamic decision layer includes: a reinforcement learning policy engine, used to establish a state space and an action space and design a reward function; and a hierarchical intervention selection tree, used to trigger corresponding intervention protocols according to the crisis level and generate intervention strategies based on the reward function; A generative interaction layer is used to generate an empathetic response that conforms to ethical norms based on the intervention strategy; wherein, the generative interaction layer includes: an ethical constraint system, used to impose ethical constraints based on an ethical rule base composed of a taboo content library, knowledge boundary constraints, privacy protection mechanisms, and an ethical knowledge graph; and an empathetic response generator, used to generate an empathetic response based on constraint decoding technology, and apply the ethical constraints to the empathetic response to generate an empathetic response that conforms to ethical norms; A brain science validation layer is used to monitor neural feedback in real time via EEG and adjust the intervention strategy. The brain science validation layer includes: a neural feedback adjustment module for monitoring changes in prefrontal alpha wave power in real time via EEG and using these changes as neural feedback; and a multimodal feedback design module for mapping neural feedback to virtual scene parameters based on the neural feedback adjustment mechanism and dynamically adjusting the intervention strategy based on these virtual scene parameters. The heterogeneous data acquisition module employs four types of physiological monitoring sensors, three types of motion sensing sensors, and two types of environmental interaction sensors. The four types of physiological monitoring sensors include heart rate variability (PPG), skin conductance response (GSR), blood oxygen saturation (SpO2), and body temperature sensors. The three types of motion sensing sensors include a triaxial accelerometer for gait analysis, a gyroscope for posture recognition, and a barometer for spatial positioning. The two types of environmental interaction sensors include a microphone for voice emotion analysis and an ambient light sensor for circadian rhythm monitoring. The heterogeneous data acquisition module adopts a layered acquisition architecture, sequentially transmitting data from the sensor hardware layer, Bluetooth 5.3 or BLE protocol stack, edge computing nodes, and data preprocessing unit to the cloud with encryption. Noise reduction of raw data is achieved through edge nodes, and a unified data description framework is established to ensure compatibility with Android Health Connect and Apple. The HealthKit protocol enables spatiotemporal alignment of data from multiple devices, employing a dynamic interpolation algorithm to balance differences in sampling rates among different sensors. The heterogeneous data acquisition module is also used for multimodal transmission and synchronization, as well as preliminary fusion of multimodal data. The multimodal transmission and synchronization are implemented based on a low-power transmission protocol and a spatiotemporal synchronization mechanism. The preliminary fusion of multimodal data is implemented based on a feature-level fusion strategy. Spatiotemporal features are extracted by inputting the original signal into a 3D-CNN network suitable for simultaneous voice and video analysis. The importance score of each modality is calculated as an energy value through a learnable feedforward neural network, and then the importance score is converted into a weight through a Softmax function.

2. The psychological healing robot based on high human emotional fit simulation analysis according to claim 1, characterized in that, The 3D-CNN-based spatiotemporal feature extraction network for processing synchronized speech and video data streams includes: A 3D-CNN architecture for spatiotemporal feature extraction is established, comprising: a data input and preprocessing layer, wherein the data input to the data input and preprocessing layer is a three-dimensional input tensor constructed from synchronously acquired time-series speech waveforms and spatial RGB data video frames, with dimensions [T×H×W×C], i.e., time × height × width × channel; a 3D convolutional kernel, which uses a three-dimensional convolutional kernel to slide synchronously on the time axis of the speech frame sequence and the spatial axis of the video frame, capturing the spatiotemporal correlation between speech fundamental frequency jitter and facial micro-expressions; a spatiotemporal feature output layer, which extracts features step by step through a 3D convolution + pooling structure based on hierarchical stacking and spatiotemporal correlation, and outputs a spatiotemporal feature map at the last layer of the spatiotemporal feature output layer; and a target optimization layer, which optimizes features based on a loss function, wherein the loss function is determined by fusing action localization error and emotion classification cross-entropy loss. Initial sentiment values ​​are obtained by processing synchronous speech and video data streams using a 3D-CNN network based on spatiotemporal feature extraction. Based on the initial sentiment value, sentiment calculation is performed using a contradiction index analysis algorithm to obtain the sentiment calculation value, wherein the contradiction index analysis algorithm includes: (1) Standardize the optimized features, including: Z-Score standardization of text sentiment polarity and speech sentiment intensity to eliminate dimensional differences; (2) Construct a contradiction index model and determine and arbitrate abnormal and contradictory data: The formula for calculating the contradiction index C in the contradiction index model is shown in equation (1) below: (1); (2); Where T text For the standardized text sentiment polarity, BERT_output represents the text sentiment polarity output by BERT, μ text For standard text sentiment polarity; T voice Pitch_Jitter represents the normalized speech emotion intensity, where μ is the fundamental frequency jitter rate. voice Standard speech emotion intensity; σ text σ is the environment vector modulation factor; voice This is the sensor state attenuation factor; when the inconsistency index C > 2.5, the difference between the two modes exceeds 2.5 times the standard deviation, triggering a manual review process, including: (A) Prioritize the use of physiological sensor data; (B) If physiological data is missing, a manual review process will be initiated; (3) Determining physiological arousal, including: calculating and classifying the physiological stress index through heart rate variability and skin conductance, wherein the formula for calculating the physiological stress index A is as follows: ; The classification is divided into three levels: low A < 0.5, medium 0.5 ≤ A ≤ 1.5, and high A > 1.5; among which, HRV LF / HF GSR indicates the lower or higher value of heart rate variability. Δ This indicates skin electrical conductance.

3. The psychological healing robot based on high human emotional fit simulation analysis according to claim 2, characterized in that, The modal weights are dynamically adjusted based on an attention mechanism or a cross-modal attention mechanism based on Transformer; wherein, the formula for dynamically adjusting the modal weights based on the attention mechanism is shown in equation (3) below: (3); in, This represents the modal weight values, totaling [number]. One modality, Attention scores are assigned to each modality; Each mode corresponds to Each sensor or feature extractor extracts a feature vector for each modality. , =1,2,..., The goal is to obtain the weights of each modality through an attention mechanism, and then calculate a weighted feature vector as the final representation. The formula for dynamically adjusting the modal weights of the Transformer-based cross-modal attention mechanism is shown in equation (4) below: (4); Where Q is the query matrix for the current sentiment state, and K... i The key matrix Key is the eigenvector matrix for each modality; d k It is the dimension of the key vector, that is, the length of the key vector in each head; Dynamically adjusting the modal weights includes: (1) Feature projection: Mapping multimodal features to the same latent space; (2) Attention aggregation: Calculate the weighted fusion features for downstream sentiment classification; (3) Introduce a spatiotemporal attention gating mechanism to suppress the weight of features in time periods with low signal-to-noise ratio.

Citation Information

Patent Citations

  • Music healing system based on brain wave emotion recognition and processing method thereof

    CN112999490A

  • Session type artificial intelligence driven personality simulation system based on context awareness and operation method

    CN117874185A