Emotional dialectical interaction system and method for traditional Chinese medicine physical robot

By using a TCM embodied robot system to collect facial images and voice signals non-contactly, and combining this with a deep reinforcement learning optimizer to generate treatment parameters, the system solves the problems of subjectivity and contact-based detection in traditional TCM emotional diagnosis. It achieves efficient TCM seven-emotion state recognition and treatment parameter generation, which is in line with the TCM principle of non-invasiveness.

CN120878094APending Publication Date: 2025-10-31YIZHIYUAN TECHNOLOGY DEVELOPMENT (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511048361.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Traditional Chinese medicine's diagnosis of emotions relies heavily on the physician's experience and is highly subjective. Existing AI emotion recognition systems cannot handle the syndrome correlation of "emotions causing disease" in traditional Chinese medicine. This results in robot therapy lacking the theoretical support of traditional Chinese medicine, and has problems such as limitations in static expression recognition, lack of guidance from the theory of the seven emotions causing disease, poor dynamic adaptability, and the burden of contact detection.

Method used

The device uses a binocular RGB-IR camera and an omnidirectional microphone array of a TCM embodied robot to collect user facial images and voice signals in a non-contact manner. It then performs cross-modal feature fusion through a TCM emotional syndrome differentiation model and generates treatment parameters by combining a deep reinforcement learning optimizer. This achieves non-contact TCM observation and auscultation quantitative fusion and establishes an automated mapping chain of 'AI features → seven emotions → syndrome type → parameters'.

Benefits of technology

It enables non-contact TCM diagnosis, improves the efficiency of facial color-micro-expression-voice feature fusion, enhances the accuracy of TCM seven emotions state recognition and the dynamic adaptability of treatment parameters, reduces user discomfort, and conforms to the non-invasive principle of TCM.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120878094A_ABST
    Figure CN120878094A_ABST
Patent Text Reader

Abstract

The emotion dialectical interaction system and method for the traditional Chinese medicine body-equipped robot are characterized by comprising the steps that S1, a user face image and a voice signal are collected; s2, extracting a forehead chromaticity feature and a facial muscle group motion dynamics feature; s3, extracting a fundamental frequency jitter rate and a formant envelope characteristic dynamic range; s4, inputting the features obtained in S2 and S3 into a traditional Chinese medicine emotion dialectical model, and outputting a traditional Chinese medicine seven-estrus state intensity vector E belonging to R; s5, mapping the E into a traditional Chinese medicine syndrome type; and S6, obtaining a diagnosis result based on the syndrome type. The invention provides a combined characterization and mapping method of facial dynamic optical flow features and voice nonlinear acoustic energy operator chaotic features oriented to traditional Chinese medicine syndrome differentiation, seven-condition state vectors are constructed, and non-contact traditional Chinese medicine observation and smelling quantification is realized; designing an artificial intelligence deep reinforcement learning parameter optimization engine based on traditional Chinese medicine syndrome type decision constraints, and generating physical treatment parameters conforming to a syndrome differentiation treatment principle; the problems that existing emotional syndrome differentiation lacks objective quantification standards, treatment parameters are disjointed with traditional Chinese medicine syndrome types, and contact type detection is relied on are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot emotional dialectics technology, specifically a TCM embodied robot emotional dialectics interaction system and method. Background Technology

[0002] Traditional Chinese medicine (TCM) diagnosis of emotional disorders relies heavily on physician experience, leading to a high degree of subjectivity. Existing AI emotion recognition systems cannot handle the syndrome correlations of "emotional illness" in TCM, resulting in a lack of theoretical support for current robotic therapy. The existing technology has the following shortcomings:

[0003] 1. Limitations of emotion recognition: It only uses static facial expression recognition and ignores the correlation between facial color and emotions in Traditional Chinese Medicine;

[0004] 2. Insufficient integration of traditional Chinese and Western medicine: Lacking a parameter mapping mechanism guided by the theory of seven emotions causing disease, it suffers from three limitations:

[0005] ① Ignoring the mapping relationship between the five tones and the five internal organs in traditional Chinese medicine, the existing technology has limitations in the specific identification of the emotion of fright;

[0006] ② A dynamic correlation between emotional state, syndrome type, and treatment parameters has not been established;

[0007] ③ It requires a contact-type heart rate sensor, which violates the non-invasive principle of traditional Chinese medicine's diagnostic methods of observation, auscultation, inquiry, and palpation;

[0008] 3. Poor dynamic adaptability: Traditional control methods cannot handle the time-varying coupling relationship between emotion and therapy;

[0009] 4. Contact-based detection burden: The need to wear physiological sensors can cause discomfort to users. Summary of the Invention

[0010] The technical problem to be solved by the present invention is to provide a TCM embodied robot emotional dialectical interaction system and method, which can effectively solve the problems mentioned in the background art.

[0011] To solve the above problems, the technical solution adopted by the present invention is: a TCM embodied robot emotional dialectical interaction system and method, comprising the following steps:

[0012] S1. Through the binocular RGB-IR camera and omnidirectional microphone array of the TCM embodied robot, non-contact synchronous acquisition of user facial images and voice signals is achieved.

[0013] S2. Extract forehead chromaticity features and facial muscle movement dynamics features based on the CIE LAB color space from the facial image. The dynamics features include the rate of change of eyebrow spacing and the tension of the orbicularis oris muscle.

[0014] S3. Extract the fundamental frequency jitter rate and formant envelope dynamic range from the speech signal;

[0015] S4. Input the features obtained in steps S2 and S3 into a traditional Chinese medicine (TCM) emotional syndrome differentiation model. This model adopts a cross-modal attention fusion mechanism with weighted regulation of zang-fu organs and qi-blood, and outputs a vector E ∈ R of the intensity of the seven emotional states in TCM 7 ;

[0016] S5. Through a TCM syndrome type determination module that integrates a fuzzy inference rule base, map E to a TCM syndrome type. The TCM syndrome type determination module uses a field-programmable gate array (FPGA) to achieve high-speed parallel inference of 47 rules;

[0017] S6. Based on the syndrome type diagnosis result, calculate initial treatment parameters through a physical therapy parameter generator constrained by the syndrome type, and adjust the parameters in real time via a deep reinforcement learning optimizer. Among them: the state space s_t includes ΔE, Δ facial muscle movement dynamics feature movement features, and Δ speech acoustic features; the action space a_t is constrained by the current syndrome type (|Δ parameter| ≤ 20%); the reward function r_t = 10 · expression relaxation degree - ∥Δparam∥2.

[0018] Preferably, the TCM emotional syndrome differentiation model described in S4 adopts a heterogeneous modal parallel processing network architecture. The network architecture includes a visual branch, a speech branch, and a feature fusion layer. The visual branch is a 3-layer convolutional neural network for processing the heat map of facial muscle movement dynamics features; the speech branch is a Bi-LSTM for processing the temporal features of speech formant envelope features; the feature fusion layer is to weighted splice bimodal features through an attention mechanism.

[0019] Preferably, the syndrome type mapping in S5 adopts a fuzzy inference mechanism. When the emotional state vector satisfies:

[0020] μanger ≥ τ1 ∧ μfear ≤ τ2 (τ1 = 0.75, τ2 = 0.3), trigger the treatment protocol for the hyperactivity of yang of the liver meridian syndrome type, where the threshold is determined by the clinical ROC curve.

[0021] An embodied robot emotional syndrome differentiation interaction system for implementing the above method, characterized in that it includes a perception unit, a calculation unit, and an execution unit. The perception unit includes a binocular RGB-IR camera and an omnidirectional microphone array (8 microphones, beamforming); the calculation unit includes a TCM syndrome differentiation calculation engine and a dynamic treatment parameter dynamic optimizer, configured to execute a parameter adaptive strategy based on deep reinforcement learning. The state observation value of the parameter adaptive strategy of the reinforcement learning includes the real-time change amount of the facial muscle movement dynamics feature movement features of the user and the real-time change amount of the speech acoustic features, and its action space is constrained by the current TCM syndrome type diagnosis result; the execution unit includes a bionic partition temperature control palm, a high-precision force control massage module (resolution better than the clinical palpation threshold), and an adjustable low-frequency electrical stimulation module.

[0022] Preferably, the binocular RGB-IR camera is equipped with near-infrared supplementary light and a polarizing filter to eliminate the influence of ambient reflection on face color analysis.

[0023] A training method for a TCM emotional syndrome differentiation model, characterized by the following input:

[0024] Bimodal dataset D = {(V i A i E i )}, where E i Annotated by TCM experts;

[0025] Model architecture:

[0026] Visual branch: Convolutional neural networks process heatmaps or optical flow feature maps of facial muscle group motion dynamics;

[0027] Speech branch: Temporal neural networks process the temporal sequence of acoustic features;

[0028] Fusion layer: A cross-modal feature dynamic weighted fusion module constructed based on the topological constraints of the graph neural network (GNN) for the transmission of the five internal organs in the "Fuxing Jue";

[0029] Loss function: L = α·KL(E pred ||E true )+β·||ΔMeridian Response||2.

[0030] Compared with the prior art, the present invention provides a TCM embodied robot emotional dialectical interaction system and method, which has the following beneficial effects:

[0031] This invention pioneers a non-contact TCM "observation and auscultation" quantitative fusion framework: deeply integrating the TCM theories of "five-color diagnosis" and "five-tone diagnosis", it establishes a mapping model from machine vision-extracted facial skin color (e.g., CIE LABa* value > 15 to judge "red face"), specific muscle group dynamics (e.g., the contraction rate of the glabella muscle), and modern acoustic features (e.g., fundamental frequency jitter rate, the five-tone tempo rate) to the TCM seven emotions state, abandoning contact physiological sensors and strictly adhering to the TCM non-invasive principle;

[0032] A three-level TCM mapping chain of "AI features → seven emotions → syndrome type → parameters" is realized: a complete automated mapping chain of "modern multimodal AI perception features → TCM seven emotions state vector → TCM visceral syndrome type diagnosis → physical therapy parameter generation and optimization" is established; and its causal validity is verified through clinical data, which deeply bridges the gap between artificial intelligence technology and traditional Chinese medicine theory and practice. Attached Figure Description

[0033] Figure 1 This is a system architecture diagram of the present invention;

[0034] Figure 2This is a flowchart of the algorithm of the present invention. Detailed Implementation

[0035] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0036] Reference Figure 1-2 This invention provides a TCM embodied robot emotional dialectical interaction system and method. The core of this invention lies in constructing a closed-loop system of "visual-voice dual-modal perception → TCM seven emotions dialectical differentiation → dynamic parameter generation → artificial intelligence reinforcement learning optimization". Its originality is reflected in:

[0037] 1. Non-contact TCM diagnosis system

[0038] This invention pioneers a quantitative mapping between forehead chromaticity in the CIE LAB color space and the "Five Colors Diagnosis" method in Traditional Chinese Medicine, establishing a causal relationship between facial redness (a*>15) and excessive heart fire: based on forehead region chromaticity detection in the CIE LAB color space (a* value>15 determines "facial redness"),

[0039] Facial muscle biomechanics: Eyebrow movement speed >15 pixels / frame indicates "surprise" emotion.

[0040] Voice and Emotion Integration: When the corners of the mouth droop more than 10° and the speech rate is less than 2 words per second, increase the weight of "worry".

[0041] 2. Dynamic optimization algorithm for disease prevention and treatment parameters

[0042] def T_calc(syndrome,E): if syndrome=="Excessive Heart Fire": #Temperature is positively correlated with anger and negatively correlated with joy.

[0043] temp=37.2+0.6*E[1]-0.4*E[0];#The current intensity decreases exponentially with the degree of fear, current=12*exp(-0.8*E[5])

[0044] return[temp,force,speed,current].

[0045] 3. Introduce a five-organ-five-color-five-tone mapping matrix in the feature fusion layer.

[0046] 4. Deep reinforcement learning optimization mechanism

[0047] State space: $s_t = (\Delta e_anger,\Delta V_{corner of the mouth},\Delta A_{speech rate})$

[0048] Action space: $a_t = (\Delta temp\in[-1,1]℃,\Delta force\in[-0.5,0.5]N)$

[0049] Reward function: $r_t = 10\cdot\text{expression relaxation} - |\Delta param|_2$.

[0050] Unified Algorithm Flow

[0051] 1. Data Acquisition Phase:

[0052] Visual: Captures the motion trajectory of 68 key facial points at 30fps.

[0053] Auditory: Acquiring speech streams at a 16kHz sampling rate

[0054] 2. Feature extraction stage:

[0055] Visual: Calculate the rate of change in eyebrow spacing and the tension of the orbicularis oris muscle.

[0056] Auditory perception: Extracting fundamental frequency jitter rate, dynamic range of speech formant envelope features, and 3. Traditional Chinese medicine diagnosis stage:

[0057] The vector of seven emotions, E = f_θ(V,A) = softmax(W_v·V + W_a·A + b).

[0058] Innovation points: Introducing the organ synergy matrix W_season to regulate feature weights. 4. Syndrome type decision stage:

[0059] Syndrome type = argmax(μ_k(E)) k∈{Liver Yang Rising, Heart and Spleen Deficiency,...}

[0060] 5. Parameter generation stage:

[0061] T = g(certificate type, E) + DQN_agent(s_t)

[0062] 6. Execution Feedback Phase:

[0063] Observe changes in user facial expressions and update the Q-value table.

[0064] Explanation of Innovation Points

[0065] 1. Pioneering contactless TCM emotional diagnosis

[0066] It pioneers a non-contact TCM "observation and auscultation" quantitative fusion framework: deeply integrating the TCM theories of "five-color diagnosis" and "five-tone diagnosis", it establishes a mapping model from facial skin color extracted by machine vision (e.g., CIE LABa* value > 15 judges "red face"), specific muscle group dynamics (e.g., the contraction rate of the glabella muscle) and modern acoustic characteristics (e.g., fundamental frequency jitter rate, the five-tone rhythm rate) to the TCM seven emotions state, abandoning contact physiological sensors and strictly adhering to the TCM non-invasive principle.

[0067] 2. Dynamic Treatment Parameter Optimization System

[0068] A three-tiered TCM mapping chain of "AI features → Seven Emotions → Syndrome Type → Parameters" is established: a complete automated mapping chain is created, consisting of "modern multimodal AI perception features → TCM Seven Emotions state vectors → TCM Zang-Fu syndrome type diagnosis → physical therapy parameter generation and optimization." Its causal validity is verified through clinical data, deeply bridging the gap between artificial intelligence technology and traditional TCM theory and practice.

[0069] Innovative parameter calculation model:

[0070]

[0071] 3. A joint optimization framework for action space with multimodal state coding and TCM syndrome type constraints.

[0072] State encoder: Vectorizes the rate of change of facial expressions

[0073] Motion constraint: Limits single-step adjustment range to ≤20%.

[0074] Unique reward mechanism: +10

[0075] Compared to traditional methods, this solution achieves:

[0076] Breakthrough in dialectical dimension: The efficiency of fusing multi-source features of facial color, micro-expression, and vocal emotion is improved by 3.2 times;

[0077] Treatment precision: The causal chain verification accuracy of 'anger → liver yang hyperactivity → low temperature stimulation' reached 94.7% (p<0.01);

[0078] The depth of TCM theory embedding: The rule base contains 12 syndrome types and 47 'emotion-organ' mapping rules, covering 80% of the emotional pathogenesis in the "Huangdi Neijing".

[0079] As a specific embodiment of the present invention:

[0080] Example 1: Intervention for Anger and Depression

[0081] 1. Visual inspection:

[0082] The contraction rate of the glabella muscle is 0.82 (reference value 0.6), and the LAB color space a* = 18.5 (facial redness threshold 15).

[0083] 2. Speech Analysis:

[0084] The five-tone tempo-rate Δ = 40% (normal <20%), and the fundamental frequency standard deviation is 52Hz.

[0085] 3. Dialectical Output:

[0086] e_anger = 0.88, e_thought = 0.63, syndrome type = "liver stagnation transforming into fire"

[0087] 4. Parameter generation:

[0088] temp=36.5+0.7*0.63-0.6*0.88=36.2℃

[0089] current = 10 * (1 - 0.4 * nonlinear decay function of fear component) = 7.2 mA 5. Execution strategy:

[0090] The bionic hand maintains low-temperature contact, limiting the massage intensity to below 3N.

[0091] Example 2: Regulation of Worrying State

[0092] 1. Visual characteristics:

[0093] Mouth corner drooping angle 22° (reference 15°), eyelid opening angle 0.65 (reference 0.8) 2. Voice characteristics:

[0094] Average speech rate: 2.3 words / second; dynamic range of formant envelope characteristics: <8dB3. Diagnostic results:

[0095] e_worry = 0.79, e_thought = 0.85, syndrome type = "deficiency of both heart and spleen"

[0096] 4. Parameter calculation:

[0097] force=2.5+1.8*ln(1+0.79)=3.8N

[0098] speed = 90 + 30 * sigmoid(0.85) = 112 times / minute 5. Execution:

[0099] Rhythmic back tapping, maintaining hand temperature at 38.2℃ - experimental data

[0100] Results of a clinical trial involving 250 patients:

[0101]

[0102] Technical effect

[0103] Diagnostic accuracy: The F1-score for recognizing the seven emotions reached 0.94;

[0104] Treatment response: Parameter adjustment cycle <0.8s;

[0105] User experience: The incidence of discomfort decreased by 62%.

[0106] The innovation of this application lies in:

[0107] This paper proposes for the first time a non-contact TCM emotional diagnosis framework based on machine vision and acoustic analysis, which realizes the first two of the TCM diagnostic methods of "observation, auscultation, inquiry and palpation" through machine vision and voice analysis.

[0108] Establish a three-level mapping system of modern AI features, traditional Chinese medicine seven emotions, and physical parameters to bridge the gap between traditional Chinese medicine theory and robot execution;

[0109] Develop a bimodal deep reinforcement learning framework to achieve dynamic closed-loop optimization of emotion-therapy.

[0110] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for diagnosing emotions and mental states using a holistic robot in Traditional Chinese Medicine, characterized in that, It includes the following steps: S1. Non-contact and synchronous acquisition of the user's facial image and voice signal through the binocular RGB-IR camera and omnidirectional microphone array of the traditional Chinese medicine embodied robot; S2. Extract the forehead chromaticity feature and facial muscle group motion dynamics feature based on the CIE LAB color space from the facial image. The dynamics feature includes the eyebrow distance change rate and the orbicularis oris muscle tension; S3. Extract the fundamental frequency jitter rate and the dynamic range of the formant envelope feature from the voice signal; S4. Input the features obtained in steps S2 and S3 into the TCM emotional differentiation model. This model adopts a cross-modal attention fusion mechanism with weighted regulation of viscera and qi and blood, and outputs the intensity vector of the seven emotions in TCM, E∈R. 7 ; S5. Map E to the traditional Chinese medicine syndrome type through the traditional Chinese medicine syndrome type determination module integrating the fuzzy inference rule base. The traditional Chinese medicine syndrome type determination module realizes high-speed parallel inference of multiple rules with a programmable logic array; S6. Based on the syndrome type diagnosis result, calculate the initial treatment parameters through the physical treatment parameter generator constrained by the syndrome type, and adjust the parameters in real time through the deep reinforcement learning optimizer. Among them, the state space s_t includes ΔE, Δ facial muscle group motion dynamics feature motion feature, and Δ voice acoustic feature; The action space a_t is constrained by the current syndrome type; the reward function r_t = 10·expression relaxation degree - ∥Δparam∥2.

2. The method for diagnosing emotional disorders using a holistic TCM robot according to claim 1, characterized in that, The traditional Chinese medicine emotional syndrome differentiation model described in S4 adopts a heterogeneous modal parallel processing network architecture. The network architecture includes a visual branch, a voice branch, and a feature fusion layer. The visual branch is a 3-layer convolutional neural network for processing the thermal map of the facial muscle group motion dynamics feature; the voice branch is a Bi-LSTM for processing the temporal feature of the voice formant envelope feature; The feature fusion layer is to weight and splice the bimodal features through the attention mechanism.

3. The method for diagnosing emotional disorders using a holistic TCM robot according to claim 1, characterized in that, The syndrome type mapping in S5 adopts a fuzzy inference mechanism. When the emotional state vector satisfies: μanger ≥ τ1 ∧ μfear ≤ τ2 (τ1 = 0.75, τ2 = 0.3), trigger the treatment protocol for the hyperactivity of liver yang syndrome type, where the threshold is determined by the clinical ROC curve.

4. An embodied robot emotional dialectical interaction system for implementing the methods of claims 1-3, characterized in that, It includes a sensing unit, a computing unit, and an execution unit. The sensing unit includes a binocular RGB-IR camera and an omnidirectional microphone array; the computing unit includes a traditional Chinese medicine syndrome differentiation computing engine and a dynamic optimization seeker for dynamic treatment parameters, configured to execute a parameter adaptive strategy based on deep reinforcement learning. The state observation value of the parameter adaptive strategy of the reinforcement learning includes the real-time change amount of the facial muscle group motion dynamics feature motion feature of the user and the real-time change amount of the voice acoustic feature, and its action space is constrained by the current traditional Chinese medicine syndrome type diagnosis result; the execution unit includes a bionic partition temperature control palm, a high-precision force control massage module, and an adjustable low-frequency electrical stimulation module.

5. The embodied robot emotional dialectical interaction system according to claim 4, characterized in that, The binocular RGB-IR camera is equipped with near-infrared supplementary light and configured with a polarization filter to eliminate the influence of environmental reflection on the facial color analysis.

6. A training method for a TCM emotional syndrome differentiation model, characterized in that, Input: Bimodal dataset D={(V i A i E i )}, where E i Annotated by TCM experts; Model architecture: Visual branch: The convolutional neural network processes the thermal map or optical flow feature map of the facial muscle group motion dynamics feature; Voice branch: The temporal neural network processes the temporal sequence of the acoustic feature; Fusion layer: A cross-modal feature dynamic weighted fusion module constructed based on the topological constraint of the graph neural network (GNN) of the five-organ transmission and transformation in the "Auxiliary Treatise on Febrile Diseases"; Loss function: L = α·KL(E pred ||E true )+β·||ΔMeridian Response||2.