Physical and psychological state assessment method and system for multi-modal biological signal fusion processing

By using a multimodal fusion method that combines features such as micro-expressions, eye movements, breathing patterns, and postural tension, the problem of insufficient stability and universality of single-modal assessment methods is solved. This achieves robust, individualized, and reliable assessment of physiological and psychological states, and enables low-latency, privacy-preserving assessment loops at the edge.

CN121817886APending Publication Date: 2026-04-10HETE (INNER MONGOLIA) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing physiological and psychological state assessment methods based on physiological micro-motion monomodal approaches lack stability and universality under conditions such as camera shake, lighting changes, occlusion, posture changes, and individual differences. They also lack individualization and long-term trend management and provide insufficient privacy protection.

Method used

Employing a multimodal fusion approach that combines features such as micro-expressions, eye-tracking behavior, breathing patterns, and postural tension, this approach achieves a robust and reliable evaluation loop through temporal alignment, quality gating, cross-modal attention fusion, and individualized calibration, combined with federated learning and differential privacy protection, and is adapted for low-latency inference at the edge.

Benefits of technology

It improves robustness in complex scenarios, supports individualized and uncertain outputs, forms an assessment-intervention-trend closed loop, is suitable for high-reliability scenarios, and protects privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121817886A_ABST
    Figure CN121817886A_ABST
Patent Text Reader

Abstract

The invention discloses a body and mind state assessment method and system based on multi-mode biological signal fusion processing. According to the method, human face and body features are collected through videos, and features such as micro-expressions, eye movements (pupils and gazes), breathing modes, postures and physiological micro-movements are extracted; generating a multi-dimensional state vector by adopting timestamp registration and a cross-modal attention and gating fusion network in combination with an individual baseline and adaptive calibration, and outputting indexes including emotion, pressure, alertness and the like and a confidence interval; and triggering a prompt or intervention suggestion when the preset threshold value is exceeded. The system supports edge end quantitative operation, and data security is guaranteed through federated learning and privacy protection technologies. The method has remarkable effects in the aspects of anti-noise robustness, individualized adaptation and credibility, and is suitable for robots and man-machine interaction scenes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence and biological signal processing, in particular to a real-time or quasi-real-time physiological and psychological state evaluation method and system based on multi-modal non-contact biological signal fusion and a computer readable storage medium. BACKGROUND

[0002] The existing method with "physiological micro-movement / sub-pixel level non-random motion features" as the core single mode is easily affected by camera shaking, light changes, occlusion, posture changes and individual differences, and has insufficient stability and universality; it mainly stays in static or one-time evaluation, lacks individualization and long-term trend management, and has insufficient support for output uncertainty and privacy protection. The market urgently needs to output multi-dimensional psychological and physiological indicators in a stable and personalized manner under the condition of general camera, non-contact, low cost and edge deployment. SUMMARY

[0003] OBJECTIVE

[0004] To solve the above problems, a method and system are proposed, which fuse micro-expression, eye movement behavior, breathing pattern, posture tension and physiological micro-movement and other multi-modal features, combine time sequence alignment, quality gating, cross-modal attention fusion, individualized calibration and uncertainty estimation, realize a robust, reliable and intervenable evaluation closed loop, and adapt to edge low-latency inference and privacy protection.

[0005] The method includes: 1) collecting video data; 2) extracting multi-modal non-contact features; 3) cross-modal time sequence alignment and quality gating; 4) cross-modal attention / gating fusion, outputting a multi-dimensional state vector; 5) individualized calibration based on individual baseline and online adaptation; 6) outputting multi-dimensional indicators and uncertainty measures, and giving intervention suggestions when the threshold is triggered; 7) running in real time or quasi-real time on the edge, using federated learning and differential privacy to protect individual data.

[0006] The system includes: a collection module, a preprocessing module, a feature extraction module, an alignment and quality module, a fusion and uncertainty module, an individualization module, an evaluation and intervention module, a privacy and deployment module.

[0007] ADVANTAGEOUS EFFECTS

[0008] Compared with the single-modal "physiological micro-movement" method, the present application significantly improves the robustness in complex scenes such as shaking / occlusion / light changes;

[0009] Supports individualization and uncertainty output, risk controllable, suitable for high reliability scenarios;

[0010] Forms an evaluation-intervention-trend closed loop, facilitating long-term management and immediate response;

[0011] ●Supports low-latency operation at the edge and protects privacy.

[0012] Differences and technical effects compared to existing technologies

[0013] ●Modal level: This invention is a "multimodal fusion" method, which includes "physiological micro-motion / subpixel-level motion spectrum features".

[0014] It is only one of the optional modes, unlike solutions that use only this mode as the core.

[0015] ●Mechanism level: Introduce a combined mechanism of cross-modal temporal alignment (coherence spectrum / DTW), quality gating, cross-modal attention fusion, individualized calibration and uncertainty estimation;

[0016] ● Output and Closed Loop: It not only provides a one-time assessment, but also outputs uncertainty and trends, triggering real-time intervention and forming an assessment-intervention-trend closed loop;

[0017] ● Deployment and Privacy: Supports real-time / near-real-time inference at the edge, combining federated learning and differential privacy to minimize raw data transmission;

[0018] ●Scalability: Applicable to extended scenarios such as robot interactive control, it transforms evaluation results into safe and controllable execution instructions through policy mapping and degradation rules. Attached Figure Description

[0019] Figure 1 System overall architecture;

[0020] Figure 2 Method and process;

[0021] Figure 3-1 Unified network architecture;

[0022] Figure 3-2 Gating-quality scoring relationship;

[0023] Figure 3-3 Uncertainty output;

[0024] Figure 4 Personalized and adaptive processes;

[0025] Figure 5 Real-time intervention and trend interface;

[0026] Figure 6 A schematic diagram of the robot strategy process. Detailed Implementation

[0027] Principle protection and non-restrictive statement

[0028] The following embodiments are used to illustrate the principles and feasible implementation paths of the present invention, aiming to meet the requirements of full disclosure and implementability. Unless otherwise expressly limited, the structures, processes, module divisions, network architectures, parameter ranges, and training strategies involved in this document are all exemplary preferred solutions and do not constitute a limitation on the scope of protection; equivalent substitutions and modifications made by those skilled in the art without departing from the core ideas of the present invention (multimodal fusion, temporal alignment and quality gating, cross-modal attention fusion, individualized calibration and uncertainty estimation) should all fall within the protection scope of the present invention.

[0029] 1. Data Acquisition and Preprocessing

[0030] ●Camera management: Resolution 720p~1080p, frame rate 25~60fps; timestamp unification and image stabilization;

[0031] ●Face / feature and pupil localization, skin / lighting normalization, multi-scale cropping and noise reduction.

[0032] 2. Feature Extraction

[0033] ● Micro-expressions: Fine-grained temporal coding based on action units (AUs);

[0034] ●Eye movement behavior: blink rate, pupil diameter change rate, gaze stability, saccade amplitude;

[0035] ● Breathing pattern: Extract the optical flow field v(x,y,t) within the ROI of the chest and abdomen, bandpass filter its amplitude sequence (e.g., 0.1-0.7Hz), identify the main periodic peak f_resp using autocorrelation / FFT / PSD or wavelet transform, and calculate the respiratory rate BR = 60·f_resp; output the variability (BRV) and respiratory irregularity (main peak bandwidth, peak ratio, etc.); when rPPG signal is present, fuse its low-frequency component consistency for robust estimation;

[0036] ● Postural tension: Based on the trajectory p_k(t) of key points of the human skeleton, the micro-motion energy E=∑_k ∑_t||Δp_k(t)||^2 or the energy of velocity / acceleration / jerk is calculated, and the shoulder-neck distance fluctuation, center of gravity drift velocity, and energy proportion of specific frequency bands (such as 3-10Hz) are extracted as tension proxy; the above indicators are individually normalized and smoothed by time window to obtain the posture tension score;

[0037] ● Physiological micro-motions (also known as "micro-vibration characteristics"): directional spectrum characteristics, frequency band energy, and coherence indices.

[0038] 3. Timing alignment and quality gating

[0039] ● Cross-modal alignment is performed using coherent spectrum analysis and / or dynamic time warping, based on frame timestamps.

[0040] ● Calculate modal quality scores (based on signal-to-noise ratio, occlusion, tracking stability, and illumination stability), and downweight or mask low-quality modes.

[0041] 4. Fusion and Uncertainty Estimation

[0042] ● Fusion Model: Employs a cross-modal attention network and gating mechanism. Specifically, for each modal input sequence Xm (m∈{micro-expression, eye movement, respiration, posture, physiological micro-movement}), the latent representation Hm is first obtained through its respective temporal encoder (such as a TCN / Transformer encoder); then, using the selected master modality (such as eye movement or an intermediate representation of comprehensive fusion) as the query Q and other modalities as key values ​​K / V, a fusion representation Fm is obtained through multi-head cross-modal attention computation; the outputs of each modality are gating and weighted according to the quality score Qm to form a weighted fusion vector Z.

[0043] Note: In this manual, "micro-vibration" is a commonly used industry term and refers only to "physiological micro-motion / subpixel-level non-involuntary motion spectrum features based on video". Its extraction methods include, but are not limited to, optical flow, rPPG correlation spectrum, coherence and directionality spectrum, etc.

[0044] ●Structural highlights: Includes (i) a modality-specific encoder layer, (ii) a cross-modality multi-head attention layer (supporting Q-KV interactions from different modalities), (iii) a gating weight allocation layer (weights are determined by quality scores and learnable parameters), and (iv) a fusion and convergence layer (such as splicing + feedforward or weighted summation).

[0045] ● Uncertainty estimation: An uncertainty branch is introduced after the fused vector Z, and the prediction variance is obtained by using MC Dropout and / or deep ensemble; the output includes point estimates and confidence intervals, and supports triggering "pause conclusion / resample / delayed observation" in high uncertainty scenarios.

[0046] ● Training: Multi-task learning, with losses including supervised loss (classification / regression), contrastive loss (cross-session / cross-domain consistency) and calibration loss (temperature calibration / ECE optimization) to improve generalization and calibration.

[0047] 5. Individualized calibration

[0048] ●Establish individual baselines (resting / normal range), update EMA within sessions, and align temperature / calibration curves or adapt the contrast learning domain between sessions.

[0049] 6. Assessment and intervention closed loop

[0050] ● Output indicators such as emotions, stress, neuroticism, self-regulation, extraversion, alertness, decisiveness, and skepticism;

[0051] ● Trigger prompts or intervention suggestions (such as breathing training / rest prompts) when the risk or uncertainty exceeds a threshold.

[0052] 7. Privacy Protection and Deployment

[0053] ● Edge inference: Model quantization / distillation reduces latency and power consumption;

[0054] ● Federated learning and differential privacy: only gradients or statistics are aggregated in the cloud, and the original video is not uploaded to the cloud.

[0055] 8. Examples

[0056] Example A (Driver fatigue / alertness monitoring): primarily eye movement, breathing, and posture; Example B (Security / Interview assessment): primarily micro-expression, eye movement, and posture; Example C (Medical triage / psychological screening): primarily breathing and rPPG.

[0057] Example D (Robot Interaction / HRI):

[0058] ●Input: Multidimensional indicators and uncertainties output by the method of this invention (such as mood, stress, alertness, individualized calibrated state vectors and their confidence intervals).

[0059] ● Strategy Mapping: The strategy generation module generates robot interaction strategy instructions based on the aforementioned indicators and uncertainties. The instructions include at least one of the following: movement speed / amplitude / path smoothing, interaction distance, language output style / speech rate / volume, and waiting time. When the uncertainty meets the threshold condition, a degraded interaction is triggered (such as pausing, requesting resampling, or delaying observation).

[0060] ●Execution: The execution module controls the robot body (mobile chassis / robotic arm / voice interaction / display, etc.) according to the strategy instructions.

[0061] ● Interface (Example, non-restricted):

[0062] ○ plan(state_Vector, unCertainty)→{motion_policy, speech_policy, wait_policy, safety_state}

[0063] ○ motion_poliCy: {speed_mode, amplitude_mode, path_smoothing}

[0064] ○ speech_poliCy: {style, tone, volume, phrases[]}

[0065] ○ wait_policy:{duration_mode}

[0066] ○ safety_state: {degraded: boolean, reason} Explanation: The above field is used to clarify the data flow and module boundaries. It is not limited to a specific enumeration / implementation, and equivalent replacements can be used.

[0067] 9. Performance Evaluation

[0068] ●Accuracy (correlation with scale / task), robustness (occlusion / jitter / lighting), real-time performance (<100ms / frame), calibration (ECE / NLL), privacy compliance (ε budget).

[0069] 10. Model Evolution and Continuous Learning

[0070] ● Online Adaptive: During the inference phase, the drift of individual feature distribution (such as mean / variance drift and confidence drift) is detected, and the output calibration parameters are updated using EMA / temperature calibration / calibration curves;

[0071] ●Feedback loop: Combine user interaction and intervention effect (such as the index falling back after intervention) to construct a weak supervision signal and update the weight or threshold strategy of multi-task loss.

[0072] ● Federated learning: The client performs small-step training / distillation on the new data locally, and uploads the gradients / statistics processed with differential privacy noise to the server for aggregation;

[0073] ●Personalized distillation: Using the global model as the teacher, small data distillation and efficient parameter fine-tuning are performed on the edge (Adapter / LoRA), and periodically merged into individualized branches;

[0074] ● Drift rollback: When uncertainty and quality score remain abnormal, roll back to the previous stable snapshot and trigger resampling.

[0075] Background Technology (Supplementary Explanation)

[0076] Besides solutions based on a single modality of "physiological micro-motion," systems based solely on facial expression recognition or heart rate variability (HRV / rPPG) also suffer from sensitivity to lighting, occlusion, camouflage, and scene domain shifts. Furthermore, they lack cross-modal verification and uncertainty control, making it difficult to maintain stable and reliable output across multiple scenarios. This invention systematically addresses the shortcomings of single-modal methods through multimodal alignment and quality gating, attention fusion, and individualized calibration.

[0077] Terminology Explanation

[0078] "Uncertainty" refers to the model's prediction variance on the current input and in-domain / out-of-domain samples, including chance and cognitive uncertainty; "Individualized baseline" refers to the reference statistics established for the long-term stable characteristics of the evaluated object.

[0079] Feasibility of implementation

[0080] This can be achieved in ordinary consumer-grade cameras and edge computing devices, and is feasible for industrialization.

Claims

1. A method for assessing physical and mental state through multimodal biological signal fusion processing, characterized in that, include: a) Collect video data containing the target person's face and upper body; b) Extract at least two of the following from the video data: micro-expression features, eye movement features, breathing pattern features, body posture features, and physiological micro-movement features; c) Perform temporal alignment and quality assessment on the features, and perform cross-modal attention and gating fusion based on the alignment results and quality assessment to obtain a multi-dimensional feature vector representing the physical and mental state of the target person; d) Perform individualized calibration of the multidimensional feature vector based on individual baseline and online adaptation; e) Output an assessment index that includes at least one of emotion, stress level, neuroticism, self-regulation, extraversion, alertness, decisiveness and skepticism index, and provide uncertainty measurement and threshold trigger prompts or intervention suggestions.

2. The method according to claim 1, characterized in that, The temporal alignment includes coherent spectral analysis and / or dynamic time warping based on timestamp registration, the quality assessment includes modal quality scoring based on signal-to-noise ratio, occlusion ratio, tracking stability and illumination stability, and the fusion is performed by downweighting or masking based on the scores.

3. The method according to claim 1 or 2, characterized in that, The uncertainty measure is obtained through MCDropout, deep ensemble and / or Bayesian estimation, and the output confidence interval is used for risk control.

4. The method according to any one of claims 1 to 3, characterized in that, The individualized calibration includes intra-session EMA updates and inter-session temperature calibration or calibration curve alignment.

5. The method according to any one of claims 1 to 4, characterized in that, The method can be executed in real time or near real time on edge devices, and achieves computational overhead control and privacy protection through model compression, federated learning and / or differential privacy.

6. A system for implementing the method according to any one of claims 1 to 5, characterized in that, include: The module includes a data acquisition module, a feature extraction module, an alignment and quality module, a fusion and uncertainty module, an individualization module, and an evaluation and intervention module.

7. A robot interactive control method based on the method described in any one of claims 1 to 5, characterized in that, include: a) Perform the assessment method described above to obtain the target person's physical and mental state indicators and uncertainty measure; b) Based on the aforementioned indicators and uncertainties, generate robot interaction strategy instructions, the instructions including at least one of motion control, language output, and waiting strategies; c) When the uncertainty exceeds a threshold, trigger a risk avoidance strategy, including at least one of degraded interaction, requesting resampling, or delayed observation.

8. A robot system, characterized in that, include: The perception and assessment module is used to perform the method described in any one of claims 1 to 5 to generate indicators of physical and mental state and uncertainty. The strategy generation module is used to generate robot interaction strategy instructions based on the aforementioned indicators and uncertainties. The execution module is used to control the movement and interactive output of the robot body according to the interaction strategy instructions.

9. A computer-readable storage medium, characterized in that, It stores a program that, when executed by a processor, implements the method described in any one of claims 1 to 5 or 7.

10. An electronic device, characterized in that, It includes at least one processor and a memory, wherein the memory stores a program that can run on the processor, and the program, when run, causes the electronic device to perform the method according to any one of claims 1 to 5 or 7.