Psychological state recognition method and system based on physiological signals

By using a quality parameter-driven dynamic weighted fusion and meta-learning framework, combined with an incremental learning module, a personalized mental state recognition model is constructed. This solves the problem of inaccurate recognition under individual differences and environmental interference, and achieves efficient and personalized closed-loop optimization.

CN121154162APending Publication Date: 2025-12-19CHENGDU WEIFU TIKE TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511550361.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing mental state recognition technologies suffer from insufficient stability, weak dynamic adaptability, and lack of closed-loop regulation capabilities when faced with significant individual differences and strong environmental interference.

Method used

A quality parameter-driven dynamic weighted fusion mechanism is adopted, combined with a meta-learning framework and an incremental learning module, to construct a personalized recognition model. Through dynamic weighted fusion and online updating of multimodal physiological signals, a closed-loop optimization system is formed.

Benefits of technology

The robustness and individualization accuracy of recognition are improved in noisy environments, enabling efficient, personalized recognition and continuous adaptation to users' psychological states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121154162A_ABST
    Figure CN121154162A_ABST
Patent Text Reader

Abstract

The invention discloses a psychological state recognition method and system based on physiological signals, and relates to the technical field of intelligent preoperative evaluation and planning of thoracic surgery, and the method comprises the steps: S1, collecting multi-modal physiological signals, and forming an original signal set; s2, performing parallel preprocessing on the original signal set to generate a denoised time sequence signal; s3, dynamically generating a modal weight coefficient based on the quality parameters, and performing weighted fusion on the multi-modal time sequence signals by using the weight coefficient; and S4, inputting the fused feature vector into a personalized recognition model, and outputting psychological state probability distribution through an adaptive feature mapping layer in the model. And S5, according to the user feedback signal or the distribution offset of the continuous monitoring data, triggering online parameter updating of the personalized recognition model, and generating a recognition model after incremental optimization. According to the psychological state recognition method and system based on the physiological signals, the problem that psychological state recognition is inaccurate due to large individual difference, strong environmental interference and weak dynamic adaptability can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multimodal physiological signal processing and psychological state recognition technology for wearable devices, specifically to a method and system for psychological state recognition based on physiological signals. Background Technology

[0002] Current psychological state recognition technologies primarily rely on single-modal physiological signal analysis, such as using EEG features to identify emotions or heart rate variability to assess stress. However, single signals are susceptible to motion artifacts and environmental noise, leading to insufficient recognition stability. While multimodal fusion methods can partially compensate for this deficiency, traditional fixed-weight fusion strategies neglect individual physiological response differences and real-time signal quality fluctuations. For example, the signal-to-noise ratio of EEG signals drops sharply during limb activity, while ECG signals are more discriminative during emotional excitement. Existing technologies lack dynamic weight adjustment mechanisms. Regarding model generalization, general models require extensive calibration data on new users, but sufficient labeled samples are difficult to obtain in real-world scenarios. Static models cannot adapt to the long-term evolution of users' psychological states, such as increased stress tolerance or the cumulative effect of fatigue.

[0003] Current solutions mitigate individual differences through fine-tuning via transfer learning, but this process damages the representational capabilities of pre-trained models and lacks a continuous evolutionary mechanism. Online learning technologies attempt to dynamically update models, but simple incremental training easily leads to catastrophic forgetting and significant loss of historical knowledge. In terms of hardware deployment, centralized cloud processing poses risks of transmission latency and privacy breaches, while edge devices are limited by computing resources and struggle to run complex models. Furthermore, the results of psychological state recognition are disconnected from actual intervention, lacking the ability to implement closed-loop adjustments based on the recognition results.

[0004] While existing patented technologies involve multi-signal fusion, they fail to address the dynamic weighting issue in quality perception. Personalized modeling largely relies on traditional machine learning, neglecting to incorporate meta-learning frameworks for rapid adaptation to small sample sizes. Furthermore, model update mechanisms often employ periodic retraining, resulting in poor real-time response and high computational overhead. Therefore, there is an urgent need for a recognition method that can dynamically perceive signal quality, adaptively fuse multimodal data, efficiently construct and continuously evolve personalized models, and support lightweight deployment to meet the demands of real-world scenarios. Summary of the Invention

[0005] In view of the shortcomings of the prior art, the purpose of this invention is to provide a method and system for psychological state recognition based on physiological signals, to solve the problems of inaccurate psychological state recognition due to large individual differences, strong environmental interference, and weak dynamic adaptability. This invention utilizes a quality parameter-driven dynamic weighted fusion mechanism to automatically optimize multimodal fusion weights based on signal reliability; it constructs a meta-learning pre-trained basic network and a user-specific adaptation layer, combined with an incremental learning module to achieve rapid personalized modeling with small samples; it triggers online updates based on distribution offset detection, enabling the model to continuously adapt to the evolution of the user's state; and finally, it forms a closed-loop optimization system to improve recognition robustness and individualized accuracy in noisy environments.

[0006] This invention provides a method for identifying psychological states based on physiological signals, comprising: S1: Collect multimodal physiological signals to form a raw signal set. The multimodal physiological signals include at least electroencephalogram (EEG) signals, electrocardiogram (ECG) signals, and electrodermal signal. S2: Perform parallel preprocessing on the original signal set to generate a denoised time-series signal, and simultaneously detect the quality parameters of each modal signal. The quality parameters include the signal-to-noise ratio and the artifact contamination index. S3: Dynamically generate modal weighting coefficients based on quality parameters, and use these weighting coefficients to perform weighted fusion of multimodal time series signals to form a spatiotemporally correlated fusion feature vector; S4: Input the fused feature vector into the personalized recognition model, and output the psychological state probability distribution through the adaptive feature mapping layer in the model. The personalized recognition model is pre-trained by the meta-learning framework to generate the basic network, realizes feature mapping through the user-specific adaptation layer, and embeds an incremental learning module to support online optimization. S5: Based on the distribution offset of user feedback signals or continuous monitoring data, trigger online updates of the parameters of the personalized recognition model to generate an incrementally optimized recognition model.

[0007] In one embodiment of the present invention, the multimodal physiological signals further include eye-tracking signals, respiratory waveform signals, and electromyography (EMG) signals; the acquisition steps are achieved through the collaboration of wearable devices and fixed sensing devices, wherein the electroencephalogram (EEG) signals are acquired using a dry electrode array, the electrocardiogram (ECG) signals are acquired through a three-lead patch, the electrodermal signal is measured by a fingertip sensor, the eye-tracking signal is captured by an embedded infrared camera to capture the pupil movement trajectory, the respiratory signal is monitored through a chest and abdominal pressure band, and the EMG signals are deployed on the facial motion unit; each signal acquisition module integrates a time synchronization device to ensure that the timestamp alignment accuracy of the cross-modal data is higher than the physiological event response threshold.

[0008] In one embodiment of the present invention, the quality parameters further include signal entropy, inter-channel consistency coefficient, and transient stability index; parallel preprocessing includes motion artifact compensation and frequency domain filtering, wherein motion artifact compensation adopts adaptive noise reference channel reduction technology, and frequency domain filtering configures independent passbands for different physiological signal characteristics: EEG signals retain the full frequency band from delta waves to gamma waves, ECG signals focus on the QRS complex characteristic frequency band, and EEG signals enhance low-frequency slowly varying components; the preprocessing process outputs a signal integrity score in real time, which is used for the calculation of weight correction factors in subsequent fusion steps.

[0009] In one embodiment of the present invention, the weighted fusion adopts a spatiotemporal dual-path attention mechanism: the temporal path extracts the temporal dependency features of each modality signal through a long short-term memory network, and the spatial path uses a graph convolutional network to model the multi-channel topological relationship; the modality weight coefficients are jointly generated by quality parameters and task context features, wherein the task context features include the current recognition scene identifier and the historical state transition probability; the fusion process introduces a gating unit to dynamically shield low-confidence signal segments, and the finally generated fusion feature vector has cross-modal spatiotemporal correlation and noise robustness.

[0010] In one embodiment of the present invention, the meta-learning framework is constructed using a model-agnostic meta-learning algorithm: during the pre-training stage of the base network, a large number of user sub-task sets are constructed to simulate individual differences, each sub-task containing a support set and a query set; the network parameters are updated through a dual-loop optimization mechanism, with the inner loop quickly adapting user characteristics on the support set and the outer loop optimizing global generalization ability on the query set; the base network outputs a universal feature encoder shared across users, and the user-specific adaptation layer is implemented by a lightweight neural network, whose parameters are fine-tuned and initialized using new user calibration data.

[0011] In one embodiment of the present invention, the user-specific adaptation layer includes a feature transformation matrix and a bias vector, the dimensions of which match the fused feature vector; the adaptation layer parameter optimization adopts a knowledge distillation constraint mechanism: the base network is used as the teacher model to output soft labels, the adaptation layer is used as the student model to fit the decision distribution of the teacher model, and the user's real labels are introduced to calculate the cross-entropy loss; the optimization process restricts the parameter change amplitude to be lower than a preset threshold to ensure that the personalization process does not destroy the generalization representation ability of the base network.

[0012] In one embodiment of the present invention, the incremental learning module includes an elastic weight solidification unit and a feature replay buffer: during online updates, the Fisher information matrix of the new parameters relative to the historical parameters is calculated, and a high penalty coefficient is applied to important parameters to mitigate the forgetting effect; the feature replay buffer stores the prototype vector of the historical fused feature vector, and during updates, the loss function is calculated by combining the new data and the replay data; the distribution offset is obtained by calculating the Mahalanobis distance between the real-time fused features and the user's historical feature library, and parameter updates are triggered when the distance exceeds an adaptive threshold.

[0013] In one embodiment of the present invention, the online parameter update adopts a sliding window regularization strategy: each update only adjusts the local parameter block of the weight matrix of the fully connected layer, and the update range is determined by the feature distribution offset direction; the update process introduces a momentum acceleration mechanism and gradient clipping constraints, and the updated model parameters can only be deployed after stability verification; the user feedback signal includes manually labeled status labels and implicit behavioral feedback, the latter indirectly generating supervision signals by analyzing user operation response delay and interaction mode changes.

[0014] In one embodiment of the present invention, the method is deployed on an edge computing-cloud platform collaborative architecture: signal acquisition and preprocessing are completed on a local embedded device, and the fused feature vector is uploaded to the cloud after encryption; the personalized recognition model adopts a layered deployment strategy, with the basic network residing in the cloud, and the user-specific adaptation layer and incremental learning module being deployed to the edge terminal; the model update instruction is uniformly issued by the cloud, and after the edge terminal performs incremental training, it feeds back the model difference parameters to the cloud for secure aggregation.

[0015] The present invention also includes a psychological state recognition system based on physiological signals, comprising: The acquisition module acquires multimodal physiological signals to form a raw signal set. The multimodal physiological signals include at least electroencephalogram (EEG) signals, electrocardiogram (ECG) signals, and electrodermal signal. The coordination module performs parallel preprocessing on the original signal set to generate a denoised time-series signal and simultaneously detects the quality parameters of each modal signal, including the signal-to-noise ratio and the artifact contamination index. The verification module dynamically generates modal weighting coefficients based on quality parameters, and uses these weighting coefficients to perform weighted fusion of multimodal time series signals to form a spatiotemporally correlated fusion feature vector. The pre-visualization module inputs the fused feature vectors into the personalized recognition model, and outputs the probability distribution of mental states through the adaptive feature mapping layer in the model. The integration module triggers online updates of the parameters of the personalized recognition model based on the distribution offset of user feedback signals or continuous monitoring data, generating an incrementally optimized recognition model.

[0016] This invention provides a psychological state recognition method and system based on physiological signals. Through a quality parameter-driven dynamic weighted fusion mechanism, it automatically optimizes the multimodal fusion weights based on signal reliability; it constructs a meta-learning pre-trained basic network and a user-specific adaptation layer, and combines an incremental learning module to achieve rapid personalized modeling with small samples; it triggers online updates based on distribution offset detection, enabling the model to continuously adapt to the evolution of the user's state; and finally forms a closed-loop optimization system to improve recognition robustness and individualization accuracy in noisy environments. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of a psychological state recognition method based on physiological signals; Figure 2 A schematic diagram illustrating the workflow of a psychological state recognition method based on physiological signals; Figure 3 This is a system architecture diagram of a psychological state recognition system based on physiological signals. Detailed Implementation

[0019] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.

[0020] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0021] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.

[0022] Please see Figure 1-3The image shows a method and system for identifying psychological states based on physiological signals according to the present invention. The method includes: S1: acquiring multimodal physiological signals to form a raw signal set, the multimodal physiological signals including at least EEG signals, ECG signals, and dermal conductance signals; S2: performing parallel preprocessing on the raw signal set to generate denoised time-series signals, and simultaneously detecting the quality parameters of each modality signal, including signal-to-noise ratio and artifact contamination index; S3: dynamically generating modality weight coefficients based on the quality parameters, and using these weight coefficients to perform weighted fusion of the multimodal time-series signals to form a spatiotemporally correlated fusion feature vector; S4: inputting the fusion feature vector into a personalized recognition model, outputting a psychological state probability distribution through an adaptive feature mapping layer in the model. The personalized recognition model is pre-trained using a meta-learning framework to generate a basic network, implementing feature mapping through a user-specific adaptation layer, and embedding an incremental learning module to support online optimization; S5: triggering online parameter updates of the personalized recognition model based on the distribution offset of user feedback signals or continuously monitored data, generating an incrementally optimized recognition model.

[0023] like Figure 1 As shown, the process includes five core steps: First, multimodal physiological signal acquisition is performed, capturing a raw signal set containing at least EEG, ECG, and SCEA signals through a biosensor array. This step forms the data foundation for subsequent processing. Next, parallel preprocessing is performed on the raw signal set, performing noise filtering and artifact removal while independently processing each modality signal to generate clean signals that meet timing requirements. Simultaneously, quality parameters reflecting signal reliability are detected, including the signal-to-noise ratio (SNR) characterizing signal purity and the artifact contamination index indicating interference intensity. Then, dynamic weight calculation is performed based on the aforementioned quality parameters, and the confidence level of each modality signal is evaluated in real time. The method generates corresponding modal weight coefficients, which are then used to adaptively weight and fuse multimodal time-series signals, forming a fused feature vector that combines temporal dynamics with spatial correlation. This fused feature vector is then input into a pre-defined personalized recognition model. This model is pre-trained using a meta-learning framework to generate a basic network structure and integrates a user-specific adaptation layer to achieve individualized feature mapping. Finally, the adaptive feature mapping layer in the model outputs a multidimensional psychological state probability distribution. Based on user-initiated feedback or automatically monitored feature distribution offsets, an online parameter update mechanism for the personalized recognition model is triggered. An incremental optimization strategy generates a performance-improved recognition model, forming a closed-loop evolutionary system. This method addresses the problem of unbalanced reliability in multimodal signals through quality-driven dynamic fusion, overcomes individual differences using a meta-learning architecture, and achieves continuous model evolution through an incremental update mechanism, significantly improving the accuracy of psychological state recognition in complex environments.

[0024] Furthermore, the types of multimodal physiological signals and their acquisition methods are limited: in addition to basic EEG, ECG, and electrodermal signals, the multimodal physiological signals must include three key physiological indicators: eye-tracking signals, respiratory waveform signals, and electromyography (EMG) signals, forming a six-modal collaborative sensing system. At the acquisition implementation level, a heterogeneous collaborative architecture of wearable devices and fixed sensing devices is adopted. For EEG signal acquisition, a non-conductive dry electrode array is used, with a ring-shaped topology covering the prefrontal cortex and motor cortex. For ECG signals, standard limb lead waveforms are obtained through a medical-grade three-lead patch to ensure the stability of R-wave peak detection. The system employs a high-sensitivity fingertip sensor to measure changes in sweat gland conductivity. Eye movement signals are acquired using a near-infrared miniature camera to capture pupil movement trajectories, and the gaze vector is calculated by combining this with corneal reflection points. Respiratory signals are monitored via an embedded pressure-sensitive band to detect chest and abdominal undulations, with temperature compensation eliminating environmental interference. Electromyography (EMG) signals are primarily deployed on the zygomaticus major and orbicularis oculi muscles to capture electrical activity related to micro-expressions. All signal acquisition modules are equipped with high-precision hardware clocks, and a bidirectional timestamp calibration protocol ensures cross-modal data time synchronization accuracy better than five milliseconds, meeting the physiological event analysis requirements for strict phase alignment of multimodal signals. This extended solution enhances the comprehensiveness of physiological representation through six-modal signal complementarity, and the heterogeneous acquisition architecture balances applicability to mobile scenarios with medical-grade data quality. The hardware-level time synchronization mechanism provides a precise alignment foundation for subsequent spatiotemporal fusion.

[0025] like Figure 1As shown, the details of quality parameter definition and preprocessing techniques are further refined: In addition to the basic signal-to-noise ratio and artifact contamination index, three advanced evaluation parameters must be added: signal entropy, inter-channel consistency coefficient, and transient stability index, forming a five-dimensional quality evaluation system; the signal entropy is calculated using a sliding window approximate entropy algorithm to quantify changes in signal complexity; the inter-channel consistency coefficient is obtained through multi-channel cross-correlation analysis to reflect signal spatial consistency; and the transient stability index is calculated based on the first-order difference variance of the signal to identify the intensity of sudden interference. The parallel preprocessing process comprises two paths: motion artifact compensation and frequency domain filtering. Motion artifact compensation employs adaptive noise reference channel reduction technology, constructing a motion noise model through independent accelerometer channels and dynamically updating filter coefficients using a least mean square algorithm. Frequency domain filtering configures independent passband parameters for different physiological signal characteristics. For EEG signals, the full frequency band from 0.5 Hz to 100 Hz is retained, covering all rhythmic components of delta waves, theta waves, alpha waves, beta waves, and gamma waves. For ECG signals, the characteristic frequency band from 0.5 Hz to 40 Hz is focused to enhance the morphological fidelity of the QRS complex. For EEG signals, a bandpass filter from 0.05 Hz to 5 Hz is used to highlight low-frequency, slowly varying emotional arousal characteristics. The preprocessing process outputs a real-time signal integrity score, which is generated by weighted calculation of five-dimensional quality parameters and serves as the input for weight correction factors in subsequent fusion steps. This scheme achieves precise quantification of signal reliability through multi-dimensional quality assessment, significantly improves the usability of data in mobile scenarios through motion artifact compensation technology, and optimizes the feature extraction basis through physiological characteristic adaptation design of frequency domain filtering.

[0026] Furthermore, the technical implementation details of the fusion mechanism are specified: the weighted fusion operation adopts a spatiotemporal dual-path attention architecture, wherein the temporal path processes the temporal dependencies of each modality signal through a long short-term memory network. This network includes forget gates, input gates, and output gate control units, and captures the long-range temporal dynamic characteristics of physiological signals through cyclic connections; the spatial path uses a graph convolutional network to model the topological relationship structure of multi-channel signals, mapping the EEG electrode positions, ECG lead layout, and skin conduction sensor placement as graph nodes, constructing an adjacency matrix based on the physiological distance between nodes, and aggregating spatial neighborhood features through graph convolutional layers; the modality weight coefficients are composed of quality parameters and any The task context features are jointly generated, including the current identification scene classification identifier and the historical state transition probability matrix. The former distinguishes application scenarios such as driving, learning, and medical care, while the latter records the user's recent psychological state transition patterns. The fusion process introduces a gating unit to dynamically shield low-confidence signal segments. This unit generates a binary mask in real time based on the signal integrity score, blocking the signal from flowing into the fusion channel when the score falls below an adaptive threshold. The final generated fused feature vector exhibits cross-modal spatiotemporal correlation and noise robustness, with its dimensionality compressed to less than 20% of the original signal space while retaining over 95% of the discriminative information. This scheme overcomes the shortcomings of traditional single-path representation through spatiotemporal dual-path processing, the task context awareness mechanism adapts the fusion weights to specific application requirements, and the dynamic gating technology effectively suppresses low-quality signal segment contamination.

[0027] like Figure 1 As shown, based on the personalized recognition model, the implementation of the meta-learning framework is as follows: The meta-learning framework uses a model-independent meta-learning algorithm to construct the basic network. In the pre-training stage, a meta-training set containing a large number of user sub-tasks is constructed to simulate individual differences. Each sub-task is divided into a support set and a query set. The network parameters are updated through a dual-loop optimization mechanism: the inner loop performs user-level fast adaptation on the support set, using one or more steps of gradient descent to adjust the network parameters to adapt to a specific user distribution; the outer loop evaluates the performance of the adapted model on the query set, calculates higher-order gradients based on cross-user generalization loss, and optimizes the initialization parameters of the basic network through backpropagation; the basic network outputs a universal feature encoder shared across users, which learns the essential features of physiological signals that are insensitive to individual differences; the user-specific adaptation layer is implemented by two fully connected neural networks, whose input dimension matches the fused feature vector, and whose output dimension is consistent with the input layer of the basic network; the adaptation layer parameters are fine-tuned and initialized using a small amount of calibration data from new users. During the fine-tuning process, the basic network parameters are frozen, and only the weight matrix and bias vector of the adaptation layer are updated. This technical solution simulates individual differences through meta-learning, enabling the basic network to quickly adapt to new users; the adaptation layer, as a lightweight conversion interface, achieves efficient personalized customization while protecting the generalization capability of the basic network.

[0028] like Figure 2 As shown, its parameter optimization mechanism is defined as follows: the user-specific adaptation layer includes a learnable feature transformation matrix and a bias vector, the dimensions of which are dynamically configured according to the size of the fused feature vector; the adaptation layer parameter optimization adopts a knowledge distillation constraint mechanism, using the basic network described in claim 5 as the teacher model, and the teacher model outputs the soft probability distribution of various psychological states as soft labels after processing calibration data; the adaptation layer, as the student model, fits the decision distribution of the teacher model, and its optimization objective includes two loss functions: one is to calculate the cross-entropy loss between the student model output and the user's real labeled label, and the other is to calculate the KL divergence loss between the student model output and the teacher model's soft label; the total loss function is the weighted sum of the two losses, where the weight of the KL divergence loss term decays exponentially with the training rounds; the optimization process restricts the change of adaptation layer parameters to be lower than a preset threshold, and ensures that the Euclidean distance after parameter update does not exceed 10% of the initial parameters through the projection gradient descent algorithm. This constraint prevents the personalization process from destroying the general representation ability of the basic network learning. This scheme transfers the essence of knowledge from the basic network through knowledge distillation, balances individualized adaptation and generalization ability protection through dual loss design, and avoids the adaptation layer from deviating excessively from the optimal solution space through parameter change constraints.

[0029] Specifically, the operation mechanism of the incremental learning module is refined as follows: The incremental learning module consists of two parts: an elastic weight solidification unit and a feature replay buffer. The elastic weight solidification unit calculates the diagonal elements of the Fisher information matrix of the new data parameters relative to the historical parameters. The element values ​​reflect the importance of the parameters to the historical tasks. During online updates, a secondary regularization penalty is applied to important parameters, and the penalty coefficient is proportional to the Fisher information content. The feature replay buffer stores the prototype vectors of the historical fused feature vectors. The prototype library is updated every 24 hours using the K-means clustering algorithm, and the storage capacity is dynamically adjusted to 5% of the user's memory. During online updates, the loss function is calculated by combining the new input data and the prototype data extracted from the replay buffer. The sampling weight of the new data is three times that of the prototype data. The distribution offset is obtained by calculating the Mahalanobis distance between the real-time fused feature vector and the user's historical feature library. The historical feature library contains the sliding window statistics of the feature vectors for the past 30 days. When the Mahalanobis distance exceeds the adaptive threshold, parameter updates are triggered. This threshold is dynamically adjusted according to the user's state stability index. This technical solution mitigates the forgetting of important historical knowledge through flexible weight solidification, maintains the model's ability to discriminate historical distributions through feature replay mechanism, and achieves precise update trigger control through Mahalanobis distance detection.

[0030] like Figure 3As shown, this invention relates to a psychological state recognition system based on physiological signals, comprising: an acquisition module that acquires multimodal physiological signals to form a raw signal set, the multimodal physiological signals including at least electroencephalogram (EEG), electrocardiogram (ECG), and electrodermal signal; a coordination module that performs parallel preprocessing on the raw signal set to generate denoised time-series signals and simultaneously detects the quality parameters of each modality signal, including signal-to-noise ratio (SNR) and artifact contamination index; a verification module that dynamically generates modality weight coefficients based on the quality parameters and uses these weight coefficients to perform weighted fusion of the multimodal time-series signals to form a spatiotemporally correlated fusion feature vector; a pre-exercise module that inputs the fusion feature vector into a personalized recognition model and outputs a psychological state probability distribution through an adaptive feature mapping layer in the model; and an integration module that triggers online parameter updates of the personalized recognition model based on the distribution offset of user feedback signals or continuous monitoring data, generating an incrementally optimized recognition model.

[0031] like Figure 3 As shown, based on online parameter updates, the specific update execution strategy and feedback signal source are defined as follows: The online parameter update adopts a sliding window regularization strategy. Each update operation only adjusts a local parameter block of the weight matrix of the fully connected layer of the personalized recognition model. The selection range of this parameter block is determined by the real-time feature distribution offset direction—when the offset principal component analysis shows that the variance ratio of the first principal component exceeds a set threshold, the parameter block corresponding to the hidden layer neuron with the highest correlation to that principal component is locked; the update process introduces a momentum acceleration mechanism and gradient clipping constraints. The momentum coefficient is set to above 0.9 to accelerate convergence, and the gradient clipping threshold is dynamically adjusted according to the historical gradient magnitude; the updated model parameters need to be verified for stability before use. The deployment and verification method involves inputting feature vector cache data from the past three days, requiring the variance of the recognition results to be lower than the baseline value before the update. The user feedback signals include two types: explicit manual labeling and implicit behavioral feedback. Manual labeling receives real-time psychological state labels through the user interface, while implicit behavioral feedback indirectly generates supervision signals by analyzing changes in user operation response delay, interaction trajectory complexity, and interface dwell time. When the operation response delay exceeds twice the normal range by two standard deviations, it is marked as an inattentive state; when the disorder of the interaction trajectory increases by more than 30%, it is marked as an anxious state. This scheme reduces computational overhead through local parameter updates, improves optimization efficiency through momentum acceleration, and enhances the coverage of supervision signals through dual-source feedback.

[0032] Specifically, the hardware deployment architecture and data security mechanism are defined as follows: the method is deployed in a collaborative architecture of edge computing nodes and cloud platform. The signal acquisition device and preprocessing module are integrated into the local embedded terminal for execution. The fused feature vector is encrypted using national cryptographic algorithms and then uploaded to the cloud server through a secure channel. The personalized recognition model adopts a layered deployment strategy. The basic network shared across users resides in the high-performance computing cluster in the cloud, while the user-specific adaptation layer and incremental learning module are deployed to the edge terminal device. The model update command is uniformly issued by the cloud with an update strategy description file. The edge terminal performs local incremental training based on this file. After training, only the model difference parameters are fed back to the cloud secure aggregation server. The transmission of difference parameters is processed using homomorphic encryption technology. The cloud aggregates at least one hundred user difference parameters before performing global model fusion. The edge terminal sets up a local model firewall to perform integrity verification and behavior auditing on the update commands issued by the cloud. When abnormal commands are detected, a rollback mechanism is initiated to restore to a trusted version. This deployment scheme alleviates the pressure on edge devices by offloading computing load, protects user privacy data from leakage through layered model deployment, and achieves collective intelligent evolution through the secure aggregation mechanism.

[0033] Furthermore, based on the output of the probability distribution of psychological states, an extended feedback intervention closed-loop system is established: when the probability of a negative psychological state exceeds the adaptive threshold, a multimodal intervention strategy library generation engine is automatically triggered. This engine includes four basic intervention resources: audio guidance for cognitive behavioral therapy, visual guidance for breathing rhythm, mindfulness meditation videos, and immersive natural scene simulations. The strategy library optimizes the push plan based on the user's historical intervention effect records, pushing interactive adjustment tasks to the user's mobile terminal. The intervention effect is quantitatively evaluated by re-collecting physiological signals, with the collection period covering five minutes before intervention to ten minutes after intervention. Evaluation indicators include the rate of decrease in skin conductance response, the degree of increase in heart rate variability, and the increase in EEG alpha wave power. The evaluation results are converted into reinforcement learning reward signals to input the recognition model optimization module. The reward function is designed as the logarithmic function of the post-intervention state improvement rate. The system establishes a user psychological state evolution map, which records the psychological state transition trajectory on a daily basis. The system uses long-term and short-term time series prediction algorithms to predict the state trend for the next 72 hours and dynamically adjusts the sensitivity threshold of the recognition model accordingly—reducing the alarm threshold by 15% when a high-stress period is predicted and increasing the threshold by 20% during stable periods to avoid false alarms. This closed-loop system realizes a complete chain from state recognition to proactive intervention. The reinforcement learning mechanism enables the model to adaptively optimize the intervention strategy, and the state prediction function provides proactive protection.

[0034] This invention discloses a psychological state recognition method and system based on physiological signals. Through a quality parameter-driven dynamic weighted fusion mechanism, it automatically optimizes the multimodal fusion weights based on signal reliability; it constructs a meta-learning pre-trained basic network and a user-specific adaptation layer, and combines an incremental learning module to achieve rapid personalized modeling with small samples; it triggers online updates based on distribution offset detection, enabling the model to continuously adapt to the evolution of the user's state; and finally forms a closed-loop optimization system to improve recognition robustness and individualization accuracy in noisy environments.

[0035] Therefore, the psychological state recognition method and system based on physiological signals of the present invention can solve the problem of inaccurate psychological state recognition due to large individual differences, strong environmental interference, and weak dynamic adaptability.

[0036] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.

Claims

1. A method for recognizing psychological states based on physiological signals, characterized in that, include: S1: Collect multimodal physiological signals to form a raw signal set, wherein the multimodal physiological signals include at least electroencephalogram (EEG) signals, electrocardiogram (ECG) signals, and electrodermal signal; S2: Perform parallel preprocessing on the original signal set to generate a denoised time-series signal, and simultaneously detect the quality parameters of each modal signal, including the signal-to-noise ratio and the artifact contamination index. S3: Based on the quality parameters, dynamically generate modal weighting coefficients, and use these weighting coefficients to perform weighted fusion of multimodal time-series signals to form a spatiotemporally correlated fusion feature vector; S4: Input the fused feature vector into the personalized recognition model, and output the psychological state probability distribution through the adaptive feature mapping layer in the model. The personalized recognition model is a basic network pre-trained by a meta-learning framework, and the feature mapping is realized through a user-specific adaptation layer. An incremental learning module is embedded to support online optimization. S5: Based on the distribution offset of user feedback signals or continuous monitoring data, trigger the online update of the parameters of the personalized recognition model to generate an incrementally optimized recognition model.

2. The psychological state recognition method based on physiological signals according to claim 1, characterized in that, The multimodal physiological signals further include eye-tracking signals, respiratory waveform signals, and electromyography (EMG) signals. The acquisition process is achieved through the collaboration of wearable devices and fixed sensing devices. Specifically, EEG signals are acquired using a dry electrode array, ECG signals are acquired through a three-lead patch, ESC signals are measured by a fingertip sensor, eye-tracking signals are captured by an embedded infrared camera to capture pupil movement trajectories, respiratory signals are monitored through a chest and abdominal pressure band, and EMG signals are deployed on a facial motion unit. Each signal acquisition module integrates a time synchronization device to ensure that the timestamp alignment accuracy of cross-modal data is higher than the physiological event response threshold.

3. The psychological state recognition method based on physiological signals according to claim 1, characterized in that, The quality parameters further include signal entropy, inter-channel consistency coefficient, and transient stability index; the parallel preprocessing includes motion artifact compensation and frequency domain filtering, wherein motion artifact compensation adopts adaptive noise reference channel reduction technology, and frequency domain filtering configures independent passbands for different physiological signal characteristics: EEG signals retain the full frequency band from delta wave to gamma wave, ECG signals focus on the QRS complex characteristic frequency band, and EEG signals enhance low-frequency slowly varying components; the preprocessing process outputs a signal integrity score in real time, which is used for the calculation of weight correction factors in subsequent fusion steps.

4. The psychological state recognition method based on physiological signals according to claim 1, characterized in that, The weighted fusion adopts a spatiotemporal dual-path attention mechanism: the temporal path extracts the temporal dependency features of each modality signal through a long short-term memory network, and the spatial path uses a graph convolutional network to model the multi-channel topological relationship; the modality weight coefficients are jointly generated by quality parameters and task context features, wherein the task context features include the current recognition scene identifier and historical state transition probabilities. The fusion process introduces a gating unit to dynamically shield low-confidence signal segments, and the final fused feature vector has cross-modal spatiotemporal correlation and noise robustness.

5. The psychological state recognition method based on physiological signals according to claim 1, characterized in that, The meta-learning framework is constructed using a model-independent meta-learning algorithm: during the pre-training stage of the base network, a large number of user sub-task sets are constructed to simulate individual differences, with each sub-task containing a support set and a query set; network parameters are updated through a dual-loop optimization mechanism, with the inner loop quickly adapting user characteristics on the support set and the outer loop optimizing global generalization ability on the query set; the base network outputs a universal feature encoder shared across users, and the user-specific adaptation layer is implemented by a lightweight neural network, whose parameters are fine-tuned and initialized using new user calibration data.

6. The psychological state recognition method based on physiological signals according to claim 1, characterized in that, The user-specific adaptation layer includes a feature transformation matrix and a bias vector, the dimensions of which match the fused feature vector. The adaptation layer parameter optimization adopts a knowledge distillation constraint mechanism: the base network is used as the teacher model to output soft labels, and the adaptation layer is used as the student model to fit the decision distribution of the teacher model. At the same time, the user's real labels are introduced to calculate the cross-entropy loss. The optimization process restricts the parameter change to be lower than a preset threshold to ensure that the personalization process does not destroy the generalization representation ability of the base network.

7. The psychological state recognition method based on physiological signals according to claim 1, characterized in that, The incremental learning module includes an elastic weight solidification unit and a feature replay buffer: during online updates, the Fisher information matrix of the new parameters relative to the historical parameters is calculated, and a high penalty coefficient is applied to important parameters to mitigate the forgetting effect; The feature replay buffer stores the prototype vector of the historical fused feature vector. When updating, the loss function is calculated by combining the new data and the replay data. The distribution offset is obtained by calculating the Mahalanobis distance between the real-time fused features and the user's historical feature library. When the distance exceeds the adaptive threshold, the parameter update is triggered.

8. The psychological state recognition method based on physiological signals according to claim 1, characterized in that, The online parameter update adopts a sliding window regularization strategy: each update only adjusts the local parameter block of the weight matrix of the fully connected layer, and the update range is determined by the feature distribution offset direction; the update process introduces a momentum acceleration mechanism and gradient clipping constraints, and the updated model parameters can only be deployed after stability verification; the user feedback signal includes manually labeled status labels and implicit behavioral feedback, the latter indirectly generating supervision signals by analyzing user operation response delay and interaction mode changes.

9. The psychological state recognition method based on physiological signals according to claim 1, characterized in that, The method is deployed in an edge computing-cloud platform collaborative architecture: signal acquisition and preprocessing are completed on a local embedded device. The fused feature vectors are encrypted and then uploaded to the cloud; the personalized recognition model adopts a layered deployment strategy, with the basic network residing in the cloud and the user-specific adaptation layer and incremental learning module being deployed to the edge terminal; the model update instructions are uniformly issued by the cloud, and after the edge terminal performs incremental training, it feeds back the model difference parameters to the cloud for secure aggregation.

10. A preoperative assessment system using the physiological signal-based psychological state recognition method according to any one of claims 1-9, characterized in that, include: The acquisition module acquires multimodal physiological signals to form a raw signal set, wherein the multimodal physiological signals include at least electroencephalogram (EEG) signals, electrocardiogram (ECG) signals, and electrodermal signal. The coordination module performs parallel preprocessing on the original signal set to generate a denoised time-series signal and simultaneously detects the quality parameters of each modal signal, including the signal-to-noise ratio and the artifact contamination index. The verification module dynamically generates modal weighting coefficients based on the quality parameters, and uses these weighting coefficients to perform weighted fusion of multimodal time-series signals to form a spatiotemporally correlated fusion feature vector. The pre-visualization module inputs the fused feature vector into the personalized recognition model and outputs the psychological state probability distribution through the adaptive feature mapping layer in the model. An integration module triggers online updates of the parameters of the personalized recognition model based on the distribution offset of user feedback signals or continuous monitoring data, generating an incrementally optimized recognition model.

Citation Information

Cited By

  • Emotion recognition method based on online cross-modal knowledge distillation

    CN121960702A

  • Power equipment auxiliary monitoring system and method based on multi-modal data fusion

    CN121980352A