Intelligent psychological state evaluation method based on multi-modal data fusion

By using multimodal data fusion and deep learning models, physiological signals, facial images, speech, and gait data are collected simultaneously, which solves the shortcomings of single-modal assessment and enables accurate and continuous monitoring and assessment of psychological state, improving the accuracy and adaptability of the assessment.

CN122163215APending Publication Date: 2026-06-09FOREIGN ECONOMIC & TRADE UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FOREIGN ECONOMIC & TRADE UNIV
Filing Date
2026-03-11
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

In existing technologies, psychological state assessment methods based on single-modal data suffer from limited information dimensions, susceptibility to environmental interference, and significant impact from individual differences. Furthermore, they lack multimodal data fusion strategies, fail to fully exploit complementary information between modalities, and cannot adapt to the dynamic changes in modal correlation under different scenarios, resulting in insufficient accuracy and reliability of assessment results.

Method used

An intelligent psychological state assessment method employing multimodal data fusion is proposed. This method simultaneously collects physiological signals, facial images, speech, behavioral actions, and gait data. It uses an adaptive fusion strategy and a deep learning model to perform feature fusion and psychological state assessment, including data preprocessing, dynamic feature extraction, feature correlation calculation, and adaptive weight adjustment.

Benefits of technology

It enables accurate, objective, and continuous monitoring of psychological states, improves the accuracy of assessment results and the robustness of the system, adapts to changes in modal correlation under different psychological state scenarios, and fully explores complementary information between multiple modalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122163215A_ABST
    Figure CN122163215A_ABST
Patent Text Reader

Abstract

The application discloses an intelligent psychological state evaluation method based on multi-modal data fusion, and relates to the technical field of psychological function testing.The method comprises the following steps: acquiring multi-modal physiological behavior data of a target object, wherein the multi-modal physiological behavior data comprises physiological signal data, facial image data, voice data, behavior action data and gait data; pre-processing the multi-modal physiological behavior data to obtain standardized data; performing dynamic feature extraction on the standardized data to obtain feature vectors corresponding to each mode; calculating the feature correlation degree between the gait mode and other modes; performing correlation fusion processing on the feature vectors corresponding to each mode to obtain a fusion feature vector; inputting the fusion feature vector into a pre-trained psychological state evaluation model to obtain a psychological state evaluation result; and through the collaborative analysis and deep fusion of multi-modal data, the problem of insufficient evaluation accuracy of single-mode data is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of psychological function testing technology, specifically to an intelligent psychological state assessment method based on multimodal data fusion. Background Technology

[0002] Mental health issues have become a significant public health challenge in today's society. Timely and accurate identification and assessment of an individual's mental state are of great importance for mental health intervention and adjustment. Traditional mental state assessment mainly relies on questionnaires and interviews with professionals, which have problems such as strong subjectivity, poor timeliness, and difficulty in continuous monitoring.

[0003] With the development of sensor technology and artificial intelligence technology, objective psychological state assessment methods based on physiological and behavioral data have gradually attracted attention. In existing technologies, some solutions use single-modal data for psychological state assessment, such as using only heart rate variability data to analyze stress levels or using only facial expressions to identify emotional states. However, single-modal data has the disadvantages of limited information dimensions, susceptibility to environmental interference, and large influence from individual differences, resulting in insufficient accuracy and reliability of assessment results.

[0004] Other solutions attempt to combine multiple modalities for psychological state assessment, but they have shortcomings in data fusion strategies: simple feature splicing methods cannot fully explore the complementary information between modalities, and fixed-weight fusion methods cannot adapt to changes in the quality of data from different modalities under different scenarios, resulting in unsatisfactory fusion results; in addition, existing solutions do not pay enough attention to the temporal dynamic changes in psychological states, making it difficult to capture the evolutionary patterns of psychological states; at the same time, existing technologies do not make sufficient use of gait information. Gait, as an important behavioral feature that can reflect an individual's psychological state, has a close relationship with other physiological behavioral modalities, but existing multimodal fusion solutions mostly adopt single gait modalities or simple fusion strategies, failing to fully utilize the correlation characteristics between gait modalities and other modalities for adaptive weight adjustment, and cannot adapt to the dynamic changes in modal correlation under different psychological states.

[0005] To address the shortcomings of existing technologies, this invention provides a method for monitoring and assessing psychological states based on multimodal data fusion. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides an intelligent psychological state assessment method based on multimodal data fusion. By collecting physiological and behavioral data from multiple modalities, an adaptive fusion strategy is adopted for feature fusion, and a deep learning model is used for psychological state assessment. This enables accurate, objective, and continuous monitoring of the psychological state of the target object, providing scientific support for psychological health intervention and adjustment.

[0007] The technical solution provided by this invention is as follows:

[0008] A method for monitoring and assessing mental state based on multimodal data fusion includes the following steps:

[0009] Acquire multimodal physiological and behavioral data of the target object, including physiological signal data, facial image data, voice data, behavioral action data, and gait data;

[0010] The multimodal physiological behavior data is preprocessed to obtain standardized data;

[0011] Dynamic feature extraction is performed on the standardized data to obtain feature vectors corresponding to each modality;

[0012] Calculate the feature correlation degree between the gait mode and other modes, and perform correlation and fusion processing on the feature vectors corresponding to each mode to obtain the fused feature vector;

[0013] The fused feature vector is input into the pre-trained psychological state assessment model to obtain the psychological state assessment result.

[0014] Furthermore, the physiological signal data includes heart rate variability data, skin conductance data, and electroencephalogram (EEG) signal data; the facial image data is a continuous sequence of image frames containing the facial region of the target object; the speech data is the speech audio signal of the target object; the behavioral action data is the limb movement trajectory data of the target object; and the gait data is the gait motion parameter data of the target object during walking, including cadence data, stride length data, plantar pressure distribution data, and body center of gravity movement trajectory data.

[0015] Furthermore, the preprocessing of the multimodal physiological behavior data to obtain standardized data includes:

[0016] The physiological signal data are subjected to bandpass filtering and baseline drift correction to remove power frequency interference and motion artifacts from the signal;

[0017] The facial image data is processed by face detection, key point localization and geometric alignment, and the facial region is normalized to a preset size (e.g. 224×224 pixels).

[0018] The speech data is subjected to endpoint detection, silence removal, and spectrum normalization.

[0019] The behavioral data is subjected to coordinate system 1 and trajectory interpolation smoothing; the gait data is subjected to gait cycle segmentation, abnormal gait point removal and temporal alignment processing to normalize the gait sequence to a uniform time scale.

[0020] Further, the step of dynamically extracting features from the standardized data to obtain feature vectors corresponding to each modality includes:

[0021] Time-domain statistical features and frequency-domain power spectrum features are extracted from the physiological signal data. The time-domain statistical features include mean, variance, and peak-to-peak value. The frequency-domain power spectrum features include the energy proportion of each frequency band.

[0022] Extract facial motion unit activation intensity features and facial expression category probability distribution features from the facial image data;

[0023] Extract Mel frequency cepstral coefficient features, fundamental frequency variation features, and speech rate features from the speech data;

[0024] Extract joint angle features, amplitude features, and rhythmic features from the behavioral data;

[0025] Gait spatiotemporal parameter features, gait symmetry features, and gait stability features are extracted from the gait data. The gait spatiotemporal parameter features include stride frequency, stride length, stride speed, and the proportion of support phase duration. The gait symmetry features include the difference between left and right stride lengths and the difference between left and right support phase times. The gait stability features include the gait variation coefficient and the amplitude of center of gravity swing.

[0026] Furthermore, the calculation of the feature correlation degree between the gait mode and other modes, and the correlation and fusion processing of the feature vectors corresponding to each mode to obtain the fused feature vector, includes:

[0027] Calculate the data quality assessment score for each modality separately;

[0028] Calculate the feature correlation degree between the gait modality feature vector and the feature vectors of other modalities, and the feature correlation degree is obtained by cosine similarity calculation;

[0029] The modal correlation adjustment factor is calculated based on the feature correlation, and the fusion weight of each modality is dynamically adjusted based on the modal correlation adjustment factor.

[0030] Based on the data quality assessment score, the preset modality base weight coefficient, and the modality correlation adjustment factor, the adaptive fusion weight of each modality feature vector is calculated;

[0031] Based on the multi-head attention mechanism, the adaptive fusion weights are used to weight and aggregate the feature vectors of each modality to obtain the fused feature vector.

[0032] Further, the step of calculating the modal correlation adjustment factor based on the feature correlation, and dynamically adjusting the fusion weights of each modality based on the modal correlation adjustment factor, includes:

[0033] The cosine similarity between the gait modal feature vector and the physiological signal modal feature vector is calculated as the gait-physiology correlation degree.

[0034] The cosine similarity between the gait modal feature vector and the facial image modal feature vector is calculated as the gait-face correlation.

[0035] If the gait-physiological correlation is greater than or equal to a preset high correlation threshold, the fusion weight of the gait mode is increased by a first preset ratio, and the fusion weight of the physiological signal mode is increased by a second preset ratio.

[0036] If the gait-face correlation is less than a preset low correlation threshold, the fusion weight of the gait modality is reduced by a third preset ratio, and the fusion weight of the speech modality is increased by a fourth preset ratio.

[0037] Wherein, the high correlation threshold is 0.7, the low correlation threshold is 0.3, the first preset ratio is 20%, the second preset ratio is 10%, the third preset ratio is 10%, and the fourth preset ratio is 15%.

[0038] Further, the calculation of data quality assessment scores for each modality includes:

[0039] Calculate the signal quality factor based on the signal-to-noise ratio of each modality data;

[0040] Calculate the data completeness factor based on the missing rate of each modality;

[0041] The signal quality factor and the data integrity factor are weighted and summed to obtain the data quality evaluation score for each modality, wherein the weight coefficient of the signal quality factor is 0.6 and the weight coefficient of the data integrity factor is 0.4.

[0042] Furthermore, the psychological state assessment model includes a feature encoding layer, a temporal modeling layer, and a classification output layer connected in sequence;

[0043] The feature encoding layer adopts a fully connected neural network structure to perform nonlinear transformation and dimensionality reduction on the fused feature vector and output the encoded feature vector.

[0044] The temporal modeling layer adopts a long short-term memory network structure to model the temporal dependency relationship of the encoded feature vectors within a continuous time window and output temporal feature vectors.

[0045] The classification output layer employs a fully connected neural network structure and a normalized exponential function to output the probability distribution of mental state categories based on the temporal feature vector.

[0046] Furthermore, the psychological state assessment result includes a psychological state category label and a corresponding confidence score; the psychological state category label is one of the following: normal state, anxious state, depressed state, and stressed state; the confidence score is the maximum category probability value output by the psychological state assessment model.

[0047] Furthermore, the method also includes:

[0048] The psychological state assessment results of multiple consecutive time periods are statistically analyzed using a preset time window length to calculate the frequency of occurrence and average confidence score of each psychological state category.

[0049] Based on the frequency of occurrence and average confidence score, a psychological state change trend curve is generated, and a psychological state trend analysis report is output.

[0050] Furthermore, the method also includes:

[0051] Based on the psychological state assessment results and the psychological state trend analysis report, an intervention recommendation plan is generated for the target group;

[0052] The proposed intervention plan includes the type of intervention, the level of intervention intensity, and specific intervention measures;

[0053] The intervention types are determined based on psychological state category labels, including relaxation training intervention, cognitive regulation intervention, behavioral activation intervention, and professional referral intervention;

[0054] The intervention intensity level is determined based on the confidence score and the trend of psychological state change, and is divided into three levels: mild intervention, moderate intervention, and severe intervention.

[0055] The beneficial effects of this invention are as follows:

[0056] First, this invention comprehensively depicts the physiological and psychological state of the target object from multiple dimensions by simultaneously collecting data from five modalities: physiological signals, facial images, voice, behavioral actions, and gait. Compared with single-modal methods, it has a richer source of information. In particular, the gait modality can reflect the changes in behavioral characteristics of an individual under different psychological states such as anxiety and depression, and can more accurately reflect the complexity of psychological states.

[0057] Second, the present invention adopts an adaptive fusion strategy based on data quality assessment and modality correlation adjustment factor, which can dynamically adjust the fusion weight according to the actual quality of each modality data and the correlation between gait modality and other modalities, effectively reducing the negative impact of low-quality data on the evaluation results. At the same time, it can adapt to the changes in modality correlation under different psychological state scenarios, improving the robustness and adaptability of the system.

[0058] Third, this invention utilizes a multi-head attention mechanism combined with modal correlation for feature fusion, which can fully explore the complementary information and correlation between different modalities. In particular, by calculating the cosine similarity between gait and modalities such as physiological signals and facial images and adjusting the weights accordingly, it achieves dynamic weight fusion that is more in line with the actual evaluation scenario, and realizes more effective information integration than simple feature splicing and fixed weight fusion.

[0059] Fourth, the present invention adopts a full-processing flow of "synchronous acquisition - dynamic features - correlation fusion", which is different from the existing technical solutions of "single step modality, simple fusion, basic features". It realizes the collaborative processing and deep fusion of multimodal data and meets the requirements of technical novelty. Attached Figure Description

[0060] Figure 1 A flowchart illustrating the overall process of psychological state monitoring and assessment based on multimodal data fusion, as provided in this embodiment of the invention.

[0061] Figure 2 This is a schematic diagram of the structure of the multimodal data acquisition module provided in an embodiment of the present invention;

[0062] Figure 3 This is a schematic diagram of the multimodal feature extraction process provided in an embodiment of the present invention;

[0063] Figure 4 This is a schematic diagram of the multimodal feature fusion module structure provided in an embodiment of the present invention;

[0064] Figure 5 This is a schematic diagram of the psychological state assessment model structure provided in an embodiment of the present invention;

[0065] Figure 6 This is a schematic diagram of the intervention suggestion generation process provided in an embodiment of the present invention;

[0066] Figure 7 This is a schematic diagram of the modal correlation calculation and weight adjustment process provided in an embodiment of the present invention. Detailed Implementation

[0067] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0068] like Figure 1 As shown in the figure, this invention provides a method for monitoring and assessing psychological state based on multimodal data fusion. The method includes five main steps: multimodal data acquisition, data preprocessing, multimodal feature extraction, feature fusion processing, and psychological state assessment.

[0069] I. Multimodal Data Acquisition

[0070] like Figure 2As shown, the multimodal data acquisition module of the present invention is a core functional component for realizing the synchronous acquisition of multimodal physiological behavior data. It is used to integrate various sensor resources and complete the synchronous acquisition, preliminary processing and output of data, providing a standardized data source for subsequent data preprocessing and feature extraction. The module is mainly composed of a sensor unit, a synchronization control unit and a data transmission unit. Each unit works together to ensure the integrity, timeliness and consistency of data acquisition.

[0071] The sensor unit includes five categories based on the type of data collected: wearable physiological sensors, image acquisition sensors, audio acquisition sensors, motion capture sensors, and gait acquisition sensors. Wearable physiological sensors are used to collect physiological data such as heart rate variability, skin conductance, and electroencephalogram (EEG) signals.

[0072] Image acquisition sensors (such as high-definition cameras) are used to capture a continuous sequence of image frames containing the facial region of a target object;

[0073] Audio acquisition sensors (such as high-fidelity microphones) are used to record the voice audio signals of the target object;

[0074] Motion capture sensors (such as depth cameras and inertial sensors) are used to acquire limb movement trajectory data;

[0075] Gait acquisition sensors (such as pressure sensors, inertial measurement units, and depth cameras) are used to collect gait parameter data such as stride frequency, stride length, and plantar pressure distribution.

[0076] The synchronization control unit has a built-in unified timestamp generation module. Through high-precision clock synchronization technology, it assigns a unique timestamp to each piece of data collected by each sensor, achieving precise alignment of physiological signals, facial images, voice, behavioral actions, and gait data in the time dimension. This ensures the temporal consistency of multimodal data and avoids data misalignment caused by acquisition delays.

[0077] The data transmission unit uses wired or wireless communication methods (such as USB, Bluetooth, Wi-Fi) to transmit synchronized multimodal physiological behavior data to the data storage or preprocessing module in real time. It also has a data caching function, which can temporarily store data when the network is interrupted or the transmission is delayed, and complete the retransmission after the communication is restored, ensuring the continuity of data collection.

[0078] This multimodal data acquisition module supports sensor type expansion and parameter configuration, and can be adapted to different types and levels of sensor devices with different precision according to actual application scenarios (such as laboratory environments and daily home scenarios), taking into account both the flexibility and practicality of data acquisition.

[0079] The multimodal data acquisition module of this invention is used to synchronously acquire multimodal physiological and behavioral data of a target object, specifically including the following five types of data:

[0080] Physiological signal data: Physiological signals of the target object are collected through wearable sensors, including heart rate variability data, skin conductance data and electroencephalogram (EEG) signal data. Heart rate variability data reflects the activity state of the autonomic nervous system and is closely related to mood and stress levels.

[0081] Electrodermal conductivity data reflects the degree of arousal of the sympathetic nervous system and is an important indicator for assessing the level of emotional activation.

[0082] Electroencephalogram (EEG) data directly reflects the brain's electrical activity and contains a wealth of information about cognitive and emotional states.

[0083] Facial image data: Facial images of the target object are captured by a camera, and a continuous sequence of image frames containing the facial region is obtained. Facial expressions are an important channel for emotional expression, and an individual's emotional state can be inferred by analyzing the movement patterns of facial muscles.

[0084] Voice data: The voice audio signal of the target object is collected through a microphone. The voice contains rich emotional information, including features such as tone, speech rate, and volume, all of which are related to emotional state.

[0085] Behavioral motion data: The limb movement trajectory data of the target object is collected through depth cameras or inertial sensors. Behavioral motion patterns can reflect the individual's psychological state. For example, in an anxious state, there may be frequent small movements, while in a depressed state, there may be slow movements.

[0086] Gait data: Gait motion parameter data of the target object during walking is collected through pressure sensors, inertial measurement units or depth cameras, including cadence data, stride length data, plantar pressure distribution data and body center of gravity movement trajectory data. Gait patterns are closely related to psychological state. For example, in an anxious state, there may be characteristics of increased cadence and shortened stride length, while in a depressed state, there may be characteristics of slowed walking speed and increased gait variability. By analyzing gait characteristics, objective physiological behavioral indicators of psychological state can be obtained.

[0087] This invention employs a synchronous acquisition mechanism, which uses a unified timestamp to align the data of each modality in time, ensuring the consistency of multimodal data in the time dimension and laying the foundation for subsequent correlation and fusion processing.

[0088] II. Data Preprocessing

[0089] After acquiring multimodal physiological and behavioral data, each modality needs to be preprocessed to eliminate noise interference and format differences, resulting in standardized data. Specific preprocessing operations include:

[0090] For physiological signal data, bandpass filtering is performed to retain the signal components of the target frequency band, while high-frequency noise and low-frequency drift are filtered out. Baseline drift correction algorithm is used to remove baseline changes caused by electrode movement or skin impedance changes. Power frequency interference and motion artifacts in the signal are removed by independent component analysis.

[0091] For facial image data, the face detection algorithm is first used to locate the facial region in the image. Then, the facial key point detection algorithm is used to obtain the position of facial feature points. Geometric transformation is then performed based on the feature point position to align and normalize the facial region to a preset size, eliminating the influence of head posture changes.

[0092] For speech data, endpoint detection algorithms are used to identify the start and end positions of the speech signal, remove silent segments from the signal, and use spectral subtraction or Wiener filtering methods for noise reduction to reduce the impact of environmental noise. The speech signal is then subjected to spectral normalization to eliminate the influence of differences in recording equipment and environment.

[0093] For action data, coordinate data from different sensors are unified into the same coordinate system, missing data points are filled by interpolation, and smoothing filters are used to eliminate jitter noise in the trajectory.

[0094] For gait data, gait cycle segmentation is first performed, dividing the continuous gait sequence into several complete gait cycles. Then, outliers in the gait data are detected and removed to eliminate abnormal data caused by factors such as sensor noise or sudden pauses of the target object. Next, time alignment and normalization are performed on each gait cycle to unify gait cycles of different durations to the same time scale, which facilitates subsequent feature extraction and fusion processing.

[0095] III. Multimodal Feature Extraction

[0096] like Figure 3 As shown, dynamic feature extraction is performed on the preprocessed standardized data to extract feature vectors that can represent psychological states from each modality of data.

[0097] For physiological signal data, time-domain statistical features and frequency-domain power spectrum features are extracted. Time-domain statistical features include basic statistical quantities such as the mean, variance, and peak-to-peak value of the signal, reflecting the overall characteristics of the signal. Frequency-domain power spectrum features are obtained by performing Fourier transform on the signal and calculating the energy proportion of each frequency band. Taking heart rate variability signals as an example, low-frequency components are related to sympathetic nerve activity, while high-frequency components are related to parasympathetic nerve activity. The ratio of low-frequency to high-frequency energy can reflect the balance state of the autonomic nervous system.

[0098] For facial image data, activation intensity features of facial action units and probability distribution features of facial expression categories are extracted. Facial action units are the basic units that describe facial muscle movements. By analyzing the activation patterns of facial action units, different emotional expressions can be identified. The probability distribution features of facial expression categories are obtained through a pre-trained expression recognition model. The expression recognition model is a convolutional neural network (CNN) model trained on a large-scale facial expression image dataset (including basic emotion category labels such as happy, sad, anxious, and calm). It can automatically identify the emotion category corresponding to the facial expression and output the category probability distribution, thereby representing the probability that the facial image belongs to each basic emotion category.

[0099] For speech data, Mel frequency cepstral coefficient features, fundamental frequency variation features, and speech rate features are extracted. Mel frequency cepstral coefficients are widely used acoustic features in speech processing, which can characterize the spectral envelope of speech. Fundamental frequency variation features reflect the fluctuation pattern of pitch and are closely related to emotional expression. Speech rate features reflect the speed of speaking, and speech rate often varies under different emotional states.

[0100] For behavioral motion data, extract joint angle features, motion amplitude features, and motion rhythm features;

[0101] Gait spatiotemporal parameters, gait symmetry features, and gait stability features are extracted from gait data. Gait spatiotemporal parameters include stride frequency, stride length, stride speed, and the proportion of the stance phase. Gait symmetry features include the difference in stride length between the left and right sides and the difference in stance time between the left and right sides. Gait stability features include the coefficient of variation of gait and the amplitude of the center of gravity swing. Joint angle features describe the degree of bending of each joint in the body and reflect body posture. Movement amplitude features measure the spatial range of limb movement. Movement rhythm features characterize the periodicity and regularity of movement.

[0102] After vectorization, the features of each modality are obtained as fixed-dimensional feature vectors, which are denoted as physiological signal feature vector, facial image feature vector, speech feature vector, and behavioral action feature vector, respectively.

[0103] For gait data, spatiotemporal gait features, gait symmetry features, and gait stability features are extracted. Spatiotemporal gait features are quantitative indicators describing the basic attributes of gait, including cadence (number of steps per unit time), stride length (distance between two adjacent steps), gait speed (distance traveled per unit time), and support duration ratio (the proportion of single-leg support time to the gait cycle). These parameters can reflect an individual's movement state and energy consumption characteristics and are correlated with psychological states such as anxiety and depression. Gait symmetry features include the difference in stride length between the left and right sides and the difference in support duration between the left and right sides, reflecting the degree of coordination of bilateral limb movements. Abnormal psychological states may lead to a decrease in gait symmetry. Gait stability features include the coefficient of variation of gait (the ratio of the standard deviation to the mean of gait parameters) and the amplitude of center of gravity swing, reflecting the regularity and stability of gait. Emotional fluctuations and psychological stress may lead to a decrease in gait stability.

[0104] After vectorization, the features of each modality are obtained as fixed-dimensional feature vectors, which are denoted as physiological signal feature vector, facial image feature vector, speech feature vector, behavioral action feature vector, and gait feature vector, respectively.

[0105] IV. Feature Fusion Processing

[0106] like Figure 4 , Figure 7 As shown, this invention employs an adaptive correlation fusion strategy based on modal correlation adjustment factors and attention mechanisms to fuse the feature vectors of each modality.

[0107] First, the data quality assessment score for each modality is calculated. For each modality, a signal quality factor is calculated based on its signal-to-noise ratio (SNR). A higher SNR indicates better signal quality and a higher signal quality factor. Simultaneously, a data integrity factor is calculated based on the missing data rate. A lower missing data rate indicates more complete data and a higher integrity factor. The signal quality factor and the data integrity factor are then weighted and summed to obtain the data quality assessment score for that modality. In this embodiment, the weighting coefficient of the signal quality factor is 0.6, and the weighting coefficient of the data integrity factor is 0.4.

[0108] The formula for calculating the data quality assessment score is as follows:

[0109] ;

[0110] in, Indicates the first Data quality assessment scores for each modality Indicates the first Signal quality factor of each mode Indicates the first Data integrity factor and signal quality factor for each modality The data integrity factor is obtained based on signal-to-noise ratio normalization and has a value range of 0 to 1. Calculated based on the data missing rate, defined as ,in For the first Data missing rate for each modality.

[0111] Secondly, the feature correlation between gait mode and other modes is calculated, using cosine similarity as the correlation metric. The calculation formula is as follows:

[0112] ;

[0113] in, Represents the gait mode feature vector. Indicates the first The feature vector of a modality Indicates gait mode and the first The feature correlation degree between the modalities ranges from -1 to 1, with a larger value indicating a higher degree of correlation.

[0114] Then, the modal correlation adjustment factor is calculated based on the feature correlation, and the fusion weights of each modality are dynamically adjusted, according to the following rules:

[0115] (1) Calculate the gait-physiological correlation If the gait-physiological correlation is greater than or equal to 0.7 (for example, in anxious scenarios, gait speed increase often occurs simultaneously with increased heart rate and skin conductance, showing a high correlation), then the base weight of the gait modality is increased by 20%, and the base weight of the physiological signal modality is increased by 10%, in order to enhance the contribution of these two highly correlated modalities to the fusion results.

[0116] (2) Calculate gait-face correlation If the gait-face correlation is less than 0.3 (for example, in a face occlusion scenario, facial image data may be missing or of poor quality, reducing the correlation with gait data), then the base weight of the gait modality will be reduced by 10%, and the base weight of the speech modality will be increased by 15% to compensate for the impact of missing facial modality information and ensure the reliability of the evaluation results.

[0117] The aforementioned dynamic weight adjustment mechanism based on modal correlation can adapt to changes in modal correlation under different psychological states. Compared with the existing technology that uses fixed weights to fuse gait with other modalities, it is more in line with actual evaluation scenarios and improves the flexibility and accuracy of the fusion strategy.

[0118] Then, based on the data quality assessment score, the preset modality base weight coefficient, and the modality correlation adjustment factor, the adaptive fusion weight of each modality feature vector is calculated. The corrected calculation formula is as follows:

[0119] ;

[0120] in, Indicates the first Adaptive fusion weights for various modalities Indicates the first The basic weighting coefficients of each modality Indicates the first Data quality assessment scores for each modality This represents the weight adjustment factor after adjustment according to the modal correlation rule. Indicates the total number of modes (in this embodiment) =5), using the above formula, the fusion weights are normalized to ensure that the sum of the weights of all modes is 1.

[0121] Next, based on the multi-head attention mechanism, the feature vectors of each modality are weighted and aggregated. The feature vectors of each modality are transformed into query vector, key vector and value vector respectively through linear transformation. Then, the attention score matrix is ​​calculated to reflect the degree of correlation between different modal features. The adaptive fusion weight is combined with the attention score, and the value vectors of each modality are weighted and summed to obtain the fused feature vector.

[0122] Multi-head attention mechanisms can learn the relationships between modalities from multiple subspaces and capture interaction information at different levels. By introducing adaptive fusion weights into the attention calculation process, dynamic fusion with data quality awareness is achieved. When the data quality of a certain modality is low, its contribution to the fusion result is automatically reduced, thereby improving the robustness of the fusion.

[0123] V. Psychological Status Assessment

[0124] like Figure 5 As shown, a psychological state assessment model pre-trained by fusing feature vector input is used for psychological state identification and assessment. The psychological state assessment model is a deep neural network model trained and optimized based on a multimodal fusion feature dataset labeled with normal state, anxiety state, depression state, and stress state. Its structure includes three parts connected in sequence: feature encoding layer, temporal modeling layer, and classification output layer. During the pre-training process, the difference between the predicted probability distribution and the true label is minimized by the cross-entropy loss function, and the model parameters are iteratively updated to improve the recognition accuracy.

[0125] The feature encoding layer adopts a multi-layer fully connected neural network structure, which includes an input layer, a hidden layer and an output layer. The input layer receives the fused feature vector, the hidden layer performs non-linear transformation on the features through a non-linear activation function, and the output layer reduces the dimensionality of the features to a pre-defined dimension (e.g., 128 dimensions) encoded feature vector. The role of the feature encoding layer is to further extract high-level semantic information related to psychological state from the fused features, while reducing the feature dimension to reduce the amount of subsequent computation.

[0126] The temporal modeling layer adopts a long short-term memory network structure to model the temporal dependencies of encoded feature vectors within a continuous time window. The long short-term memory network controls the flow of information through a gating mechanism, which can effectively capture dependencies over long time spans. It is suitable for processing data with temporal dynamic characteristics, such as psychological states. The temporal modeling layer receives encoded feature vectors from multiple consecutive time points as input sequences and outputs temporal feature vectors that integrate temporal context information.

[0127] The classification output layer employs a fully connected neural network structure, mapping temporal feature vectors to predicted scores for mental state categories. Then, a normalized exponential function (Softmax function) is used to convert the predicted scores into a probability distribution. The formula for calculating the normalized exponential function is as follows:

[0128] ;

[0129] in, Indicates the prediction is the first The probability of a mental state. Indicates the classification output layer for the first Predicted scores for the class This represents the total number of psychological state categories.

[0130] In this embodiment, the psychological state categories include four types: normal state, anxious state, depressed state, and stressed state. The psychological state assessment results include psychological state category labels and corresponding confidence scores. The psychological state category label is the category with the highest probability, and the confidence score is the probability value corresponding to that category.

[0131] The psychological state assessment model is pre-trained using labeled training data. During training, the cross-entropy loss function is used to measure the difference between the predicted probability distribution and the true label. The model parameters are updated through the backpropagation algorithm, enabling the model to accurately identify different psychological states.

[0132] VI. Trend Analysis

[0133] To provide more comprehensive psychological state assessment information, this invention also supports a trend analysis function for psychological states, which performs sliding statistics on the psychological state assessment results of multiple consecutive time periods with a preset time window length, and calculates the frequency of occurrence and average confidence score of each psychological state category within the time window.

[0134] Based on the frequency of occurrence and the average confidence score, a psychological state change trend curve is generated to intuitively show the change pattern of the target object's psychological state over time. When a continuous abnormality or obvious deterioration trend in the psychological state is detected, the intelligent psychological state assessment system (hereinafter referred to as the "assessment system") constructed based on the method of this invention can generate early warning information and output a psychological state trend analysis report. The assessment system integrates a multimodal data acquisition module, a data preprocessing module, a feature extraction module, a feature fusion module, a psychological state assessment module, a trend analysis module, and an intervention suggestion generation module, which can realize the fully automated processing from data acquisition to assessment report output, providing a reference for mental health intervention.

[0135] VII. Generation of Intervention Recommendations

[0136] like Figure 6 As shown, based on the psychological state assessment results and trend analysis report, this invention further provides an intelligent intervention suggestion generation function, realizing a complete closed loop from assessment to intervention.

[0137] First, the intervention type is determined based on the psychological state category label. In this embodiment, the correspondence between the four intervention types and the psychological state categories is set as follows:

[0138] Psychological state categories Corresponding intervention type Intervention Target Anxiety Relaxation training intervention Lower physiological arousal levels and alleviate anxiety symptoms Depressive state Behavioral activation intervention Increase participation in positive activities to improve emotional state stress state Cognitive adjustment intervention Adjusting cognitive patterns and improving stress coping skills Continuous abnormal state Professional referral intervention It is recommended to seek professional psychological counseling or medical help.

[0139] Then, the intervention intensity level is determined based on the confidence score and the trend of psychological state changes. The rules for determining the intervention intensity level are as follows:

[0140] Mild intervention: When the confidence score of the psychological state assessment result is below 0.6, or when the trend of psychological state changes shows that the state is improving, mild intervention is adopted. It mainly provides self-help intervention suggestions, such as recommending relevant relaxation audio, popular science knowledge on mental health, etc.

[0141] Moderate intervention: When the confidence score of the psychological state assessment results is between 0.6 and 0.8, and the trend of psychological state changes shows that the state is continuously stable or slightly deteriorating, a moderate intervention is adopted, providing a structured intervention program, such as a daily relaxation practice plan, cognitive restructuring practice tasks, etc., and setting up intervention effect tracking reminders.

[0142] Severe intervention: When the confidence score of the psychological state assessment result is higher than 0.8, or when the trend of psychological state changes shows that the state continues to deteriorate beyond the preset threshold, severe intervention is adopted. In addition to providing immediate crisis intervention guidance, it also generates professional referral suggestions, prompting users to seek help from professional psychological counselors or psychiatrists in a timely manner, and can send early warning notifications to emergency contacts with user authorization.

[0143] The formula for calculating the intervention intensity level is as follows:

[0144] ;

[0145] in, Indicates the intervention intensity score. This represents the confidence score. A factor representing the degree of trend deterioration (with a value ranging from 0 to 1, and a larger value for a more severe trend). and These are the weighting coefficients, as shown in this embodiment. =0.5, =0.5, when A value <0.4 indicates mild intervention; a value ≤0.4 indicates moderate intervention. A value <0.7 indicates moderate intervention; when A value ≥0.7 indicates severe intervention.

[0146] Finally, based on the intervention type and intensity level, specific intervention measures are matched from a pre-set intervention measure library to generate intervention suggestion plans. The intervention measure library stores standardized intervention measure templates for different intervention types and intensity levels, including information such as intervention measure name, intervention content description, suggested implementation frequency, and expected effects. The generated intervention suggestion plans can be pushed to the target subject through terminal devices to guide them in self-psychological adjustment.

[0147] The above description is merely a specific embodiment of the present invention, and the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An intelligent psychological state assessment method based on multimodal data fusion, characterized in that, Includes the following steps: Acquire multimodal physiological and behavioral data of the target object, including physiological signal data, facial image data, voice data, behavioral action data, and gait data; The multimodal physiological behavior data is preprocessed to obtain standardized data; Dynamic feature extraction is performed on the standardized data to obtain feature vectors corresponding to each modality; Calculate the feature correlation degree between the gait mode and other modes, and perform correlation and fusion processing on the feature vectors corresponding to each mode to obtain the fused feature vector; The fused feature vector is input into the pre-trained psychological state assessment model to obtain the psychological state assessment result.

2. The intelligent psychological state assessment method based on multimodal data fusion according to claim 1, characterized in that: The physiological signal data includes heart rate variability data, skin conductance data, and electroencephalogram (EEG) signal data; the facial image data is a sequence of continuous image frames containing the facial region of the target object; the speech data is the speech audio signal of the target object; the behavioral action data is the limb movement trajectory data of the target object; and the gait data is the gait motion parameter data of the target object during walking, including cadence data, stride length data, plantar pressure distribution data, and body center of gravity movement trajectory data.

3. The intelligent psychological state assessment method based on multimodal data fusion according to claim 1, characterized in that: The multimodal physiological behavior data is preprocessed to obtain standardized data, including: The physiological signal data are subjected to bandpass filtering and baseline drift correction to remove power frequency interference and motion artifacts from the signal; The facial image data is processed for face detection, key point localization and geometric alignment, and the facial region is normalized to a preset size; The speech data is subjected to endpoint detection, silence removal, and spectrum normalization. The behavioral action data is subjected to coordinate system 1 and trajectory interpolation smoothing processing; The gait data is processed by gait cycle segmentation, abnormal gait point removal and temporal alignment to normalize the gait sequence to a uniform time scale.

4. The intelligent psychological state assessment method based on multimodal data fusion according to claim 1, characterized in that: Dynamic feature extraction is performed on the standardized data to obtain feature vectors corresponding to each modality, including: Time-domain statistical features and frequency-domain power spectrum features are extracted from the physiological signal data. The time-domain statistical features include mean, variance, and peak-to-peak value. The frequency-domain power spectrum features include the energy proportion of each frequency band. Extract facial motion unit activation intensity features and facial expression category probability distribution features from the facial image data; Extract Mel frequency cepstral coefficient features, fundamental frequency variation features, and speech rate features from the speech data; Extract joint angle features, amplitude features, and rhythmic features from the behavioral data; Gait spatiotemporal parameter features, gait symmetry features, and gait stability features are extracted from the gait data. The gait spatiotemporal parameter features include stride frequency, stride length, stride speed, and the proportion of support phase duration. The gait symmetry features include the difference between left and right stride lengths and the difference between left and right support phase times. The gait stability features include the gait variation coefficient and the amplitude of center of gravity swing.

5. The intelligent psychological state assessment method based on multimodal data fusion according to claim 1, characterized in that: Calculate the feature correlation degree between gait mode and other modes, and perform correlation and fusion processing on the feature vectors corresponding to each mode to obtain a fused feature vector, including: Calculate the data quality assessment score for each modality separately; Calculate the feature correlation degree between the gait modality feature vector and the feature vectors of other modalities, and the feature correlation degree is obtained by cosine similarity calculation; The modal correlation adjustment factor is calculated based on the feature correlation, and the fusion weight of each modality is dynamically adjusted based on the modal correlation adjustment factor. Based on the data quality assessment score, the preset modality base weight coefficient, and the modality correlation adjustment factor, the adaptive fusion weight of each modality feature vector is calculated; Based on the multi-head attention mechanism, the adaptive fusion weights are used to weight and aggregate the feature vectors of each modality to obtain the fused feature vector.

6. The intelligent psychological state assessment method based on multimodal data fusion according to claim 5, characterized in that: The calculation of data quality assessment scores for each modality includes: Calculate the signal quality factor based on the signal-to-noise ratio of each modality data; Calculate the data completeness factor based on the missing rate of each modality; The signal quality factor and the data integrity factor are weighted and summed to obtain the data quality evaluation score for each modality, wherein the weight coefficient of the signal quality factor is 0.6 and the weight coefficient of the data integrity factor is 0.

4.

7. The intelligent psychological state assessment method based on multimodal data fusion according to claim 1, characterized in that: The psychological state assessment model includes a feature encoding layer, a temporal modeling layer, and a classification output layer connected in sequence. The feature encoding layer adopts a fully connected neural network structure to perform nonlinear transformation and dimensionality reduction on the fused feature vector and output the encoded feature vector. The temporal modeling layer adopts a long short-term memory network structure to model the temporal dependency relationship of the encoded feature vectors within a continuous time window and output temporal feature vectors. The classification output layer employs a fully connected neural network structure and a normalized exponential function to output the probability distribution of mental state categories based on the temporal feature vector.

8. The intelligent psychological state assessment method based on multimodal data fusion according to claim 1, characterized in that: The psychological state assessment results include psychological state category labels and corresponding confidence scores; the psychological state category labels are one of the following: normal state, anxious state, depressed state, and stressed state; the confidence score is the maximum category probability value output by the psychological state assessment model.

9. The intelligent psychological state assessment method based on multimodal data fusion according to claim 1, characterized in that: The method further includes: The psychological state assessment results of multiple consecutive time periods are statistically analyzed using a preset time window length to calculate the frequency of occurrence and average confidence score of each psychological state category. Based on the frequency of occurrence and average confidence score, a psychological state change trend curve is generated, and a psychological state trend analysis report is output.

10. The intelligent psychological state assessment method based on multimodal data fusion according to claim 9, characterized in that: The method further includes: Based on the psychological state assessment results and the psychological state trend analysis report, an intervention recommendation plan is generated for the target group; The proposed intervention plan includes the type of intervention, the level of intervention intensity, and specific intervention measures; The intervention types are determined based on psychological state category labels, including relaxation training intervention, cognitive regulation intervention, behavioral activation intervention, and professional referral intervention; The intervention intensity level is determined based on the confidence score and the trend of psychological state change, and is divided into three levels: mild intervention, moderate intervention, and severe intervention.