Method for constructing depression risk prediction model based on multi-modal sample information

By constructing a depression risk prediction model based on multimodal sample information, using nonlinear dynamics, manifold modeling and multi-scale signal decomposition technology to extract features, and fusion of features through self-attention mechanisms, the problem of insufficient prediction accuracy of depression risk prediction in the existing technology is solved, and higher prediction accuracy and adaptability are achieved.

CN119943384APending Publication Date: 2025-05-06JUSTIFIED (SHANGHAI) NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 12 Cited by

Patent Information

Application Number
CN202510012694.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing depression risk prediction techniques are less accurate and robust, especially in real-time and dynamic risk prediction, and the multimodal data fusion effect is not good.

Method used

The depression risk prediction model construction method based on multimodal sample information is adopted. By obtaining multimodal data of physiological signals, speech features and micro-expression actions, the features are extracted using nonlinear dynamics theory, manifold modeling and multi-scale signal decomposition technology, and feature fusion is performed through self-attention mechanism, and dynamic modeling and risk prediction are finally performed based on the state space model.

Benefits of technology

It improves the accuracy and reliability of the prediction of depression risk, enhances the adaptability and robustness of the model to multimodal data, and can more accurately reflect the dynamic changes in depression risk.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943384A_ABST
    Figure CN119943384A_ABST
Patent Text Reader

Abstract

The invention relates to the field of mental health assessment, and discloses a depression risk prediction model construction method based on multi-modal sample information, and the method comprises the following steps: collecting and preprocessing physiological signals, voice features and micro-expression action data; performing phase-space reconstruction on the physiological signal; constructing a manifold model, and extracting local geometric features of the multi-modal signals; the short-term dynamic change and the long-term trend of the signals are separated through a multi-scale variational mode decomposition method; performing dynamic weighted fusion on the multi-modal features by adopting a self-attention mechanism to generate a unified high-quality feature vector; and constructing a nonlinear state space model based on the fusion features, optimizing a state transition and output relationship through extended Kalman filtering, and generating a real-time emotional state and depression risk score. The method can accurately capture the dynamic characteristics of the multi-modal signal, significantly improves the accuracy and robustness of depression risk prediction, and provides efficient technical support for mental health assessment and intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of mental health assessment, and in particular to a method for constructing a depression risk prediction model based on multimodal sample information. Background Art

[0002] In recent years, with the increasing prominence of mental health problems, depression has become an important public health issue that needs to be addressed worldwide. Early identification and intervention of depression are the key to reducing its incidence and preventing serious consequences. However, due to the complex and hidden manifestations of depression, traditional diagnostic methods mainly rely on psychological questionnaires and clinical interviews, which are highly dependent on subjective factors and are easily affected by the patient's subjective will or cognitive bias, resulting in inaccurate diagnostic results.

[0003] In the existing technology, some methods try to assess the risk of depression through single-modal data (such as physiological signals, voice features or facial expressions). However, such methods have the following main shortcomings: first, the feature dimension of single-modal data is low, it is difficult to fully reflect the multidimensional physiological and psychological characteristics of depression, and the feature information is incomplete; second, there are strong nonlinearities and time series in the signal, and it is difficult to capture these complex dynamic characteristics using only a single analysis method; in addition, the existing methods lack an effective mechanism for the fusion of multimodal data, usually adopting simple splicing or static weight allocation, and cannot dynamically adjust the weights of different modal features, resulting in the prediction model being difficult to adapt to the complexity of data distribution.

[0004] Understandably, these defects lead to low accuracy and robustness of existing depression risk prediction technologies, especially in terms of real-time and dynamic risk prediction. Therefore, there is an urgent need for a technical method that can fully integrate multimodal data characteristics and dynamically model signal change rules to improve the comprehensiveness, accuracy and adaptability of depression risk prediction. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention provides a method for constructing a depression risk prediction model based on multimodal sample information, which solves the problems in the prior art of insufficient accuracy in depression risk prediction, incomplete single-modal feature information, and poor multimodal data fusion effect.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method for constructing a depression risk prediction model based on multimodal sample information, comprising the following steps:

[0007] Acquire multimodal data of physiological signals, voice features, and micro-expression movements, and pre-process the multimodal data;

[0008] Use nonlinear dynamics theory to reconstruct the phase space of physiological signals and extract nonlinear features;

[0009] Based on differential geometry theory, multimodal signal distribution is modeled and local geometric features of the signal are extracted;

[0010] Extract modal features at different time scales through multi-scale signal decomposition technology;

[0011] Construct a multimodal feature fusion model and use the self-attention mechanism to perform weighted fusion of multimodal data;

[0012] Based on the fusion features, the depression risk prediction model is trained, the state space transfer relationship is established, and the depression risk prediction results are output.

[0013] Preferably, the collection and preprocessing of the multimodal data includes:

[0014] Use non-contact photoplethysmography to obtain physiological signals of heart rate variability, blood pressure fluctuations, and blood oxygen saturation;

[0015] Extract speech features, including pitch, speech rate, and speech energy;

[0016] Capture micro-expression movement features, including dynamic changes of the frontalis and masseter muscles;

[0017] Dynamic time warping is used to align the time series of signals of different modalities.

[0018] Preferably, the phase space reconstruction is based on Takens embedding theorem, and the physiological signal is constructed into a high-dimensional dynamic feature space through time delay and setting of embedding dimension, so as to reveal the nonlinear dynamic law of the signal.

[0019] Preferably, the nonlinear characteristics include:

[0020] Lyapunov exponent of the system's sensitivity to initial conditions;

[0021] Fractal dimension for characterizing signal complexity;

[0022] Dynamic evolution characteristics of system attractors.

[0023] Preferably, the manifold modeling includes:

[0024] Construct a nearest neighbor graph based on multimodal signals and calculate the weighted distance matrix between samples;

[0025] The Laplace matrix of the manifold is generated, and the spectral characteristics of the signal are extracted based on the Laplace-Beltrami operator to characterize the local geometric characteristics of the signal on the manifold.

[0026] Preferably, the multiscale signal decomposition is based on a multiscale variational mode decomposition method, which decomposes the signal into multiple intrinsic mode components and extracts features on different time scales, wherein:

[0027] High-frequency components reflect short-term signal changes;

[0028] The low-frequency component reflects the long-term trend characteristics.

[0029] Preferably, the multimodal feature fusion is achieved through a self-attention mechanism, by constructing query, key and value matrices, dynamically weighting the multimodal features, and generating a fused feature representation based on the weighted feature importance.

[0030] Preferably, the depression risk prediction model is implemented by constructing a nonlinear state space relationship, and the state space model includes:

[0031] State transition equations that describe the dynamic changes of the system;

[0032] An output equation that maps feature inputs to a depression risk score.

[0033] Preferably, the state space model is optimized by an extended Kalman filter, wherein:

[0034] Initialize state transition parameters using historical data;

[0035] The state variables are adjusted dynamically using the observed data to optimize the prediction model.

[0036] Preferably, the output of the depression risk prediction model includes:

[0037] Real-time emotional state results, using a two-dimensional emotional model to represent arousal and valence;

[0038] Depression risk score, a quantitative assessment of a patient's depression risk based on a trained nonlinear model.

[0039] The present invention provides a method for constructing a depression risk prediction model based on multimodal sample information.

[0040] It has the following beneficial effects:

[0041] 1. The present invention uses nonlinear dynamics analysis, manifold modeling and multiscale signal decomposition technology to deeply mine the dynamic characteristics and multidimensional distribution characteristics of multimodal data. By extracting the nonlinear characteristics of physiological signals, the local geometric characteristics of signals, and the short-term and long-term time scale characteristics, the model can more comprehensively capture the relevant information of depression risk, thereby improving the accuracy and reliability of prediction.

[0042] 2. The present invention uses a self-attention mechanism to dynamically weighted fuse multimodal data, making full use of the complementarity of different modal features while suppressing the influence of redundant features and noise. Combined with the dynamic characteristics of state-space modeling, the model can achieve real-time updates of risk status in time series, thereby more accurately reflecting the dynamic changes of depression risk.

[0043] 3. This invention solves the asynchrony and noise interference problems of multimodal data in the time dimension through technologies such as dynamic time warping and extended Kalman filtering. Whether it is high-dimensional, nonlinearly distributed data or physiological and behavioral signals with random fluctuations, the model can adapt to the data characteristics in different scenarios and ensure the prediction stability and robustness in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a schematic diagram of the construction method of the present invention;

[0045] Figure 2 It is a schematic diagram of the structure of the depression risk prediction model of the present invention. DETAILED DESCRIPTION

[0046] The following will be combined with the drawings of the present invention specification to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0047] Please see attached Figure 1 The present invention provides a method for constructing a depression risk prediction model based on multimodal sample information. Through the collection and preprocessing of multimodal data, nonlinear dynamic analysis, manifold modeling, multi-scale signal decomposition, feature fusion and dynamic modeling, a model that can accurately predict the risk of depression is constructed.

[0048] like Figure 1 As shown, the method for constructing a depression risk prediction model based on multimodal sample information may include the following steps:

[0049] S1, obtain multimodal data and preprocess it;

[0050] S2, phase space reconstruction using nonlinear dynamics theory;

[0051] S3, manifold modeling based on differential geometry theory;

[0052] S4, extracting time scale features through multi-scale signal decomposition;

[0053] S5. Build a multimodal feature fusion model;

[0054] S6. Train the prediction model and output the results.

[0055] Each step of the method of the present invention will be described in detail below.

[0056] For step S1, in this embodiment, the purpose of this step is to provide standardized multimodal data input for subsequent steps, including three types of core data: physiological signals, voice features, and micro-expression movements, and ensure data quality and consistency through preprocessing.

[0057] As an option, the collection of physiological signals uses non-contact photoplethysmography (PPG / IPPG) technology. Specifically, this technology obtains key physiological characteristic signals of the human body by detecting changes in the absorption and reflection of light in blood and tissues. These signals include but are not limited to the following:

[0058] Heart rate variability (HRV) is obtained by analyzing the changes in the RR intervals in the ECG signal and is used to reflect the regulatory ability of the autonomic nervous system.

[0059] Blood pressure fluctuations are estimated through pulse wave signals, and arterial blood pressure changes are obtained by combining pulse wave transmission time with model fitting.

[0060] Blood oxygen saturation is calculated by the changes in the absorption degree of light of different wavelengths in the blood.

[0061] In a possible implementation, the sampling frequency of the physiological signal is set between 50 Hz and 100 Hz. The specific selection is related to the performance of the device to ensure high-resolution capture of rapid physiological changes.

[0062] Exemplarily, the collection of speech features is completed by recording speech in natural language conversations. It should be noted that the collected speech signal is recorded at a sampling frequency of more than 16kHz to retain high-frequency detail information. Speech feature extraction is based on short-time Fourier transform (STFT), which performs time-frequency domain analysis on the signal to extract emotion-related features, including but not limited to:

[0063] Pitch, obtained by fundamental frequency detection method;

[0064] Speech rate (Speed), calculated by the time duration characteristics of the speech signal;

[0065] Speech energy (Energy) reflects the strength of the sound through short-time energy calculation.

[0066] In some embodiments, the collection of micro-expression movements is completed by capturing facial video sequences, and the video resolution can be set to 720p or higher to ensure clear recording of facial muscle details. Specifically, the focus of micro-expression detection is the dynamic changes of local facial muscles, especially the frontalis, masseter, and periocular muscles. As an implementation method, the displacement information of facial key points can be extracted by optical flow method to characterize the characteristics of micro-expression movements.

[0067] It should be noted that in order to ensure the uniformity and consistency of multimodal data, the following data preprocessing steps are adopted in this embodiment:

[0068] As an option, the dynamic time warping (DTW) algorithm is used to align the time scales of multimodal signals. The basic principle of DTW is to find the optimal matching path between time series through dynamic programming to minimize the distance between the two series. Specifically:

[0069]

[0070] Where X = {x1, x2, …, x m} and Y={y1,y2,…,y n} is a two-modal time series; π is a matching path; d(x, y) represents the distance between two points (e.g., Euclidean distance).

[0071] It should be understood that the purpose of time alignment is to align the data points of different modes with the time axis to ensure that the signals at the same time point have a corresponding relationship. This is especially important for subsequent multimodal fusion.

[0072] Specifically, in order to reduce the impact of noise on data quality, wavelet transform denoising technology is used for physiological signals. As an implementation method, wavelet transform removes high-frequency noise and retains low-frequency features by decomposing the signal into different frequency components. Exemplarily, the denoising steps based on the soft threshold filter are as follows:

[0073] Decomposing physiological signals into multiple frequency bands;

[0074] Apply the set threshold on each frequency band to filter out the noise components smaller than the threshold;

[0075] The denoised signal is reconstructed by inverse wavelet transform.

[0076] In a possible implementation, the noise reduction process of the micro-expression video signal adopts a median filtering technique to eliminate random noise points that may be introduced by the camera device.

[0077] As an option, in this embodiment, all data features need to be normalized. Normalization methods include minimum-maximum normalization and Z-score normalization. The specific selection can be determined according to the characteristics of the data.

[0078] Exemplarily, for physiological signals, minimum-maximum normalization is used; for speech and micro-expression features, Z-score normalization is used.

[0079] It should be noted that the purpose of data preprocessing is to ensure that multimodal data has a unified format, clear features, and high-quality signals, and all processing steps have provided a sufficient basis for the implementation of subsequent steps.

[0080] It can be understood that the data acquisition and preprocessing process in this embodiment is not only applicable to the construction of a depression risk prediction model, but can also provide effective support for other multimodal data-driven models.

[0081] As for step S2, in this embodiment, the physiological signal is analyzed based on nonlinear dynamics theory, the dynamic characteristics of the physiological signal are revealed by phase space reconstruction technology, and important nonlinear features describing the complex behavior of the signal are extracted therefrom.

[0082] It should be noted that in this embodiment, nonlinear dynamic analysis is applicable to time series data in physiological signals, such as heart rate variability (HRV) and blood pressure fluctuation signals. These signals can be modeled through nonlinear dynamic theory to characterize the complex regulatory mechanism of the patient's physiological state and provide key features for subsequent model construction.

[0083] As an alternative, phase space reconstruction uses Takens embedding theorem, the basic principle of which is to map a one-dimensional time series into a high-dimensional phase space through a time delay embedding method, thereby revealing the intrinsic dynamic structure of the signal. Specifically, assuming that the time series is {x1, x2, …, x N}, the process of phase space reconstruction can be described as:

[0084] X t =[x t ,x t+τ ,x t+2τ ,…,x t+(m-1)τ ]

[0085] in:

[0086] X t represents the reconstructed phase space vector;

[0087] τ is the time delay, which is usually calculated and determined by the autocorrelation function or mutual information method;

[0088] m is the embedding dimension, which is determined by the False Nearest Neighbors (FNN) method.

[0089] In a possible implementation, the selection of the time delay τ can be determined by calculating the first zero crossing position of the signal autocorrelation function to ensure that the selection of the time delay can fully reflect the periodic characteristics of the signal. Exemplarily, for the heart rate variability signal, the time delay is generally between 0.1 and 0.5 seconds.

[0090] It should be understood that the choice of embedding dimension m determines the dimension of the reconstructed phase space. In order to avoid the loss of phase space information due to too low embedding dimension, or the introduction of redundant information due to too high embedding dimension, the FalseNearest Neighbors (FNN) method is used in this embodiment to optimize the dimension. Specifically, the FNN method detects the authenticity of the neighboring relationship between points in the high-dimensional space and gradually increases the embedding dimension until the proportion of false neighbors is lower than the set threshold.

[0091] In some embodiments, in order to effectively characterize the dynamic behavior of the signal in the phase space, the following nonlinear features are further extracted:

[0092] Specifically, the Lyapunov index is used to describe the sensitivity of the system to the initial conditions and is the core indicator for measuring the chaos of the system. The process of calculating the Lyapunov index includes the following steps:

[0093] Select the initial point X0 and its neighboring point X0 ′ , and calculate its initial distance δ0=∥X0 ′ -X0∥;

[0094] Evolving over time, calculate the distance δ between two trajectories t ;

[0095] Calculate the Lyapunov exponent after taking the logarithm and normalizing:

[0096]

[0097] It should be noted that a positive Lyapunov exponent indicates that the system has chaotic characteristics, that is, the trajectory is sensitive to the initial conditions. For the physiological signals of patients with depression, the Lyapunov exponent is usually lower, reflecting the weakening of their system's regulatory ability.

[0098] As another feature, the fractal dimension is used to quantify the complexity and distribution characteristics of the signal. Exemplarily, this embodiment uses the Grassberger-Procaccia algorithm to calculate the fractal dimension, and the process includes:

[0099] Select any pair of points in phase space (Xi ,X j ), calculate its Euclidean distance r = ∥X i -X j ∥;

[0100] Define the fraction of point pairs within a certain radius r:

[0101]

[0102] Where H is the Heaviside function, when r-∥X i -X j ∥ takes the value 1, otherwise it takes the value 0;

[0103] Calculate the fractal dimension D:

[0104]

[0105] It is important to understand that the higher the fractal dimension, the greater the complexity of the signal. For signals from people with depression, the fractal dimension is usually low, indicating that the dynamic behavior of the system tends to be simpler.

[0106] In a possible implementation, in order to further reveal the dynamic pattern of the signal, the present embodiment also calculates Kolmogorov complexity to quantify the unpredictability of the signal. Kolmogorov complexity is based on the minimum description of the signal compression length and reflects the information content and randomness in the signal.

[0107] It is understandable that the high-dimensional dynamic features obtained through phase space reconstruction and nonlinear feature extraction can fully describe the changing rules of the signal in the time dimension. Specifically, the features after phase space reconstruction can not only capture the dynamic patterns of the signal (such as periodicity and nonlinear relationships), but also reveal the chaos, complexity and stability of the signal through the feature extraction process.

[0108] The Lyapunov index provides an important indicator for measuring the dynamic behavior of the system. The larger the value, the more sensitive the system is to the initial conditions, reflecting the divergent characteristics of the trajectory in the phase space. For the heart rate variability (HRV) signal of healthy individuals, its Lyapunov index is usually in a reasonable range, indicating that the dynamic behavior of the system has a certain adaptability. However, the Lyapunov index value of the HRV signal of patients with depression tends to be lower, indicating that the regulatory ability of their cardiovascular system is weakened. Therefore, this index is an important feature for judging risk in the model.

[0109] It is further explained that fractal dimension reveals the diversity of system dynamic behavior by quantifying signal complexity. High fractal dimension usually corresponds to more complex signal characteristics. For example, HRV signals of healthy individuals show more complex dynamic patterns, reflecting the health of their autonomic nervous system. The fractal dimension of HRV signals of patients with depression is usually low, indicating that their dynamic regulation ability is inhibited. Therefore, fractal dimension can provide stable and robust features for the model.

[0110] In one possible implementation, Kolmogorov complexity reflects the randomness and regularity of a signal by evaluating its compressibility. For HRV signals of healthy individuals, their Kolmogorov complexity is in a moderate range, with both certain regularity and appropriate randomness, reflecting the flexible adjustment ability of the system. However, HRV signals of patients with depression often show the characteristics of low Kolmogorov complexity, reflecting the characteristics of increased regularity and reduced randomness of the signal, further confirming the weakening of the system's dynamic adjustment ability.

[0111] It should be noted that these extracted nonlinear features do not exist in isolation, but are interrelated. For example, the changes in Lyapunov exponents and fractal dimensions are usually consistent, and the combination of the two can more comprehensively describe the chaos and complexity of the signal. In some embodiments, through the joint analysis of features, the dynamic characteristics of physiological signals of healthy individuals and patients with depression can be more accurately distinguished.

[0112] In one possible implementation, these extracted high-dimensional dynamic features (including Lyapunov exponents, fractal dimensions, and Kolmogorov complexity) can be combined into feature vectors as input data for subsequent model training and prediction. It should be understood that this high-dimensional feature representation can effectively reduce the redundant information of time series signals and enhance the model's ability to predict the risk of depression.

[0113] It should be noted that the method described in this embodiment is not only applicable to the nonlinear feature extraction of heart rate variability (HRV) signals, but can also be applied to the dynamic analysis of other physiological signals (such as blood pressure fluctuations and blood oxygen saturation). As an option, for different types of physiological signals, the calculation methods of time delay and embedding dimension can be flexibly adjusted to adapt to the characteristic differences of the signals.

[0114] For step S3, in this embodiment, the high-dimensional distribution of the multimodal signal is modeled by manifold modeling through differential geometry theory to extract the local geometric features of the signal and provide accurate input data for the subsequent depression risk prediction model.

[0115] It should be noted that the basic idea of ​​manifold modeling is to map high-dimensional data to a low-dimensional embedding space to capture the intrinsic structure of the data. Multimodal signals (such as physiological signals, voice features, and micro-expression movements) present nonlinear distributions in high-dimensional space, which can be viewed as low-dimensional manifolds nested in a higher-dimensional Euclidean space. Therefore, extracting the geometric properties on the manifold can reveal the essential characteristics of multimodal data.

[0116] As an option, this embodiment adopts a manifold learning method based on a graph structure to characterize the local similarity relationship between multimodal data by constructing a weighted nearest neighbor graph. Specifically, the steps for constructing the nearest neighbor graph are as follows:

[0117] First, select any data point x in the multimodal sample data i , and define its neighborhood as a set of k nearest neighbor points It should be understood that k here is a preset parameter, and its size affects the representation ability of the neighborhood.

[0118] Then, for each pair of data points x i and x j The Euclidean distance between i -x j ∥ 2 In one possible implementation, the similarity between data points is defined as the Gaussian kernel weights:

[0119]

[0120] Among them, W ij is the weight matrix element, and σ is the distance scale parameter, which controls the steepness of the weight distribution.

[0121] It should be noted that the weight matrix W represents the similarity between multimodal sample points. The larger its value is, the closer the distance between the two points is and the higher the similarity is.

[0122] In some embodiments, in order to further extract the geometric characteristics of the signal distribution, the Laplacian matrix L of the graph is constructed using the weight matrix W. Specifically, the Laplacian matrix of the graph is defined by the following formula:

[0123] L=DW

[0124] Where D is the degree matrix, and its diagonal elements D ii Defined as i The sum of the associated weights:

[0125]

[0126] It can be understood that the Laplace matrix L characterizes the local smoothness of the graph structure, and further analysis of L can reveal the geometric properties on the manifold.

[0127] Specifically, this embodiment uses the Laplace-Beltrami operator to perform spectral decomposition on the manifold, thereby extracting the local features of the signal on the manifold. It should be noted that the eigenvalues ​​and eigenvectors of the Laplace-Beltrami operator reflect the local and global geometric structures of the manifold, and its mathematical form is:

[0128] Δf(x)=λf(x)

[0129] in:

[0130] Δ is the Laplace-Beltrami operator;

[0131] λ is the eigenvalue of the manifold, representing the global geometric characteristics;

[0132] f(x) is the corresponding eigenvector, representing the local geometric characteristics.

[0133] In one possible implementation, the low-dimensional manifold features of the signal can be extracted by calculating the eigenvalues ​​and eigenvectors of the Laplacian matrix L. These features can be used to describe the intrinsic structure of the multimodal signal in the high-dimensional space, while reducing the redundant information in the original high-dimensional data.

[0134] For example, for physiological signal and micro-expression feature data, the manifold feature extraction process can be specifically implemented as follows:

[0135] First, the high-dimensional sample data is standardized to ensure the scale consistency between different modal data;

[0136] Then, the weighted distance matrix W is calculated through the nearest neighbor graph;

[0137] Finally, the Laplacian matrix L of the graph is spectrally decomposed, and the eigenvectors corresponding to the first d smallest non-zero eigenvalues ​​are selected to construct a low-dimensional embedding space, where d is the target dimension of the low-dimensional manifold.

[0138] It can be understood that this method can effectively reduce the data dimension while preserving the global and local geometric structures of multimodal signals, providing high-quality input features for subsequent risk prediction models.

[0139] It should be noted that the manifold modeling method of this embodiment is suitable for processing signal data with nonlinear distribution characteristics in multimodal samples. In particular, for high-dimensional, nonlinear modal data that presents local structural differences in a specific space (such as speech features and micro-expression features), the manifold modeling method can effectively extract its inherent distribution law and provide a unified geometric basis for the low-dimensional representation of features.

[0140] Specifically, different modal data have differences in their characteristic distributions, for example:

[0141] The distribution of physiological signals (such as heart rate variability and blood pressure fluctuations) usually has strong periodicity and timing;

[0142] The characteristic distribution of speech signals is affected by emotions and semantic content, and it exhibits complex nonlinear changes in frequency and time domains;

[0143] Micro-expression characteristics are mainly composed of dynamic changes in local facial muscles, which usually present short-term, drastic changes and are random.

[0144] Based on the above modal characteristics, the manifold modeling method needs to flexibly adjust the nearest neighbor parameter k and the distance scale parameter σ according to the data distribution characteristics to optimize the feature extraction effect.

[0145] It is important to understand that the setting of the nearest neighbor parameter k directly determines the size of the sample point neighborhood, thereby affecting the ability to represent local structures.

[0146] As an option, the value of k is usually in the range of 5 to 15. Specifically:

[0147] A too small k value may result in an insufficient neighborhood range and fail to fully reflect the local characteristics of the data;

[0148] If the k value is too large, redundant information may be introduced, resulting in dilution of local characteristics.

[0149] For example, for speech features, since their distribution is usually dense and continuous, it is appropriate to select a smaller k value (such as k=5) to accurately capture their local changes. For micro-expression features, since their dynamic changes are fast and the sample distribution may be sparse, a larger k value is usually selected to enhance the coverage of the sample point neighborhood.

[0150] The distance scale parameter σ controls the decay rate of the similarity in the weight matrix W, and its setting needs to be adjusted according to the distribution span of the modal features. Specifically:

[0151] When the distance between data points is small, a smaller σ value can more accurately reflect the local similarity of adjacent sample points;

[0152] When the distance between data points is large and the distribution range is wide, a larger σ value needs to be selected to avoid most elements in the weight matrix approaching zero.

[0153] In one possible implementation, for physiological signals (such as heart rate variability), whose sample points are relatively evenly distributed and have small distances, σ can be set to a range of 0.5 to 1.0; while for speech features or micro-expression features, due to their larger distribution span, σ is usually selected to be in the range of 1.0 to 2.0.

[0154] It should be noted that the value of σ can be optimized through cross-validation or empirical formulas. For example, the mean or median of the global sample point distance can be set as the initial estimate of σ, and then gradually tuned in the experiment.

[0155] In one possible implementation, based on the distribution characteristics of different modal signals, the manifold modeling process includes the following optimization steps:

[0156] For speech features, the frequency distribution of the signal is represented by log spectrum or Mel-frequency cepstral coefficients (MFCC), and the local characteristics of the signal in the time and frequency domains are extracted by constructing a Gaussian kernel weight matrix.

[0157] For micro-expression features, we focus on capturing the spatial distribution differences of facial muscle movements and enhance the representation ability of sparsely distributed signals by increasing the nearest neighbor parameter k and the distance scale σ.

[0158] For physiological signals, combined with their temporal characteristics, a sliding window approach is used to construct a nearest neighbor graph to ensure the temporal consistency of the local structure of the signal.

[0159] It can be understood that the above optimization steps ensure that the manifold modeling method can adapt to the distribution characteristics of different modal data and provide high-quality low-dimensional embedding representation for the fusion of multimodal features.

[0160] It can be understood that, through the above manifold modeling steps, this embodiment extracts low-dimensional geometric features of multimodal signals in high-dimensional space, providing sufficient data support for the training of depression risk prediction models in subsequent steps. Feature extraction not only improves the robustness of the model, but also enhances the model's adaptability to nonlinear signals.

[0161] For step S4, in this embodiment, modal features of different time scales are extracted from multimodal data by multiscale signal decomposition technology. The core goal of the multiscale decomposition method is to decompose the signal into multiple intrinsic mode components (IMF) to capture the dynamic characteristics of the signal on short-term and long-term time scales, and provide rich feature input for the construction of subsequent models.

[0162] It should be noted that multimodal signals have different characteristics on the time scale, for example:

[0163] Physiological signals (such as heart rate variability and blood pressure fluctuations) often contain low-frequency trends and high-frequency fluctuations;

[0164] Speech rate variations and energy fluctuations in speech signals may appear as rapid changes on short time scales;

[0165] Micro-expression characteristics often show drastic changes in a very short period of time.

[0166] Through multi-scale signal decomposition, these modal features can be distinguished into different time scale components, thereby capturing the intrinsic dynamics of the signal more accurately.

[0167] As an option, in this embodiment, the signal is decomposed using the Multiscale Variational Mode Decomposition (MVMD) technique. MVMD is a signal decomposition method based on variational theory, which captures the multi-scale characteristics of the signal by constructing multiple intrinsic mode components (IMFs).

[0168] Specifically, the goal of MVMD is to decompose the input signal f(t) into a set of band-limited signals {u k (t)} and optimize the center frequency ω of each mode k , to minimize the following objective function:

[0169]

[0170] in:

[0171] u k (t) is the kth intrinsic modal component;

[0172] ω k is the corresponding center frequency;

[0173] δ(t) is the Dirac function;

[0174] * indicates a convolution operation.

[0175] It can be understood that MVMD optimizes the center frequency ω by iterative optimization. k and the band-limited signal u k (t), realizes adaptive decomposition of the input signal and can effectively capture the characteristics of the signal at different time scales.

[0176] In a possible implementation, for multimodal data, the decomposition step of MVMD includes the following process:

[0177] Specifically, for physiological signals (such as heart rate variability), MVMD is able to separate modal components at the following time scales:

[0178] High-frequency components: reflect short-term fluctuations, such as instantaneous changes in heart rate;

[0179] Mid-frequency components: reflect mid-term regulatory behaviors, such as the response of the sympathetic nervous system;

[0180] Low-frequency components: reflect long-term trends, such as changes in heart rate as emotions fluctuate.

[0181] For example, after decomposition, the heart rate variability signal f(t) can be divided into three main modal components:

[0182] u1(t): contains high-frequency features (such as changes caused by breathing and sympathetic nerve activity);

[0183] u2(t): contains intermediate frequency features (such as blood pressure fluctuations and parasympathetic nerve regulation signals);

[0184] u3(t): Contains low-frequency features (such as changes caused by long-term mood swings).

[0185] For speech signals, MVMD can decompose its changes in frequency domain and time domain. For example:

[0186] The high-frequency components capture short-term energy fluctuations in the speech signal;

[0187] The mid-frequency component captures the pitch and rate changes of speech;

[0188] The low-frequency component reflects the long-term changing trend of the overall voice emotion.

[0189] In micro-expression features, since the dynamic changes of facial muscles are usually fast, MVMD can separate the following modes:

[0190] High-frequency components: describe the rapid movements of local facial muscles, such as subtle changes around the eyes or at the corners of the mouth;

[0191] Mid-frequency components: describe slower changes in facial expressions, such as the combined movements of eyebrows and masseter muscles;

[0192] Low-frequency component: describes the trend of overall facial expression changes, such as the slow changes in expression during emotional transitions.

[0193] It should be noted that, in order to ensure that the decomposed modal components can retain the original characteristics of the signal, the MVMD technology in this embodiment is optimized by adjusting the following key parameters:

[0194] Mode number K: indicates the number of modes after the signal is decomposed. Usually, K = 3 to 5 is selected according to the complexity of the signal.

[0195] Center frequencyω kInitialization: The spectrum information of the signal can be obtained through fast Fourier transform (FFT) as a reference for the initial center frequency.

[0196] Iteration step size and termination condition: Set a smaller step size to ensure the accuracy of decomposition, and set the termination condition based on the residual of the decomposed modal component.

[0197] In some embodiments, to further analyze the modal components, statistical characteristics of each mode may be calculated. For example:

[0198] Energy distribution: Analyze the energy proportion of different modal components and determine the contribution of the signal on different time scales;

[0199] Spectral characteristics: Extract frequency distribution characteristics through spectrum analysis of modal components;

[0200] Time-varying characteristics: Analyze the time-varying characteristics of modal components through sliding windows to reveal the dynamic change rules of the signal.

[0201] It can be understood that this embodiment realizes multi-scale decomposition of multi-modal signals through MVMD technology, and extracts the dynamic characteristics of signals at different time scales. These modal components can effectively distinguish short-term and long-term change characteristics, providing high-quality input data for subsequent feature fusion and model training.

[0202] It should be noted that MVMD technology is not only applicable to physiological signals, but can also be flexibly adapted to the decomposition and analysis of speech features and micro-expression features. Through the multi-scale decomposition of different modal signals, the present invention achieves a comprehensive characterization of complex signal characteristics and provides key support for the depression risk prediction model.

[0203] For step S5, in this embodiment, a multimodal feature fusion model is constructed to perform weighted fusion of data from different modalities (such as physiological signal features, voice features, and micro-expression features) through a self-attention mechanism. The core goal of this step is to generate a unified high-quality feature vector by fully exploring the correlation and importance between multimodal data to provide input for subsequent models.

[0204] It should be noted that multimodal data has differences in feature dimensions, time scales, and signal distribution. For example, physiological signals are mainly based on time series features, speech features include frequency domain characteristics, and micro-expression movement features are manifested as short-term dynamic changes in space. To this end, this embodiment achieves dynamic weighting and unified representation of multimodal features by constructing a feature fusion model based on the self-attention mechanism.

[0205] As an option, the multimodal feature fusion model of this embodiment is based on the self-attention mechanism, the basic idea of ​​which is to assign dynamic weights to each modal feature, and give more important features higher weights by calculating the correlation between modal features, thereby generating a unified feature representation.

[0206] Specifically, the core calculation of the self-attention mechanism includes the following three steps:

[0207] First, the input multimodal features are mapped into query matrix Q, key matrix K and value matrix V respectively. The calculation form of these matrices is:

[0208] Q=XW Q ,K=XW K ,V=XW V

[0209] in:

[0210] X represents multimodal feature input;

[0211] W Q , W K , W V The mapping weight matrices for query, key, and value, respectively.

[0212] Then, the attention score is obtained by calculating the dot product of the query matrix Q and the key matrix K and normalizing it:

[0213]

[0214] in:

[0215] d k Indicates the dimension of the key vector, which is used for normalization and scaling to avoid excessive inner product values;

[0216] The softmax function is used to convert the scores into a probability distribution.

[0217] Finally, the attention scores are used to perform weighted summation on the value matrix V to obtain the fused multimodal feature representation.

[0218] It can be understood that the above self-attention mechanism effectively enhances the complementarity between features by calculating the correlation between modal features, while suppressing the influence of redundant features and noise.

[0219] In a possible implementation, this embodiment uses physiological signal features, voice features, and micro-expression features as input, and the specific steps are as follows:

[0220] Specifically, physiological signal features include Lyapunov exponents and fractal dimensions extracted from phase space reconstruction and nonlinear dynamics analysis; speech features include pitch, speech rate and short-term energy extracted from time domain and frequency domain analysis of speech signals; micro-expression features include local dynamic features extracted based on facial muscle movements.

[0221] During the feature fusion process, these features are input into the self-attention mechanism according to the following process:

[0222] First, each modal feature is linearly transformed to unify the feature dimension;

[0223] Then, the features are mapped into query matrix Q, key matrix K and value matrix V, and the weights are calculated through the attention mechanism;

[0224] Finally, the fused feature vector is normalized to generate a unified multimodal feature representation.

[0225] It should be noted that the attention mechanism in this embodiment can dynamically assign importance weights to different modal features. For example, in some scenarios, physiological signal features may be more reflective of the risk of depression. At this time, the attention mechanism will automatically give it a higher weight and reduce the weight of other modal features. Through this dynamic weighting mechanism, the robustness and adaptability of the model can be effectively enhanced.

[0226] For example, in certain emotional states (such as depression), the heart rate variability features in physiological signals are often more important than speech features, and the attention mechanism can capture this pattern and assign higher weights to physiological features. In addition, when there are abnormalities in the speech signal such as lowered intonation and slower speech speed, the attention mechanism will increase the weight of the speech features.

[0227] In some embodiments, in order to further optimize the effect of feature fusion, this embodiment also adopts a multi-head attention mechanism, the basic principle of which is to capture different correlation patterns of modal features by calculating multiple attention heads in parallel. Specifically:

[0228] First, each modality feature is divided into multiple subspaces, and the attention score is calculated for each subspace;

[0229] Then, the attention outputs of all subspaces are concatenated to generate the final fused feature representation.

[0230] The mathematical expression of the multi-head attention mechanism is:

[0231] MultiHead(Q,K,V)=Concat(head1,…,head h )W O

[0232] in:

[0233] head i =Attention(QWi i Q ,KW i K ,VW i V );

[0234] W O is the output mapping matrix.

[0235] It is important to understand that the multi-head attention mechanism improves the expressiveness of fused features by capturing different feature correlation patterns, and is particularly suitable for scenarios where there are characteristic differences in multimodal data.

[0236] It can be understood that this embodiment not only realizes efficient and unified representation of multimodal data by constructing a feature fusion model of the self-attention mechanism, but also enhances the dynamic adjustment ability of features to importance weights. This dynamic weighting mechanism can effectively suppress the influence of irrelevant features and noise, and provide robust and high-quality input features for subsequent models.

[0237] It should be noted that the feature fusion model of the present invention is applicable to different types of multimodal data, such as physiological signals, voice features, and micro-expression features. Through the combination of self-attention mechanism and multi-head attention mechanism, the present invention can improve the fusion effect of multimodal features and provide important technical support for depression risk prediction.

[0238] For step S6, in this embodiment, based on the multimodal fusion features, the depression risk prediction model is trained, a nonlinear state space relationship is established, and the depression risk prediction result is output. The core of this step is to capture the temporal dependency of the multimodal fusion features through a dynamic modeling method to achieve accurate prediction of the depression risk state.

[0239] Training a depression risk prediction model

[0240] As an option, this embodiment uses a nonlinear state space model to model depression risk. The state space model can effectively capture the dynamic changes of multimodal features over time, and establish a mathematical description of depression risk prediction through state transfer equations and output equations.

[0241] Specifically, the nonlinear state-space model consists of the following two parts:

[0242] State transfer equation: describes the dynamic change relationship of the system state over time, in the following form:

[0243] x t+1 =f(x t ,u t )+η t

[0244] in:

[0245] x t represents the hidden state of the system (implicit level of depression risk);

[0246] u t represents the fused feature vector;

[0247] f(x t ,u t ) is the state transfer function;

[0248] η t is the process noise.

[0249] Output equation: Mapping hidden states to observations, in the form of:

[0250] y t =g(x t )+v t

[0251] in:

[0252] y t represents the depression risk prediction results;

[0253] g(x t ) is the output function;

[0254] v t is the observation noise.

[0255] It should be noted that the state-space model characterizes the dynamic dependency of multimodal fusion features through the state transfer equation, and realizes the mapping of depression risk through the output equation.

[0256] State transfer and optimization process

[0257] In a possible implementation, to ensure the accuracy of the state space model, this embodiment optimizes the state transfer process by using an extended Kalman filter (EKF). EKF is a state estimation method for nonlinear systems that can dynamically correct state estimation values ​​through iterative calculations.

[0258] Specifically, the optimization process of EKF includes the following steps:

[0259] Status prediction:

[0260]

[0261] P t|t-1 =F t P t-1|t-1 F t +Q t

[0262] in:

[0263] is the predicted state at time t;

[0264] P t|t-1 is the covariance matrix of state prediction;

[0265] F t is the Jacobian matrix of the state transfer function f(·);

[0266] Q t is the process noise covariance matrix.

[0267] Status Update:

[0268]

[0269] P t|t =(IK t H t ) t|t-1

[0270] in:

[0271] K t is the Kalman gain;

[0272] H t is the Jacobian matrix of the output function g(·);

[0273] R t is the observation noise covariance matrix.

[0274] It should be understood that EKF can effectively reduce the uncertainty in the state transfer process and improve the accuracy of the depression risk prediction model by correcting the state estimate at each time step.

[0275] In one possible implementation, the model training process includes the following steps:

[0276] First, the multimodal fusion features and their corresponding depression risk label data are divided into a training set and a validation set;

[0277] Then, the gradient descent optimization algorithm is used to optimize the state transfer function f(x t ,u t) and the output function g(x t ) parameters are trained to minimize the error between the predicted value and the true label;

[0278] For example, the loss function can be defined as Mean Squared Error (MSE):

[0279]

[0280] Among them, y i and represent the true value and the predicted value respectively.

[0281] It should be noted that, in order to avoid overfitting of the model, this embodiment adopts regularization technology, such as adding L2 norm regularization term:

[0282]

[0283] Among them, λ is the regularization coefficient and θ is the model parameter.

[0284] Exemplarily, after model training, the state space model of this embodiment can output the following two results:

[0285] Real-time emotional state: The current emotional state is described through a two-dimensional emotional model (arousal and valence). For example, emotional states can be divided into "low arousal-negative valence" and "high arousal-positive valence".

[0286] Depression risk score: Outputs a continuous quantitative score to assess an individual's risk level of depression. For example, the score ranges from 0 to 1, where higher values ​​indicate greater risk.

[0287] Understandably, this continuous risk score can provide a more nuanced description of the degree of depression risk, facilitating the development of medical interventions and subsequent treatment plans.

[0288] This embodiment can capture the dynamic change characteristics of multimodal features through state space modeling, and use EKF to improve the dynamic adaptability of the prediction model. Compared with the traditional static prediction model, the method of the present invention can update the depression risk prediction results in real time, significantly enhancing the ability to capture complex emotional states and physiological signals.

[0289] It should be noted that the method described in the present invention is not only applicable to the prediction of depression risk, but can also be extended to other mental health status assessment scenarios. Through the combination of dynamic modeling and parameter optimization, this embodiment provides reliable technical support for the accurate prediction of depression risk.

[0290] In general, the present invention collects and preprocesses multimodal data, uses nonlinear dynamics theory to extract the complex dynamic characteristics of physiological signals, constructs a manifold model based on differential geometry theory to extract the local geometric features of the signal, and extracts the short-term and long-term temporal characteristics of the modality through multi-scale signal decomposition. Subsequently, the self-attention mechanism is used to dynamically weight the fusion of multimodal features, and finally the fusion features are dynamically modeled based on the state space model, and the depression risk score and emotional state are output. The present invention can accurately capture the nonlinear dynamic changes of multimodal data, significantly improve the robustness and accuracy of depression risk prediction, and provide efficient technical support for mental health assessment and intervention.

[0291] Please see attached Figure 2 The present invention also provides a depression risk prediction model constructed by the above construction method, and the model structure includes: a data input module, a nonlinear feature extraction module, a manifold feature extraction module, a multi-scale analysis module, a feature fusion module and a prediction module. The function and specific implementation of each module are described in detail below.

[0292] Data input module

[0293] In this embodiment, the data input module is used to receive multimodal data, including physiological signals (such as heart rate variability, blood pressure fluctuations and blood oxygen saturation), voice features (such as intonation, speech speed and voice energy) and micro-expression action data (such as dynamic changes in facial muscles). This module is connected to the acquisition device and transmits the collected signals to the subsequent modules in a unified format.

[0294] As an option, the data input module may include time alignment and signal preprocessing functions, such as time synchronization processing of multimodal data based on the dynamic time warping (DTW) method, removing signal noise and implementing standardization processing to ensure the integrity and consistency of the input data.

[0295] It should be noted that this module serves as the input port of the model and provides high-quality raw data for subsequent feature extraction.

[0296] Nonlinear feature extraction module

[0297] The nonlinear feature extraction module is used to extract dynamic characteristics from physiological signals. In this embodiment, the module extracts nonlinear characteristics of physiological signals, such as Lyapunov exponents and fractal dimensions, based on the phase space reconstruction method. These characteristics can effectively characterize the chaos, complexity and dynamic regulation ability of physiological signals.

[0298] Manifold feature extraction module

[0299] In this embodiment, the manifold feature extraction module is used to construct a nearest neighbor graph based on the high-dimensional manifold structure of the multimodal signal, and extract the geometric features of the signal on the manifold by calculating the Laplacian matrix.

[0300] It should be noted that the manifold feature extraction module can efficiently process the nonlinear distribution characteristics of multimodal data, map high-dimensional data to low-dimensional space, and retain the local structure and global geometric characteristics of the data. By extracting manifold features, the module can enhance the ability to describe complex signal characteristics and reduce the interference of redundant information on model performance.

[0301] Multiscale Analysis Module

[0302] The multi-scale analysis module extracts modal features of different time scales by using the multi-scale variational mode decomposition method (MVMD). In this embodiment, the module decomposes the input multi-modal signal to obtain the short-term dynamic changes and long-term trends of the signal respectively.

[0303] As an option, the module can flexibly adjust the parameters of modal decomposition, such as the number of modes and decomposition accuracy, to adapt to the characteristics of different signals. By performing statistical analysis on the decomposed modal components, such as calculating energy distribution and spectral characteristics, the module can provide rich information on the time scale for subsequent feature fusion.

[0304] It should be understood that the decomposition process of the signal by the multi-scale analysis module realizes the time scale alignment of different modal features, ensuring the fusion and modeling of short-term and long-term features in a unified framework.

[0305] Feature fusion module

[0306] The feature fusion module realizes weighted fusion of multimodal features through the self-attention mechanism to generate a unified feature vector. In this embodiment, the module dynamically calculates the correlation weights of different modal features and assigns greater importance to features with higher weights, thereby improving the accuracy and robustness of multimodal feature fusion.

[0307] It should be noted that the design of the feature fusion module fully considers the characteristic differences of multimodal data and effectively solves the conflict of feature priorities between different modalities through a dynamic weighting mechanism.

[0308] Prediction Module

[0309] The prediction module performs dynamic modeling based on the fusion features to generate real-time emotional states and depression risk scores. In this embodiment, the module uses a nonlinear state space model to capture the dynamic changes of multimodal features through state transfer equations and output equations, and updates the prediction results in real time.

[0310] As an option, the prediction module integrates an extended Kalman filter (EKF) algorithm for optimizing state estimation and risk score calculation. In the output of the risk score, the module is able to provide a continuous risk quantification value ranging from 0 to 1 based on the model's prediction results, where higher scores indicate a higher risk of depression.

[0311] Exemplarily, the module can also output emotional state results, describing the patient's current psychological state based on a two-dimensional emotional model (such as arousal and valence), providing a basis for medical intervention and emotional management.

[0312] It can be understood that the depression risk prediction model described in the present invention realizes efficient processing and feature modeling of multimodal data through the collaborative work of multiple modules. The data input module ensures the quality and consistency of the input signal; the nonlinear feature extraction module, the manifold feature extraction module and the multi-scale analysis module respectively conduct in-depth mining of signal characteristics from different angles; the feature fusion module realizes the unified expression of multimodal features through a dynamic weighting mechanism; the prediction module generates real-time risk scores and emotional states based on state space modeling. The model has a clear design structure and clear functions, providing an efficient and accurate technical implementation path for depression risk prediction.

[0313] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for constructing a depression risk prediction model based on multimodal sample information, characterized in that: The following steps are involved: Acquire multimodal data of physiological signals, voice features, and micro-expression movements, and pre-process the multimodal data; Use nonlinear dynamics theory to reconstruct the phase space of physiological signals and extract nonlinear features; Based on differential geometry theory, multimodal signal distribution is modeled and local geometric features of the signal are extracted; Extract modal features at different time scales through multi-scale signal decomposition technology; Construct a multimodal feature fusion model and use the self-attention mechanism to perform weighted fusion of multimodal data; Based on the fusion features, the depression risk prediction model is trained, the state space transfer relationship is established, and the depression risk prediction results are output.

2. The method for constructing a depression risk prediction model based on multimodal sample information according to claim 1, characterized in that: The collection and preprocessing of the multimodal data includes: Use non-contact photoplethysmography to obtain physiological signals of heart rate variability, blood pressure fluctuations, and blood oxygen saturation; Extract speech features, including pitch, speech rate, and speech energy; Capture micro-expression movement features, including dynamic changes of the frontalis and masseter muscles; Dynamic time warping is used to align the time series of signals of different modalities.

3. The method for constructing a depression risk prediction model based on multimodal sample information according to claim 1, characterized in that: The phase space reconstruction is based on Takens embedding theorem, which constructs the physiological signal into a high-dimensional dynamic feature space through time delay and embedding dimension setting, so as to reveal the nonlinear dynamic law of the signal.

4. The method for constructing a depression risk prediction model based on multimodal sample information according to claim 3, characterized in that: The nonlinear characteristics include: Lyapunov exponent of the system's sensitivity to initial conditions; Fractal dimension for characterizing signal complexity; Dynamic evolution characteristics of system attractors.

5. The method for constructing a depression risk prediction model based on multimodal sample information according to claim 1, characterized in that: The manifold modeling includes: Construct a nearest neighbor graph based on multimodal signals and calculate the weighted distance matrix between samples; The Laplace matrix of the manifold is generated, and the spectral characteristics of the signal are extracted based on the Laplace-Beltrami operator to characterize the local geometric characteristics of the signal on the manifold.

6. The method for constructing a depression risk prediction model based on multimodal sample information according to claim 1, characterized in that: The multiscale signal decomposition is based on a multiscale variational mode decomposition method, which decomposes the signal into multiple intrinsic mode components and extracts features on different time scales, wherein: High-frequency components reflect short-term signal changes; The low-frequency component reflects the long-term trend characteristics.

7. The method for constructing a depression risk prediction model based on multimodal sample information according to claim 1, characterized in that: The multimodal feature fusion is achieved through a self-attention mechanism. By constructing query, key and value matrices, the multimodal features are dynamically weighted, and a fused feature representation is generated based on the weighted feature importance.

8. The method for constructing a depression risk prediction model based on multimodal sample information according to claim 1, characterized in that: The depression risk prediction model is implemented by constructing a nonlinear state space relationship, and the state space model includes: State transition equations that describe the dynamic changes of the system; An output equation that maps feature inputs to a depression risk score.

9. The method for constructing a depression risk prediction model based on multimodal sample information according to claim 8, characterized in that: The state-space model is optimized by extended Kalman filtering, where: Initialize state transition parameters using historical data; The state variables are adjusted dynamically using the observed data to optimize the prediction model.

10. The method for constructing a depression risk prediction model based on multimodal sample information according to claim 1, characterized in that: The outputs of the depression risk prediction model include: Real-time emotional state results, using a two-dimensional emotional model to represent arousal and valence; Depression risk score, a quantitative assessment of a patient's depression risk based on a trained nonlinear model.

Citation Information

Cited By

  • Method for predicting depression risk of chronic disease patient

    CN120299720A

  • A method for predicting depression risk in patients with chronic diseases

    CN120299720B

  • Cardiotoxicity prediction system for children with leukemia based on electrocardio dynamic evolution characteristics

    CN120319497A

  • Wearable double-layer adaptive Kalman filtering ambulatory blood pressure monitoring system

    CN120501395A

  • User state evaluation method and system based on heart rate characteristics

    CN120527007A