Dynamic prediction system for mild depression based on deep learning of multi-modal time series data

CN121506501BActive Publication Date: 2026-08-07NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV OF INFORMATION SCI & TECH
Filing Date
2025-12-25
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0002]抑郁症作为全球主要精神障碍疾病之一,其早期诊断与动态干预对改善患者预后具有关键作用,传统诊断依赖临床量表评估与医生经验判断,存在主观性强、敏感性不足的局限,随着生物传感技术与人工智能的发展,基于多模态生理信号的客观检测成为研究热点,眼动追踪、皮肤电反应、脑电、语音及面部表情数据等生物标志物被证实与抑郁状态密切相关,但单一模态数据易受噪声干扰且信息碎片化,深度学习技术虽能提取高维特征,却难以直接建模生物信号与临床表征间的复杂关联,此外,抑郁症治疗需兼顾疗效、安全性与经济性,但现有方案多依赖经验性选择,缺乏多目标优化的科学决策框架,因此,亟需一种整合多模态数据、动态预测病情演变并生成个性化治疗方案的智能化技术

Benefits of technology

[0058]一、本发明通过多模态数据融合技术,整合眼动追踪、皮肤电反应、脑电、语音及面部表情数据,突破单一生物信号的局限性,形成更全面的抑郁症特征表征,临床-数据双驱权重融合算法结合生理信号与临床标注信息,动态优化特征权重,提升特征与疾病状态的关联性;创新性地将特征划分为生理信号与情绪表达超节点集,通过动态约束系数实时更新超图结构,解决了传统静态模型无法捕捉病情动态变化的缺陷,超图卷积神经网络模型进一步引入临床先验知识,强化模型对轻度抑郁症核心特征的识别能力,实现高精度的早期预测与病情演变跟踪,为个性化干预提供科学依据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121506501B_ABST
    Figure CN121506501B_ABST
Patent Text Reader

Abstract

The application discloses a mild depression dynamic prediction system based on multi-modal time series data deep learning, which comprises a data acquisition module, which acquires relevant data of suspected patients with mild depression; a data processing module, which pre-processes the collected data, and generates a multi-modal time series feature set by adopting a clinical-data double-drive weight fusion algorithm; a key feature set acquisition module, which obtains an optimal time scale according to the time resolution relationship of the multi-modal time series feature set; a time series feature vector set is obtained based on the optimal time scale, and the set is corrected to obtain a multi-modal key feature set; a feature vector acquisition module, which constructs an initial hypergraph based on the multi-modal key feature set and in combination with clinical priori, and obtains a mild depression prediction feature vector; and an optimal scheme acquisition module, which generates a Pareto optimal scheme set based on the mild depression prediction feature vector. The application realizes high-precision prediction and disease tracking, and provides an intelligent tool for precise medical treatment of depression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and more specifically to a dynamic prediction system for mild depression based on deep learning of multimodal time-series data. Background Technology

[0002] Depression, as one of the major mental disorders worldwide, relies heavily on early diagnosis and dynamic intervention to improve patient prognosis. Traditional diagnosis depends on clinical scale assessments and physician experience, which are limited by strong subjectivity and insufficient sensitivity. With the development of biosensor technology and artificial intelligence, objective detection based on multimodal physiological signals has become a research hotspot. Biomarkers such as eye tracking, skin conductance response, electroencephalography (EEG), speech and facial expression data have been proven to be closely related to depressive states. However, single-modal data is easily affected by noise and is fragmented. Although deep learning technology can extract high-dimensional features, it is difficult to directly model the complex relationship between biosignals and clinical manifestations. In addition, the treatment of depression needs to balance efficacy, safety and economy, but existing solutions mostly rely on empirical selection and lack a scientific decision-making framework for multi-objective optimization. Therefore, there is an urgent need for an intelligent technology that integrates multimodal data, dynamically predicts the evolution of the disease, and generates personalized treatment plans.

[0003] Traditional depression prediction technologies suffer from several shortcomings: First, single-modal biosignal analysis cannot fully reflect the heterogeneity of the disease; for example, EEG data can only capture neural activity while ignoring nonverbal cues related to emotional expression. Second, static prediction models struggle to adapt to dynamic changes in the condition; feature extraction at fixed time scales cannot capture the correlation between short-term fluctuations and long-term trends. Third, treatment plan formulation lacks quantitative basis, requiring doctors to manually weigh efficacy, side effects, and costs, potentially leading to suboptimal decisions. Existing dynamic prediction research often focuses on optimizing a single objective, failing to consider conflicting multi-dimensional goals simultaneously and lacking embedded clinical prior knowledge, resulting in insufficient model interpretability. Furthermore, personalized recommendation systems rely solely on simple matching based on patient preferences, without dynamically adjusting based on demographic characteristics and the urgency of the condition, limiting the accuracy and feasibility of the proposed solutions. Summary of the Invention

[0004] The purpose of this invention is to provide a dynamic prediction system for mild depression based on deep learning of multimodal time series data, so as to achieve high-precision early prediction and tracking of disease progression, and provide a scientific basis for personalized intervention.

[0005] This invention adopts the following technical solution: a dynamic prediction system for mild depression based on deep learning of multimodal time series data, comprising:

[0006] The module includes a data acquisition module, a data processing module, a key feature set acquisition module, a feature vector acquisition module, and an optimal solution acquisition module.

[0007] The data acquisition module is used to collect eye-tracking data, skin conductance data, electroencephalogram (EEG) data, voice data, and facial expression data from suspected patients with mild depression.

[0008] The data processing module is used to preprocess the data collected in the data acquisition module and extract the modal features of the data. The modal features are fused using a clinical-data dual-drive weighted fusion algorithm to generate a multimodal time series feature set.

[0009] The key feature set acquisition module is used to construct a candidate set of time scales based on the temporal resolution relationship of the multimodal temporal feature sets, and to select the optimal time scale of modal features from the candidate set of time scales using a three-dimensional optimized temporal scale selection algorithm. Based on the optimal time scale, a set of temporal feature vectors is obtained, and the set of temporal feature vectors is corrected by a feature redundancy correction algorithm to obtain the multimodal key feature set.

[0010] The feature vector acquisition module is used to divide the multimodal key feature set into a physiological signal supernode set and an emotion expression supernode set, construct an initial hypergraph by combining clinical priors, update the initial hypergraph by time segments through dynamic constraint coefficients, and perform convolution operation on the updated hypergraph through a hypergraph convolutional neural network model to obtain the predictive feature vector of mild depression.

[0011] The optimal solution acquisition module is used to analyze the key influencing dimensions of a patient's depressive state based on the predictive feature vector of mild depression. It constructs a multi-objective function with treatment efficiency, side effects, and cost as objectives, and uses a non-dominated sorting genetic algorithm to solve the multi-objective function to generate a Pareto optimal solution set.

[0012] Furthermore, the data acquisition module is configured to perform the following actions:

[0013] Eye-tracking data was obtained by recording fixation point data and pupil change data when the patient viewed pictures and read text using a 120-240Hz eye tracker; skin conduction level and response peak were collected using fingertip electrodes to obtain skin conduction response data; EEG data was obtained by placing electrodes according to international standards using a 64-channel EEG device to collect resting-state and task-state signals of the brain; speech data was obtained by recording audio of the patient reading aloud and speaking freely using a 48kHz microphone; and facial expression data was obtained by capturing the patient's facial dynamics and key point changes using a 25fps camera.

[0014] Furthermore, the data processing module is configured to perform the following actions:

[0015] For eye-tracking data, abnormal fixation points caused by device shake or blinking are removed, and missing data segments are filled in using interpolation; pupil change data is smoothed; and fixation point data is aligned with the presentation time of the corresponding stimulus based on timestamps.

[0016] For the skin conductance response data, abnormal jump values ​​were removed, and baseline drift was filtered out using the sliding window method; skin conductance signals were extracted within the task period, and resting state and reactive state data segments were distinguished.

[0017] For EEG data, interference from eye movement and electromyography was removed using independent component analysis; the 64-channel signal was grouped according to the international standard brain region division criteria, and data segments of resting state and task state were extracted and filtered.

[0018] For speech data, environmental noise and silent segments are removed from the audio, and the data is stored according to the task type of reading aloud and free narration. The duration and location of pauses in the speech data are marked, and the audio sampling rate is standardized.

[0019] For facial expression data, a facial key point detection algorithm is used to correct the tilted image caused by camera angle deviation and align the facial reference points; image noise is removed; facial dynamic changes in consecutive frames are extracted at a frame rate of 25fps, and the start and end frames of expression changes are marked.

[0020] The specific formula for generating a multimodal temporal feature set is as follows:

[0021] ;

[0022] in, Represents the multimodal temporal characteristics at time step t. This represents the total number of modal features. Represents the balance coefficient. Represents the mutual information function. This represents the feature of the m-th modal feature at time step t. This represents clinically labeled data. Represents the similarity function. This represents the clinical standard value corresponding to the m-th modal feature.

[0023] Furthermore, the key feature set acquisition module is configured to perform the following actions:

[0024] The specific formula for obtaining the optimal time scale is:

[0025] ;

[0026] in, The optimal time scale represents the m-th modal feature. Represents the candidate set for the time scale. This represents the m-th modal feature on the time scale. Features This represents the time-scale penalty coefficient. The clinical recommendation scale represents the m-th modality feature;

[0027] The collection of time-series feature vectors includes:

[0028] Based on the optimal time scale, the modal features are divided into several consecutive and equally long time windows. Within each time window, the basic statistics of the modal features are calculated, including the mean, variance, maximum, minimum, and median, to capture the central tendency and dispersion of the modal features at that time scale. For modal features with dynamic characteristics, dynamic statistics within the time window are additionally calculated, including the slope, rate of change, number of peaks, and duration. The basic and dynamic statistics obtained from all time windows are summarized to obtain a set of time-series feature vectors.

[0029] Among them, the modal features with dynamic change characteristics are those with a normalized rate of change standard deviation of not less than 0.15, a slope mean of not less than 0.08, or a number of peaks of not less than 2 within the time window corresponding to the optimal time scale.

[0030] The specific formula for obtaining the multimodal key feature set is as follows:

[0031] ;

[0032] in, This represents the key feature of the modified m-th modal feature. This represents the feature statistics extraction function. This represents the redundancy penalty coefficient. Represents an exponential function. This represents the function for calculating characteristic redundancy. This represents the feature of the m-th modal feature at the optimal time scale.

[0033] Furthermore, the feature vector acquisition module is configured to perform the following actions:

[0034] Features related to EEG and TENS data in the multimodal key feature set are categorized into the physiological signal supernode set, while features related to eye-tracking, speech, and facial expression data are categorized into the emotion expression supernode set. A feature association information database for mild depression in clinical diagnosis is collected to form clinical prior association rules. Using the physiological signal supernode set and the emotion expression supernode set as nodes, hyperedges are established according to the clinical prior association rules to obtain an initial hypergraph.

[0035] Dynamic constraint coefficient The expression is:

[0036] ;

[0037] in, This represents the function that takes the maximum value. Represents the cosine similarity function. Indicates the initial hypergraph. This indicates a clinical prior hypergraph. Indicates the similarity threshold;

[0038] The feature embedding matrices of the physiological signal supernode set and the emotion expression supernode set in the updated hypergraph are extracted respectively. Using a hypergraph convolutional neural network model, the matrices are sequentially subjected to inverse square / inverse operations on the hypergraph association matrix, hyperedge degree matrix, and supernode degree matrix, followed by matrix multiplication. The LeakyReLU activation function is then applied to obtain the feature representations of the physiological signal supernodes and the emotion expression supernodes. The specific formulas are as follows:

[0039] ;

[0040] in, The supernode feature representation at time step t. This represents the LeakyReLU activation function. Represents the degree matrix of supernodes. express The inverse square root matrix, Represents the hypergraph incidence matrix. Represents the hypermarginality matrix. express transpose, Indicates the convolution kernel weights;

[0041] The feature representations of physiological signals and emotional expression are fused to obtain a fused feature representation. The dimensionality of the fused feature representation is then reduced to obtain a predictive feature vector for mild depression.

[0042] Furthermore, clinical prior association rules include internal association rules for physiological signals, internal association rules for emotional expression, and cross-domain association rules between physiological signals and emotional expression.

[0043] Furthermore, the optimal solution acquisition module is configured to perform the following actions:

[0044] The expression for the multi-objective function is:

[0045] ;

[0046] in, Represents a multi-objective function; Indicates the treatment plan; This represents a function that takes the minimum value. This represents the objective function for treatment efficiency. , This represents the predictive feature vector for mild depression. This represents the feature vector set for predicting mild depression. Indicates weight, Indicates responsiveness. Represents the exponentially decaying term. This represents the first correction factor. express right The adaptation distance; Describe the objective function for side effects. , Indicates the type of side effect. Represents a set of side effect types. Indicates the risk value. This represents the second correction factor. Indicates sensitivity, This indicates a suspected case of mild depression. Indicates the exponential enhancement term; Represents the cost objective function, , Indicates the cost type. Represents a set of cost types. Indicates the cost value. This represents the third correction factor. Indicates the urgency of the treatment;

[0047] The generation of the Pareto optimal solution set includes:

[0048] Step 1: Each individual in the initial population corresponds to a set of treatment plan parameters, including the frequency of psychotherapy, drug dosage, and intensity of physical therapy; the individual is a suspected patient with mild depression.

[0049] Step 2: Calculate the multi-objective function value for each individual in the initial population;

[0050] Step 3: Perform non-dominated sorting on the individuals in the initial population, and divide the individuals into different front layers according to their dominance relationship. Individuals in the same front layer do not dominate each other.

[0051] Step 4: Obtain the distribution density of each individual in its frontal layer to obtain the crowding degree of the individual;

[0052] Step 5: Based on the non-dominated ranking results and crowding, select individuals from the initial population to enter the next generation using the roulette wheel selection method;

[0053] Step 6: Perform crossover on the individuals entering the next generation, with a crossover probability of 0.7-0.9, and generate new individuals through single-point crossover;

[0054] Step 7: Perform mutation operation on the new individual, with the mutation probability set to 0.05-0.15, and generate mutated individuals by randomly changing the parameter values;

[0055] Step 8: Merge the individuals in the initial population with the mutated individuals, and repeat steps 3-4 to perform non-dominated sorting and crowding calculation to obtain a new population;

[0056] Step 9: Repeat steps 4-8. After iterating for 40-60 generations, extract individuals in the first frontier layer of the new population to form the Pareto optimal solution set.

[0057] Compared with the prior art, the present invention, employing the above technical solution, has the following technical effects:

[0058] I. This invention integrates eye-tracking, electrodermal response, electroencephalography (EEG), speech, and facial expression data through multimodal data fusion technology, overcoming the limitations of single biological signals to form a more comprehensive characterization of depression. The clinical-data dual-drive weight fusion algorithm combines physiological signals and clinical annotation information to dynamically optimize feature weights and improve the correlation between features and disease state. It innovatively divides features into sets of physiological signals and emotional expression supernodes, and updates the hypergraph structure in real time through dynamic constraint coefficients, solving the deficiency of traditional static models in capturing dynamic changes in the condition. The hypergraph convolutional neural network model further introduces clinical prior knowledge, strengthening the model's ability to identify the core features of mild depression, achieving high-precision early prediction and tracking of disease evolution, and providing a scientific basis for personalized intervention.

[0059] Second, this invention generates a Pareto optimal solution set through a non-dominated sorting genetic algorithm, avoiding the one-sidedness of traditional single-objective optimization. This algorithm maintains the diversity of the solution set through non-dominated sorting and crowding, ensuring coverage of different treatment needs. It introduces a preference-attribute dual-dimensional matching technology, combining the priority ranking of treatment goals by patients and their families with demographic characteristics to accurately match the optimal solution. This breaks through the experience-driven decision-making model, realizes the quantitative matching of treatment parameters, expected effects and risk control, improves the adherence to the treatment plan and the clinical conversion rate, and provides a replicable intelligent tool. Attached Figure Description

[0060] Figure 1 This is an overall structural diagram of the present invention.

[0061] Figure 2 This is a training result diagram of the hypergraph convolutional neural network model of this invention.

[0062] Figure 3 This is a graph showing the working characteristics of the test subjects on the test set in this embodiment of the invention.

[0063] Figure 4This is a confusion matrix diagram of the prediction results in an embodiment of the present invention.

[0064] Figure 5 This is a radar comparison chart of performance indicators in an embodiment of the present invention. Detailed Implementation

[0065] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.

[0066] To achieve the above objectives, this invention proposes a dynamic prediction system for mild depression based on deep learning of multimodal temporal data, such as... Figure 1 As shown, it includes:

[0067] The system includes a data acquisition module, a data processing module, a key feature set acquisition module, a feature vector acquisition module, an optimal solution acquisition module, and a solution output module.

[0068] The data acquisition module is used to collect eye-tracking data, skin conductance data, electroencephalogram (EEG) data, voice data, and facial expression data from suspected patients with mild depression; specifically:

[0069] Eye-tracking data was obtained by recording fixation point data and pupil change data when the patient viewed pictures and read text using a 120-240Hz eye tracker; skin conduction level and response peak were collected using fingertip electrodes to obtain skin conduction response data; EEG data was obtained by placing electrodes according to international standards using a 64-channel EEG device to collect resting-state and task-state signals of the brain; speech data was obtained by recording audio of the patient reading aloud and speaking freely using a 48kHz microphone; and facial expression data was obtained by capturing the patient's facial dynamics and key point changes using a 25fps camera.

[0070] The data processing module preprocesses the data acquired by the data acquisition module, extracts the modal features of the data, and uses a clinical-data dual-drive weighted fusion algorithm to fuse the modal features to generate a multimodal time-series feature set; specifically:

[0071] For eye-tracking data, abnormal fixation points caused by device shaking or blinking are removed, and missing data segments are filled in using interpolation; pupil change data is smoothed to eliminate high-frequency noise interference; and fixation point data is aligned with the presentation time of the corresponding stimulus based on timestamps to mark the effective fixation interval.

[0072] For the skin conductance response data, abnormal jump values ​​caused by poor electrode contact were removed, and the sliding window method was used to filter baseline drift. According to the experimental task timeline, skin conductance signals within the task period were extracted to distinguish between resting state and reactive state data segments. The skin conductance level values ​​were standardized to eliminate the magnitude deviation caused by individual physiological differences.

[0073] For EEG data, independent component analysis was used to remove artifacts such as eye movement and electromyography, and retain effective EEG signals. The 64-channel signal was grouped according to the international standard brain region division standard, and data segments of resting state and task state were extracted and filtered to retain characteristic frequency bands related to emotions.

[0074] For speech data, environmental noise and silent segments are removed from the audio, and effective speech segments are extracted; the data is classified and stored according to the task type of reading aloud and free narration, and the pause duration and location in the speech data are marked; the audio sampling rate is standardized to the standard required for analysis to ensure data format consistency.

[0075] For facial expression data, a facial key point detection algorithm is used to correct the tilted image caused by camera angle deviation and align the facial reference points; image noise caused by uneven lighting is removed to enhance the feature contrast of the facial expression area; facial dynamic changes in continuous frames are extracted at a frame rate of 25fps, and the start and end frames of expression changes are marked.

[0076] The specific formula for generating a multimodal temporal feature set is as follows:

[0077] ;

[0078] in, Represents the multimodal temporal characteristics at time step t. This represents the total number of modal features. Represents the balance coefficient. Represents the mutual information function. This represents the feature of the m-th modal feature at time step t. This represents clinically labeled data. Represents the similarity function. This represents the clinical standard value corresponding to the m-th modal feature.

[0079] The key feature set acquisition module is used to construct a candidate set of time scales based on the temporal resolution relationship of multimodal temporal feature sets, and to select the optimal time scale for modal features from the candidate set using a three-dimensional optimized temporal scale selection algorithm; based on the optimal time scale, a set of temporal feature vectors is obtained, and the set of temporal feature vectors is corrected by a feature redundancy correction algorithm to obtain the multimodal key feature set; specifically:

[0080] The specific formula for obtaining the optimal time scale is:

[0081] ;

[0082] in, The optimal time scale represents the m-th modal feature. Represents the candidate set for the time scale. This represents the m-th modal feature on the time scale. Features This represents the time-scale penalty coefficient. The clinical recommendation scale represents the m-th modality feature;

[0083] The collection of time-series feature vectors includes:

[0084] Based on the optimal time scale, the modal features are divided into several consecutive and equally long time windows. Within each time window, the basic statistics of the modal features are calculated, including the mean, variance, maximum, minimum, and median, to capture the central tendency and dispersion of the modal features at that time scale. For modal features with dynamic characteristics, dynamic statistics within the time window are additionally calculated, including the slope, rate of change, number of peaks, and duration. The basic and dynamic statistics obtained from all time windows are summarized to obtain a set of time-series feature vectors.

[0085] Among them, the modal features with dynamic change characteristics are those with a normalized rate of change standard deviation of not less than 0.15, a slope mean of not less than 0.08, or a number of peaks of not less than 2 within the time window corresponding to the optimal time scale.

[0086] The specific formula for obtaining the multimodal key feature set is as follows:

[0087] ;

[0088] in, This represents the key feature of the modified m-th modal feature. This represents the feature statistics extraction function. This represents the redundancy penalty coefficient. Represents an exponential function. This represents the function for calculating characteristic redundancy. This represents the feature of the m-th modal feature at the optimal time scale.

[0089] The feature vector acquisition module is used to divide the multimodal key feature set into a physiological signal supernode set and an emotional expression supernode set, construct an initial hypergraph based on clinical priors, and update the initial hypergraph according to time segments using dynamic constraint coefficients; then, a hypergraph convolutional neural network model is used to perform convolution operations on the updated hypergraph to obtain the predictive feature vector for mild depression; specifically:

[0090] Features related to EEG and TENS data in the multimodal key feature set are categorized into a physiological signal supernode set, while features related to eye-tracking, speech, and facial expression data are categorized into an emotion expression supernode set. A feature association information database for mild depression in clinical diagnosis is collected to form clinical prior association rules. Using the physiological signal and emotion expression supernode sets as nodes, hyperedges are established according to the clinical prior association rules to obtain an initial hypergraph. Directly associated feature pairs correspond to core hyperedges, and indirectly associated feature pairs correspond to ordinary hyperedges. Core hyperedges are assigned an initial weight of 1.2-1.5, and ordinary hyperedges are assigned an initial weight of 0.8-1.0.

[0091] Based on clinical diagnostic guidelines for mild depression, authoritative treatment manuals, and high-impact journal articles, we extracted clearly documented correlation descriptions between physiological signal characteristics and emotional expression characteristics to establish a preliminary correlation information database. We conducted expert interviews with psychiatrists and psychological assessors to record characteristic correlation phenomena observed in clinical practice, updating the preliminary correlation information database. We collected multimodal clinical data from past patients with mild depression, analyzed case reviews, statistically determined feature co-occurrence frequencies, and selected strongly correlated feature pairs with frequencies higher than 70%, incorporating them into the updated preliminary correlation information database. We then deduplicated and verified the correlation information in this database, eliminating contradictory or low-confidence entries to obtain a strongly correlated feature correlation information database for mild depression.

[0092] Clinical prior association rules include internal association rules for physiological signals, internal association rules for emotional expression, and cross-domain association rules between physiological signals and emotional expression;

[0093] Dynamic constraint coefficient The expression is:

[0094] ;

[0095] in, This represents the function that takes the maximum value. Represents the cosine similarity function. Indicates the initial hypergraph. This indicates a clinical prior hypergraph. Indicates the similarity threshold;

[0096] when Remodeling is triggered at this time;

[0097] The feature embedding matrices of the physiological signal supernode set and the emotion expression supernode set in the updated hypergraph are extracted respectively. Using a hypergraph convolutional neural network model, the matrices are sequentially subjected to inverse square / inverse operations on the hypergraph association matrix, hyperedge degree matrix, and supernode degree matrix, followed by matrix multiplication. The LeakyReLU activation function is then applied to obtain the feature representations of the physiological signal supernodes and the emotion expression supernodes. The specific formulas are as follows:

[0098] ;

[0099] in, The supernode feature representation at time step t. This represents the LeakyReLU activation function. Represents the degree matrix of supernodes. express The inverse square root matrix, Represents the hypergraph incidence matrix. Represents the hypermarginality matrix. express transpose, Indicates the convolution kernel weights;

[0100] The feature representations of physiological signals and emotional expression are fused to obtain a fused feature representation. The dimensionality of the fused feature representation is then reduced to obtain a predictive feature vector for mild depression.

[0101] The optimal solution acquisition module is used to analyze the key influencing dimensions of a patient's depressive state based on the predicted feature vector of mild depression. A multi-objective function is constructed with treatment efficiency, side effects, and cost as objectives. A non-dominated sorting genetic algorithm is used to solve the multi-objective function to generate a Pareto optimal solution set. Specifically:

[0102] The expression for the multi-objective function is:

[0103] ;

[0104] in, This represents a multi-objective function that comprehensively measures the performance of treatment plan x across three objective dimensions: treatment efficiency, side effects, and cost. This represents a function that takes the minimum value. This represents the objective function for treatment efficiency. , This represents the predictive feature vector for mild depression. This represents the feature vector set for predicting mild depression. Indicates weight, Indicates responsiveness. Represents the exponentially decaying term. This represents the first correction factor. express right The adaptation distance; The objective function represents the side effects. , Indicates the type of side effect. Represents a set of side effect types. Indicates the risk value. This represents the second correction factor. Indicates sensitivity, This indicates a suspected case of mild depression. Indicates the exponential enhancement term; Represents the cost objective function, , Indicates the cost type. Represents a set of cost types. Indicates the cost value. This represents the third correction factor. This indicates the urgency of the treatment, determined based on the urgency of the patient's condition.

[0105] The generation of the Pareto optimal solution set includes:

[0106] Step 1: Set the initial population size to 80-120 individuals. Each individual corresponds to a set of treatment plan parameters, including the frequency of psychotherapy, drug dosage, and intensity of physical therapy. The individuals are suspected patients with mild depression.

[0107] Step 2: Calculate the multi-objective function value for each individual in the initial population, i.e., the quantitative results of treatment efficiency, side effects, and cost;

[0108] Step 3: Perform non-dominated sorting on the individuals in the initial population, and divide the individuals into different front layers according to their dominance relationship. Individuals in the same front layer do not dominate each other.

[0109] Step 4: Obtain the distribution density of each individual in its frontal layer to obtain the crowding degree of the individual, which is used to maintain population diversity;

[0110] Step 5: Based on the non-dominated ranking results and crowding, select individuals from the initial population to enter the next generation using the roulette wheel selection method;

[0111] Step 6: Perform crossover on the individuals entering the next generation, with a crossover probability of 0.7-0.9, and generate new individuals through single-point crossover;

[0112] Step 7: Perform mutation operation on the new individual, with the mutation probability set to 0.05-0.15, and generate mutated individuals by randomly changing the parameter values;

[0113] Step 8: Merge individuals from the initial population with mutated individuals, repeat steps 3-4 to perform non-dominated sorting and crowding calculation, and retain the better individuals to form a new population;

[0114] Step 9: Repeat steps 4-8. After iterating for 40-60 generations, extract individuals in the first frontier layer of the new population to form the Pareto optimal solution set.

[0115] The solution output module is used to collect patients' and their families' preference weights for treatment efficiency, side effects, and costs, as well as patients' demographic information, through structured questionnaires; the preference-attribute dual-dimensional matching technology selects the optimal solution from the Pareto optimal solution set as the final personalized treatment plan, and organizes it into a structured report output.

[0116] Example 1:

[0117] Develop prediction and personalized treatment plans for mild depression in adolescents.

[0118] Eye-tracking data, electrodermal response (EDS) data, electroencephalogram (EEG) data, speech data, and facial expression data were collected from adolescents aged 13-18 years suspected of having mild depression. An optimal timescale of 500ms was selected for EEG data, and an optimal timescale of 1s was selected for facial expression data. A database of correlations between clinically diagnosed characteristics of mild depression was compiled, such as the correlation between specific EEG bands and symptoms of depressed mood, and the correlation between EDS intensity and anxiety levels. Treatment efficiency targets should consider the adolescents' acceptance and response to different treatment methods; side effect targets should focus on the potential impact of treatment on adolescents' growth and development and cognitive function; and cost targets should cover examination fees, treatment costs, and time costs during the treatment process.

[0119] A structured questionnaire was designed, covering expectations of treatment efficiency, tolerance for side effects, and acceptable treatment costs. Demographic information on adolescents was also collected, including age, gender, presence of underlying medical conditions, and family economic status. This questionnaire was used to collect the preference weights of adolescents and their families regarding treatment efficiency, side effects, and costs. A preference-attribute two-dimensional matching technique was employed. First, the collected preference weights were categorized as high, medium, and low, forming a preference weight matrix. Then, the core attributes of each option in the Pareto optimal solution set were extracted, including expected treatment efficiency, side effect risk level, and treatment cycle cost, constructing a solution attribute matrix. The dimensional matching degree between preferences and solutions was calculated, matching solution attributes with preference weights. Expected efficiency was positively correlated with efficiency preference weights, side effect risk level was negatively correlated with side effect preference weights, and treatment cycle cost was negatively correlated with cost preference weights, resulting in the individual matching degree for each option across the three dimensions. The individual matching degrees were then adjusted based on the adolescents' demographic information to obtain the adjusted overall matching degree. The overall matching degree of all the plans is ranked, and the plan with the highest overall matching degree is selected as the final personalized treatment plan for the adolescent. The plan is then compiled into a structured report output that includes specific treatment methods, treatment parameters, expected results, and precautions.

[0120] Example 2:

[0121] Data from a 6-month clinical trial conducted by the Department of Psychology at a top-tier hospital was obtained from a dataset website. The trial recruited 234 patients suspected of having mild depression, aged 18-65, including 102 men and 132 women. In this clinical trial, an eye tracker with a frequency of 120 Hz was used to record the eye movements of each patient while viewing 20 emotion-related images (including positive, negative, and neutral images). Each image was viewed for 10 seconds, and parameters such as scan rate, fixation point position, fixation duration, mean pupil diameter, and variance of pupil diameter were obtained. Finger electrode sampling was used to collect the patients' electrodermal (EDS) signals while completing an emotion Stroop task. The sampling frequency was 1000 Hz, and the duration was 5 minutes. Parameters such as EDS conduction level, peak EDS response, EDS response frequency, and EDS response amplitude were obtained. A 64-channel EEG was used, with electrodes placed according to the international 10-20 system standard, to collect 3 minutes of resting-state EEG signals (eyes closed) and 5 minutes of task-oriented EEG signals. The sampling frequency was 1000 Hz, with a bandpass filter of 0.5-45 Hz. Wave (0.5-4Hz), Wave (4-8Hz), Wave (8-13Hz), Parameters such as average power and power percentage of each frequency band in the 13-30Hz wave were collected. Audio recordings of patients reading standard emotional sentences and freely narrating recent life events were obtained using a professional microphone with a 48kHz sampling rate. The reading task included 10 sentences, and the free narration lasted approximately 3 minutes. Parameters such as the fundamental frequency mean, fundamental frequency variation coefficient, speech rate, pause duration, and formant characteristics were acquired. Facial dynamics of patients during conversations with doctors and while watching emotional videos were captured using a 25fps high-definition camera, with a recording duration of approximately 10 minutes. 68 facial key points were extracted, and parameters such as smile duration, brow furrowing frequency, eye opening degree, and mouth corner movement amplitude were obtained. Based on all collected parameters, two psychiatrists with over 10 years of clinical experience independently diagnosed the patients, using this data as clinical annotation data. Ultimately, 156 patients were identified as having mild depression, and 78 were identified as a healthy control group.

[0122] Set balance coefficient .

[0123] Based on a database of fixation patterns in normal populations, clinical standard values ​​corresponding to the modal characteristics of eye-tracking data were determined; based on normal psychological assessment tools and the range of skin conductance response, clinical standard values ​​corresponding to the modal characteristics of skin conductance response data were determined; based on the power distribution of each frequency band in healthy populations, clinical standard values ​​corresponding to the modal characteristics of electroencephalogram (EEG) data were determined; based on the prosodic characteristics of normal speech, clinical standard values ​​corresponding to the modal characteristics of speech data were determined; based on the frequency of facial expression changes in healthy populations, clinical standard values ​​corresponding to the modal characteristics of facial expression data were determined.

[0124] The candidate timescales for eye-tracking data are defined as follows: T1 = {50ms, 100ms, 200ms, 500ms}; T2 = {100ms, 200ms, 500ms, 1s}; T3 = {100ms, 200ms, 500ms, 1s, 2s}; T4 = {100ms, 200ms, 500ms}; and T5 = {200ms, 500ms, 1s, 2s}. A timescale penalty coefficient is also defined. The clinically recommended timescales for eye-tracking data are 200ms, for electrodermal response (EDR) data 500ms, for electroencephalogram (EEG) data 1s, for speech data 200ms, and for facial expression data 1s. Ultimately, the optimal timescales for modal features of eye-tracking data, EEG data, speech data, facial expression data, and ESR data are found to be 200ms, 1s, 200ms, 1s, and 500ms respectively.

[0125] Set redundancy penalty coefficient .

[0126] Assign an initial weight of 1.3 to the core superedge. Abnormal wave power indicates a depressed mood; weakened skin conductance indicates a flat tone of voice; and decreased pupil diameter indicates inattention.

[0127] Assign an initial weight of 0.9 to the ordinary hyperedge. Abnormal wave power indicates changes in facial expression; abnormal gaze pattern indicates changes in speech rhythm.

[0128] Set similarity threshold When updating the initial supermap, set the initial... Training Round 5 Training round 10 Training round 15 ;when When this is triggered, remodeling is performed, making... Restored to 0.92.

[0129] The network structure parameters for the LeakyReLU activation function (negative slope = 0.01) are: 3 hypergraph convolutional layers, [64, 32, 16] hidden units per layer, Dropout ratio of 0.3, learning rate of 0.001, optimizer of Adam, and batch size of 32.

[0130] The dataset, consisting of preprocessed eye-tracking data, skin conductance response data, electroencephalogram data, speech data, and facial expression data, was divided into training set, validation set, and test set in a 14:3:3 ratio.

[0131] Training configuration for the hypergraph convolutional neural network model:

[0132] Maximum number of training rounds: 100 rounds;

[0133] Early stopping strategy: Stop training if the validation set loss does not improve for 10 consecutive rounds;

[0134] Loss function: Cross-entropy loss;

[0135] Evaluation metrics: accuracy, precision, recall, and F1 score.

[0136] To verify the generalization ability and stability of the system proposed in this invention, 5-fold cross-validation was performed on all samples.

[0137] Cross-validation configuration: 5 folds; each fold contains approximately 187 training samples and approximately 47 validation samples; each fold uses the same network structure and training strategy; evaluation metrics are accuracy, precision, recall, and F1 score. The results of cross-validation are shown in Tables 1 and 2.

[0138] Table 1 Results of cross-validation at each fold

[0139]

[0140] Table 2 Summary of Cross-Validation Results

[0141]

[0142] As can be seen from Tables 1 and 2, the average accuracy reached 96.6%, indicating that the system proposed in this invention maintains high performance on different data subsets; the standard deviation is small (<5%), indicating that the system proposed in this invention has good stability and strong generalization ability; the average recall rate is 97.8%, and the false negative rate is low, which meets the requirements for clinical application; the perfect prediction in the fourth fold may be due to the typical characteristics of the data in that fold.

[0143] Figure 2 (a) is a curve showing the change in training loss during the training process of the hypergraph convolutional neural network model. It can be seen that as the number of training rounds increases, the training loss shows a steep downward trend, rapidly decreasing from about 0.6 at the beginning to about 0.0523 in the 28th round, indicating that the model can quickly learn data features and converge. Figure 2 (b) is a curve showing the change in validation loss during the training process of the hypergraph convolutional neural network model. It can be seen that the validation loss decreased rapidly in the early stage, but stabilized and fluctuated slightly after the 18th round. No significant improvement was observed by the 28th round, which triggered the preset early stopping mechanism and effectively prevented the model from overfitting. Figure 2 (c) is the training accuracy curve of the hypergraph convolutional neural network model. It can be seen that the training accuracy steadily increases and eventually reaches a high level of 98.77%. Figure 2(d) is the curve of the validation accuracy of the hypergraph convolutional neural network model. It can be seen that the validation accuracy increases synchronously with the increase of training rounds and stabilizes at around 82.86%; indicating that the hypergraph convolutional neural network model maintains good robustness while learning key features.

[0144] Figure 3 The receiver operating characteristic (ROC) curve on the test set shows that the ROC curve is extremely close to the upper left corner, and the area under the curve (AUC) reaches 0.993. This indicates that the system proposed in this invention has extremely high sensitivity and specificity in distinguishing patients with mild depression from healthy controls, and its overall classification performance is excellent.

[0145] Figure 4 The confusion matrix obtained by using the system proposed in this invention shows that, in 36 test samples, the system correctly identified 23 depressed patients (true positives) and 11 healthy subjects (true negatives), with only 1 missed diagnosis (false negative) and 1 misdiagnosis (false positive). This indicates that the system proposed in this invention has an extremely low missed diagnosis rate (4.2%) and misdiagnosis rate in clinical applications, effectively assisting doctors in early screening and demonstrating high safety.

[0146] Figure 5 The radar chart comparing the performance metrics of the system proposed in this invention shows that, across the five dimensions of accuracy, precision, recall, F1 score, and area under the curve (AUC), the performance of the test set (red solid line) highly overlaps with the average performance of 5-fold cross-validation (blue solid line), and the enclosed areas are similar. The average accuracy of cross-validation reaches 96.6%, indicating that the system proposed in this invention not only performs excellently on a single test set but also exhibits extremely low standard deviation in cross-validation. This demonstrates that the system proposed in this invention has extremely strong generalization ability and stability across different data distributions.

[0147] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A dynamic prediction system for mild depression based on deep learning of multimodal time-series data, characterized in that, include: Data acquisition module, data processing module, key feature set acquisition module, feature vector acquisition module, and optimal solution acquisition module; The data acquisition module is used to collect eye-tracking data, skin conductance data, electroencephalogram (EEG) data, voice data, and facial expression data from suspected patients with mild depression. The data processing module is used to preprocess the data collected in the data acquisition module, extract the modal features of the data, and use the clinical-data dual-drive weighted fusion algorithm to fuse the modal features to generate a multimodal time series feature set. The key feature set acquisition module is used to construct a candidate set of time scales based on the temporal resolution relationship of the multimodal temporal feature sets, and to select the optimal time scale of the modal features from the candidate set of time scales using a three-dimensional optimized temporal scale selection algorithm; based on the optimal time scale, a set of temporal feature vectors is obtained, and the set of temporal feature vectors is corrected by a feature redundancy correction algorithm to obtain the multimodal key feature set; The feature vector acquisition module is used to divide the multimodal key feature set into a physiological signal supernode set and an emotional expression supernode set, construct an initial hypergraph based on clinical priors, and update the initial hypergraph according to time segments using dynamic constraint coefficients; then, a hypergraph convolutional neural network model is used to perform convolution operations on the updated hypergraph to obtain the predictive feature vector for mild depression; specifically: Features related to EEG and TENS data in the multimodal key feature set are categorized into the physiological signal supernode set, while features related to eye-tracking, speech, and facial expression data are categorized into the emotion expression supernode set. A feature association information database for mild depression in clinical diagnosis is collected to form clinical prior association rules. Using the physiological signal supernode set and the emotion expression supernode set as nodes, hyperedges are established according to the clinical prior association rules to obtain an initial hypergraph. Dynamic constraint coefficient The expression is: ; in, This represents the function that takes the maximum value. Represents the cosine similarity function. Indicates the initial hypergraph. This indicates a clinical prior hypergraph. Indicates the similarity threshold; The feature embedding matrices of the physiological signal supernode set and the emotion expression supernode set in the updated hypergraph are extracted respectively. Using a hypergraph convolutional neural network model, the matrices are sequentially subjected to inverse square / inverse operations on the hypergraph association matrix, hyperedge degree matrix, and supernode degree matrix, followed by matrix multiplication. The LeakyReLU activation function is then applied to obtain the feature representations of the physiological signal supernodes and the emotion expression supernodes. The specific formulas are as follows: ; in, The supernode feature representation at time step t. This represents the LeakyReLU activation function. Represents the degree matrix of supernodes. express The inverse square root matrix, Represents the hypergraph incidence matrix. Represents the hypermarginality matrix. express transpose, Indicates the convolution kernel weights; The feature representations of physiological signals and emotional expression are fused to obtain a fused feature representation. The dimensionality of the fused feature representation is then reduced to obtain a predictive feature vector for mild depression. The optimal solution acquisition module is used to analyze the key influencing dimensions of a patient's depressive state based on the predicted feature vector of mild depression. A multi-objective function is constructed with treatment efficiency, side effects, and cost as objectives. A non-dominated sorting genetic algorithm is used to solve the multi-objective function to generate a Pareto optimal solution set. Specifically: Step 1: Each individual in the initial population corresponds to a set of treatment plan parameters, including the frequency of psychotherapy, drug dosage, and intensity of physical therapy; the individual is a suspected patient with mild depression. Step 2: Calculate the multi-objective function value for each individual in the initial population; Step 3: Perform non-dominated sorting on the individuals in the initial population, and divide the individuals into different front layers according to their dominance relationship. Individuals in the same front layer do not dominate each other. Step 4: Obtain the distribution density of each individual in its frontal layer to obtain the crowding degree of the individual; Step 5: Based on the non-dominated ranking results and crowding, select individuals from the initial population to enter the next generation using the roulette wheel selection method; Step 6: Perform crossover on the individuals entering the next generation, with a crossover probability of 0.7-0.9, and generate new individuals through single-point crossover; Step 7: Perform mutation operation on the new individual, with the mutation probability set to 0.05-0.15, and generate mutated individuals by randomly changing the parameter values; Step 8: Merge the individuals in the initial population with the mutated individuals, and repeat steps 3-4 to perform non-dominated sorting and crowding calculation to obtain a new population; Step 9: Repeat steps 4-8. After iterating for 40-60 generations, extract individuals in the first frontier layer of the new population to form the Pareto optimal solution set.

2. The dynamic prediction system for mild depression based on deep learning of multimodal time series data according to claim 1, characterized in that, The data acquisition module is configured to perform the following actions: Eye-tracking data was obtained by recording fixation point data and pupil change data when the patient viewed pictures and read text using a 120-240Hz eye tracker; skin conduction level and response peak were collected using fingertip electrodes to obtain skin conduction response data; EEG data was obtained by placing electrodes according to international standards using a 64-channel EEG device to collect resting-state and task-state signals of the brain; speech data was obtained by recording audio of the patient reading aloud and speaking freely using a 48kHz microphone; and facial expression data was obtained by capturing the patient's facial dynamics and key point changes using a 25fps camera.

3. The dynamic prediction system for mild depression based on deep learning of multimodal time series data according to claim 2, characterized in that, The data processing module is configured to perform the following actions: For eye-tracking data, abnormal fixation points caused by device shake or blinking are removed, and missing data segments are filled in using interpolation; pupil change data is smoothed; and fixation point data is aligned with the presentation time of the corresponding stimulus based on timestamps. For the skin conductance response data, abnormal jump values ​​were removed, and baseline drift was filtered out using the sliding window method; skin conductance signals were extracted within the task period, and resting state and reactive state data segments were distinguished. For EEG data, interference from eye movement and electromyography was removed using independent component analysis; the 64-channel signal was grouped according to the international standard brain region division criteria, and data segments of resting state and task state were extracted and filtered. For speech data, environmental noise and silent segments are removed from the audio, and the data is stored according to the task type of reading aloud and free narration. The duration and location of pauses in the speech data are marked, and the audio sampling rate is standardized. For facial expression data, a facial key point detection algorithm is used to correct the tilted image caused by camera angle deviation and align the facial reference points; image noise is removed; facial dynamic changes in consecutive frames are extracted at a frame rate of 25fps, and the start and end frames of expression changes are marked. The specific formula for generating a multimodal temporal feature set is as follows: ; in, Represents the multimodal temporal characteristics at time step t. This represents the total number of modal features. Represents the balance coefficient. Represents the mutual information function. This represents the feature of the m-th modal feature at time step t. This represents clinically labeled data. Represents the similarity function. This represents the clinical standard value corresponding to the m-th modal feature.

4. The dynamic prediction system for mild depression based on deep learning of multimodal time series data according to claim 1, characterized in that, The key feature set acquisition module is configured to perform the following actions: The specific formula for obtaining the optimal time scale is: ; in, The optimal time scale represents the m-th modal feature. Represents the candidate set for the time scale. This represents the m-th modal feature on the time scale. Features This represents the time-scale penalty coefficient. This represents the clinical recommendation scale for the m-th modality feature. Represents the mutual information function. This represents clinically labeled data; The collection of time-series feature vectors includes: Based on the optimal time scale, the modal features are divided into several consecutive and equally long time windows. Within each time window, the basic statistics of the modal features are calculated, including the mean, variance, maximum, minimum, and median, to capture the central tendency and dispersion of the modal features at that time scale. For modal features with dynamic characteristics, dynamic statistics within the time window are additionally calculated, including the slope, rate of change, number of peaks, and duration. The basic and dynamic statistics obtained from all time windows are summarized to obtain a set of time-series feature vectors. Among them, the modal features with dynamic change characteristics are those with a normalized rate of change standard deviation of not less than 0.15, a slope mean of not less than 0.08, or a number of peaks of not less than 2 within the time window corresponding to the optimal time scale. The specific formula for obtaining the multimodal key feature set is as follows: ; in, This represents the key feature of the modified m-th modal feature. This represents the feature statistics extraction function. This represents the feature of the m-th modal feature at time step t. This represents the redundancy penalty coefficient. Represents an exponential function. This represents the function for calculating characteristic redundancy. This represents the feature of the m-th modal feature at the optimal time scale.

5. The dynamic prediction system for mild depression based on deep learning of multimodal time series data according to claim 1, characterized in that, Clinical prior association rules include internal association rules for physiological signals, internal association rules for emotional expression, and cross-domain association rules between physiological signals and emotional expression.

6. The dynamic prediction system for mild depression based on deep learning of multimodal time series data according to claim 1, characterized in that, The optimal solution acquisition module is configured to perform the following actions: The expression for the multi-objective function is: ; in, Represents a multi-objective function; Indicates the treatment plan; This represents the function that takes the maximum value. This represents a function that takes the minimum value. This represents the objective function for treatment efficiency. , This represents the predictive feature vector for mild depression. This represents the feature vector set for predicting mild depression. Indicates weight, Indicates responsiveness. Represents the exponentially decaying term. This represents the first correction factor. express right The adaptation distance; The objective function represents the side effects. , Indicates the type of side effect. Represents a set of side effect types. Indicates the risk value. This represents the second correction factor. Indicates sensitivity, This indicates a suspected case of mild depression. Indicates the exponential enhancement term; Represents the cost objective function, , Indicates the cost type. Represents a set of cost types. Indicates the cost value. This represents the third correction factor. Indicates the urgency of the treatment.

Citation Information

Patent Citations

  • Intelligent system and method for early screening of depression based on electroencephalogram-eye movement multi-modal data fusion

    CN120227030A

  • Multi-modal fusion-based depression classification method and system

    CN120656721A