Depression onset risk prediction method and system based on causal representation learning
By acquiring multi-source latent monitoring data and utilizing causal hierarchy partitioning and causal variational autoencoder networks, a causal link graph model is constructed, which solves the accuracy and reliability problems of predicting the risk of depression in existing technologies and achieves ultra-early and accurate prediction of the risk of depression.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 四川互慧软件有限公司
- Filing Date
- 2026-04-09
- Publication Date
- 2026-06-09
AI Technical Summary
Existing methods for predicting the risk of developing depression rely on patient subjective reports and simple psychological tests, which cannot accurately capture subtle changes in an individual's daily behavior and physiological state, and fail to effectively utilize causal relationships from multiple sources, resulting in low accuracy and reliability of predictions.
By acquiring multi-source latent monitoring data, extracting causal latent representation vectors using causal hierarchy partitioning rules and causal variational autoencoder networks, constructing a causal link graph model, and performing causal weighted fusion processing, the probability prediction results of depression incidence risk are generated.
It has achieved ultra-early and accurate prediction of the risk of developing depression, improved the accuracy and reliability of prediction, revealed the intrinsic causal mechanism of depression, and improved the effectiveness and efficiency of prevention and treatment.
Smart Images

Figure CN121983323B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart healthcare technology, and more specifically, to a method and system for predicting the risk of developing depression based on causal representation learning. Background Technology
[0002] Depression, a common mental disorder that seriously affects human physical and mental health and social functioning, has a complex and not fully understood pathogenesis. Traditional methods for predicting the risk of developing depression mainly rely on patient-reported symptoms, clinical interviews, and simple psychological assessment scales. However, these methods have many limitations. On the one hand, patients may provide inaccurate or incomplete information due to insufficient awareness of their own symptoms, subjective concealment, or limited expressive ability, thus affecting the reliability of the prediction results. On the other hand, traditional methods often only focus on current overt symptoms, making it difficult to capture potential early signs of depression before its onset, especially subtle changes hidden in an individual's daily behavior, physiological state, and environmental factors.
[0003] With the rapid development of information and sensor technologies, multi-source data collection has become possible. Wearable devices and smart terminals can continuously collect multi-source implicit monitoring data such as an individual's autonomic nervous function and touch behavior. This data, combined with environmental spatiotemporal big data and non-depression-related medical visit trajectory data from electronic health records, contains a wealth of information related to the onset of depression. However, current data utilization methods mostly involve simple statistical analysis or machine learning modeling, without fully considering the causal relationships between data points. This fails to effectively eliminate the bias of confounding variables and accurately reveal the intrinsic causal mechanisms of depression, resulting in low accuracy and reliability in predicting the risk of developing depression. Consequently, it cannot meet the practical needs for ultra-early and precise prevention and intervention of depression. Summary of the Invention
[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a method for predicting the risk of developing depression based on causal representation learning, the method comprising:
[0005] Acquire a multi-source latent monitoring data set of the target object, wherein the multi-source latent monitoring data set includes continuously collected nonlinear micro-feature data units of autonomic nervous function with time sequence identification, micro-feature data units of touch behavior of smart terminals, spatiotemporal big data units of environment, and non-depression-related medical trajectory data units in electronic health records;
[0006] According to the preset causal hierarchy classification rules, the multi-source latent monitoring data set is subjected to causal variable hierarchy labeling processing to obtain the exogenous exposure factor hierarchy label, endophenotypic abnormality mediating hierarchy label and depression onset outcome hierarchy label corresponding to each data unit in the multi-source latent monitoring data set.
[0007] The pre-built causal variational autoencoder network is invoked to perform causal latent representation extraction processing on the multi-source latent monitoring data set carrying the hierarchical label of the causal variable, generating a set of causal latent representation vectors that are causally associated with the onset of depression. Each vector dimension in the set of causal latent representation vectors corresponds to a causal feature unit after eliminating confounding variable bias.
[0008] Based on the set of causal latent representation vectors and the hierarchical labels of causal variables, a directed acyclic graph learning process with a penalty term is performed to construct a causal link graph model that includes the causal transmission path between the hierarchy of exogenous exposure factors, the hierarchy of mediating endophenotypic abnormalities, and the hierarchy of depressive onset outcomes.
[0009] Based on the causal effect weights corresponding to each causal feature unit in the causal link graph model, the real-time collected individual multi-source time-series monitoring data are subjected to causal weighted fusion processing to obtain an individual causal weighted fusion feature time-series matrix. The individual causal weighted fusion feature time-series matrix is then input into the early-onset risk assessment model for depression to perform probability prediction processing, generating a probability output result of depression onset risk containing prediction time window identifiers.
[0010] Furthermore, embodiments of the present invention also provide a system for predicting the risk of developing depression based on causal representation learning, comprising:
[0011] A processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to perform the above-described method for predicting the incidence risk of depression based on causal representation learning by executing the machine-executable instructions.
[0012] Based on the above, by acquiring a multi-source latent monitoring data set containing various types of time-series identifiers for the target object, various potential information related to the onset of depression was collected. Using a pre-defined causal hierarchy classification rule, the multi-source latent monitoring data was processed with causal variable hierarchical labeling, which defined the position of different data units in the causal chain of depression. A pre-constructed causal variational autoencoder network was invoked to extract a set of causal latent representation vectors, effectively eliminating confounding variable bias and accurately extracting features causally associated with depression, thus improving the effectiveness and reliability of the features. Based on the causal latent representation vector set and causal variable hierarchical labeling, a causal link graph model was constructed, intuitively presenting the causal transmission path between exogenous exposure factors, endophenotypic abnormalities, and the outcome of depression, deeply revealing the intrinsic causal mechanism of depression. Finally, based on the causal effect weights in the causal link graph model, the real-time collected individual multi-source time-series monitoring data was causally weighted and fused, and input into a depression ultra-early onset risk assessment model for probability prediction. This generated a depression onset risk probability output containing prediction time window identifiers, achieving ultra-early and accurate prediction of depression risk, greatly improving the effectiveness and efficiency of depression prevention and treatment. Attached Figure Description
[0013] Figure 1 This is a schematic diagram of the execution flow of the method for predicting the risk of depression based on causal representation learning provided in an embodiment of the present invention.
[0014] Figure 2 This is a schematic diagram of exemplary hardware and software components of the depression incidence risk prediction system based on causal representation learning provided in an embodiment of the present invention. Detailed Implementation
[0015] Figure 1 This is a flowchart illustrating a method for predicting the risk of developing depression based on causal representation learning, provided in one embodiment of the present invention. A detailed description follows.
[0016] Step S110: Obtain a multi-source latent monitoring data set of the target object. The multi-source latent monitoring data set includes continuously collected nonlinear micro-feature data units of autonomic nervous function with time-series identification, micro-feature data units of smart terminal touch behavior, environmental spatiotemporal big data units, and non-depression-related medical trajectory data units in electronic health records.
[0017] In this embodiment, the collection process of multi-source covert monitoring data is completed through compliant authorization. All data involving personal privacy is immediately de-identified after collection. Personal identification is anonymized using an irreversible hash algorithm. End-to-end encryption technology is used in all data storage and processing stages to prevent the leakage of sensitive information.
[0018] The nonlinear micro-feature data unit of autonomic nervous function is continuously collected through wearable physiological monitoring devices, including nonlinear derived features of physiological signals such as heart rate variability, skin conductance response, and respiratory rate fluctuations. Each data unit is accompanied by a time sequence identifier accurate to the second. The micro-feature data unit of smart terminal touch behavior is collected through the user-authorized smart terminal operating system interface, including operation behavior features such as touch pressure, swipe speed, click interval, and page dwell time. All data collection operations are only initiated after the user actively authorizes them and are only used for risk prediction calculations in this method. The environmental spatiotemporal big data unit is obtained through open and compliant environmental data interfaces, including public environmental features such as air quality, meteorological conditions, and regional population activity density corresponding to spatiotemporal location, without involving any individual location trajectory tracking. The non-depression-related medical visit trajectory data unit in electronic health records is obtained through authorized medical data platforms, including medical records from departments other than psychiatry, medication records, physical examination indicators, and other health data. All data are used after removing direct identity identifiers.
[0019] Step S120: Perform causal variable hierarchical labeling processing on the multi-source latent monitoring data set according to the preset causal hierarchy division rules to obtain the exogenous exposure factor hierarchical label, endophenotypic abnormality mediating hierarchical label, and depression onset outcome hierarchical label corresponding to each data unit in the multi-source latent monitoring data set.
[0020] Step S121: Parse the original acquisition source identifier and data structure type identifier of each data unit in the multi-source latent monitoring data set. Based on the original acquisition source identifier and data structure type identifier, classify the environmental spatiotemporal big data unit and the non-depression-related medical trajectory data unit in the electronic health record to the exogenous exposure factor level, and add an exogenous exposure factor level label to each data unit classified to the exogenous exposure factor level.
[0021] In this embodiment, the metadata segment of each data unit includes a predefined original data source identifier field and a data structure type identifier field. The original data source identifier field is used to mark the data generation channel, and the data structure type identifier field is used to mark whether the data is time-series continuous, discrete event-type, or static attribute-type. The two identifier fields of each data unit are read sequentially. When the original data source identifier corresponds to an environmental data interface or a medical data platform, the corresponding data unit is classified into the exogenous exposure factor level. This exogenous exposure factor level includes all variables that may have external impacts on an individual's physiological and psychological state. After classification, a fixed coded value corresponding to the exogenous exposure factor level is written to the level marker field of the data unit.
[0022] Step S122: Parse the original acquisition source identifier and data structure type identifier of each data unit in the multi-source latent monitoring data set. Based on the original acquisition source identifier and data structure type identifier, classify the nonlinear micro-feature data unit of autonomic nervous function and the micro-feature data unit of smart terminal touch behavior into the endophenotype abnormality mediating level, and add an endophenotype abnormality mediating level label to each data unit classified into the endophenotype abnormality mediating level.
[0023] In this embodiment, when the original data collection source identifier of a data unit corresponds to a wearable physiological monitoring device or a smart terminal operating system interface, the corresponding data unit is classified into the endophenotype abnormality mediating level. This endophenotype abnormality mediating level includes all intermediate variables that can reflect changes in an individual's internal physiological and behavioral states. It represents the intrinsic response characteristics generated after external exposure factors act on the individual and is also a direct mediating variable pointing to the subsequent outcome of depression. After classification, a fixed code value corresponding to the endophenotype abnormality mediating level is written into the level label field of the data unit. This code value is mutually exclusive with the code value of the exogenous exposure factor level to avoid level confusion in subsequent processing.
[0024] Step S123: Obtain the clinical diagnosis record of depression of the target object corresponding to the multi-source latent monitoring data set within the preset follow-up time window, generate a depression onset outcome hierarchical marker with binary status identifier based on the clinical diagnosis record of depression, and perform time-series association binding processing on the depression onset outcome hierarchical marker and the data unit at the corresponding time point in the multi-source latent monitoring data set.
[0025] In this embodiment, the length of the follow-up time window is set according to general standards for clinical research. For target subjects with a clear clinical diagnosis of depression, a positive depression outcome level label is generated; for target subjects without a record of a diagnosis of depression during the follow-up period, a negative depression outcome level label is generated. After the label is generated, all data units within the corresponding time range are associated and bound to the outcome label according to the time sequence identifier of each data unit, ensuring that the monitoring data of each time segment corresponds to a clear outcome label.
[0026] Step S124: Perform data integrity scanning on the multi-source latent monitoring data set after hierarchical labeling is completed, identify data units with missing hierarchical labels or conflicting hierarchical labels, and re-execute the hierarchical classification judgment operation according to the original acquisition source identifier and data structure type identifier of the data unit until all data units carry a unique and conflict-free hierarchical label.
[0027] In this embodiment, the integrity scan process sequentially checks whether the hierarchical label field of each data unit is empty and whether multiple hierarchical label values exist simultaneously. For units with missing labels, their original acquisition source identifier and data structure type identifier are reread, and the classification process in step S121 or S122 is repeated to supplement the label; for units with multiple conflicting labels, all existing labels are cleared and the classification judgment is re-executed to ensure that each data unit carries only one valid hierarchical label, avoiding label noise in the subsequent causal extraction process.
[0028] Step S125: Arrange and integrate the multi-source latent surveillance data set carrying complete hierarchical labels according to the time axis to generate a hierarchical structured time series dataset containing hierarchical data subsets of exogenous exposure factors, hierarchical data subsets of mediating endophenotypic abnormalities, and hierarchical labels of depression onset and outcome for each time point.
[0029] In this embodiment, all data units are sorted according to the order of their time series identifiers. Data units belonging to the exogenous exposure factor level at the same time point are grouped into an exogenous exposure factor level subset, and data units belonging to the endophenotypic abnormality mediating level are grouped into an endophenotypic abnormality mediating level subset. Both data subsets at each time point are associated with the corresponding time period's depression onset outcome level marker. In the integrated dataset, the entry structure at each time point remains consistent, facilitating batch processing by subsequent time series algorithms.
[0030] Step S126: Extract the correlation features between data units within each level in the hierarchical structured time series dataset, construct an initial causal path hypothesis set across levels based on the correlation features, and use the initial causal path hypothesis set as a priori constraint for subsequent causal latent representation extraction processing.
[0031] In this embodiment, the correlation characteristics are obtained by statistically analyzing the temporal synchronization fluctuations of different data units within the same level. When the numerical changes of two data units show a synchronous upward or downward trend at multiple consecutive time points, it is determined that the two units are correlated. Based on the correlation characteristics, initial causal path hypotheses are constructed, pointing from the exogenous exposure factor hierarchy to the endophenotype abnormality mediating hierarchy, and from the endophenotype abnormality mediating hierarchy to the depression onset outcome hierarchy. These hypotheses will serve as prior constraints for subsequent causal learning, narrowing the search space of the causal structure and improving learning efficiency and the reasonableness of the results.
[0032] Step S130: Call the pre-built causal variational autoencoder network to perform causal latent representation extraction processing on the multi-source latent monitoring data set carrying the hierarchical label of the causal variable, and generate a set of causal latent representation vectors that are causally associated with the onset of depression. Each vector dimension in the set of causal latent representation vectors corresponds to a causal feature unit after eliminating confounding variable bias.
[0033] Step S131: Input the multi-source latent monitoring data set carrying the hierarchical label of the causal variable into the encoder module of the causal variational autoencoder network, and perform dimensionality compression processing on the input data through the multi-layer nonlinear mapping network of the encoder module to generate initial latent representation vector distribution parameters, which include a mean vector and a log-variance vector.
[0034] In this embodiment, the encoder module of the causal variational autoencoder network consists of three stacked fully connected neural networks. Each layer is followed by a nonlinear activation function to perform nonlinear transformation on the high-dimensional raw monitoring data. The input data first enters the first fully connected network for preliminary feature transformation, outputting intermediate features with a dimension smaller than the input dimension. These intermediate features then enter the next two fully connected networks for further compression, ultimately outputting two vectors of the same dimension, which serve as the mean vector and log-variance vector of the initial latent representation vector distribution, respectively. The dimensions of these two vectors are consistent with the preset latent representation space dimension.
[0035] Step S132: Obtain a preset set of promiscuous variables, which includes age identifier, gender identifier, underlying disease identifier and season identifier. Perform one-hot encoding on the set of promiscuous variables to obtain a promiscuous variable encoding vector. Then, concatenate the promiscuous variable encoding vector with the initial latent representation vector distribution parameters to obtain extended latent representation distribution parameters carrying promiscuous information.
[0036] In this embodiment, confounding variables are variables known to simultaneously affect exposure factors, endophenotypic traits, and the risk of developing depression. Failure to control for these variables can lead to bias in causal effect estimation. First, each confounding variable is one-hot encoded, converting discrete category labels into vectors of fixed dimensions. The encoded vectors of all confounding variables are concatenated to form a unified confounding variable encoding vector. This confounding variable encoding vector is then concatenated along the feature dimensions with the mean vector and log-variance vector of the initial latent representation, respectively, to obtain extended latent representation distribution parameters that include information about the confounding variables.
[0037] Step S133: Call the distribution alignment module in the causal variational autoencoder network, and use the promiscuous variable encoding vector as the grouping basis to perform maximum mean difference distance calculation on the extended latent representation distribution parameters to generate distribution alignment loss values that represent the distribution differences of latent representations under different promiscuous levels.
[0038] In this embodiment, the distribution alignment module divides all samples into multiple different confounding groups based on the values of the confounding variable encoding vectors, such as grouping by age range or by gender. For each latent representation dimension, the maximum mean difference distance of the latent representation distribution under different confounding groups is calculated. This maximum mean difference distance is used to measure the degree of difference in latent representation distributions between different groups. The sum of the distance values of all dimensions yields the overall distribution alignment loss value. The larger the loss value, the greater the difference in latent representation distributions under different confounding groups, and the more severe the confounding bias.
[0039] Step S134: Based on the distribution alignment loss value, perform distribution alignment constraint optimization on the extended latent representation distribution parameters, and adjust the network weight parameters of the encoder module through the gradient backpropagation algorithm to make the latent representation distributions corresponding to different promiscuous levels gradually become consistent, thereby obtaining the causal latent representation vector distribution parameters that eliminate promiscuous variable bias.
[0040] In this embodiment, during the backpropagation phase of model training, the distribution alignment loss is used as part of the optimization objective. The weight parameters of the encoder module are adjusted using the gradient descent algorithm, so that the distribution differences of the latent representations under different profanity groups gradually decrease. When the distribution alignment loss is reduced to a preset range, the distributions of latent representations under different profanity levels basically overlap. At this point, the latent representations have eliminated the bias caused by profanity variables, and their changes are driven only by true causal relationships.
[0041] Step S135: Perform random sampling from the causal latent representation vector distribution parameters to obtain an initial causal latent representation vector sample. Input the initial causal latent representation vector sample into the decoder module of the causal variational autoencoder network for data reconstruction processing to generate a reconstructed data set that matches the dimension of the input data.
[0042] In this embodiment, the random sampling operation is implemented using a reparameterization technique. A sampling vector following a corresponding distribution is generated based on the mean vector and log-variance vector of the causal latent representation vector distribution. This sampling process is differentiable and supports gradient backpropagation. The initial causal latent representation vector samples obtained from the sampling are input to the decoder module. The decoder module consists of a three-layer fully connected neural network, whose network structure is symmetrical to the encoder module. After multiple layers of nonlinear mapping, it outputs reconstructed data with the same dimension as the original input data, used to verify whether the latent representation retains the key information of the original data.
[0043] Step S136: Calculate the reconstruction error loss value between the reconstructed data set and the original input data, and simultaneously calculate the KL divergence loss value between the causal latent representation vector distribution parameters and the standard normal prior distribution. Perform a weighted summation operation on the reconstruction error loss value, the KL divergence loss value, and the distribution alignment loss value to obtain the total optimized loss function value of the causal variational autoencoder network.
[0044] In this embodiment, the reconstruction error loss is obtained by calculating the mean square error between the reconstructed data and the original input data, and is used to measure the degree to which the latent representation retains the original information; the KL divergence loss is used to constrain the distribution of the latent representation to be as close as possible to the standard normal distribution, thereby improving the generalization ability of the latent representation; and the distribution alignment loss is used to eliminate confounding bias. The three loss values are multiplied by preset weighting coefficients and then summed to obtain the total optimized loss function value. The weighting coefficients are preset according to the importance of different loss terms and remain fixed during training.
[0045] Step S137: Based on the total optimization loss function value, perform joint iterative optimization training on the encoder module and decoder module of the causal variational autoencoder network until the total optimization loss function value converges to a preset threshold range, and extract the causal latent representation vector corresponding to each input data unit from the output layer of the trained encoder module.
[0046] Step S1371: Divide the multi-source latent monitoring data set into a training data subset and a validation data subset.
[0047] In this embodiment, all samples are randomly divided into a training data subset and a validation data subset according to a preset ratio. The training data subset is used for iterative updates of model parameters, while the validation data subset is used to evaluate the model's generalization ability and avoid overfitting. During the partitioning process, the ratio of positive to negative samples in both subsets is kept consistent with the original dataset to ensure the consistency of dataset distribution.
[0048] Step S1372: Initialize all trainable weight parameters of the encoder module and decoder module in the causal variational autoencoder network to random small values, and set the optimizer type identifier, initial learning rate parameter, and batch training sample quantity parameter.
[0049] In this embodiment, the weight parameters are initialized using the He initialization method, which sets all weights to random small values following a specific distribution, avoiding instability during training due to excessively large or small initial weights. The optimizer is an adaptive moment estimation optimizer, with the initial learning rate parameter pre-set according to the network size and the batch training sample size parameter set according to hardware computing power, ensuring efficient and stable training.
[0050] Step S1373: Randomly select a batch of multi-source latent monitoring data units carrying hierarchical labels from the training data subset, input the selected data units into the encoder module of the current iteration round for forward propagation calculation, and obtain the causal latent representation vector distribution parameters of the current batch.
[0051] In this embodiment, each iteration first randomly selects samples of a batch size from a subset of the training data to form the current training batch. The sample selection process uses sampling without replacement to ensure that each sample is selected only once in a training cycle. After the selected samples are input into the encoder module, they are sequentially processed through the encoder's three fully connected layers for forward propagation calculation, outputting the mean vector and log-variance vector of the latent representation distribution corresponding to each sample in the current batch.
[0052] Step S1374: Calculate the distribution alignment loss value, the reconstruction error loss value, and the KL divergence loss value based on the causal latent representation vector distribution parameters of the current batch, and calculate the total optimization loss function value of the current batch according to the preset weighting coefficient combination.
[0053] In this embodiment, based on the latent representation distribution parameters output by the current batch, the calculation process of distribution alignment loss, reconstruction error loss, and KL divergence loss is executed sequentially. The three loss values are multiplied by their corresponding weighting coefficients and then summed to obtain the total optimized loss function value for the current batch. Gradient information is retained in all intermediate results during the calculation process for subsequent backpropagation updates.
[0054] Step S1375: Based on the total optimization loss function value of the current batch, calculate the gradient value of each trainable weight parameter in the encoder module and decoder module through an automatic differentiation mechanism, and update all trainable weight parameters according to the calculated gradient value using a preset optimization algorithm.
[0055] In this embodiment, the automatic differentiation mechanism is implemented through the built-in function of the deep learning framework. The gradient of each weight parameter is calculated by backpropagation based on the total optimization loss function value. After the gradient calculation is completed, the adaptive moment estimation optimizer is used to update the value of each weight parameter according to the gradient value. The momentum mechanism is introduced during the update process to accelerate the convergence speed and suppress the fluctuation of the training process.
[0056] Step S1376: After completing a batch parameter update, input all data units in the verification data subset into the updated encoder module and decoder module for forward propagation calculation, calculate the total optimized loss function value corresponding to the verification data subset, and record the verification loss value.
[0057] In this embodiment, after the parameters of each batch are updated, the generalization ability of the current model is evaluated immediately using a subset of validation data. During the validation process, gradient calculation and parameter updates are not performed; only forward propagation is performed to obtain the total optimized loss function value of the validation set.
[0058] Step S1377: Determine whether the validation loss value of multiple consecutive iterations shows a downward trend or whether the fluctuation range exceeds the preset stability threshold. If the validation loss value continues to decrease or the fluctuation range exceeds the stability threshold, continue to perform batch extraction and forward propagation calculation operations for the next round of iteration training.
[0059] In this embodiment, after each complete training cycle iteration, the changes in the validation loss value of the most recent training cycles are statistically analyzed. If the validation loss value is still in a continuous decreasing state, or the fluctuation of the validation loss value in adjacent cycles exceeds the preset stability threshold, it indicates that the model has not yet converged, and the next round of training is continued.
[0060] Step S1378: If the verification loss value stops decreasing and the fluctuation amplitude is lower than the preset stability threshold, then the model training is determined to have reached the convergence state, the iterative optimization process is stopped, and the weight parameters of the encoder module and decoder module corresponding to the current iteration are saved as the model parameters after training is completed.
[0061] In this embodiment, when the validation loss value no longer decreases after multiple consecutive training cycles and the fluctuation range is lower than the preset stability threshold, it indicates that the model has fitted the core pattern of the data. Continuing to train will not improve the model performance and may even lead to overfitting. At this time, training is stopped and the current model weight parameters are saved for subsequent inference calculations.
[0062] Step S1379: Input all data units in the multi-source latent monitoring data set into the trained encoder module in sequence for forward propagation calculation, and extract the causal latent representation vector corresponding to each data unit from the output of the last hidden layer of the encoder module.
[0063] In this embodiment, after the model training is completed, all multi-source latent monitoring data units are input into the encoder module. After the encoder performs forward propagation calculation, the mean vector of the last layer output of the encoder is extracted as the causal latent representation vector corresponding to the data unit. There is no need to perform random sampling operations, which ensures that the extracted latent representation has determinism and consistency.
[0064] Step S13710: Perform dimensional standardization on the extracted causal latent representation vectors to generate a final causal latent representation vector set with a uniform scale representation. Each vector dimension in the final causal latent representation vector set corresponds to a causal feature unit after eliminating confounding variable bias.
[0065] In this embodiment, the dimensionality standardization process calculates the mean and standard deviation of all samples in each latent representation dimension. The mean of each sample in that dimension is subtracted from the mean and then divided by the standard deviation, so that the latent representations of all dimensions follow a distribution with a mean of zero and a standard deviation of one, eliminating scale differences between different dimensions and facilitating subsequent causal structure learning.
[0066] Step S138: Arrange and integrate all the extracted causal latent representation vectors according to time order and hierarchical label to generate a set of causal latent representation vectors containing temporal relationship and hierarchical attribution information. Each dimension unit of the causal latent representation vector set corresponds to a causal feature unit after eliminating confounding variable bias.
[0067] In this embodiment, the causal latent representation vectors are sorted according to the temporal identifier of the data units. Simultaneously, a hierarchical label corresponding to the original data unit is appended to each vector, ensuring that each vector clearly belongs to either the exogenous exposure factor hierarchy or the endophenotypic anomaly mediation hierarchy. In the integrated set of causal latent representation vectors, the value changes of each dimension are driven solely by the true causal relationship and are unaffected by confounding variables.
[0068] Step S140: Based on the set of causal latent representation vectors and the hierarchical labels of causal variables, perform directed acyclic graph learning with a penalty term to construct a causal link graph model that includes the causal transmission path between the exogenous exposure factor hierarchy, the endophenotypic abnormality mediation hierarchy, and the depression onset outcome hierarchy.
[0069] Step S141: The causal latent representation vector set is split into a latent representation matrix of exogenous exposure factors, a latent representation matrix of mediating endophenotypic abnormalities, and a latent representation vector of depression onset outcome according to the hierarchical labeling of the causal variables. The rows of the latent representation matrix of exogenous exposure factors correspond to time points and the columns correspond to the feature units of exposure factors. The rows of the latent representation matrix of mediating endophenotypic abnormalities correspond to time points and the columns correspond to the feature units of endophenotypic abnormalities.
[0070] In this embodiment, the splitting process is performed based on the hierarchical labeling of each causal latent representation vector. All latent representation vectors labeled as exogenous exposure factor level are stacked in chronological order to form an exogenous exposure factor latent representation matrix. The number of rows in the matrix equals the number of time points, and the number of columns equals the number of latent representation dimensions corresponding to the exogenous exposure factor. Similarly, latent representation vectors labeled as mediating endophenotype abnormalities are stacked to form an endophenotype abnormality mediating latent representation matrix. The latent representation vector of depression's onset outcome is obtained by arranging the depression onset outcome labels corresponding to all time points in chronological order, with each element corresponding to the outcome state at a given time point.
[0071] Step S142: Perform time-lag processing on the latent characterization matrix of the exogenous exposure factors and the latent characterization matrix of the mediating endophenotypic abnormalities to generate an extended feature matrix containing the current time point and multiple lag time points. The extended feature matrix is used to capture the delayed causal effects of exposure factors and endophenotypic abnormalities on the disease outcome.
[0072] In this embodiment, the time-series lag processing, for each feature dimension, concatenates the value of the current time point of that dimension with the values of multiple preceding lag time points to form an extended feature containing time-series lag information. For example, if the preset number of lag time points is 3, then the extended feature of each feature dimension includes the values of the current time point, the previous 1 time point, the previous 2 time points, and the previous 3 time points. The extended features of all dimensions are concatenated to form an extended feature matrix. The number of columns in the matrix is the original number of feature dimensions multiplied by the number of lag time points plus one, which can capture delayed causal effects between variables and avoid missing causal relationships with long time spans.
[0073] Step S143: Initialize an adjacency matrix with a dimension equal to the sum of the total number of all causal feature units and the total number of disease outcome units. Each element in the adjacency matrix represents whether there is a causal edge between two corresponding variables and the direction of the causal edge. All diagonal elements are set to zero to ensure that there are no self-loops.
[0074] In this embodiment, each row and column of the adjacency matrix corresponds to a variable. These variables include the latent representation dimensions of all exogenous exposure factors, the latent representation dimensions of all endophenotypic abnormalities' mediating latent representation dimensions, and the outcome variable of depression. Therefore, the dimension of the adjacency matrix equals the total number of these three types of variables. Each element of the matrix is a real number; a positive number indicates a positive causal effect between the row variable and the column variable, a negative number indicates a negative causal effect, and zero indicates no causal relationship. All diagonal elements are set to zero to exclude self-loop relationships between variables, which conforms to the basic definition of causality.
[0075] Step S144: Construct a smooth linear structural equation model with the adjacency matrix as the variable. The smooth linear structural equation model represents the value of each variable as a linear combination of other variables plus a noise term. The coefficients of the linear combination are determined by the non-zero elements of the corresponding rows of the adjacency matrix.
[0076] In this embodiment, the smooth linear structural equation model assumes that the value of each outcome variable is composed of a linear combination of all causal variables pointing to it, plus an independent noise term. The coefficients of the linear combination are the weights of the corresponding edges in the adjacency matrix. The model uses a smoothed continuous representation, which allows the optimization process of the adjacency matrix to be implemented using the gradient descent algorithm, avoiding the combinatorial explosion problem of traditional discrete graph structure search.
[0077] Step S145: Substitute the extended feature matrix and the latent representation vector of the depression outcome into the smooth linear structural equation model, calculate the sum of squared residuals between the predicted and actual values of each variable using the least squares method, and sum the sums of squared residuals of all variables to obtain the model fitting loss function value.
[0078] In this embodiment, firstly, based on the current adjacency matrix, the predicted value of each variable is calculated according to the linear structural equation model. The predicted value is the weighted sum of all other variables pointing to that variable, with the weights being the values of the corresponding elements in the adjacency matrix. Then, the residual between the predicted value and the actual observed value of each variable is calculated. The sum of the squares of the residuals is the fitting error of that variable. The sum of the fitting errors of all variables yields the overall model fitting loss function value. The smaller the model fitting loss function value, the better the model fits the observed data.
[0079] Step S146: Add a sparsity penalty term and a directed acyclic constraint term based on the adjacency matrix to the model fitting loss function value. The sparsity penalty term is used to control the number of causal edges to prevent overfitting, and the directed acyclic constraint term is used to ensure that the learned graph structure does not have directed cycles.
[0080] In this embodiment, the sparsity penalty term is implemented by calculating the sum of the absolute values of all non-zero elements in the adjacency matrix. The larger the proportion of this sparsity penalty term in the total loss function, the fewer causal edges are learned, which can effectively avoid the model learning too many spurious causal relationships and prevent overfitting. The directed acyclic constraint term is constructed based on the trace operation of the matrix, which can transform the acyclic constraint of the graph structure into a continuously differentiable numerical constraint, ensuring that the graph structure corresponding to the optimized adjacency matrix does not have directed cycles, which conforms to the basic properties of causality.
[0081] Step S147: The augmented Lagrange method combined with the gradient descent algorithm is used to iteratively optimize the joint objective function with added penalty and constraint terms. In each iteration, the element values in the adjacency matrix are updated, and the Lagrange multipliers and penalty coefficients are updated to force the satisfaction of the directed acyclic constraint.
[0082] In this embodiment, the joint objective function is obtained by weighted summation of the model fitting loss, the sparsity penalty term, and the directed acyclic constraint term. An augmented Lagrange method is used to handle the directed acyclic equality constraints. By iteratively updating the values of the adjacency matrix, Lagrange multipliers, and penalty coefficients, the optimal solution satisfying the constraints is gradually approximated. Each iteration first fixes the Lagrange multipliers and penalty coefficients, updates the elements of the adjacency matrix using gradient descent, and then updates the Lagrange multipliers and penalty coefficients based on the current constraint satisfaction status of the adjacency matrix, ensuring that the directed acyclic constraints are gradually satisfied.
[0083] Step S148: During the iterative optimization process, a hierarchical structure constraint is applied to the adjacency matrix according to the hierarchical label of the causal variable, prohibiting causal edges from the mediating level of endophenotypic abnormality to the level of exogenous exposure factors, and prohibiting causal edges from the depression onset outcome level to any other level, ensuring that the causal direction conforms to the preset hierarchical order.
[0084] In this embodiment, after each gradient descent update of the adjacency matrix, a hierarchical structure constraint check is performed. Elements in the adjacency matrix that point to exogenous exposure factor hierarchical variables corresponding to endophenotype abnormality mediating hierarchical variables are forcibly set to zero. At the same time, elements that point to any other hierarchical variable corresponding to depression outcome are also forcibly set to zero. This ensures that the causal direction can only be exogenous exposure factor hierarchical pointing to endophenotype abnormality mediating hierarchical or depression outcome hierarchical, and endophenotype abnormality mediating hierarchical pointing to depression outcome hierarchical, which conforms to the preset causal hierarchy logic and avoids causal directions that do not conform to common sense.
[0085] Step S149: When the iterative optimization reaches the maximum number of iterations or the change in the adjacency matrix is lower than the preset convergence threshold, the optimization process is stopped, and the final adjacency matrix is thresholded. Adjacency matrix elements with absolute values lower than the preset significance threshold are set to zero, and elements with absolute values higher than the threshold are retained as valid causal edges.
[0086] In this embodiment, a maximum number of iterations is preset as one of the termination conditions to prevent the optimization process from running indefinitely. A convergence threshold is also set; when the sum of the absolute values of the differences between elements in the adjacency matrix obtained from two consecutive iterations is lower than this threshold, it indicates that the adjacency matrix has stabilized, and optimization stops. After optimization stops, the saliency of each element in the adjacency matrix is assessed. Elements with absolute values lower than the preset saliency threshold are set to zero, and only causal edges with sufficiently large weights are retained, filtering out weak spurious edges caused by noise.
[0087] Step S1410: Construct a causal link graph model containing variable nodes and directed edges based on the adjacency matrix after thresholding. Each node in the causal link graph model corresponds to a causal feature unit or disease outcome unit. Each directed edge points from the cause node to the result node. The weight of the edge corresponds to the element value in the adjacency matrix. The causal link graph model fully presents the causal transmission path from the exogenous exposure factor level through the endophenotypic abnormality mediation level to the depression disease outcome level.
[0088] In this embodiment, when constructing the causal link graph model, firstly, a corresponding node is created for each variable. The node attributes include information such as the variable's hierarchical affiliation and the meaning of its corresponding original feature. Then, the thresholded adjacency matrix is traversed, and for each non-zero element, a directed edge is created from the node corresponding to the row variable to the node corresponding to the column variable. The edge attributes include information such as weight value and causal effect direction. The completed causal link graph model can clearly show the causal relationship between all causal features, as well as the complete transmission path from external exposure to disease outcome.
[0089] Step S150: Perform causal weighted fusion processing on the real-time collected individual multi-source time-series monitoring data according to the causal effect weight corresponding to each causal feature unit in the causal link graph model to obtain the individual causal weighted fusion feature time-series matrix, and input the individual causal weighted fusion feature time-series matrix into the early onset risk assessment model of depression for onset probability prediction processing to generate an output result of the onset risk probability of depression containing the prediction time window identifier.
[0090] Step S151: Obtain individual multi-source time-series monitoring data of the target object in real time during multiple consecutive collection cycles. The individual multi-source time-series monitoring data includes real-time sequences of nonlinear micro-features of autonomic nervous function corresponding to nodes in the causal link graph model, real-time sequences of micro-features of smart terminal touch behavior, real-time units of environmental spatiotemporal big data, and real-time records of non-depression-related medical trajectories in electronic health records.
[0091] In this embodiment, the acquisition channels and data types of the real-time individual multi-source time-series monitoring data are completely consistent with the multi-source implicit monitoring data during the model training phase. Each data unit also comes with a precise time-series identifier, ensuring the consistency of the distribution between real-time data and training data. Privacy protection processing is also performed during the acquisition process. All sensitive data is anonymized and encrypted at the acquisition end before being transmitted to the processing module to avoid data leakage.
[0092] Step S152: Perform the same causal latent representation extraction process as in the model training stage on each feature real-time sequence in the individual multi-source time-series monitoring data, call the encoder module of the trained causal variational autoencoder network to map each feature real-time sequence into an individual causal latent representation real-time vector, and obtain the individual causal latent representation real-time matrix.
[0093] In this embodiment, the causal latent representation extraction process of real-time data is completely consistent with the training phase. First, the real-time data is preprocessed in the same way as in the training phase. Then, it is input into the encoder module of the trained causal variational autoencoder. After forward propagation calculation, the corresponding causal latent representation vector is extracted. All vectors are stacked in chronological order to form an individual causal latent representation real-time matrix. The rows of the matrix correspond to the collection time point, and the columns correspond to the causal feature dimensions.
[0094] Step S153: Extract the total causal effect value corresponding to each causal feature unit from the causal link graph model, and normalize the total causal effect value to obtain the causal fusion weight coefficient of each causal feature unit. The magnitude of the causal fusion weight coefficient is positively correlated with the total causal effect value.
[0095] In this embodiment, the total causal effect value refers to the overall influence of each causal feature unit on the outcome of depression, including the sum of direct and all indirect effects. The total causal effect values of all feature units are normalized so that the sum of all weight coefficients is one. Features with larger total causal effect values have larger fusion weight coefficients, ensuring that features with a greater impact on the risk of developing depression occupy a higher proportion in the subsequent fusion process.
[0096] Step S154: Multiply each column vector in the real-time matrix of individual causal latent representation with the causal fusion weight coefficient of the corresponding causal feature unit element by element to obtain the weighted real-time matrix of individual causal latent representation. Perform in-row concatenation operation on the weighted matrix according to time points to generate the individual causal weighted fusion feature vector corresponding to each time point.
[0097] In this embodiment, the element-wise multiplication operation refers to multiplying the values of all time points for each feature dimension by the causal fusion weight coefficient corresponding to that dimension, so that the weighted feature value reflects the causal importance of the feature. After weighting, the values of all feature dimensions corresponding to each time point are concatenated into a one-dimensional vector, which is the individual causal weighted fusion feature vector for that time point. The dimension of the vector is equal to the total number of causal features.
[0098] Step S155: Stack the individual causal weighted fusion feature vectors of each time point in chronological order to construct an individual causal weighted fusion feature time series matrix containing time dimension information and feature dimension information. The row index of the individual causal weighted fusion feature time series matrix corresponds to the continuous collection time point, and the column index corresponds to the causal feature dimension after weighted fusion.
[0099] In this embodiment, the stacking operation is performed in chronological order. The feature vector of the earliest collected time point is used as the first row of the matrix, and the feature vector of the latest collected time point is used as the last row of the matrix. The constructed time series matrix fully preserves the temporal evolution information of the features, which is convenient for subsequent time series prediction models to capture the dynamic change patterns of the features.
[0100] Step S156: Input the individual causal weighted fusion feature time series matrix into the pre-trained early-onset risk assessment model for depression. The early-onset risk assessment model for depression includes a bidirectional long short-term memory network layer and a Cox proportional hazards output layer. The bidirectional long short-term memory network layer performs forward and backward contextual information capture processing on the time series matrix to extract the dynamic evolution pattern vector of individual causal features.
[0101] In this embodiment, the input layer dimension of the early-onset risk assessment model for depression is perfectly matched with the dimension of the individual causal weighted fusion feature vector. The input temporal matrix first enters the bidirectional long short-term memory network layer, which contains forward long short-term memory units and backward long short-term memory units. The forward unit processes temporal data in chronological order from front to back, capturing the influence of historical features on the current state; the backward unit processes temporal data in reverse chronological order from back to front, capturing the implicit influence of future features on the current state. The outputs of the two units are concatenated to obtain the context-aware features at each time point. The context-aware features at the last time point are the dynamic evolution pattern vector of the individual causal features, which fully encodes the temporal evolution law of the features.
[0102] Step S157: Input the dynamic evolution pattern vector into the Cox proportional hazards output layer. The Cox proportional hazards output layer calculates the cumulative morbidity function value of an individual within multiple preset future time windows based on a linear combination of a preset baseline morbidity function and the dynamic evolution pattern vector.
[0103] In this embodiment, the Cox proportional hazards output layer is constructed based on survival analysis theory. After inputting the dynamic evolution pattern vector, it first obtains the risk coefficient through linear transformation, and then combines the risk coefficient with the preset baseline risk function to calculate the instantaneous risk of disease in an individual at different future time points. The instantaneous risk is integrated on the time axis to obtain the cumulative risk function value of disease in different time windows. The cumulative risk function value of disease reflects the cumulative risk level of disease in an individual within the corresponding time window.
[0104] Step S158: Perform probability transformation processing on the cumulative incidence risk function value to generate the individual's incidence probability value of depression in a future preset first time window, the incidence probability value of depression in a future preset second time window, and the incidence probability value of depression in a future preset third time window, thus forming an output result of the incidence probability of depression containing time window identifiers.
[0105] In this embodiment, the probability transformation process converts the cumulative incidence risk function value into a probability value between 0 and 1. The three preset time windows correspond to different prediction spans, and each probability value is accompanied by a corresponding time window identifier, so that the output results can clearly show the incidence risk level in different future time periods and meet the prediction needs of different scenarios.
[0106] Step S159: Compare the output result of the probability of developing depression with the preset probability division threshold, and generate the corresponding probability interval identifier based on the comparison result. The probability interval identifier includes a first probability interval identifier, a second probability interval identifier, a third probability interval identifier, and a fourth probability interval identifier.
[0107] In this embodiment, the preset risk probability classification threshold is set based on clinical guidelines and large-scale sample statistical results, dividing the probability range of 0 to 1 into four intervals, each corresponding to a different risk level. After each probability value is compared with the threshold, it is matched with the corresponding risk probability interval identifier, which makes it easy for non-professionals to quickly understand the risk level and facilitates subsequent risk stratification management.
[0108] Step S210: Perform causal effect value quantification on each directed edge in the causal link graph model, and extract the adjacency matrix element value corresponding to each directed edge as the direct causal effect coefficient. The direct causal effect coefficient represents the expected change of the result variable when the cause variable changes by one unit.
[0109] In this embodiment, the direct causal effect coefficient of each directed edge is directly taken from the value of the corresponding element in the adjacency matrix. The sign of the coefficient indicates the direction of the causal effect, and the absolute value indicates the magnitude of the effect. It directly reflects the degree of direct influence of the cause variable on the result variable and is the basis for subsequent calculation of indirect effects and total effects.
[0110] Step S220: For a multi-step causal path starting from the exogenous exposure factor level node, passing through the endophenotypic abnormality mediating level node, and reaching the depression pathogenesis outcome level node, the indirect causal effect value of each multi-step path is calculated using the path coefficient product method. The indirect effect value of the path is obtained by multiplying all the direct causal effect coefficients on the path.
[0111] In this embodiment, the indirect effect of a multi-step causal path is obtained by multiplying the direct causal effect coefficients of all edges on the path. For example, for a path containing two edges, the indirect effect is equal to the coefficient of the first edge multiplied by the coefficient of the second edge. This value reflects the degree of indirect influence of exogenous exposure factors on the disease outcome through mediating variables, and can identify transmission paths hidden outside the direct effects.
[0112] Step S230: Summate all indirect causal effect values that pass through the same endophenotype abnormality mediation node to obtain the total mediation effect of the endophenotype abnormality mediation node between the exposure factor and the disease outcome. The total mediation effect quantifies the mediating strength of the endophenotype abnormality in the causal transmission process.
[0113] In this embodiment, each mediating node of endophenotypic abnormality may belong to multiple indirect causal paths. The total mediating effect of the node is obtained by summing all the indirect effect values that pass through the node. The larger the value, the more critical the role of the mediating node is in the transmission process from external exposure to disease outcome, and it is a potential important intervention target.
[0114] Step S240: Calculate the total causal effect for each exogenous exposure factor node in the causal link graph model. Add the direct causal effect value of the node pointing to the depression outcome level node to the total causal effect value through all mediating paths to obtain the total causal effect value of the exposure factor node on the outcome.
[0115] In this embodiment, the total causal effect is the sum of the direct effects and all indirect effects, comprehensively reflecting the overall impact of each exogenous exposure factor on the outcome of depression. It is the core basis for subsequent feature weighting and risk factor ranking.
[0116] Step S250: Sort the nodes in the causal link graph model according to the total causal effect value corresponding to each causal feature unit, and generate a core risk driver ranking list. The core risk driver ranking list is used to identify the key feature units that contribute the most to the onset of depression.
[0117] In this embodiment, all causal feature units are sorted from largest to smallest according to the total causal effect value. The features ranked at the top are the core risk drivers that contribute the most to the risk of disease. This list of core risk drivers can provide clear target directions for clinical research and intervention program development.
[0118] Step S260: Perform stability assessment on the causal link graph model. Use a bootstrap resampling method to extract multiple sample subsets from the original data. Repeat the directed acyclic graph learning operation with a penalty term on each sample subset to obtain multiple adjacency matrix samples.
[0119] In this embodiment, the self-service resampling method extracts multiple sample subsets of the same size as the original dataset from the original dataset through sampling with replacement. Each sample subset independently performs causal structure learning to obtain corresponding adjacency matrix samples. Multiple adjacency matrix samples can reflect the fluctuation of the causal structure learning results and are used to evaluate the stability of the original causal link graph model.
[0120] Step S270: Calculate the frequency of occurrence and standard deviation of weight coefficient of each causal edge in multiple adjacency matrix samples, mark the causal edges with a frequency lower than a preset frequency threshold or a standard deviation of weight coefficient higher than a preset stability threshold as unstable edges, and remove the unstable edges from the causal link graph model.
[0121] In this embodiment, the frequency of occurrence of a causal edge refers to the proportion of non-zero values of that edge in multiple adjacency matrix samples. The lower the frequency, the more unstable the existence of the edge. The larger the standard deviation of the weight coefficient, the greater the fluctuation of the effect value of the edge, and the lower the reliability. Edges that simultaneously satisfy or individually satisfy both removal conditions are marked as unstable edges and removed, while only highly stable causal edges are retained, thereby improving the reliability of the causal link graph model.
[0122] Step S280: Simplify and optimize the causal link graph model after removing unstable edges, retain the top core nodes whose cumulative contribution rate of causal effect value reaches the preset ratio and the causal edges between nodes, and generate a simplified core causal link graph model.
[0123] In this embodiment, all retained nodes are sorted from largest to smallest according to their total causal effect value, and their effect values are accumulated sequentially. When the accumulated value reaches a preset proportion of the total effect value of all nodes, the accumulation stops, these nodes and the causal edges between them are retained, and the remaining nodes and related edges are deleted. The simplified model retains only the core causal relationships that contribute the most, reducing the complexity of subsequent calculations and avoiding interference from redundant information.
[0124] Step S290: The causal feature unit identifier, node hierarchy, and causal effect coefficient between nodes of each node in the simplified core causal link graph model are stored in a structured manner to construct an interpretable knowledge base for the causal causes of depression.
[0125] In this embodiment, the structured storage is implemented using a graph database, where the attributes of each node and edge are stored as independent fields, supporting quick querying of the hierarchical affiliation, causal effects, and transmission paths between any two nodes.
[0126] Step S310: Construct an early onset risk assessment model for initial depression. The early onset risk assessment model for initial depression consists of an input layer, a bidirectional long short-term memory network layer, a fully connected mapping layer, and a Cox proportional hazards output layer connected in sequence. The dimension of the input layer matches the dimension of the individual causal weighted fusion feature vector. The bidirectional long short-term memory network layer includes forward long short-term memory units and backward long short-term memory units.
[0127] In this embodiment, the dimension of the input layer is set according to the total number of causal features to ensure that it can directly receive the temporal matrix of individual causal weighted fusion features as input. The dimension of the hidden layer of the bidirectional long short-term memory network is preset according to the complexity of the temporal data. The forward and backward hidden states of each time step are concatenated and output to the fully connected mapping layer. The dimension of the fully connected mapping layer is lower than the output dimension of the bidirectional long short-term memory network layer, and it is used to compress and nonlinearly transform the temporal features. Finally, it is connected to the Cox proportional hazards output layer to output the risk prediction result.
[0128] Step S320: Obtain the sample queue training dataset, which includes the causal weighted fusion feature time series matrix of multiple diagnosed depression patients at multiple consecutive time points before diagnosis and the corresponding diagnosis time point identifier, and also includes the causal weighted fusion feature time series matrix of multiple healthy control individuals within the same follow-up period and the undiagnosed identifier.
[0129] In this embodiment, the training dataset for the sample cohort comes from a large-scale prospective follow-up cohort. All samples have been screened using strict inclusion and exclusion criteria. The diagnosis time point of the patient samples is recorded accurately to the month. The follow-up duration of the healthy control samples is consistent with the monitoring duration before diagnosis of the patient samples, ensuring that the time series lengths of the two groups of samples match and avoiding model bias caused by differences in time series lengths.
[0130] Step S330: Take the causal weighted fusion feature time series matrix of each individual in the sample queue training dataset as input, and take the diagnosis time point identifier or non-diagnosis identifier of the corresponding individual as supervision label to construct a model training sample pair set.
[0131] In this embodiment, each sample pair contains a time series matrix and a corresponding survival label. The survival label consists of two fields: event status and survival time. Event status indicates whether depression was diagnosed during the follow-up period, and survival time indicates the length of time from the start of follow-up to the occurrence of the event or the end of follow-up. The label format conforms to the training requirements of the Cox proportional hazards model.
[0132] Step S340: Input the time series matrix in the model training sample pair set into the initial early-onset risk assessment model of depression in sequence. Pass the time series matrix to the bidirectional long short-term memory network layer through the input layer. The bidirectional long short-term memory network layer fuses the forward hidden state and the backward hidden state at each time step to generate the context-aware feature vector of each time step.
[0133] In this embodiment, the forward long short-term memory unit processes each row of the timing matrix sequentially from the first time step to the last time step, and outputs the corresponding forward hidden state at each time step; the backward long short-term memory unit processes the timing matrix in reverse from the last time step to the first time step, and outputs the corresponding backward hidden state at each time step; the two hidden states are concatenated at each time step to obtain the context-aware feature vector of that time step, which contains both the timing information before and after that time point.
[0134] Step S350: Input the context-aware feature vector of the last time step into the fully connected mapping layer, and generate a compressed temporal evolution pattern feature vector through the linear transformation and nonlinear activation function of the fully connected mapping layer. The dimension of the temporal evolution pattern feature vector is smaller than the dimension of the original context-aware feature vector.
[0135] In this embodiment, the fully connected mapping layer first performs a linear transformation on the context-aware feature vector of the last time step, compressing its dimension to a preset smaller dimension. Then, it transforms it through a nonlinear activation function, introducing nonlinear expressive power. The resulting temporal evolution pattern feature vector can more compactly represent the dynamic evolution law of individual features, reducing the computational complexity of subsequent output layers.
[0136] Step S360: Input the feature vector of the time-series evolution pattern into the Cox proportional hazards output layer. The Cox proportional hazards output layer calculates the incidence risk function value of each individual at each time point during the observation period according to the preset baseline risk function. Integrate the incidence risk function value on the time axis to obtain the cumulative risk function value.
[0137] In this embodiment, the baseline risk function is pre-fitted using survival data from a large population and does not change with individual characteristics, but is only related to time. The feature vector of the time-series evolution pattern is linearly transformed to obtain the individual's risk coefficient. The risk coefficient is multiplied by the baseline risk function to obtain the individual's instantaneous risk of disease at each time point. The instantaneous risk is integrated over time to obtain the cumulative risk function value, which reflects the individual's cumulative risk of disease during the observation period.
[0138] Step S370: Construct a partial likelihood loss function based on the cumulative risk function value of each individual and the actual time point of diagnosis or the undiagnosed status. The partial likelihood loss function is used to measure the degree of agreement between the risk ranking predicted by the model and the actual situation. Individuals with correct risk function value ranking are given a positive contribution, while individuals with incorrect ranking are given a negative contribution.
[0139] In this embodiment, the partial likelihood loss function is constructed based on the partial likelihood theory of survival analysis. For all individuals who have an event at the same time point, their predicted risk values are compared and ranked with the risk values of other individuals who have not had an event before that time point. The loss value of the correctly ranked samples is lower, and the loss value of the incorrectly ranked samples is higher. This loss function can effectively handle the censoring problem in survival data, that is, the label information of undiagnosed healthy control samples.
[0140] Step S380: Minimize the partial likelihood loss function using the gradient descent optimization algorithm. Calculate the gradients of all trainable parameters in the bidirectional long short-term memory network layer, the fully connected mapping layer, and the Cox proportional risk output layer through backpropagation. Update the parameter values based on the gradients and iteratively perform forward and backward propagation operations until the loss function converges.
[0141] In this embodiment, the optimization algorithm uses an adaptive moment estimation optimizer. In each iteration, forward propagation is performed once to calculate the loss value, and then backpropagation is performed to calculate the gradient of all trainable parameters. The parameter values are updated according to the gradient. During the iteration process, the loss value gradually decreases. When the change in the loss value is lower than a preset threshold, the model is determined to have converged.
[0142] Step S390: Introduce an early stopping mechanism during model training. Divide the sample queue training dataset into a training subset and a validation subset. After each training round, calculate the partial likelihood loss value on the validation subset. Stop training when the partial likelihood loss value on the validation subset no longer decreases for several consecutive rounds. Save the model parameters of the current round as a pre-trained risk assessment model for the very early onset of depression.
[0143] In this embodiment, the early stopping mechanism is used to prevent model overfitting. The sample queue training dataset is divided into a training subset and a validation subset according to a preset ratio. The training subset is used for parameter updates, and the validation subset is used to evaluate the model's generalization ability. When the validation set loss value no longer decreases after several consecutive training epochs, it indicates that the model has shown an overfitting trend. At this point, training is stopped, the current model parameters are saved, and the model is ensured to have optimal generalization ability.
[0144] Step S410: Obtain the individual causal weighted fusion feature time series matrix of the target object within the past continuous preset historical time window range, and input the individual causal weighted fusion feature time series matrix within the past continuous preset historical time window range into the pre-constructed dynamic abnormal trajectory recognition module for abnormal trajectory matching processing.
[0145] In this embodiment, the length of the preset historical time window is set based on the clinical research results of the prodromal period of depression to ensure that it can cover the period of characteristic changes in the prodromal period. All causal weighted fusion feature vectors of the target object within this time period are extracted, stacked to form a time series matrix, and input into the dynamic abnormal trajectory recognition module for prodromal depression for matching processing to identify whether there are abnormal characteristic changes in the prodromal period.
[0146] Step S420: The dynamic abnormal trajectory identification module for the prodromal phase of depression has a pre-stored template library of multiple prodromal abnormal trajectory subtypes obtained by clustering sample queue data. The abnormal trajectory subtype template library includes a first prodromal abnormal trajectory template, a second prodromal abnormal trajectory template, and a third prodromal abnormal trajectory template. Each abnormal trajectory template includes the mean sequence of causal features and the abnormal threshold sequence of the corresponding subtype at each time point before the onset of the disease.
[0147] In this embodiment, the prodromal abnormal trajectory subtype template library is obtained by clustering analysis of the characteristic time series data of a large number of confirmed patients before the onset of the disease. Each subtype corresponds to a different prodromal characteristic evolution mode. The mean sequence in the template is the characteristic mean of all samples of the subtype at each time point, and the abnormal threshold sequence is the normal fluctuation range boundary of the subtype characteristics, which is used to determine whether the individual characteristics conform to the abnormal pattern of the subtype.
[0148] Step S430: Calculate the similarity distance metric between the individual causal weighted fusion feature time series matrix and each abnormal trajectory template using the dynamic time warping algorithm, and simultaneously calculate the similarity distance metric between the individual causal weighted fusion feature time series matrix and the normal fluctuation baseline template of healthy people.
[0149] In this embodiment, the dynamic time warping algorithm can handle the problem of inconsistent lengths between two time series. By flexibly adjusting the alignment of the time axes, it calculates the optimal matching distance between the two sequences. The smaller the optimal matching distance, the more similar the evolution patterns of the two sequences are. The similarity distance between the individual time series matrix and each abnormal trajectory template and healthy baseline template is calculated separately.
[0150] Step S440: Subtract the similarity distance metric between the individual causal weighted fusion feature time series matrix and the abnormal trajectory template from the minimum similarity distance metric between the individual and the normal fluctuation baseline template of the healthy population to obtain the individual abnormal trajectory deviation score. The individual abnormal trajectory deviation score is used to quantify the degree to which the individual time series features deviate from the normal range.
[0151] In this embodiment, the abnormal trajectory deviation score reflects the degree of proximity of an individual's characteristics to the most similar abnormal trajectory minus the degree of proximity to the normal baseline. The higher the score, the closer the individual's characteristics are to the abnormal trajectory pattern and the greater the degree of deviation from the normal range. It is the core indicator for determining whether the individual is in the prodromal period.
[0152] Step S450: If the individual abnormal trajectory deviation score exceeds the preset prodromal abnormal judgment threshold, the target object is determined to be in the abnormal trajectory state of the prodromal depression stage, and the subtype identifier corresponding to the abnormal trajectory template with the smallest similarity distance to the individual causal weighted fusion feature time series matrix is identified as the matched abnormal trajectory subtype.
[0153] In this embodiment, the threshold for determining prodromal abnormalities is determined through ROC curve analysis of a large-scale sample, which balances the sensitivity and specificity of the determination. When the deviation score exceeds this threshold, the individual is determined to be in a prodromal abnormal state, and the subtype corresponding to the abnormal trajectory template with the lowest similarity is taken as the individual's matching subtype.
[0154] Step S460: Based on the historical onset time window distribution data corresponding to the matched abnormal trajectory subtype, predict the expected onset time window for the target object to progress from the current state to a clinically diagnosed depression within a specific time period in the future, and generate a prodromal progression warning information containing the expected start time point and the expected end time point.
[0155] In this embodiment, each abnormal trajectory subtype corresponds to the time distribution data of historical samples from the appearance of the trajectory pattern to clinical diagnosis. Based on the quantiles of the distribution data, the start and end times of the expected onset time window are determined, and the generated early warning information can clearly indicate the time range of risk progression.
[0156] Step S470: The output results of the probability of developing depression and the early warning information of the prodromal period are correlated and integrated to generate a comprehensive risk assessment report that includes the risk probability value, probability interval identifier, prodromal period trajectory subtype identifier and expected onset time window.
[0157] In this embodiment, the comprehensive risk assessment report integrates quantified risk probability, risk level, prodromal subtype information, and expected onset time window, which can not only indicate the level of risk, but also clarify the type of risk and the time of progression.
[0158] For example, after generating a comprehensive risk assessment report that includes risk probability values, probability interval identifiers, prodromal trajectory subtype identifiers, and expected onset time windows, the method further includes:
[0159] Step S510: Obtain a subset of stable health baseline data of the target object within a continuously preset stable baseline time window during the initial collection phase. The stable health baseline data subset has no record of major stress events and no self-reported emotional abnormalities during the time period corresponding to the stable health baseline data subset.
[0160] In this embodiment, the stable health baseline time window is set at the beginning of the monitoring period. During this period, the target object's health status is stable, there are no major life events or self-reports of abnormal emotions, and the corresponding monitoring data can represent the target object's normal health status.
[0161] Step S520: Input the stable health baseline data subset into a pre-trained causal generative adversarial network. The causal generative adversarial network includes a generator module and a discriminator module. The generator module is used to learn the joint distribution of causal features under the health state of the target object. The discriminator module is used to distinguish the feature sequence generated by the generator from the real health baseline feature sequence.
[0162] In this embodiment, the causal generative adversarial network is trained based on the stable health baseline data of the target object. The generator module learns the joint distribution and temporal evolution law among multiple causal features under the health state, and can generate feature sequences that conform to the health state of the object. The discriminator module learns to distinguish between the real health baseline sequence and the generated sequence. Through adversarial training between the generator and the discriminator, the generator can finally generate highly realistic individual health feature sequences.
[0163] Step S530: The generator module of the causal generative adversarial network performs multiple random sampling and generation operations to generate a set of individualized health counterfactual feature sequences of the target object under the assumption of no exposure to the risk of depression within a future continuous preset counterfactual time window. The set of individualized health counterfactual feature sequences contains multiple possible health state evolution trajectory samples.
[0164] In this embodiment, the counterfactual time window has the same length as the subsequent real-time monitoring window. The generator module generates multiple feature sequences that conform to the distribution of the object's health status through multiple random samplings. These sequences represent the possible health status evolution trajectory of the object without additional exposure to depression risk, and constitute a set of counterfactual feature sequences.
[0165] Step S540: Perform kernel density estimation on the feature values at each time point in the individualized health counterfactual feature sequence set, calculate the percentile distribution of the feature values at each time point, and extract the feature values corresponding to the first preset percentile and the second preset percentile as the boundary of the individualized anomaly determination confidence interval. The boundary of the individualized anomaly determination confidence interval includes the first confidence upper limit and the second confidence upper limit.
[0166] In this embodiment, kernel density estimation is used to fit the value distribution of each feature at each time point. Based on the fitted distribution, the feature values corresponding to two preset percentiles are calculated as individualized anomaly judgment thresholds. The first confidence upper limit corresponds to the mild anomaly threshold, and the second confidence upper limit corresponds to the severe anomaly threshold. The two thresholds can achieve anomaly stratification judgment of different degrees of severity.
[0167] Step S550: Compare the feature values in the time series matrix of individual causal weighted fusion features collected in real time for the target object with the boundaries of the individualized anomaly judgment confidence interval at the corresponding time point, and count the number of consecutive days exceeding the first confidence limit and the number of consecutive days exceeding the second confidence limit.
[0168] In this embodiment, for each feature at each time point, the value is compared with the corresponding two confidence limits to determine whether it exceeds the threshold. Then, the number of consecutive days that each feature exceeds the threshold is counted. The longer the consecutive days, the stronger the persistence of the anomaly and the lower the possibility of random fluctuations.
[0169] Step S560: If the duration parameter satisfies either continuously exceeding the first confidence limit or the first preset duration threshold, or continuously exceeding the second confidence limit or the second preset duration threshold, and the current abnormal feature is determined by the causal link graph model to be mainly driven by the endophenotypic abnormality mediating level rather than the temporary fluctuation of the exogenous exposure factor level, then the current abnormal state is determined to be a pathological precursor abnormality.
[0170] In this embodiment, the first preset number of days threshold and the second preset number of days threshold are set according to the time standard for clinical abnormality judgment. When the duration of the abnormality exceeds the threshold, the source of the abnormal feature is first determined according to the causal link diagram model. If the abnormal feature mainly belongs to the mediating level of endophenotypic abnormality and its change is not caused by short-term fluctuations of exogenous exposure factors, it is determined to be a pathological precursor abnormality, excluding false positive abnormalities caused by temporary fluctuations of external factors.
[0171] Step S570: If the duration parameter does not meet the above conditions, or if the current abnormality is determined by the causal link diagram model to be mainly dominated by temporary fluctuations at the level of exogenous exposure factors and no continuous change has occurred at the mediating level of endophenotypic abnormality, then the current abnormal state is determined to be normal emotional fluctuation, a false positive rejection label is generated, and the corresponding warning information is removed from the preliminary abnormality determination result.
[0172] In this embodiment, abnormalities that do not last long enough to reach the threshold, or abnormalities that are caused by short-term changes in exogenous exposure factors and do not show persistent abnormalities in endophenotypic characteristics, are all judged as normal emotional fluctuations. The corresponding preliminary warning information is marked as a false positive and removed, effectively reducing the false positive rate of the warning.
[0173] Step S580: For individuals who still retain abnormal markers after the pathological abnormality determination is completed, counterfactual intervention simulation processing is performed through the causal link diagram model. The core risk factor characteristic values of the individual whose causal effect values are ranked in the preset preorder position are adjusted one by one to the normal level of healthy people. The values are then re-entered into the early onset risk assessment model for depression to calculate the probability of onset risk after intervention.
[0174] In this embodiment, the counterfactual intervention simulation first selects several core risk factors with the highest total causal effect values, adjusts the feature values of these factors to the average level of healthy people, simulates the state after intervention on these risk factors, and then inputs the adjusted feature sequence into the risk assessment model to calculate the probability of disease incidence after intervention and evaluate the potential effect of intervention.
[0175] Step S590: Calculate the ratio of the difference between the probability of disease onset before intervention and the probability of disease onset after intervention to the probability of disease onset before intervention as the percentage of risk reduction. If the percentage of risk reduction is lower than the preset effective intervention threshold, the current preliminary abnormal judgment result is determined to be an uninterventional false positive result, a false positive removal flag is generated and removed from the final output result.
[0176] In this embodiment, the percentage of risk reduction reflects the degree to which the intervention measures reduce the risk. If the reduction is lower than the preset effective intervention threshold, it means that the current abnormal characteristics do not contribute much to the risk of disease, and even if intervention is implemented, the risk cannot be effectively reduced. This result is judged as a false positive and is removed, thereby further improving the accuracy of the final warning result.
[0177] Step S5100: Output the final risk assessment results of the individual, which have been verified by both pathological abnormality determination and counterfactual intervention simulation, to the clinical decision support system. At the same time, output the causal effect contribution value of each core risk factor in the causal link diagram model to the risk of disease as reference information for clinical intervention targets.
[0178] In this embodiment, the risk results after double verification have high accuracy. After being output to the clinical decision support system, they can provide risk alerts to clinicians. At the same time, the output core risk factor contribution values can clearly indicate which factors are the main causes of increased risk, providing doctors with clear target references for developing individualized intervention plans.
[0179] For example, after constructing a causal link graph model containing variable nodes and directed edges based on the thresholded adjacency matrix, the method further includes:
[0180] Step S610: Obtain the time series set of causal latent representation vectors of the target object on multiple consecutive time slices, and divide the time series set of causal latent representation vectors into multiple time slice subsets according to the preset time slice length. Each time slice subset corresponds to a causal latent representation vector sequence within a consecutive time period.
[0181] In this embodiment, the preset time slice length is set according to the evolution speed of causal relationships. Each time slice corresponds to a continuous time period of fixed length. The entire time series is divided into multiple non-overlapping time slice subsets in chronological order. Each subset contains all causal latent representation vectors within the corresponding time period, which are used to analyze local causal relationships in different time periods.
[0182] Step S620: For each time slice subset, the directed acyclic graph learning process with penalty term is repeatedly performed to generate a temporary causal link graph model corresponding to the time slice subset. The temporary causal link graph model contains the local causal transmission relationship between each causal feature unit within the time slice.
[0183] In this embodiment, a directed acyclic graph learning process with a penalty term is independently executed for each subset of time slices to learn the local causal structure within that time period, thereby obtaining a temporary causal link graph model corresponding to each time slice. These models can reflect the dynamic changes in causal relationships within different time periods.
[0184] Step S630: Arrange the temporary causal link graph models corresponding to multiple consecutive time slices in chronological order to construct the temporal evolution sequence of the causal link graph model. Each element in the temporal evolution sequence is a graph structure object with a set of nodes and a set of edges.
[0185] In this embodiment, the temporary causal link graph models corresponding to all time slices are arranged in chronological order to form a temporal evolution sequence of causal structures. Adjacent graph structures in the sequence correspond to adjacent time periods, which can fully demonstrate the evolution of causal relationships over time.
[0186] Step S640: Perform node matching and alignment processing on the graph structure objects in the temporal evolution sequence to ensure that the same causal feature unit in different time slices corresponds to the same node identifier, and generate a temporal evolution graph sequence after node alignment.
[0187] In this embodiment, the node matching and alignment process uses the node identifier of the initial global causal link graph model as a benchmark. It matches the nodes in the temporary causal link graph model of each time slice with the global nodes to ensure that the same causal feature unit has the same node identifier in the graph model of all time slices, eliminates the difference in node identifiers between graph models of different time slices, and facilitates subsequent time series comparison and analysis.
[0188] Step S650: Extract the occurrence frequency and weight coefficient change trajectory of each directed causal edge in the time-series evolution graph sequence at different time slices, and construct the time-varying feature curve of each directed causal edge. The horizontal axis of the time-varying feature curve is the time slice sequence identifier, and the vertical axis is the weight coefficient value.
[0189] In this embodiment, for each valid causal edge in the global causal link graph, we count whether it exists in the temporary graph model of each time slice and its corresponding weight coefficient. Using the time slice as the horizontal axis and the weight coefficient as the vertical axis, we plot the time-varying characteristic curve of the edge. This time-varying characteristic curve can show the change law of the strength of the causal edge over time.
[0190] Step S660: Perform trend analysis processing on the time-varying characteristic curve, calculate the slope parameter and curvature parameter of the weight coefficient of each directed causal edge as a function of time, and classify the causal edges into stable causal edges, enhanced causal edges, weakened causal edges, and fluctuating causal edges according to the slope parameter and curvature parameter.
[0191] In this embodiment, the slope parameter reflects the overall trend of the weight coefficients changing over time. A positive slope indicates that the effect gradually strengthens, a negative slope indicates that the effect gradually weakens, and a slope close to zero indicates that the effect is basically stable. The curvature parameter reflects the rate of change of the trend; the greater the curvature, the faster the trend changes. Based on the range of slope and curvature values, causal edges are divided into four types to facilitate the identification of causal paths with different dynamic characteristics.
[0192] Step S670: Identify from the time-series evolution graph sequence enhanced causal edges that exist continuously in multiple consecutive time slices and whose weight coefficients show a monotonically increasing trend, and mark the enhanced causal edges as key transmission paths for risk accumulation.
[0193] In this embodiment, the effect of the enhanced causal edge continues to increase over time, indicating that the corresponding causal transmission path is constantly being strengthened, reflecting the accumulation process of risk factors. Marking the above edge as the key transmission path of risk accumulation can identify the core causal path that drives the risk to gradually increase.
[0194] Step S680: Identify mutational causal edges that suddenly appear in a specific time slice and whose weight coefficients exceed a preset mutation threshold from the time-series evolution graph sequence, and mark the mutational causal edges as risk-triggered key transmission paths.
[0195] In this embodiment, a mutation-type causal edge that did not exist or had a very low weight in previous time slices suddenly appears in a certain time slice and has a weight exceeding a preset mutation threshold indicates that the corresponding causal relationship was suddenly established, which usually corresponds to the occurrence of a risk-triggered event. Marking the above edge as a key transmission path for risk triggering can identify the triggering factors that lead to a sudden increase in risk.
[0196] Step S690: Obtain the data subsets of exogenous exposure factors and mediating data subsets of endophenotypic abnormalities corresponding to each time slice of the target object, calculate the correlation coefficient between the cumulative effect value of the risk accumulation key transmission path and the cumulative exposure amount of exogenous exposure factors, and at the same time calculate the correlation coefficient between the mutation effect value of the risk triggering key transmission path and the mutation magnitude of endophenotypic abnormalities in the same period.
[0197] In this embodiment, the correlation between the cumulative effect of the risk accumulation path and the cumulative amount of exogenous exposure can be calculated to determine whether the path is driven by the continuous accumulation of external exposure factors; the correlation between the mutation effect of the risk triggering path and the magnitude of endophenotypic abnormal mutation can be calculated to determine whether the path is driven by the sudden abnormality of the endophenotype.
[0198] Step S6100: Based on the significance level of the correlation coefficient, select causal edges that are highly synchronized with changes in exposure factors as environmental response causal edges, and select causal edges that are highly synchronized with abnormal changes in endophenotypes as endophenotype-driven causal edges.
[0199] In this embodiment, when the significance level of the correlation coefficient is lower than the preset threshold, it is determined to be highly synchronous, and the corresponding causal edges are classified as environmental response type or endophenotypic driven type. Different types of causal edges correspond to different intervention strategies. Environmental response type edges need to be intervened by reducing external exposure, while endophenotypic driven edges need to be intervened by regulating internal physiological state.
[0200] Step S6110: The environmental response causal edge, the internal surface type driving causal edge, the risk accumulation key transmission path, and the risk triggering key transmission path are superimposed and fused to generate an enhanced causal link graph model containing dynamic evolution information. Each edge in the enhanced causal link graph model carries a time-varying feature type label and a key path type label in addition to the weight coefficient.
[0201] In this embodiment, the enhanced causal link graph model adds time-varying feature type labels and critical path type labels to each edge on the basis of the original static causal link graph. It can fully display the dynamic evolution characteristics and functional types of causal edges, and can not only explain the existence and strength of causal relationships, but also explain their evolutionary laws and driving factors.
[0202] For example, after generating the enhanced causal link graph model containing dynamic evolution information, the method further includes:
[0203] Step S710: Obtain the time series set of causal latent representation vectors of multiple time slices of patients diagnosed with multiple different subtypes of depression before diagnosis. Repeat the time slice division operation, the temporary causal link graph model construction operation, and the time series evolution graph sequence generation operation for each subtype patient group to obtain the group average causal link time series evolution template library corresponding to each subtype. The group average causal link time series evolution template library includes the time series template of the first depression subtype, the time series template of the second depression subtype, and the time series template of the third depression subtype.
[0204] In this embodiment, the pathogenesis and causal evolution path of different subtypes of depression are different. For each subtype of patient population, a population average temporal evolution template corresponding to that subtype is constructed. The template contains the typical causal structure evolution law before the onset of the disease in patients of that subtype, which serves as the matching benchmark for subsequent subtype identification.
[0205] Step S720: Perform graph similarity matching calculation between the enhanced causal link graph model of the target object and each subtype time series template in the population average causal link time series evolution template library. The graph similarity matching calculation includes similarity measurement values in three dimensions: node set overlap degree calculation, edge set overlap degree calculation, and edge weight coefficient distribution similarity calculation.
[0206] In this embodiment, graph similarity matching comprehensively measures the similarity between two graph structures from three dimensions: node set overlap, edge set overlap, and edge weight coefficient distribution similarity. The three-dimensional measurement values can fully reflect the similarity between two causal graph structures.
[0207] Step S730: Perform a weighted summation operation on the node set overlap degree, the edge set overlap degree, and the edge weight coefficient distribution similarity to obtain the comprehensive graph similarity score between the target object and each subtype time series template. Use the subtype identifier corresponding to the subtype time series template with the highest comprehensive graph similarity score as the predicted depression subtype identifier of the target object.
[0208] In this embodiment, the three similarity measures are multiplied by preset weighting coefficients and then summed to obtain a comprehensive similarity score. The weighting coefficients are preset according to the importance of the three dimensions. The higher the score, the more similar the causal evolution pattern of the target object is to the subtype. The subtype with the highest score is taken as the predicted subtype of the target object, thereby realizing individualized subtype identification.
[0209] Step S740: Retrieve the corresponding subtype-specific intervention strategy knowledge base based on the predicted depression subtype identifier. The subtype-specific intervention strategy knowledge base contains a priority ranking list of intervention targets for different causal transmission pathways and an identifier of intervention method type.
[0210] In this embodiment, the subtype-specific intervention strategy knowledge base is built based on clinical research and expert consensus. Each depression subtype corresponds to a set of specific intervention strategies, including the intervention priorities and corresponding intervention methods for different causal pathways under that subtype, which can provide targeted intervention suggestions for individuals with different subtypes.
[0211] Step S750: Identify all direct causal edges and indirect causal paths pointing to the level nodes of the depression outcome from the enhanced causal link graph model of the target object, and calculate the percentage of the causal effect contribution value of each direct causal edge and each indirect causal path relative to the total causal effect of all nodes pointing to the outcome as the contribution percentage.
[0212] In this embodiment, the contribution percentage reflects the relative contribution of each causal edge or path to the risk of disease. It is obtained by dividing the effect value of each edge or path by the sum of all effects pointing to the disease outcome, which can quantify the contribution ratio of different paths to the individual's disease risk.
[0213] Step S760: Mark the direct causal edges and indirect causal paths whose contribution percentage exceeds the preset contribution threshold as core intervention target paths, and extract the starting node, intermediate node and ending node on the path as a set of intervention target candidate nodes from the core intervention target paths.
[0214] In this embodiment, a preset contribution threshold is used to screen core paths with significant contributions. Paths with a contribution exceeding this threshold are marked as core intervention target paths. These paths contribute significantly to the risk of disease development, and intervening in these paths can more effectively reduce the risk of disease development. All nodes on the path are extracted as candidate intervention targets.
[0215] Step S770: For each candidate intervention target node, query the intervention type identifier and expected effect decay time parameter that match the causal feature unit corresponding to the exogenous exposure factor node in the subtype-specific intervention strategy knowledge base, and generate an individualized intervention suggestion entry containing the intervention target node identifier, the matching intervention type identifier, and the expected effect decay time parameter.
[0216] In this embodiment, each candidate node of the intervention target corresponds to a causal feature unit. The intervention method corresponding to the feature unit is queried in the subtype-specific intervention strategy knowledge base. Each intervention method has a corresponding expected effect decay time parameter, that is, the duration of the effect after the intervention is implemented. Combining this information, individualized intervention suggestion items for the target are generated to ensure the targeting and operability of the suggestions.
[0217] Step S780: Based on the causal transmission direction of the core intervention target path, sort multiple individualized intervention suggestion items in order from upstream to downstream to generate a hierarchical intervention path sequence. The intervention suggestion items at the beginning of the hierarchical intervention path sequence correspond to intervention targets at the level of exogenous exposure factors, and the intervention suggestion items at the end correspond to intervention targets at the level of mediating endophenotypic abnormalities.
[0218] In this embodiment, the causal transmission direction is from the exogenous exposure factor level to the mediating level of endophenotypic abnormalities and then to the disease outcome. Therefore, intervention recommendations for upstream exogenous exposure factors are placed first, and intervention recommendations for downstream endophenotypic abnormalities are placed later, forming a hierarchical intervention path sequence, which conforms to the logic of causal transmission and can gradually block the risk transmission path from the root cause to the result.
[0219] Step S790: Combine and encapsulate the hierarchical intervention path sequence with the predicted depression subtype identifier, the risk probability interval identifier, and the expected onset time window to generate an individualized intervention plan report containing multi-dimensional information. Each intervention suggestion item in the individualized intervention plan report is associated with a specific causal feature unit and an operable intervention action descriptor.
[0220] In this embodiment, the individualized intervention program report integrates multi-dimensional information such as intervention path, subtype information, risk level, and expected onset time window. Each intervention recommendation corresponds to a specific actionable action, allowing clinicians or individuals to directly implement the intervention based on the recommendations in the report. At the same time, it clarifies the causal target and expected effect corresponding to each intervention action, thereby improving the scientific nature and effectiveness of the intervention.
[0221] Based on the same inventive concept, please refer to Figure 2 The diagram shows a schematic block diagram of a depression incidence risk prediction system 100 based on causal representation learning provided in an embodiment of this application. The depression incidence risk prediction system 100 based on causal representation learning may include a communication unit 110, a machine-readable storage medium 120, and a processor 130.
[0222] In this embodiment, the machine-readable storage medium 120 can also be integrated into the processor 130 and can communicate and interact with external systems through the communication unit 110. The machine-readable storage medium 120 stores machine-executable instructions for executing the scheme of this application, and the processor 130 executes the machine-executable instructions stored in the machine-readable storage medium 120 to implement the depression incidence risk prediction method based on causal representation learning provided in the aforementioned method embodiments.
[0223] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.
Claims
1. A method for predicting the risk of developing depression based on causal representation learning, characterized in that, The method includes: Acquire a multi-source latent monitoring data set of the target object, wherein the multi-source latent monitoring data set includes continuously collected nonlinear micro-feature data units of autonomic nervous function with time sequence identification, micro-feature data units of touch behavior of smart terminals, spatiotemporal big data units of environment, and non-depression-related medical trajectory data units in electronic health records; According to the preset causal hierarchy classification rules, the multi-source latent monitoring data set is subjected to causal variable hierarchy labeling processing to obtain the exogenous exposure factor hierarchy label, endophenotypic abnormality mediating hierarchy label and depression onset outcome hierarchy label corresponding to each data unit in the multi-source latent monitoring data set. The pre-built causal variational autoencoder network is invoked to perform causal latent representation extraction processing on the multi-source latent monitoring data set carrying the hierarchical label of the causal variable, generating a set of causal latent representation vectors that are causally associated with the onset of depression. Each vector dimension in the set of causal latent representation vectors corresponds to a causal feature unit after eliminating confounding variable bias. Based on the set of causal latent representation vectors and the hierarchical labels of causal variables, a directed acyclic graph learning process with a penalty term is performed to construct a causal link graph model that includes the causal transmission path between the hierarchy of exogenous exposure factors, the hierarchy of mediating endophenotypic abnormalities, and the hierarchy of depressive onset outcomes. Based on the causal effect weights corresponding to each causal feature unit in the causal link graph model, the real-time collected individual multi-source time-series monitoring data are subjected to causal weighted fusion processing to obtain an individual causal weighted fusion feature time-series matrix. The individual causal weighted fusion feature time-series matrix is then input into the early-onset risk assessment model for depression to perform probability prediction processing, generating a probability output result of depression onset risk containing prediction time window identifiers.
2. The method for predicting the risk of depression based on causal representation learning according to claim 1, characterized in that, The step involves performing causal variable hierarchical labeling on the multi-source latent surveillance dataset according to a preset causal hierarchy classification rule, resulting in exogenous exposure factor hierarchical labels, endophenotypic abnormality mediating hierarchical labels, and depression outcome hierarchical labels corresponding to each data unit in the multi-source latent surveillance dataset, including: The original acquisition source identifier and data structure type identifier of each data unit in the multi-source latent monitoring data set are analyzed. Based on the original acquisition source identifier and data structure type identifier, the environmental spatiotemporal big data unit and the non-depression-related medical trajectory data unit in the electronic health record are classified into the exogenous exposure factor level. An exogenous exposure factor level label is added to each data unit classified into the exogenous exposure factor level. The original acquisition source identifier and data structure type identifier of each data unit in the multi-source latent monitoring data set are analyzed. Based on the original acquisition source identifier and data structure type identifier, the nonlinear micro-feature data unit of the autonomic nervous function and the micro-feature data unit of the smart terminal touch behavior are classified into the endophenotype abnormality mediating level. An endophenotype abnormality mediating level label is added to each data unit classified into the endophenotype abnormality mediating level. Obtain the clinical diagnosis record of depression of the target object corresponding to the multi-source latent monitoring data set within a preset follow-up time window, generate a depression onset outcome hierarchical marker with binary status identifier based on the clinical diagnosis record of depression, and perform time-series association binding processing on the depression onset outcome hierarchical marker and the data unit at the corresponding time point in the multi-source latent monitoring data set; After the hierarchical labeling is completed, a data integrity scan is performed on the multi-source latent monitoring data set to identify data units with missing hierarchical labels or conflicting hierarchical labels. The hierarchical classification judgment operation is re-executed according to the original acquisition source identifier and data structure type identifier of the data unit until all data units carry a unique and conflict-free hierarchical label. The multi-source latent surveillance dataset with complete hierarchical labels is arranged and integrated according to the time axis to generate a hierarchical structured time series dataset containing hierarchical data subsets of exogenous exposure factors, hierarchical data subsets of mediating endophenotypic abnormalities, and hierarchical labels of depression onset and outcome for each time point. Extract the correlation features between data units within each level of the hierarchical structured time series dataset, construct an initial causal path hypothesis set across levels based on the correlation features, and use the initial causal path hypothesis set as a priori constraint for subsequent causal latent representation extraction processing.
3. The method for predicting the risk of depression based on causal representation learning according to claim 1, characterized in that, The process involves invoking a pre-built causal variational autoencoder network to extract causal latent representations from a multi-source latent monitoring dataset carrying the hierarchical labels of the causal variables, generating a set of causal latent representation vectors that are causally associated with the onset of depression, including: The multi-source latent monitoring data set carrying the hierarchical label of the causal variable is input into the encoder module of the causal variational autoencoder network. The input data is subjected to dimensionality compression processing through the multi-layer nonlinear mapping network of the encoder module to generate initial latent representation vector distribution parameters, which include a mean vector and a log-variance vector. Obtain a preset set of promiscuous variables, which includes age identifier, gender identifier, underlying disease identifier, and season identifier. Perform one-hot encoding on the set of promiscuous variables to obtain a promiscuous variable encoding vector. Then, concatenate the promiscuous variable encoding vector with the initial latent representation vector distribution parameters to obtain an extended latent representation distribution parameter carrying promiscuous information. The distribution alignment module in the causal variational autoencoder network is invoked to perform maximum mean difference distance calculation on the extended latent representation distribution parameters based on the promiscuous variable encoding vector, thereby generating a distribution alignment loss value that represents the distribution difference of latent representations under different promiscuous levels. Based on the distribution alignment loss value, the distribution alignment constraint optimization process is performed on the extended latent representation distribution parameters. The network weight parameters of the encoder module are adjusted by the gradient backpropagation algorithm so that the latent representation distributions corresponding to different promiscuous levels gradually tend to be consistent, and the causal latent representation vector distribution parameters that eliminate the bias of promiscuous variables are obtained. Random sampling is performed on the causal latent representation vector distribution parameters to obtain an initial causal latent representation vector sample. The initial causal latent representation vector sample is then input into the decoder module of the causal variational autoencoder network for data reconstruction processing to generate a reconstructed data set that matches the dimension of the input data. Calculate the reconstruction error loss value between the reconstructed data set and the original input data, and simultaneously calculate the KL divergence loss value between the causal latent representation vector distribution parameters and the standard normal prior distribution. Perform a weighted summation operation on the reconstruction error loss value, the KL divergence loss value, and the distribution alignment loss value to obtain the total optimized loss function value of the causal variational autoencoder network. Based on the total optimization loss function value, the encoder module and decoder module of the causal variational autoencoder network are jointly iteratively optimized and trained until the total optimization loss function value converges to a preset threshold range. Then, the causal latent representation vector corresponding to each input data unit is extracted from the output layer of the trained encoder module. All extracted causal latent representation vectors are arranged and integrated according to time order and hierarchical label to generate a set of causal latent representation vectors containing temporal relationship and hierarchical attribution information. The dimension unit of each vector in the set of causal latent representation vectors corresponds to a causal feature unit after eliminating confounding variable bias.
4. The method for predicting the risk of depression based on causal representation learning according to claim 3, characterized in that, The process involves jointly iteratively optimizing and training the encoder and decoder modules of the causal variational autoencoder network based on the total optimized loss function value until the total optimized loss function value converges to a preset threshold range. Then, the causal latent representation vector corresponding to each input data unit is extracted from the output layer of the trained encoder module, including: The multi-source latent monitoring dataset is divided into a training data subset and a validation data subset; Initialize all trainable weight parameters of the encoder and decoder modules in the causal variational autoencoder network to random small values, and set the optimizer type identifier, initial learning rate parameter, and batch training sample quantity parameter. A batch of multi-source latent monitoring data units carrying hierarchical labels is randomly selected from the training data subset. The selected data units are input into the encoder module of the current iteration round for forward propagation calculation to obtain the causal latent representation vector distribution parameters of the current batch. The distribution alignment loss value, the reconstruction error loss value, and the KL divergence loss value are calculated based on the causal latent representation vector distribution parameters of the current batch, and the total optimization loss function value of the current batch is calculated according to the preset weighting coefficient combination. Based on the total optimization loss function value of the current batch, the gradient value of each trainable weight parameter in the encoder module and decoder module is calculated through an automatic differentiation mechanism, and all trainable weight parameters are updated according to the calculated gradient value using a preset optimization algorithm. After completing a batch parameter update, all data units in the validation data subset are input into the updated encoder and decoder modules for forward propagation calculation. The total optimization loss function value corresponding to the validation data subset is calculated and the validation loss value is recorded. Determine whether the validation loss value of multiple consecutive iterations shows a downward trend or whether the fluctuation range exceeds the preset stability threshold. If the validation loss value continues to decrease or the fluctuation range exceeds the stability threshold, continue to perform batch extraction and forward propagation calculation operations for the next round of iteration training. If the verification loss value stops decreasing and the fluctuation range is lower than the preset stability threshold, the model training is determined to have reached the convergence state, the iterative optimization process is stopped, and the weight parameters of the encoder module and decoder module corresponding to the current iteration are saved as the model parameters after training is completed. All data units in the multi-source latent monitoring dataset are sequentially input into the trained encoder module for forward propagation calculation, and the causal latent representation vector corresponding to each data unit is extracted from the output of the last hidden layer of the encoder module. The extracted causal latent representation vectors are subjected to dimensional standardization to generate a final causal latent representation vector set with a uniform scale representation. Each vector dimension in the final causal latent representation vector set corresponds to a causal feature unit after eliminating confounding variable bias.
5. The method for predicting the risk of depression based on causal representation learning according to claim 1, characterized in that, The method involves performing penalized directed acyclic graph learning based on the set of causal latent representation vectors and the hierarchical labels of causal variables to construct a causal link graph model that includes the causal transmission path between the exogenous exposure factor hierarchy, the endophenotypic abnormality mediation hierarchy, and the depression onset outcome hierarchy. This includes: The set of causal latent representation vectors is split into latent representation matrix of exogenous exposure factors, latent representation matrix of mediating endophenotypic abnormalities, and latent representation vector of depression pathogenesis outcome according to the hierarchical labeling of causal variables. The rows of the latent representation matrix of exogenous exposure factors correspond to time points and the columns correspond to the feature units of exposure factors. The rows of the latent representation matrix of mediating endophenotypic abnormalities correspond to time points and the columns correspond to the feature units of endophenotypic abnormalities. The latent characterization matrix of the exogenous exposure factors and the latent characterization matrix of the mediating endophenotypic abnormalities are subjected to time-lag processing to generate an extended feature matrix containing the current time point and multiple lag time points. The extended feature matrix is used to capture the delayed causal effects of exposure factors and endophenotypic abnormalities on the disease outcome. Initialize an adjacency matrix with a dimension equal to the sum of the total number of all causal feature units and the total number of disease outcome units. Each element in the adjacency matrix represents whether there is a causal edge between two corresponding variables and the direction of the causal edge. All diagonal elements are set to zero to ensure that there are no self-loops. A smooth linear structural equation model is constructed with the adjacency matrix as the variable. The smooth linear structural equation model represents the value of each variable as a linear combination of other variables plus a noise term. The coefficients of the linear combination are determined by the non-zero elements of the corresponding rows of the adjacency matrix. Substitute the extended feature matrix and the latent representation vector of the depression outcome into the smooth linear structural equation model, calculate the sum of squared residuals between the predicted and actual values of each variable using the least squares method, and sum the sums of squared residuals of all variables to obtain the model fitting loss function value. Based on the model fitting loss function value, a sparsity penalty term and a directed acyclic constraint term based on the adjacency matrix are added. The sparsity penalty term is used to control the number of causal edges to prevent overfitting, and the directed acyclic constraint term is used to ensure that the learned graph structure does not have directed cycles. An augmented Lagrange method combined with gradient descent algorithm is used to iteratively optimize the joint objective function with added penalty and constraint terms. In each iteration, the element values in the adjacency matrix are updated, and the Lagrange multipliers and penalty coefficients are updated to force the satisfaction of the directed acyclic constraint. During the iterative optimization process, a hierarchical structure constraint is applied to the adjacency matrix based on the hierarchical label of the causal variables, prohibiting causal edges from the mediating level of endophenotypic abnormalities to the level of exogenous exposure factors, and prohibiting causal edges from the level of depression onset outcome to any other level, ensuring that the causal direction conforms to the preset hierarchical order. When the iterative optimization reaches the maximum number of iterations or the change in the adjacency matrix is lower than the preset convergence threshold, the optimization process is stopped, and the final adjacency matrix is thresholded. Adjacency matrix elements with absolute values lower than the preset significance threshold are set to zero, and elements with absolute values higher than the threshold are retained as valid causal edges. A causal link graph model containing variable nodes and directed edges is constructed based on the adjacency matrix after thresholding. Each node in the causal link graph model corresponds to a causal feature unit or a disease outcome unit. Each directed edge points from the cause node to the result node, and the weight of the edge corresponds to the element value in the adjacency matrix. The causal link graph model fully presents the causal transmission path from the exogenous exposure factor level through the endophenotypic abnormality mediation level to the depression disease outcome level.
6. The method for predicting the risk of depression based on causal representation learning according to claim 5, characterized in that, After constructing a causal link graph model containing variable nodes and directed edges based on the thresholded adjacency matrix, the method further includes: The causal effect value is quantified for each directed edge in the causal link graph model, and the adjacency matrix element value corresponding to each directed edge is extracted as the direct causal effect coefficient. The direct causal effect coefficient represents the expected change of the result variable when the cause variable changes by one unit. For a multi-step causal path starting from the exogenous exposure factor level node, passing through the endophenotypic abnormality mediating level node, and reaching the depression pathogenesis outcome level node, the indirect causal effect value of each multi-step path is calculated using the path coefficient product method. The indirect effect value of the path is obtained by multiplying all the direct causal effect coefficients on the path. The total mediating effect of the endophenotype abnormality mediating node between the exposure factor and the disease outcome is obtained by summing all the indirect causal effect values that pass through the same endophenotype abnormality mediating node. The total mediating effect quantifies the mediating strength of the endophenotype abnormality in the causal transmission process. For each exogenous exposure factor node in the causal link graph model, the total causal effect is calculated. The direct causal effect value of the exogenous exposure factor node pointing to the depression outcome level node is added to the indirect causal effect value through all mediating paths to obtain the total causal effect value of the exposure factor node on the outcome. The nodes in the causal link graph model are ranked according to the total causal effect value corresponding to each causal feature unit, and a core risk driver ranking list is generated. The core risk driver ranking list is used to identify the key feature units that contribute the most to the onset of depression. The stability assessment of the causal link graph model is performed by using a bootstrap resampling method to extract multiple sample subsets from the original data. The directed acyclic graph learning operation with a penalty term is repeatedly performed on each sample subset to obtain multiple adjacency matrix samples. Calculate the frequency of occurrence and standard deviation of weight coefficient of each causal edge in multiple adjacency matrix samples, mark causal edges whose frequency of occurrence is lower than a preset frequency threshold or whose standard deviation of weight coefficient is higher than a preset stability threshold as unstable edges, and remove the unstable edges from the causal link graph model; The causal link graph model after removing unstable edges is simplified and optimized, retaining the top core nodes whose cumulative contribution rate of causal effect value reaches a preset ratio and the causal edges between nodes, and generating a simplified core causal link graph model. The causal feature unit identifier, node hierarchy, and causal effect coefficients between nodes of each node in the simplified core causal link graph model are stored in a structured manner to construct an interpretable knowledge base of the causal causes of depression.
7. The method for predicting the risk of depression based on causal representation learning according to claim 1, characterized in that, The real-time collected individual multi-source time-series monitoring data is causally weighted and fused according to the causal effect weight corresponding to each causal feature unit in the causal link graph model to obtain an individual causal weighted fusion feature time-series matrix. This matrix is then input into a depression early-onset risk assessment model for probability prediction, generating a depression incidence risk probability output containing a prediction time window identifier, including: Acquire individual multi-source time-series monitoring data of the target object in real time during multiple consecutive collection cycles. The individual multi-source time-series monitoring data includes real-time sequences of nonlinear micro-features of autonomic nervous function corresponding to nodes in the causal link graph model, real-time sequences of micro-features of touch behavior of smart terminals, real-time units of environmental spatiotemporal big data, and real-time records of non-depression-related medical trajectories in electronic health records. The same causal latent representation extraction process as in the model training stage is performed on each feature real-time sequence in the individual multi-source time-series monitoring data. The encoder module of the trained causal variational autoencoder network is called to map each feature real-time sequence into an individual causal latent representation real-time vector, thus obtaining the individual causal latent representation real-time matrix. The total causal effect value corresponding to each causal feature unit is extracted from the causal link graph model. The total causal effect value is normalized to obtain the causal fusion weight coefficient of each causal feature unit. The magnitude of the causal fusion weight coefficient is positively correlated with the total causal effect value. Each column vector in the real-time matrix of individual causal latent representation is multiplied element-wise with the causal fusion weight coefficient of the corresponding causal feature unit to obtain the weighted real-time matrix of individual causal latent representation. The weighted matrix is then concatenated in rows according to time points to generate the individual causal weighted fusion feature vector corresponding to each time point. The individual causal weighted fusion feature vectors at each time point are stacked in chronological order to construct an individual causal weighted fusion feature time series matrix containing time dimension information and feature dimension information. The row index of the individual causal weighted fusion feature time series matrix corresponds to the continuous collection time point, and the column index corresponds to the causal feature dimension after weighted fusion. The individual causal weighted fusion feature time series matrix is input into a pre-trained early-onset risk assessment model for depression. The early-onset risk assessment model for depression includes a bidirectional long short-term memory network layer and a Cox proportional risk output layer. The bidirectional long short-term memory network layer performs forward and backward contextual information capture processing on the time series matrix to extract the dynamic evolution pattern vector of individual causal features. The dynamic evolution pattern vector is input into the Cox proportional hazards output layer, which calculates the cumulative morbidity function value of an individual within multiple preset future time windows based on a linear combination of a preset baseline morbidity function and the dynamic evolution pattern vector. The cumulative incidence risk function value is subjected to probability transformation processing to generate the individual's incidence probability value of depression in a future preset first time window, the incidence probability value of depression in a future preset second time window, and the incidence probability value of depression in a future preset third time window, which constitutes the incidence probability output result of depression including time window identifier. The output result of the probability of developing depression is compared with a preset risk probability division threshold. Based on the comparison result, a corresponding risk probability interval identifier is generated. The risk probability interval identifier includes a first probability interval identifier, a second probability interval identifier, a third probability interval identifier, and a fourth probability interval identifier.
8. The method for predicting the risk of depression based on causal representation learning according to claim 7, characterized in that, Before inputting the individual causal weighted fusion feature time series matrix into the pre-trained early-onset risk assessment model for depression, the method further includes: A risk assessment model for the very early onset of initial depression is constructed. The model consists of an input layer, a bidirectional long short-term memory network layer, a fully connected mapping layer, and a Cox proportional hazards output layer connected in sequence. The dimension of the input layer matches the dimension of the individual causal weighted fusion feature vector. The bidirectional long short-term memory network layer includes forward long short-term memory units and backward long short-term memory units. Obtain a sample cohort training dataset, which includes a causal weighted fusion feature time series matrix of multiple diagnosed depression patients at multiple consecutive time points before diagnosis and the corresponding diagnosis time point identifier, and also includes a causal weighted fusion feature time series matrix of multiple healthy control individuals within the same follow-up period and undiagnosed identifier. The causal weighted fusion feature time series matrix of each individual in the sample queue training dataset is used as input, and the diagnosis time point identifier or non-diagnosis identifier of the corresponding individual is used as supervision label to construct a model training sample pair set. The time series matrix in the training sample pair set of the model is sequentially input into the initial early-onset risk assessment model of depression. The time series matrix is passed to the bidirectional long short-term memory network layer through the input layer. The bidirectional long short-term memory network layer fuses the forward hidden state and the backward hidden state at each time step to generate the context-aware feature vector of each time step. The context-aware feature vector of the last time step is input into the fully connected mapping layer. The compressed temporal evolution pattern feature vector is generated through the linear transformation and nonlinear activation function of the fully connected mapping layer. The dimension of the temporal evolution pattern feature vector is smaller than that of the original context-aware feature vector. The feature vector of the time-series evolution pattern is input into the Cox proportional hazards output layer. The Cox proportional hazards output layer calculates the incidence risk function value of each individual at each time point during the observation period according to the preset baseline risk function. The incidence risk function value is integrated on the time axis to obtain the cumulative risk function value. A partial likelihood loss function is constructed based on the cumulative risk function value of each individual and the actual time point of diagnosis or the undiagnosed status. The partial likelihood loss function is used to measure the degree of agreement between the risk ranking predicted by the model and the actual situation. Individuals with correct risk function value ranking are given a positive contribution, while individuals with incorrect ranking are given a negative contribution. The partial likelihood loss function is minimized by using the gradient descent optimization algorithm. The gradients of all trainable parameters in the bidirectional long short-term memory network layer, the fully connected mapping layer, and the Cox proportional risk output layer are calculated by backpropagation. The parameter values are updated according to the gradients. The forward and backward propagation operations are iteratively executed until the loss function converges. An early stopping mechanism is introduced during model training. The sample queue training dataset is divided into a training subset and a validation subset. After each training round, the partial likelihood loss value on the validation subset is calculated. Training is stopped when the partial likelihood loss value on the validation subset no longer decreases for several consecutive rounds. The model parameters of the current round are saved as a pre-trained risk assessment model for the very early onset of depression.
9. The method for predicting the risk of depression based on causal representation learning according to claim 1, characterized in that, After generating the output result of the probability of developing depression including the prediction time window identifier, the method further includes: Obtain the individual causal weighted fusion feature time series matrix of the target object within a continuous preset historical time window, and input the individual causal weighted fusion feature time series matrix within the continuous preset historical time window into the pre-constructed dynamic abnormal trajectory recognition module for abnormal trajectory matching processing; The dynamic abnormal trajectory recognition module for the prodromal phase of depression has a pre-stored template library of multiple prodromal abnormal trajectory subtypes obtained by clustering sample queue data. The abnormal trajectory subtype template library includes a first prodromal abnormal trajectory template, a second prodromal abnormal trajectory template, and a third prodromal abnormal trajectory template. Each abnormal trajectory template includes the mean sequence of causal characteristics and the abnormal threshold sequence of the corresponding subtype at each time point before the onset of the disease. The similarity distance metric between the individual causal weighted fusion feature time series matrix and each abnormal trajectory template is calculated using a dynamic time warping algorithm. At the same time, the similarity distance metric between the individual causal weighted fusion feature time series matrix and the normal fluctuation baseline template of healthy people is also calculated. The individual abnormal trajectory deviation score is obtained by subtracting the minimum similarity distance metric between the individual causal weighted fusion feature time series matrix and the abnormal trajectory template from the similarity distance metric between the template and the normal fluctuation baseline template of the healthy population. The individual abnormal trajectory deviation score is used to quantify the degree to which the individual time series features deviate from the normal range. If the deviation score of the individual abnormal trajectory exceeds the preset threshold for judging the abnormality of the prodromal period, it is determined that the target object is currently in the abnormal trajectory state of the prodromal period of depression, and the subtype identifier corresponding to the abnormal trajectory template with the smallest similarity distance to the individual causal weighted fusion feature time series matrix is identified as the matched abnormal trajectory subtype. Based on the historical onset time window distribution data corresponding to the matched abnormal trajectory subtypes, the expected onset time window for the target object to progress from the current state to a clinically diagnosed depression within a specific time period in the future is predicted, and a prodromal progression warning information containing the expected start time point and the expected end time point is generated. The output results of the probability of developing depression and the early warning information of the prodromal period are correlated and integrated to generate a comprehensive risk assessment report that includes the risk probability value, probability interval identifier, prodromal trajectory subtype identifier, and expected onset time window.
10. A system for predicting the risk of developing depression based on causal representation learning, characterized in that, include: processor; A machine-readable storage medium for storing machine-executable instructions of the processor; The processor is configured to execute the method for predicting the risk of developing depression based on causal representation learning as described in any one of claims 1 to 9 by executing the machine-executable instructions.
Citation Information
Patent Citations
Disease prediction method and system based on causal inference and dynamic integration of multiple tags
CN116364274A
Disease risk assessment method and system based on clinical scale and causal discovery
CN117711624A