Subspace balanced division-based depression classification model
Depression data is processed based on subspace equalization division, and a multi-screening model is constructed in combination with machine learning algorithms, which solves the problem of failure to capture the heterogeneous characteristics of depression in the existing technology, and realizes high-precision depression screening and personalized treatment.
Patent Information
- Application Number
- CN202510246920.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art fails to fully reveal the potential information in the data when dealing with the complex data structures of depression, especially failing to effectively capture the heterogeneous characteristics of depression, resulting in poor diagnostic accuracy and therapeutic effectiveness.
The data is processed based on subspace equalization division, and multiple subsample spaces are constructed through factor analysis, K-means clustering and SMOTE methods. Multiple screening models are constructed in combination with machine learning algorithms such as RF, GBDT, SVM, etc., and the diagnostic accuracy is improved through feature selection and model interpretation modules.
It has achieved rapid detection of high accuracy, high sensitivity and high specificity for patients with depression, significantly improved screening accuracy and efficiency, and can identify depression symptoms and potential causes in different subgroups, providing patients with personalized treatment.
Smart Images

Figure CN120280170A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent medical detection, and particularly to a depression classification model based on subspace equilibrium partitioning. Background Art
[0002] Depression is a widespread mental illness that profoundly affects an individual's emotional state, cognitive processes, and behavioral manifestations.
[0003] Depression exhibits significant heterogeneity, leading to diverse symptom manifestations and making it difficult to unify the diagnostic criteria. Some patients may mainly present with low mood, while others may mainly have somatic symptoms or anxiety. This diversity of symptoms increases the complexity of clinical diagnosis and easily leads to misdiagnosis or missed diagnosis. The heterogeneity of depression also means that the same diagnostic criteria may not be applicable to all patients, thereby reducing the accuracy of diagnosis. The heterogeneity of depression also results in different responses of patients to treatment, making it difficult for treatment plans based on unified diagnosis to meet the individual needs of each patient, thus affecting the treatment effect and the prognosis of patients. Therefore, accurately identifying depression patients and effectively solving the problem of depression are of great significance for improving public health levels, reducing social medical costs, and saving lives.
[0004] In recent years, machine learning technology has been increasingly widely applied in the medical field, which can effectively improve the ability to identify depression patients. However, the current systematic research on the heterogeneity of depression and the development of its screening models are still very limited. In particular, traditional data partitioning methods have deficiencies in dealing with the complex data structure of depression, failing to fully reveal the potential information in the data, especially failing to effectively capture the heterogeneous characteristics of depression, which limits the screening performance of machine learning models trained based on these data. Summary of the Invention
[0005] The present invention aims to solve the deficiencies of the prior art and provides a depression classification model based on subspace equilibrium partitioning to achieve accurate identification and screening of depression patients, thereby significantly improving the accuracy of diagnosis and the possibility of personalized treatment.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A depression classification model based on subspace equilibrium partitioning, comprising:
[0008] A data set processing and preparation unit, which collects the Revised BDC Scale, conducts statistics and depression classification, obtains complete sample data through outlier and duplicate value processing, and encodes the Revised BDC Scale using LabelEncoder to ensure the correct conversion and format adaptation of the model input data, meeting the requirements of the algorithm for data format;
[0009] Data space partitioning and data augmentation unit, which uses the SBP method to partition the encoded data to form multiple different sub-sample spaces, gradually divides the sub-sample spaces into training sub-spaces and test sub-spaces, performs data augmentation on the training sub-spaces to complete data balancing, and finally combines the training sub-spaces and test sub-spaces into a training set and a test set;
[0010] Depressive symptom screening factor module, which takes the partitioned data as input, uses feature selection techniques in machine learning combined with statistical chi-square tests to target the best combination of depressive symptom screening factors for a specific significance level to obtain the best set of depressive symptom screening factors;
[0011] Screening model module, which takes the depressive symptom screening factors that select the best feature subset as input, constructs multiple screening models including RF, GBDT, Bagging, SVM, SBP-RF, SBP-GBDT, SBP-Bagging, and SBP-SVM, uses grid search and cross-validation to establish an optimal model for the training set, and tests with the sample data in the test set to obtain a depression screening model;
[0012] Model evaluation module, which uses the accuracy, sensitivity, specificity, area under the receiver operating characteristic curve, and F1 score of the screening model as evaluation indicators to evaluate the application performance of the model with multi-dimensional indicators;
[0013] Model interpretation module, which uses the RF and GBDT algorithms to deeply analyze the contribution degree of depressive symptom screening factors, and realizes the efficient interpretation of depression screening results by calculating SHAP values.
[0014] Specifically, the dataset in the dataset processing and preparation unit consists of social psychology information, sociodemographic information, and the BDC scale. The revised BDC scale uses multi-dimensional information to enrich the content covered by the scale and improve data availability and model interpretability.
[0015] Specifically, the BDC scale is used for binary classification to reduce the complexity of screening for depressive patients, which specifically includes the following steps:
[0016] S1. On the premise of ensuring the anonymity of participants and obtaining their informed consent, each participant is required to fill out a questionnaire according to their feelings in the past week;
[0017] S2. Process the questionnaire data and perform binary classification according to the total BDC score and the evaluated depressive state;
[0018] Set 10 as the threshold, which is used as the standard for judging the presence or absence of depression. If the total BDC score of the participant exceeds this threshold, it is assigned a value of 1, indicating the presence of depression; if the score is equal to or lower than the threshold, it is assigned a value of 0, indicating no depression.
[0019] Specifically, in the data space division and data augmentation unit, a sample subspace is constructed by the SBP method, and data division is completed using the FA, K-means, and SMOTE methods, and virtual samples are built to expand the sample space of the training set, which specifically includes the following steps:
[0020] P1. Execute the FA algorithm on the original data set to extract several hidden factors FAC i , explaining the variance in the original variables. The internal correlation between the original feature variables and the hidden factor FAC i is represented by Equation (1);
[0021] X p×1 = A p×m FAC m×1 + ε p×1 (p < m); (1)
[0022] In the formula, A is the loading matrix, ε is the special factor, X p represents the original feature, p is the number of original features, and m represents the number of hidden factors;
[0023] The covariance Cov(FAC, ε) indicates that FAC i and ε are independent of each other;
[0024] The variance D(FAC) = I m , indicating that FAC i are uncorrelated with each other to ensure the accuracy and effectiveness of factor analysis. I m is the standard matrix;
[0025] P2. Use the extracted hidden factor FAC i as the input data set and adopt the K-means clustering algorithm to accurately divide K unique subspaces S i , and the formation of each S i strictly follows the carefully set optimization function, that is, Equation (2), to ensure that it can gradually converge to the optimal state during the iteration process;
[0026]
[0027] In the formula, x ij represents the sample in the subspace S i , i represents the subspace number, j represents the index of the sample in the subspace, u i , …, u k are the clustering centroids of S i , where K is the number of clusters or the number of subspaces, and n i represents the number of samples in the subspace S i ;
[0028] P3. Map the clustering labels to the original data. According to a specific ratio p:1-p (0 < p < 1), randomly sample a training subspace Str i and a test subspace Ste i from each subspace S of the original data with the corresponding labels; i of the instances;
[0029] P4. Apply the SMOTE method to correct the problem of sample imbalance for Str i . Equation (3) represents the minority sample generation function. During the sample generation stage, the feature information of the new sample x' new strictly depends on the neighboring samples x' i and x' ij in the training subspace Str ijt ;
[0030] x' new = x' ij +(x' ijt -x' ij ); (3)
[0031] In the formula, x' ij represents the j-th sample of the minority class in the training subspace Str i , x' ijt represents the t-th adjacent minority class sample of x' ij , and r is a random number within the range of [0, 1];
[0032] P5. Combine the training subspace Str i and the test subspace Ste i after sampling and correction processing into the final training set Str and test set Ste respectively.
[0033] Specifically, in the depression symptom screening factor module, use the SelectKBest feature mining method in machine learning to obtain the depression symptom screening factors related to depression. Select the features strongly correlated with the target variable through the SelectKBest method. Based on univariate analysis, select k optimal subsets from the feature set. Determine the best feature set by using the chi-square test, that is, Equation (4), (χ 2 , significance level p < 0.01);
[0034]
[0035] In the formula, n represents the total number of features, O i' represents the actual occurrence frequency of the i'-th feature, and E i' represents the screening occurrence frequency of the i'-th feature.
[0036] Specifically, in the screening model module, a multi-model depression screening system is established, which specifically includes the following steps:
[0037] N1. Based on the training set and validation set generated by using the traditional stratified random division method, on this basis, using the established depression symptom screening factors as input variables, and using RF classifier, Bagging classifier, SVM, and GBDT, a composite multi-classifier screening model is constructed;
[0038] N2. Based on the training set and validation set generated by using the SBP method, using the same depression symptom screening factors as input, and using RF classifier, Bagging classifier, SVM, and GBDT, construct SBP-RF, SBP-Bagging, SBP-SVM, and SBP-GBDT multi-classifier screening models;
[0039] N3. Determine the optimal model parameters for each classifier screening model using grid search and 5-fold cross-validation, and then use these optimal parameters to train the training set;
[0040] N4. Test the multiple classifier screening models on the test set. For the same test sample, each model independently outputs the screening result to complete the screening and auxiliary diagnosis of depression patients and healthy people.
[0041] Specifically, the evaluation index of the model evaluation module is defined by the confusion matrix, and the confusion matrix H is expressed as:
[0042]
[0043] Among them, TP: judged as positive, actually positive, that is, true positive; FN: judged as negative, actually positive, that is, false negative; FP: judged as positive, actually negative, that is, false positive; TN: judged as negative, actually negative, that is, true negative;
[0044] The expressions of accuracy, sensitivity, specificity, and F1 score are as follows:
[0045]
[0046] The beneficial effects of the present invention are:
[0047] In the process of depression screening, the present invention ingeniously incorporates the consideration of depression heterogeneity. Through a refined data partitioning strategy, it can achieve rapid batch detection of depression with high accuracy, high sensitivity, and high specificity, significantly improving the screening precision and efficiency, solving the problem that it is difficult to capture the heterogeneous characteristics of depression in existing data partitioning, and the deficiencies of low accuracy, sensitivity, and specificity of detection indicators in screening and detecting depression, thus realizing reliable intelligent depression detection; this model can more acutely identify one or more specific depressive symptoms and their potential causes in different sub-groups, so as to provide targeted treatment for patients. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is the overall working framework diagram of the model of the present invention;
[0049] Figure 2 It is the screening flow chart of the revised BDC scale of the present invention
[0050] Figure 3 It is the working flow chart of SBP of the present invention;
[0051] Figure 4 It is the detailed working flow chart of SBP of the present invention;
[0052] Figure 5 It is the overall working flow chart of the model of the present invention;
[0053] Figure 6 It is the importance ranking diagram of depression predictors of the present invention;
[0054] Figure 7 It is the SHAP value ranking diagram of the present invention;
[0055] The following will be described in detail with reference to the embodiments of the present invention and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] The present invention will be further described below in conjunction with the embodiments:
[0057] As Figure 1 shown, a depression classification model based on subspace balanced partitioning includes a data set processing and preparation unit, a data space partitioning and data augmentation unit, a depressive symptom screening factor module, a screening model module, a model evaluation module, and a model interpretation module.
[0058] As Figure 2As shown in the figure, the data set processing and preparation unit collects the revised BDC scale, conducts statistics and depression classification, and obtains the complete sample data through outlier and duplicate value processing. The data set consists of social psychology information, sociodemographic information, and the BDC scale. The revised BDC scale uses multi-dimensional information to enrich the content covered by the scale and improve the data availability and model interpretability. The BDC scale is used for binary classification to reduce the complexity of screening for depression patients, and the specific steps are as follows:
[0059] S1. On the premise of ensuring the anonymity of participants and obtaining their informed consent, each participant is required to fill out a questionnaire according to their feelings in the past week;
[0060] S2. Process the questionnaire data and perform binary classification according to the total BDC score and the evaluated depression status;
[0061] Set 10 as the threshold, which is used as the standard for judging the presence or absence of depression. If the total BDC score of the participant exceeds this threshold, it is assigned a value of 1, indicating the presence of depression; if the score is equal to or lower than the threshold, it is assigned a value of 0, indicating no depression.
[0062] Subsequently, on this basis, the LabelEncoder is used to encode the revised BDC scale to ensure the correct conversion and format adaptation of the model input data and meet the requirements of the algorithm for the data format.
[0063] As Figure 3 、 Figure 4 shown in the figure, the data space division and data augmentation unit uses the SBP method to divide the encoded data to form multiple different sub-sample spaces, gradually divides the sub-sample spaces into training sub-spaces and test sub-spaces, performs data augmentation on the training sub-spaces, generates intra-group similar samples, and realizes data balance with the minimum noise. Finally, the training set and test set are merged by the training sub-space and the test sub-space;
[0064] Specifically, the sample sub-space is constructed by the SBP method, and the data division is completed by using factor analysis (FA), K-means, and SMOTE methods, and virtual samples are built to expand the sample space of the training set. The specific steps are as follows:
[0065] P1. Execute the FA algorithm on the original data set to extract several hidden factors FAC i , explain the variance in the original variables, and the internal correlation between the original feature variables and the hidden factor FAC i is represented by Equation (1);
[0066] X p×1 =A p×m FAC m×1 +ε p×1(p <m); (1)
[0067] Where A is the load matrix, ε is the special factor, and X p represents the original features, p is the number of original features, and m represents the number of hidden factors;
[0068] Covariance Cov(FAC,ε), which represents FAC i and ε are independent of each other;
[0069] Variance D(FAC) = I m , indicating FAC i There is no correlation between them to ensure the accuracy and effectiveness of factor analysis. m is a standard matrix;
[0070] P2, extract the hidden factor FAC i As the input data set, the K-means clustering algorithm is used to accurately divide K unique subspaces S i , each S i The formation of strictly follows the carefully set optimization function, namely, formula (2), to ensure that it can gradually converge to the optimal state during the iteration process;
[0071]
[0072] In the formula, x ij Represents the subspace S i The samples in the subspace, i represents the number of the subspace, j represents the index of the sample in the subspace, and u i ,…,u k YesS i The cluster centroids of , where K is the number of clusters or the number of subspaces, n i Denotes the subspace S i The number of samples in ;
[0073] P3, correspond the cluster labels to the original data, according to a specific ratio p:1-p(0 <p<1),从每个对应好标签的原始数据的子空间S i Randomly extract training subspace Str i and the test subspace Ste i Examples of
[0074] P4, for Str i The SMOTE method is used to correct the problem of sample imbalance. Formula (3) represents the minority sample generation function. In the sample generation stage, the new sample x' new The feature information strictly depends on the training subspace Str i The neighboring sample x' in ij and x' ijt ;
[0075] x' new = x' ij +(x' ijt - x' ij ); (3)
[0076] Wherein, x' ij represents the j-th sample of the minority class in the training subspace Str i , x' ijt represents the t-th adjacent minority class sample of x' ij , and r is a random number within the range of [0, 1];
[0077] P5. Combine the training subspace Str i and the test subspace Ste i that have undergone sampling and calibration processing into the final training set Str and test set ste respectively.
[0078] The depressive symptom screening factor module takes the partitioned data as input, and uses the feature selection technology in machine learning combined with the statistical chi-square test to optimize the best set of depressive symptom screening factors for a specific significance level;
[0079] Specifically, use the SelectKBest feature mining method in machine learning to obtain the depressive symptom screening factors related to depression. Select the features with strong correlation with the target variable through the SelectKBest method. Based on univariate analysis, select k optimal subsets from the feature set. Determine the best feature set by using the chi-square test, i.e., Equation (4), (χ 2 , significance level p < 0.01);
[0080]
[0081] Wherein, n represents the total number of features, O i' represents the actual occurrence frequency of the i'-th feature, and E i' represents the screening occurrence frequency of the i'-th feature.
[0082] Compared with the traditional data generation method, through this generation method, the new sample x' new not only maintains a high degree of consistency with the original sample, but also further enhances its similarity within the subspace Str i .
[0083] Such as Figure 5As shown in the figure, the screening model module takes the screening factors of depressive symptoms that select the best feature subset as input, constructs multiple screening models including RF, GBDT, Bagging, SVM, SBP-RF, SBP-GBDT, SBP-Bagging, and SBP-SVM. It establishes the optimal model for the training set using grid search and cross-validation, and uses the sample data in the test set for testing to obtain the depression screening model;
[0084] The screening model module establishes a multi-model depression screening system, which specifically includes the following steps:
[0085] N1. Based on the training set and validation set generated by using the traditional stratified random division method, on this basis, using the established screening factors of depressive symptoms as input variables, and using the RF classifier, Bagging classifier, SVM, and GBDT, a composite multi-classifier screening model is constructed;
[0086] N2. Based on the training set and validation set generated by using the SBP method, using the same screening factors of depressive symptoms as input, and using the RF (Random Forest) classifier, Bagging classifier, SVM (Support Vector Machine), and GBDT (Gradient Boosting Decision Tree), construct the SBP-RF, SBP-Bagging, SBP-SVM, and SBP-GBDT multi-classifier screening models;
[0087] N3. Use grid search and 5-fold cross-validation to determine the best model parameters for each classifier screening model, and then use these optimal parameters to train the training set;
[0088] N4. Conduct test set tests on multiple classifier screening models. For the same test sample, each model independently outputs the screening results to complete the screening and auxiliary diagnosis of depression patients and healthy people.
[0089] The model evaluation module uses the accuracy, sensitivity, specificity, the area under the receiver operating characteristic curve (ROC) (AUC), and the F1 score of the screening model as evaluation indicators to evaluate the application performance of the model with multi-dimensional indicators;
[0090] The evaluation indicators of the model evaluation module are defined by the confusion matrix, and the confusion matrix H is expressed as:
[0091]
[0092] Among them, TP: judged as positive, actually positive, that is, true positive; FN: judged as negative, actually positive, that is, false negative; FP: judged as positive, actually negative, that is, false positive; TN: judged as negative, actually negative, that is, true negative;
[0093] The indicators in the field of disease detection include accuracy, sensitivity, specificity, AUC, and F1-score; accuracy refers to the probability of correct diagnosis among all screened and diagnosed samples; sensitivity refers to the probability of no missed diagnosis during screening and diagnosis; specificity refers to the probability of no misdiagnosis during screening and diagnosis; AUC refers to the area under the ROC curve formed by sensitivity and (1 - specificity) at different classification thresholds; F1-score is the harmonic mean of accuracy and sensitivity.
[0094] The expressions of accuracy, sensitivity, specificity, and F1-score are as follows:
[0095]
[0096] The model interpretation module uses the RF and GBDT algorithms to deeply analyze the contribution of depression screening factors, and realizes the efficient interpretation of depression screening results by calculating SHAP values.
[0097] Example 1:
[0098] 1) Dataset processing and preparation unit:
[0099] There are 397 depression samples and 207 healthy person samples. After removing outliers and duplicate data, a total of 594 valid samples are obtained; since all the features involved in this experimental study are categorical features, LabelEncoder is used to encode the categorical variables, and appropriate conversions are made for each column of data that can be classified to meet the input requirements of the model.
[0100] 2) Data space partitioning and data augmentation unit:
[0101] The encoded sample data is processed using the SBP method. First, the FA algorithm is executed on the original encoded data set to extract several hidden factors that can explain most of the variance in the original variables;
[0102] After that, the extracted hidden factors are used as the input data set, and the k-means clustering algorithm is used to accurately divide multiple unique sample subspaces and obtain corresponding labels;
[0103] Subsequently, the clustering labels are mapped to the original data, and according to a ratio of 4:1, instances of the training subspace and the test subspace are randomly selected from each subspace of the original data with corresponding labels;
[0104] Then, the SMOTE method is used to perform data augmentation on the training subspace, and data balance is achieved by forming more similar samples within the group, maximizing the reduction of noise generated by data augmentation;
[0105] Finally, the trained subspaces and test subspaces after sampling and calibration are respectively combined into the final training set and test set.
[0106] 3) Depression symptom screening factor module:
[0107] Using the SelectKBest feature mining method in machine learning combined with the chi-square test in statistics, the depression screening factors with a significance level p < 0.01 are preferentially selected, and the statistical characteristics of the depression screening factors in the training set are analyzed to reduce the noise input.
[0108] 4) Screening model module:
[0109] For the screening of depression, using the tested samples as input information, multiple depression detection models are constructed by RF, GBDT, Bagging, SVM, SBP-RF, SBP-GBDT, SBP-Bagging, and SBP-SVM based on the training set combined with grid search and 5-fold cross-validation to complete the screening and auxiliary diagnosis of depression patients and healthy people;
[0110] More specifically, among the best hyperparameters obtained by grid search and 5-fold cross-validation based on the training set,
[0111] The set numbers of decision trees selected by SBP-RF and RF models sum to {500; 200};
[0112] The loss functions of SBP-GBDT and GBDT are set to logarithmic likelihood loss, which is applicable to binary classification, and the set numbers of decision trees selected sum to {200; 150};
[0113] SBP-Bagging and Bagging models use the entropy of information gain to select the best split, and the set numbers of decision trees selected sum to {140; 500};
[0114] SBP-SVM and SVM use the radial basis function kernel, and the selected regularization parameter arrays sum to {100, 100}.
[0115] 5) Model evaluation module:
[0116] The evaluation index of the model evaluation module of the depression classification model can be defined by the confusion matrix, and the confusion matrix H is expressed as:
[0117]
[0118] where TP: judged as positive, actually positive, i.e., true positive; FN: judged as negative, actually positive, i.e., false negative; FP: judged as positive, actually negative, i.e., false positive; TN: judged as negative, actually negative, i.e., true negative;
[0119] The indicators in the field of disease detection include accuracy, sensitivity, specificity, AUC, and F1 score; accuracy refers to the probability of correct diagnosis among all screened and diagnosed samples; sensitivity refers to the probability of no missed diagnosis during screening and diagnosis; specificity refers to the probability of no misdiagnosis during screening and diagnosis; AUC refers to the area under the ROC curve formed by sensitivity and (1 - specificity) at different classification thresholds; F1 score is the harmonic mean of accuracy and sensitivity. The expressions of accuracy, sensitivity, specificity, and F1 score are as follows:
[0120]
[0121]
[0122] In the evaluation of the depression screening and diagnosis model, multiple evaluation indicators such as accuracy, sensitivity, specificity, AUC, and F1 score are used to objectively evaluate the application performance of the model from multiple dimensions.
[0123] The results of the evaluation indicators in this Example 1 are shown in Table 1:
[0124] Table 1 - Results of evaluation indicators (accuracy, sensitivity, specificity, AUC, and F1 score)
[0125]
[0126] As can be seen from Table 1, the four models based on the SBP method, namely SBP - GBDT, SBP - RF, SBP - Bagging, and SBP - SVM, are generally superior to the traditional GBDT, RF, Bagging, and SVM models in terms of accuracy, sensitivity, specificity, AUC, and F1 score, and their performance is stable. The SBP method has better application potential in constructing a depression screening model.
[0127] 6) Model interpretation module:
[0128] The RF and GBDT algorithms are used to deeply analyze the contribution degree of depression screening factors. By calculating the SHAP value, the efficient interpretation of the depression screening results is realized.
[0129] The contribution degree of depression screening factors in the RF and GBDT algorithms is as Figure 6 shown, and the SHAP value is as Figure 7 shown.
[0130] In the process of depression screening, the present invention ingeniously incorporates the consideration of the heterogeneity of depression. Through a refined data partitioning strategy, it can achieve rapid batch detection of depression with high accuracy, high sensitivity, and high specificity, significantly improving the screening precision and efficiency. It solves the defect that it is difficult to capture the heterogeneous characteristics of depression in the existing data partitioning, and the low accuracy, sensitivity, and specificity of the detection indicators in screening and detecting depression, realizing reliable intelligent depression detection. This model can more sensitively identify a specific one or more depressive symptoms and their potential causes in different sub-groups, so as to provide targeted treatment for patients.
[0131] The above is an exemplary description of the present invention. Obviously, the specific implementation of the present invention is not limited by the above methods. As long as various improvements are made by adopting the method concept and technical solution of the present invention, or directly applied to other occasions without improvement, they are all within the protection scope of the present invention.
Claims
1. A depression classification model based on subspace equilibrium partitioning, characterized in that, Including: A data set processing and preparation unit that collects the revised BDC scale, conducts statistics and depression classification, obtains complete sample data through outlier and duplicate value processing, and encodes the revised BDC scale using LabelEncoder to ensure the correct conversion and format adaptation of the model input data and meet the requirements of the algorithm for data format; A data space division and data augmentation unit that uses the SBP method to divide the encoded data to form multiple different sub-sample spaces, gradually divides the sub-sample spaces into training sub-spaces and test sub-spaces, augments the data in the training sub-spaces to complete data balancing, and finally combines the training sub-spaces and test sub-spaces into a training set and a test set; A depression symptom screening factor module that takes the divided data as input and uses feature selection techniques in machine learning combined with statistical chi-square tests to select the best set of depression symptom screening factors for a specific significance level; A screening model module that takes the depression symptom screening factors that select the best feature subset as input, constructs multiple screening models such as RF, GBDT, Bagging, SVM, SBP-RF, SBP-GBDT, SBP-Bagging, and SBP-SVM, establishes an optimal model for the training set using grid search and cross-validation, and tests using the sample data in the test set to obtain a depression screening model; A model evaluation module that uses the accuracy, sensitivity, specificity, area under the receiver operating characteristic curve, and F1 score of the screening model as evaluation indicators to evaluate the application performance of the model with multi-dimensional indicators; A model interpretation module that uses the RF and GBDT algorithms to deeply analyze the contribution degree of the depression screening factors and realizes the efficient interpretation of the depression screening results by calculating the SHAP value.
2. The classification model for depression based on subspace equilibrium partitioning according to claim 1, wherein The data set in the data set processing and preparation unit consists of social psychology information, sociodemographic information, and the BDC scale. The revised BDC scale uses multi-dimensional information to enrich the content covered by the scale and improve the data availability and model interpretability.
3. The classification model for depression based on subspace equilibrium partitioning according to claim 2, wherein, The BDC scale is used for binary classification to reduce the complexity of screening for depression patients, specifically including the following steps: S1. On the premise of ensuring the anonymity of participants and obtaining their informed consent, each participant is required to fill out a questionnaire according to their feelings in the past week; S2. Process the questionnaire data and conduct binary classification according to the total BDC score and the evaluated depression status; Set 10 points as the threshold, which is used as the standard for judging the presence or absence of depression. If the total BDC score of the participant exceeds this threshold, it is assigned a value of 1, indicating the presence of depression; if the score is equal to or lower than the threshold, it is assigned a value of 0, indicating no depression.
4. The classification model for depression based on subspace equilibrium partitioning according to claim 1, characterized in that In the data space division and data augmentation unit, construct sample sub-spaces through the SBP method, complete data division using the FA, K-means, and SMOTE methods, and build virtual samples to expand the sample space of the training set, specifically including the following steps: P1. Execute the FA algorithm on the original data set to extract several hidden factors FAC i , and explain the variance in the original variables. The internal correlation between the original feature variables and the hidden factor FAC i is represented by Equation (1); X p×1 = A p×m FAC m×1 + ε p×1 (p < m); (1) where A is the load matrix, ε is the special factor, and X p represents the original features, p is the number of original features, and m represents the number of hidden factors; The covariance Cov(FAC, ε) indicates that FAC i and ε are independent of each other; The variance D(FAC) = I m , indicating that FAC i are uncorrelated with each other to ensure the accuracy and effectiveness of factor analysis, and I m is the standard matrix; P2. Use the extracted hidden factor FAC i as the input data set, and adopt the K-means clustering algorithm to accurately divide K unique subspaces S i , where each S i is formed in strict accordance with the carefully set optimization function, i.e., Equation (2), to ensure that it can gradually converge to the optimal state during the iteration process; where x ij represents a sample in subspace S i , i represents the subspace number, j represents the index of the sample within that subspace, and u i , …, u k are the cluster centroids of S i , where K is the number of clusters or the number of subspaces, and n i represents the number of samples in subspace S i ; P3. Map the cluster labels to the original data, and randomly select training subspace Str i and test subspace Ste i from each subspace S of the original data with the corresponding labels according to a specific ratio p:1-p (0 < p < 1); i instance. P4. Regarding Str i Apply the SMOTE method to correct the problem of sample imbalance. Equation (3) represents the minority sample generation function. During the sample generation stage, the feature information of the new sample x' new strictly depends on the neighboring samples x' i and x' ij in the training subspace Str ijt ; x' new = x' ij +(x' ijt - x' ij ); (3) where x' ij represents the j-th sample of the minority class in the training subspace Str i , x' ijt represents the t-th adjacent minority class sample of x' ij , and r is a random number within the range of [0, 1]; P5. Combine the trained subspace Str i and the test subspace Ste i that have undergone sampling and calibration processes into the final training set Str and test set Ste respectively.
5. The depression classification model based on subspace equilibrium partitioning according to claim 1, characterized in that, In the depressive symptom screening factor module, the SelectKBest feature mining method in machine learning is used to obtain the depressive symptom screening factors related to depression. The SelectKBest method is used to select the features with strong correlation with the target variable. Based on univariate analysis, k optimal subsets are selected from the feature set. By using the chi-square test, that is, Equation (4), (χ 2 , significance level p < 0.01) to determine the best feature set; where n represents the total number of features, O i' represents the actual occurrence frequency of the i'-th feature, and E i' represents the screening occurrence frequency of the i'-th feature.
6. The classification model for depression based on subspace equilibrium partitioning according to claim 1, wherein In the screening model module, establish a multi-model depression screening system, specifically including the following steps: N1. Based on the training set and validation set generated by using the traditional stratified random partitioning method, on this basis, using the established depression symptom screening factors as input variables, and using RF classifier, Bagging classifier, SVM, and GBDT, a composite multi-classifier screening model is constructed; N2. Based on the training set and validation set generated by using the SBP method, using the same depression symptom screening factors as input, and using RF classifier, Bagging classifier, SVM, and GBDT, SBP-RF, SBP-Bagging, SBP-SVM, and SBP-GBDT multi-classifier screening models are constructed; N3. Determine the optimal model parameters for each classifier screening model using grid search and 5-fold cross-validation, and then use these optimal parameters to train the training set; N4. Test the multiple classifier screening models on the test set. For the same test sample, each model independently outputs the screening result to complete the screening and auxiliary diagnosis of depression patients and healthy individuals.
7. A depression classification model based on subspace equilibrium partitioning according to claim 1, characterized in that, The evaluation indicators of the model evaluation module are defined by the confusion matrix. The confusion matrix H is expressed as: Among them, TP: judged as positive, actually positive, that is, true positive; FN: judged as negative, actually positive, that is, false negative; FP: judged as positive, actually negative, that is, false positive; TN: judged as negative, actually negative, that is, true negative; The expressions of accuracy, sensitivity, specificity, and F1 score are as follows: