An Unsupervised Method for Depression Classification and Stage Inference Based on Structural Features

CN122575722APending Publication Date: 2026-08-14UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-26
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0005]为了解决现有抑郁症临床分型过于依赖回顾性临床描述、缺乏客观神经病理学支撑,以及现有分析方法无法同时实现个体化亚型识别与病理进展量化的问题,本发明提出一种基于结构特征的无监督抑郁症分型与阶段推理方法

Benefits of technology

[0013]本发明的有益效果:本发明通过全脑尺度的无监督建模框架,实现了对抑郁症脑结构异质性的深度解析与病程演进的客观量化。首先,本发明克服了传统回顾性临床分类(如首发/复发划分)的粗糙性,从数据驱动的神经解剖学视角识别出具有生物学解释力的客观亚型(如小脑起始型、扣带回起始型、基底节起始型),为抑郁症的异质性机制探索提供了客观依据。其次,本发明突破了传统聚类方法仅能识别离散亚型的局限,通过引入分段线性轨迹模型,在同一个数学框架下实现了对患者“所属亚型”与“病理阶段”的双重定位,能够精准刻画脑结构异常随病程连续演进的动态过程。此外,通过 16 个核心 ROI 的映射聚合策略,本发明在保留核心解剖信息的同时显著降低了高阶生成式模型在排列空间搜索时的计算开销,并结合多源干扰信号回归校正,确保了分型结果在多中心大型数据集上的稳健性。最后,本发明通过自动化输出亚型归属概率与阶段评分,有效解决了传统全群体平均分析因忽视个体间差异化受损路径以及病程阶段混叠而导致的病理特征“平减”问题,成功从复杂异质的临床样本中剥离出具有高度特异性的脑结构受损演进信号,为抑郁症的病情监测、严重程度量化及个性化诊疗决策提供了科学、客观的数据支持。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575722A_ABST
    Figure CN122575722A_ABST
Patent Text Reader

Abstract

This invention discloses an unsupervised method for classifying and inferring stages of depression based on structural features, applied to the field of neuroimaging information processing. It addresses the problems of existing clinical classifications of depression relying too heavily on retrospective clinical descriptions, lacking objective neuropathological support, and the inability of existing analytical methods to simultaneously achieve individualized subtype identification and quantification of pathological progression. This invention first acquires imaging data and extracts brain region features, mapping them to 16 core regions of interest. Next, it performs regression correction for multi-source interference signals, defining the abnormal evolution of biomarkers as a piecewise linear function changing over time. Then, it uses an evolutionary modeling algorithm based on piecewise linear functions, combined with information criteria and likelihood assessment, to determine the optimal number of subtypes, and obtains globally optimal evolutionary path parameters through a sampling strategy. Finally, it fits the test samples to the optimal evolutionary model, determines the subtype classification by calculating posterior probabilities, and outputs the pathological stage score of the sample on the corresponding trajectory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of neuroimaging information processing technology, and specifically relates to a magnetic resonance imaging classification technique for depression. Background Technology

[0002] Major Depressive Disorder (MDD) is a leading cause of functional disability worldwide and exhibits high clinical heterogeneity. Current clinical research and assessment primarily rely on retrospective clinical phenotypic descriptions, such as classifying patients into adolescent / adult onset groups or first-episode / relapse subtypes based on age of onset and frequency of illness. While this classification method has some reference value in clinical management, its criteria are rather crude, often ignoring the complex spatial distribution of brain structural damage caused by depression, and lacking data-driven objective neuropathological support. This makes it difficult for traditional classifications to accurately characterize the essential neurobiological differences among MDD patients, and also fails to explain why patients with similar clinical labels exhibit significant deviations in disease progression trajectories.

[0003] Advances in magnetic resonance imaging (MRI) have provided an opportunity to reveal the objective neurobiological heterogeneity of malignant dysplasia (MDD). Large-scale neuroimaging studies have shown that MDD patients commonly exhibit structural damage in key brain regions such as the hippocampus, anterior cingulate cortex, insula, and frontal lobe. However, most existing neuroimaging analysis paradigms are based on the traditional case-control model, focusing on extracting static average differences between groups. This approach forcibly categorizes patients at different pathological stages and with different damage patterns into the same group, obscuring the individualized trajectory of structural damage. More importantly, existing unsupervised analysis techniques often separate "subtype identification" from "stage classification." When dealing with a disease like depression, where damage characteristics are subtle and atypically distributed, traditional discrete classification methods cannot capture the dynamic process of continuous changes in brain structural abnormalities throughout the disease course, resulting in an inability to objectively quantify the degree of pathological progression in patients.

[0004] Meanwhile, in the trend of multi-center collaborative research, systematic biases from different scanning sites often interfere with real biological signals. When processing high-dimensional anatomical spatial data encompassing the whole brain, how to effectively eliminate site bias while achieving robust trajectory search and parameter estimation with limited computational resources remains a pressing challenge in the current technological field. Therefore, an analytical scheme that integrates cross-site correction, high-order evolutionary trajectory modeling, and automated staged inference is needed to reveal the deep-seated subtyping patterns and evolutionary logic of depression from a data-driven perspective, providing a scientific basis for accurately assessing pathological states. Summary of the Invention

[0005] To address the problems of existing clinical classifications of depression relying too heavily on retrospective clinical descriptions and lacking objective neuropathological support, as well as the inability of existing analytical methods to simultaneously achieve individualized subtype identification and quantification of pathological progression, this invention proposes an unsupervised depression classification and stage reasoning method based on structural features.

[0006] The technical solution adopted in this invention is: an unsupervised depression classification and staged reasoning method based on structural features, comprising:

[0007] S1. Extraction of primitive brain region features: Obtain structural magnetic resonance imaging data of the subject and use multimodal brain atlas to extract the primitive gray matter volume features of 122 primitive brain regions, including 120 standard brain regions and the nucleus accumbens.

[0008] S2, Multi-level Region of Interest Mapping and Aggregation: 122 original brain regions are mapped into 16 core regions of interest covering the cortex and subcortical regions, and the mean gray matter volume of non-zero voxels in each region of interest is calculated.

[0009] S3. Multi-source interference signal regression correction: A generalized linear model is constructed based on the mean gray matter volume of non-zero voxels in each region of interest. The systematic interference of subject gender, age, age squared, scanning site and total intracranial volume on brain structure data is eliminated through regression analysis to obtain the corrected residual characteristic value.

[0010] S4. Construction of Standardized Feature Matrix: Using the healthy control group as a benchmark, perform Z-score transformation on the corrected residual feature values ​​to obtain the standardized deviation matrix that quantifies the degree of brain structural damage;

[0011] S5. Construct a subtype evolution model based on piecewise linear functions, and obtain the order of damage to each region of interest on the time axis based on the standardized deviation matrix, thus obtaining the evolution order.

[0012] S6. Based on the corrected residual characteristic values ​​of the test subjects obtained in step S3 and the evolution sequence of the test subjects obtained in step S5, calculate the posterior probability to determine the subtype classification.

[0013] The beneficial effects of this invention are as follows: This invention achieves in-depth analysis of the heterogeneity of brain structure in depression and objective quantification of disease progression through a whole-brain-scale unsupervised modeling framework. First, this invention overcomes the crudeness of traditional retrospective clinical classifications (such as first-episode / relapse classification), identifying objective subtypes with biological explanatory power (such as cerebellar origin, cingulate gyrus origin, and basal ganglia origin) from a data-driven neuroanatomical perspective, providing objective evidence for exploring the heterogeneity mechanisms of depression. Second, this invention breaks through the limitation of traditional clustering methods that can only identify discrete subtypes. By introducing a piecewise linear trajectory model, it achieves dual localization of the patient's "subtype" and "pathological stage" within the same mathematical framework, accurately depicting the dynamic process of continuous evolution of brain structural abnormalities throughout the disease. Furthermore, through a mapping and aggregation strategy of 16 core ROIs, this invention significantly reduces the computational overhead of higher-order generative models in permutation space search while preserving core anatomical information, and combined with multi-source interference signal regression correction, ensures the robustness of the classification results on large multi-center datasets. Finally, by automatically outputting subtype attribution probabilities and stage scores, this invention effectively solves the problem of "flattening" of pathological features caused by the neglect of individual differences in damaged pathways and overlapping disease stages in traditional population average analysis. It successfully extracts highly specific brain structural damage evolution signals from complex and heterogeneous clinical samples, providing scientific and objective data support for the monitoring of depression, the quantification of its severity, and personalized diagnosis and treatment decisions. Attached Figure Description

[0014] Figure 1 This is the overall network flowchart of the present invention.

[0015] Figure 2 This is a diagram of the method architecture of the present invention.

[0016] Figure 3 This is a schematic diagram illustrating the evaluation process for determining the optimal number of subtypes in an embodiment of the present invention;

[0017] Among them, (a) is the cross-validation information criterion (CVIC) curve, (b) is the log-likelihood distribution plot, (c) is the MCMC sampling convergence trajectory, and (d) is the histogram of the likelihood values ​​of the optimal subtype evolution model.

[0018] Figure 4 is a topographic map showing the pathological evolution trajectory of the three depression subtypes identified in this invention across 16 core regions of interest (ROIs) in the whole brain. Detailed Implementation

[0019] To facilitate understanding of the technical content of this invention by those skilled in the art, the following description, in conjunction with the accompanying drawings, further illustrates the invention.

[0020] like Figures 1 to 4As shown, this invention discloses an unsupervised method for depression classification and staged reasoning based on structural features. Among them, Figure 1 This is an overall flowchart of the present invention; Figure 2 This is a schematic diagram of the method framework of the present invention; Figure 3 This is a parameter evaluation diagram used in the embodiments of the present invention to determine the optimal number of subtypes and verify the convergence of the model; Figure 4 This is a pathological evolution trajectory diagram of the three depression subtypes identified in this invention across 16 core regions of interest (ROIs). The implementation process of the method of this invention includes the following steps:

[0021] 1. This embodiment uses structural imaging data from the first phase of the Rest-meta-MDD (Chinese Depression Brain Imaging Consortium Large-Scale Multicenter Study) involving depressed and healthy subjects. To ensure the homogeneity of the data and the reliability of the analysis, strict automated and manual quality control was performed on the original structural imaging data. Finally, 1265 patients with depression (MDD group) and 1089 healthy controls (HC group) were included after quality control. The original structural imaging data were preprocessed using the Statistical Parameter Mapping (SPM) plugin and the Computational Anatomy Toolkit (CAT12), specifically including: (1) segmenting brain structures and tissues using CAT12 to obtain the tissue component maps, specifically dividing the structural imaging data into gray matter, white matter and cerebrospinal fluid; (2) registering individual brain images to the MNI (Montreal Neuroscience Institute) standard space by calling the SPM underlying algorithm through CAT12. Here, the individual brain images include the original structural imaging data and the segmented tissue component maps; (3) performing head motion correction and artifact removal by calling the SPM underlying algorithm through CAT12.

[0022] Those skilled in the art will know that SPM and CAT12 are professional neuroimaging processing software based on MATLAB.

[0023] Subsequently, gray matter volume (GMV) features of 122 primitive brain regions were extracted based on the hybrid anatomical atlas of the second and third editions of anatomical autolabelling (AAL2 and AAL3).

[0024] To balance computational complexity and anatomical resolution, the 122 original brain regions were mapped and aggregated into 16 core regions of interest (ROIs). The mapping rules are as follows: the cortical region was aggregated into 9 ROIs: frontal lobe, temporal lobe, parietal lobe, occipital lobe, insula, cerebellum, central region, cingulate gyrus, and parahippocampal gyrus; the subcortical region retained 7 ROIs: hippocampus, amygdala, caudate nucleus, putamen, globus pallidus, thalamus, and nucleus accumbens (NAC).

[0025] The arithmetic mean of the gray matter volumes of all voxels within a Region of Interest (ROI) is extracted as the ROI feature, and a structural feature vector is constructed based on the features of each ROI.

[0026] 2. For the extracted ROI features, a generalized linear model is constructed to eliminate covariate interference. Specifically, for subject j at the i-th ROI feature value... The following regression equation is established:

[0027]

[0028] in, , , , , , These are the regression coefficients; These are the subject j's age, sex, and total intracranial volume; This is the random error term of the regression equation, used to characterize the differences in individual brain structural features after removing the interference of the above covariates; The s-th component in the scan site vector using one-hot encoding is used to characterize whether subject j is scanned at the s-th site, where s represents the scan site number. The total number of sites S is determined based on the actual number of medical centers or scanner centers included in the multi-center database. In this embodiment, S corresponds to the number of sites 23 in the Rest-meta-MDD dataset.

[0029] Calculate the residual eigenvalues ​​after removing covariates using this generalized linear regression model. (i.e., the random error term) The observed value), its calculation formula is:

[0030]

[0031] in, , , , , , The regression coefficients are estimated values, obtained by fitting using ordinary least squares (OLS). It is by The remaining portion after removing variations caused by covariates such as age, sex, site, and total intracranial volume.

[0032] To eliminate systematic bias in cross-center data, this invention employs the ComBat algorithm of Empirical Bayesian methods to remove batch effects from the residual feature matrix. Subsequently, a standardization transformation is performed using a healthy control group as a baseline. To ensure that the obtained Z-score values ​​positively characterize the degree of brain structural damage, the Z-score value for the i-th ROI is calculated using the following formula:

[0033]

[0034] in, and Let be the arithmetic mean and standard deviation of the voxel gray matter volumes of the healthy control group within the i-th ROI, respectively. Through this transformation, the obtained... This represents the standardized deviation of subject j at the i-th ROI, which is calculated by using the residual eigenvalues. Distribution relative to the healthy control group (mean) with standard deviation The value is obtained by standardizing the deviation of the gray matter volume, and a positive value indicates a higher degree of abnormality in the gray matter volume of the brain region relative to the normal level. Population-level characteristic statistical analysis revealed that, after correction for false discovery rate (FDR), the Cohen's d index of brain regions showing significant inter-group differences at the population level was less than 0. This confirms that the MDD subjects generally exhibited a neuropathological phenotype of abnormal gray matter volume atrophy, providing theoretical support for subsequent evolutionary trajectory modeling and subtype classification.

[0035] 3. This invention defines disease progression as the progression over time t ( Continuously varying piecewise linear functions It comprises N Z-score events, representing the sum of abnormal threshold events across all regions of interest (ROIs). The value of this Z-score is determined by the product of the number of ROIs involved in the modeling and the number of progression stage thresholds set for each ROI. For example, when there are M ROIs and K progression stage thresholds set for each ROI, the total number of events N = M × K. In this case, the overall disease progression is divided into N+1 discrete stages, including a baseline stage where no events occur. In practical applications, the number of progression stage thresholds K for each ROI is typically set to 2 to 5 based on clinical granularity requirements, with each stage threshold... A stepped threshold is set based on the standard deviation of the threshold from the healthy control group distribution to quantify the degree of abnormal damage at different levels. N is the total number of events allowed in the entire model, i.e., the sum of the threshold points in all ROI regions. In this embodiment, the current ROI region K=3 is set, and the threshold points for the three progression stages are as follows: (Slight deviation) (Moderate damage) (Severe abnormality).

[0036] For each ROI i, its anomaly accumulation trajectory The mathematical expression is as follows:

[0037]

[0038] Where t is the normalized pathological evolution time, and its value ranges from [0, 1]. Let represent the abnormal cumulative feature value of the i-th ROI over disease evolution time t. This serves as the theoretical benchmark function for subsequent iterative optimization and probabilistic inference, achieved by combining the actual observed characteristics of the subjects with a piecewise linear function. The predicted theoretical abnormality level is fitted and compared to calculate the probability that the subject belongs to a specific subtype and a specific stage of progression. This represents the maximum abnormal threshold of the biomarker throughout its evolution (in this embodiment, the value is 2). , … This represents the Z-score threshold corresponding to the first, second, ..., R level abnormal states. , … This represents the expected time point when the i-th ROI reaches the level 1, 2, ..., R abnormal state (i.e., triggering the corresponding event). Its value is determined by data-driven parameter optimization based on the sample data distribution through maximum likelihood estimation or Markov chain Monte Carlo iteration (MCMC) algorithm.

[0039] The Z-score event refers to the standardized deviation value of a specific brain region, i.e., the Z-score value calculated in step 2. Reaching a specific stage threshold Discrete inflection points triggered by time.

[0040] The abnormality level R of a specific brain region in an individual is automatically mapped and determined based on the cascade intervals into which the Z-score falls; as evolutionary time t progresses... Crossing the threshold At this time, the corresponding discrete Z-score event is triggered at the theoretical level of the model (the Z-score value of the actual individual observation falls into the cascade interval determined by the threshold).

[0041] This disease progression model allows different brain regions to exhibit a heterogeneous order of damage over time based on data-driven patterns, thereby effectively capturing the dynamic trajectory of the evolution of regions of interest from normal to severely abnormal and thus obtaining the corresponding evolutionary sequence.

[0042] 4. Use the Expectation Maximization (EM) algorithm to iteratively search for the optimal subtype solution.

[0043] Step E: Given the current subtype evolution model parameters (evolutionary order) and weight ), calculate the posterior probability that each subject j belongs to subtype c. The evolution order is as follows. The sequence of all discrete Z-score events in the whole brain, categorized as subtype c, is used to characterize the dynamic evolutionary trajectory of cross-brain region damage. This sequence is input as a known conditional parameter into the likelihood function in the current iteration step; the weights... Characterizes the mixed proportion of subtype c in the disease population, and satisfies In the first iteration after the algorithm is launched, the weights... Assign a uniformly distributed initial value as the calculation baseline, such as setting 3 subtypes, initial... = = =1 / 3, and is solved dynamically by M steps in subsequent iterations.

[0044] Calculate the posterior probability that each subject j belongs to subtype c. The formula is as follows:

[0045]

[0046] The posterior probability is based on the full-stage overall likelihood under the current subtype. Combined with the weights And it is calculated using Bayes' theorem normalization. The overall likelihood probability across all stages is given by... The nested calculation corresponding to the brain region dimension and the stage dimension is obtained by traversing the N+1 disease progression stages contained in the subtype evolutionary model and summing the likelihood values ​​for each stage. Specifically, given an evolutionary sequence... When the subtype evolution model reaches a specific stage of disease progression, the algorithm first calculates the conditional probability density (i.e., the actual observed feature vector) of all independent brain regions of the subject. Deviation from piecewise linear function The probability of the predicted theoretical level of abnormality is multiplied together to obtain the single-stage whole-brain likelihood value for the subject at that specific stage. Then, all single-stage whole-brain likelihood values ​​calculated for all discrete stages are summed to finally obtain the overall likelihood across all stages. .

[0047] M-step: Update the subtype proportion weights based on the posterior probability. Furthermore, Markov chain Monte Carlo (MCMC) random sampling is used to find the optimal evolution sequence in the event permutation space that maximizes the overall likelihood function. .in, This represents the feature vector of the j-th subject, which is the standardized deviation of each ROI corresponding to the j-th subject. The vector formed by J represents the total number of subjects, which in this embodiment is the total number of subjects participating in the calculation.

[0048] In this embodiment, the MCMC independent run point is set to 25 times, and the formal sampling quantity is... To ensure global convergence of the subtype evolution model, the expectation-maximization algorithm combined with the Bayesian information criterion is used to iterate through the number of subtypes (2-6), and the optimal solution of the model is determined by combining the cross-validation information criterion (CVIC) and log-likelihood. As shown in Figure 3, when the number of subtypes is set to 3, the cross-validation information criterion (CVIC) reaches its minimum value, and the log-likelihood of the test set reaches its highest level and remains stable. Therefore, 3 is established as the optimal number of subtypes in this embodiment. At the same time, the MCMC sampling trajectory quickly enters the stable range after the initial fluctuations, proving the global convergence and robustness of the model parameter estimation.

[0049] 5. Based on the converged optimal subtype evolution model, such as Figure 4 As shown, this invention identifies three subtype trajectories with different neuropathological significance:

[0050] Subtype 1 (posterior-subcortical initiation): This subtype accounted for approximately 35% (440 cases) of the subjects. Its pathological progression showed that the abnormalities initially began in the cerebellar structures; subsequently affecting the medial temporal lobe structures, including the hippocampus, parahippocampal gyrus, and amygdala; and finally progressing to the temporal cortex and insula. This trajectory reflects a spatiotemporal progression of pathological damage from posterior to anterior.

[0051] Subtype 2 (cingulate-extensive cortical type): This subtype accounted for approximately 34% (433 cases) of the subjects. Its pathological abnormalities begin in the cingulate gyrus; early on, it exhibits involvement of a large area of ​​the cortex, including the frontal, parietal, and occipital lobes; in later stages, the pathological damage further extends to the subcortical system. This trajectory reflects a pattern of damage centered on higher cognitive regulatory networks.

[0052] Subtype 3 (basal ganglia origin): This subtype accounted for approximately 31% (392 cases) of the subjects. Its structural abnormalities first appeared in the globus pallidus; then rapidly spread to the caudate nucleus, putamen, and nucleus accumbens; and further affected the limbic system and the entire cerebral cortex. This trajectory demonstrates a typical evolutionary path radiating from the basal ganglia to the whole brain.

[0053] Through the above classification, this invention successfully outlines the spatiotemporal evolution profile of different biological subtypes of depression at the core region of interest scale. Combined with the mathematical verification shown in Figure 3, it is demonstrated that this invention can stably identify pathological subtypes with biological heterogeneity when processing large-scale, multi-center clinical imaging data, providing quantitative support for understanding the structural damage patterns of depression.

[0054] 6. Through an automated inference engine, this invention can receive sMRI image data from newly admitted subjects, process it through steps 1 and 2 above to obtain its standardized features, and output its subtype classification probability and progression stage score. For example... Figure 2 As shown in step 8, the output interface presents the quantified diagnostic reference information. In this embodiment, taking 3 subtypes and 10 fine evolution stages as an example: the output results include a topographic map of the abnormal brain distribution of the subject, the posterior probability percentage of belonging to subtypes 1 to 3 (e.g., A%, B%, C%), and the subject's position in each pathological stage on their respective trajectory (…). to The probability distribution score (e.g., a%, b%) is used to determine the risk level of depression. Through this quantitative reasoning process, the present invention can not only identify the biological subtype to which the subject belongs, but also accurately locate the severity of its pathological damage (i.e., stage score) on its specific evolutionary path, thereby providing multi-dimensional objective data support for the heterogeneity assessment of depression.

[0055] Those skilled in the art will recognize that the embodiments described herein are for the purpose of helping to understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Various modifications and variations can be made to the invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the invention should be included within the scope of the claims of the invention.

Claims

1. An unsupervised method for depression classification and staged reasoning based on structural features, characterized in that, include: S1. Extraction of primitive brain region features: Obtain structural magnetic resonance imaging data of the subject and use multimodal brain atlas to extract the primitive gray matter volume features of 122 primitive brain regions, including 120 standard brain regions and the nucleus accumbens. S2, Multi-level Region of Interest Mapping and Aggregation: 122 original brain regions are mapped into 16 core regions of interest covering the cortex and subcortical regions, and the mean gray matter volume of non-zero voxels in each region of interest is calculated. S3. Multi-source interference signal regression correction: A generalized linear model is constructed based on the mean gray matter volume of non-zero voxels in each region of interest. The systematic interference of subject gender, age, age squared, scanning site and total intracranial volume on brain structure data is eliminated through regression analysis to obtain the corrected residual characteristic value. S4. Construction of standardized feature matrix: Using the healthy control group as a benchmark, perform Z-score transformation on the corrected residual feature values ​​to obtain the standardized deviation matrix that quantifies the degree of brain structural damage. S5. Construct a subtype evolution model based on piecewise linear functions, and obtain the order of damage to each region of interest on the time axis based on the standardized deviation matrix, thus obtaining the evolution order. S6. Based on the corrected residual characteristic values ​​of the test subjects obtained in step S3 and the evolution sequence of the test subjects obtained in step S5, calculate the posterior probability to determine the subtype classification.

2. The unsupervised depression classification and staged reasoning method based on structural features according to claim 1, characterized in that, In step S2, 122 original brain regions are mapped into 16 core regions of interest covering both the cortex and subcortex. Specifically, the cortical region is aggregated into 9 regions of interest: frontal lobe, temporal lobe, parietal lobe, occipital lobe, insula, cerebellum, central region, cingulate gyrus, and parahippocampal gyrus. The subcortical region retains 7 regions of interest: hippocampus, amygdala, caudate nucleus, putamen, globus pallidus, thalamus, and nucleus accumbens. Thus, 16 core regions of interest are mapped.

3. The unsupervised depression classification and staged reasoning method based on structural features according to claim 2, characterized in that, The generalized linear model expression for step S3 is as follows: ; in, , , , , , These are the regression coefficients; These are the subject j's age, gender, and total intracranial volume; This is the random error term; The s-th component in the scan site vector using one-hot encoding is used to characterize whether subject j is scanned at the s-th site, where s represents the scan site number and S is the total number of sites.

4. The unsupervised depression classification and staged reasoning method based on structural features according to claim 3, characterized in that, The residual eigenvalues ​​in step S3 are expressed as follows: ; in, , , , , , These are the estimated values ​​of the regression coefficients.

5. The unsupervised depression classification and staged reasoning method based on structural features according to claim 4, characterized in that, The elements of the standardized deviation matrix in step S4 are the standardized deviation values ​​of each subject in the specific region of interest, denoted as Z-scores. The Z-score of subject j in the i-th region of interest is expressed as: ; in, and , respectively, represent the arithmetic mean and standard deviation of the gray matter volume of all voxels in the i-th region of interest for the healthy control group.

6. The unsupervised depression classification and staged reasoning method based on structural features according to claim 5, characterized in that, Step S5 details Includes the following processes: Based on the standard deviation of the subject's deviation from the distribution of the healthy control group in the current region of interest, the number of progression stages and the corresponding progress stage thresholds of the current region of interest are set in a stepwise manner. Based on the progress stage threshold, the cumulative anomaly trajectory of each region of interest is represented by a piecewise linear function; Based on the heterogeneous order of damage presented by the abnormal accumulation trajectory on the time axis, the evolution sequence that the subject is currently interested in is obtained.

7. The unsupervised depression classification and staged reasoning method based on structural features according to claim 6, characterized in that, Step S6 uses the expectation-maximization algorithm to iteratively search for the optimal subtype and obtain the subtype assignment, including the following steps: First, set the number of subtype categories C, and then construct C independent evolutionary order variables based on step S5; Step E: Given the current subtype evolution model parameters: evolution order and the corresponding weights Calculate the posterior probability that each subject j belongs to subtype c. ; Let c be the evolutionary order variable corresponding to the c-th subtype; The posterior probability is based on the full-stage overall likelihood under the current subtype. Combining weights And it was obtained by normalization using Bayes' theorem; The overall likelihood probability throughout the entire stage This is obtained by traversing N+1 disease progression stages and summing the likelihood values ​​for each stage, where N is the sum of the number of progression stages across all regions of interest for the subject; specifically: given an evolutionary sequence When traversing to a specific developmental stage, the conditional probability densities of all independent regions of interest of the subject at that specific developmental stage are multiplied together to obtain the single-stage whole-brain likelihood value for that subject at that specific stage. Subsequently, all single-stage whole-brain likelihood values ​​calculated for all developmental stages are summed to finally obtain the overall likelihood across all stages. ; The conditional probability density of an independent region of interest is the deviation of the actual residual eigenvalue of that region of interest from the piecewise linear function. The probability of the predicted theoretical level of abnormality; M-step: Update the subtype proportion weights based on the posterior probability, and use Markov chain Monte Carlo random sampling to find the optimal evolution order in the event permutation space that maximizes the overall likelihood function.