Machine learning-based obsessive-compulsive disorder subtype classifier and classification method

By integrating multimodal neuroanatomical features and cluster analysis, an OCD subtype classifier was constructed, which solved the problems of insufficient classification accuracy and treatment prediction of existing models and achieved high-precision OCD subtype division and individualized treatment guidance.

CN120804849APending Publication Date: 2025-10-17QIQIHAR MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510839596.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing machine learning models can only distinguish OCD patients from healthy controls, and lack the ability to explore disease subtypes unsupervised and predict treatment responses, resulting in approximately 40%-60% of patients not responding to first-line drug treatments.

Method used

A machine learning-based OCD subtype classifier was used, integrating multimodal neuroanatomical features such as gray matter volume, cortical thickness, and functional connectivity. The COMBAT algorithm was used to eliminate batch effects. Feature screening was performed using the t-test and LASSO regression to construct a classification model. Principal component analysis and k-means clustering were combined to divide neurobiological subtypes and construct a logistic regression prognostic model.

Benefits of technology

The classification accuracy of OCD patients and healthy controls was improved to AUC ≥ 0.85, significantly reducing the risk of misdiagnosis or missed diagnosis, achieving a stable division of neurobiological subtypes, and the accuracy of logistic regression in predicting non-response to SSRI drug treatment reached 78.3%, supporting individualized treatment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804849A_ABST
    Figure CN120804849A_ABST
Patent Text Reader

Abstract

The invention discloses a machine learning-based obsessive-compulsive disorder subtype classifier and a classification method, and belongs to the field of obsessive-compulsive disorder classification. The problems that an existing machine learning model can only distinguish an OCD patient from a healthy control and lacks unsupervised exploration and treatment response prediction on disease subtypes are solved. Comprising the steps that a data input module collects multi-modal data of a patient suffering from obsessive-compulsive disorder; the preprocessing module performs preprocessing based on voxel and surface morphology measurement on the sMRI data, and eliminates a multi-site data batch effect by using a COMBAT algorithm; the feature selection module is used for screening different brain region features between the obsessive-compulsive disorder group and the healthy control group and carrying out standardization processing; the model construction module is connected with the feature selection module and comprises a supervised classification unit and an unsupervised clustering unit; the prognosis prediction module inputs the subtype labels output by the unsupervised clustering unit and clinical feature data into a logistic regression model, and predicts the individual treatment alleviation probability; and the output module generates a subtype classification result. The method is mainly used in medical field.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of obsessive-compulsive disorder classification, and particularly relates to an obsessive-compulsive disorder subtype classifier based on machine learning. BACKGROUND

[0002] Obsessive-compulsive disorder (OCD) is a highly heterogeneous mental disorder, and current clinical diagnosis mainly relies on the symptom criteria of the Diagnostic and Statistical Manual of Mental Disorders, Fifth Edition (DSM-5) and subjective evaluation tools such as the Yale-Brown Obsessive-Compulsive Scale (Y-BOCS), lacking objective biological diagnostic basis. The traditional symptom-based classification method has significant limitations: the classification results are poor in stability, and the association with neurophysiological mechanisms is unclear, making it difficult to guide individualized treatment.

[0003] Although neuroimaging studies have revealed some brain circuit abnormalities (such as orbitofrontal-striatal circuit dysfunction), there are still two major bottlenecks: on the one hand, existing studies mainly focus on single modality image data (such as only analyzing structural or functional images) or isolated brain regions, failing to systematically integrate multi-modal neuroanatomical features such as gray matter volume, cortical thickness, and functional connectivity; on the other hand, the development of machine learning-based models mainly focuses on the binary classification task of distinguishing OCD patients from healthy controls (accuracy about 70%-80%), and has not realized the unsupervised subtype exploration of the internal heterogeneity of the disease, and lacks the ability to predict treatment response. This technical defect directly leads to the dilemma in clinical practice: about 40%-60% of patients do not respond to first-line drug (such as 5-hydroxytryptamine reuptake inhibitors) treatment. SUMMARY

[0004] Therefore, the present application aims to provide an obsessive-compulsive disorder subtype classifier based on machine learning and a classification method to solve the problem that existing machine learning models can only distinguish OCD patients from healthy controls, lack unsupervised exploration of disease subtypes and treatment response prediction function, and about 40%-60% of patients do not respond to first-line drug treatment.

[0005] To achieve the above-mentioned purpose, the present application adopts the following technical solution: an obsessive-compulsive disorder subtype classifier based on machine learning, the classifier comprising: a data input module, a preprocessing module, a feature selection module, a model construction module, a prognosis prediction module, and an output module; The data input module is used to collect multi-modal data of obsessive-compulsive disorder patients, including structural magnetic resonance imaging (sMRI) data, functional magnetic resonance imaging (fMRI) data, and clinical feature data. The preprocessing module is in communication connection with the data input module, and is used for preprocessing of voxel-based morphometry and surface-based morphometry of structural magnetic resonance imaging (sMRI) data and eliminating batch effects of multi-site data by using a COMBAT algorithm; The feature selection module is in communication connection with the preprocessing module, and is used for checking and screening features of different brain regions between the obsessive-compulsive disorder group and the healthy control group and performing standardization processing on the features; The model construction module is in communication connection with the feature selection module, and comprises: A supervised classification unit adopts a support vector machine or a random forest algorithm to optimize parameters by using a grid search and constructs a binary classifier of the obsessive-compulsive disorder patients and the healthy controls; An unsupervised clustering unit adopts a k-means algorithm to perform principal component analysis dimension reduction on the difference features, determines an optimal clustering number by using a silhouette coefficient and an elbow diagram distortion value, and divides the obsessive-compulsive disorder neurobiological subtypes; The prognosis prediction module is used for inputting the subtype labels output by the unsupervised clustering unit and clinical feature data into a logistic regression model to predict individual treatment remission probability; An output module is used for generating a subtype classification result.

[0006] Further, a preferred mode is also provided, and the preprocessing of the structural magnetic resonance imaging (sMRI) data based on voxel-based morphometry comprises: Bias field correction, noise elimination and skull stripping are performed on the sMRI data; A DARTEL algorithm is used to generate a gray matter volume map; Gray matter volume values of each brain region are extracted based on an AAL90 atlas.

[0007] Further, a preferred mode is also provided, and the preprocessing of the structural magnetic resonance imaging (sMRI) data based on surface-based morphometry comprises: An interface between gray matter and white matter in the sMRI data is extracted as a cortical surface, and basic cortical indicators are calculated according to the cortical surface; Cortical thickness, sulcal depth, local gyrus index and fractal dimension are calculated based on the basic cortical indicators.

[0008] Further, a preferred mode is also provided, and the determination method of the optimal clustering number in the unsupervised clustering unit comprises: A silhouette coefficient is calculated:

[0009] Wherein, a is an average distance of a sample to all other points in the same cluster (intra-cluster distance), bis the average distance to all points in the other cluster closest to it; A elbow plot is drawn according to the profile coefficients and a slope of the adjacent cluster number distortion value is calculated.

[0010] Further, a preferred mode is also proposed, wherein the classifier further comprises a verification module for evaluating model performance, including: using Jaccard similarity coefficient to evaluate clustering stability.

[0011] Further, a preferred mode is also proposed, wherein the feature selection module normalizes the features, including:

[0012] wherein, min ( X ) is the minimum value of the feature X , max ( X ) is the maximum value of the feature X .

[0013] Based on the same inventive concept, the present application also proposes an obsessive-compulsive disorder subtype classification method, which is realized based on the classifier of any one of the above-mentioned embodiments, and the method comprises the following steps: acquiring multi-modal data of obsessive-compulsive disorder patients; eliminating batch effects and extracting voxel-based morphometric features and surface-based morphometric features by using COMBAT algorithm; screening differential brain region features based on t-test and performing normalization processing; binary classifier of obsessive-compulsive disorder patients and healthy controls; dividing patient subtypes by PCA dimension reduction and k-means clustering; integrating subtype labels and clinical features to predict treatment response and outputting subtype classification results.

[0014] Based on the same inventive concept, the present application also proposes a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, the processor executes the obsessive-compulsive disorder subtype classification method according to the above-mentioned embodiments.

[0015] Based on the same inventive concept, the present application also proposes a computer readable storage medium, which stores a computer program, and when the computer program is run by a processor, the processor executes the steps of the obsessive-compulsive disorder subtype classification method as described above.

[0016] Compared with the prior art, the present application has the following beneficial effects: The classifier provided by the application integrates the multi-modal neuroanatomical features (sMRI / fMRI) of gray matter volume, cortical thickness, functional connection and clinical variables, adopts the COMBAT algorithm to eliminate the batch effect of multi-center data, combines t-test and LASSO regression for feature screening, and constructs a classification model with an AUC of 0.85 or more (five-fold cross-validation) when distinguishing the patients with obsessive-compulsive disorder from the healthy controls, which is more than 25% higher than the diagnostic accuracy of the traditional scale, and significantly reduces the misdiagnosis and missed diagnosis risks caused by subjective evaluation. The application combines principal component analysis (retaining 95% variance) with k-means clustering, optimizes the cluster number selection through the double indexes of profile coefficient and distortion slope, makes the Jaccard similarity coefficient of the obsessive-compulsive disorder subtype division stable at more than 0.92, and solves the technical problem of large fluctuation of the traditional symptom classification results. The application constructs a logistic regression prognosis model based on the neurobiological subtype labels (such as "extensive cortical atrophy type" and "local functional abnormality type"), integrates the clinical characteristics such as disease duration and drug compliance, and the prediction accuracy of the non-response to SSRI drugs reaches 78.3%, which realizes the closed-loop decision support from "diagnostic classification" to "treatment prediction" for the first time, and makes it possible to intervene in the patients with non-response to first-line drugs. BRIEF DESCRIPTION OF DRAWINGS

[0017] The accompanying drawings, which form a part of the present application, are used to provide further understanding of the present application, and the illustrative embodiments of the present application and their description are used to explain the present application, and do not constitute improper limitations on the present application. In the drawings: Figure 1 A schematic diagram of an obsessive-compulsive disorder subtype classifier based on machine learning according to the application. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the application. It should be explained that, in the case of no conflict, the embodiments in the application and the features in the embodiments can be combined with each other, and the described embodiments are only a part of the embodiments of the application, but not all the embodiments.

[0019] Embodiment one, see Figure 1 The application discloses an obsessive-compulsive disorder subtype classifier based on machine learning, which comprises: a data input module, a preprocessing module, a feature selection module, a model construction module, a prognosis prediction module and an output module; The data input module is used to collect multi-modal data of patients with obsessive-compulsive disorder, including structural magnetic resonance imaging sMRI data, functional magnetic resonance imaging fMRI data and clinical feature data. The preprocessing module is in communication connection with the data input module, and is used for preprocessing voxel-based morphometry and surface-based morphometry of structural magnetic resonance imaging (sMRI) data and eliminating batch effects of multi-site data by using a COMBAT algorithm; The feature selection module is in communication connection with the preprocessing module, and is used for checking and screening features of different brain regions between the OCD group and the healthy control group, and performing standardization processing on the features. The model construction module is in communication connection with the feature selection module, and includes: A supervised classification unit adopts a support vector machine or a random forest algorithm to optimize parameters by grid search, and constructs a binary classifier of OCD patients and healthy controls. An unsupervised clustering unit adopts a k-means algorithm to perform principal component analysis (PCA) dimension reduction on the difference features, determines the optimal clustering number by using a silhouette coefficient and an elbow diagram distortion value, and divides OCD neurobiological subtypes. The prognosis prediction module is used for inputting the subtype labels output by the unsupervised clustering unit and clinical feature data into a logistic regression model, and predicting individual treatment remission probability. An output module is used for generating a subtype classification result.

[0020] The current machine learning model is mostly used for binary classification of OCD patients and healthy controls. The embodiment introduces an unsupervised clustering technology to realize refined classification of OCD patient subtypes. In the embodiment, multi-modal data including sMRI data, fMRI data and clinical feature data are used, which is helpful to comprehensively capture multi-level features of the brain of OCD patients, and thus improves the accuracy of subtype division.

[0021] The COMBAT algorithm is used to eliminate batch effects of multi-site data, which ensures consistency of data from different laboratories and equipment. This can effectively reduce bias introduced by differences in data collection environments, and improve the generalization ability and reliability of the model. The feature selection module can accurately screen out brain region features with significant differences between the OCD group and the healthy control group, and perform standardization processing on the selected features, thereby enhancing the robustness and accuracy of the model. By using the k-means clustering algorithm and PCA dimension reduction, the application can perform unsupervised subtype division on OCD patients, thereby identifying different neurobiological subtypes.

[0022] Embodiment two, the embodiment is a further limitation of the OCD subtype classifier based on machine learning in embodiment one, and the voxel-based morphometry preprocessing of the sMRI data includes: Bias field correction, noise removal and skull stripping are performed on the sMRI data; DARTEL algorithm is used to generate the gray matter volume map; Based on the AAL90 atlas, the gray matter volume values of each brain region are extracted.

[0023] The structural magnetic resonance imaging (sMRI) data has a bias field effect caused by magnetic field inhomogeneity. The bias field correction step eliminates this effect, making the boundaries of gray matter, white matter and other brain tissues clearer and improving the data quality. Through the noise removal method, unnecessary interference components are removed, improving the clarity and usability of the data. By removing the skull part, accurate measurement of brain tissue is ensured. This is the basis for accurate differentiation of gray matter, white matter and cerebrospinal fluid regions, reducing false signal interference.

[0024] DARTEL algorithm is used to align the brain structure and estimate the gray matter volume, making the brain images of different subjects more accurately aligned. Further, based on the AAL90 atlas, the gray matter volume values of each brain region are extracted. The AAL90 atlas covers 90 brain regions. By extracting the gray matter volume values of each brain region through this standard atlas, not only detailed information of individual brain regions can be provided, but also differences in gray matter volume of different brain regions can be compared.

[0025] Embodiment three, this embodiment is a further limitation of the obsessive-compulsive disorder subtype classifier based on machine learning of embodiment one, the preprocessing of structural magnetic resonance imaging (sMRI) data based on surface morphometry includes: Extracting the gray-white matter interface in sMRI data as the cortical surface, and calculating the basic cortical indicators based on the cortical surface; Based on the basic cortical indicators, calculate the cortical thickness, sulcal depth, local gyrus index and fractal dimension.

[0026] By extracting the gray-white matter interface in sMRI data as the cortical surface, the structure of the cerebral cortex can be more accurately modeled. By calculating the basic cortical indicators of the cortical surface, the overall structure of the cerebral cortex can be more comprehensively understood from a macroscopic perspective. These indicators include cortical thickness, sulcal depth, local gyrus index, etc. These features are crucial for understanding the differences between different subtypes of obsessive-compulsive disorder (OCD). Each indicator reflects different neural structural characteristics, which helps to analyze the neural mechanisms of obsessive-compulsive disorder from multiple perspectives.

[0027] Based on these basic cortical indicators (such as cortical thickness, sulcal depth, etc.), further calculations of cortical thickness, sulcal depth, local gyrus index and fractal dimension, etc. can provide more dimensional fine-grained data for machine learning models. This multi-dimensional feature combination can improve the recognition accuracy of different OCD subtypes by the classification model, thereby more effectively performing individualized diagnosis.

[0028] Embodiment four, the embodiment is further limited to the OCD subtype classifier based on machine learning of embodiment one, and the determination method of the optimal cluster number in the unsupervised clustering unit comprises: Calculate the contour coefficient:

[0029] Wherein, a is the average distance of the sample to all other points in the same cluster (intra-cluster distance), b is the average distance to all points in the nearest other cluster; Draw the elbow diagram according to the contour coefficient and calculate the slope of the adjacent cluster number distortion value.

[0030] Embodiment five, the embodiment is further limited to the OCD subtype classifier based on machine learning of embodiment one, and the classifier further comprises a verification module, which is used to evaluate the model performance, including: using Jaccard similarity coefficient to evaluate the clustering stability.

[0031] Embodiment six, the embodiment is further limited to the OCD subtype classifier based on machine learning of embodiment one, and the feature selection module in the embodiment comprises:

[0032] Wherein, min ( X ) is the minimum value of the feature X , max ( X ) is the maximum value of the feature X .

[0033] Embodiment seven, the OCD subtype classification method comprises: Obtaining multi-modal data of OCD patients; Eliminating batch effect by COMBAT algorithm and extracting voxel-based morphometric features and surface-based morphometric features; Filtering differential brain region features based on t-test and performing standardization processing; Binary classifier of OCD patients and healthy controls; Subtype classification of patients by PCA dimension reduction and k-means clustering; Integrate subtype labels and clinical features to predict treatment response, and output subtype classification results.

[0034] Embodiment eight, a computer device according to the embodiment, comprising a memory and a processor, the memory stores a computer program, when the processor runs the computer program stored in the memory, the processor executes the subtype classification method of obsessive-compulsive disorder according to embodiment seven.

[0035] Embodiment nine, a computer readable storage medium according to the embodiment, the computer readable storage medium stores a computer program, the computer program is run by the processor to execute the steps of the subtype classification method of obsessive-compulsive disorder according to embodiment seven.

[0036] Embodiment ten, the embodiment is a specific embodiment of the subtype classification method of obsessive-compulsive disorder based on machine learning according to embodiment one, and also used to explain embodiments two to six, specifically: A subtype classification method of obsessive-compulsive disorder based on machine learning, the classifier comprises: Data input module, preprocessing module, feature selection module, model construction module, prognosis prediction module and output module; The data input module is used to collect multi-modal data of obsessive-compulsive disorder patients and healthy controls, including: Structural magnetic resonance imaging sMRI data: gray matter volume, cortical thickness, sulcal depth; Functional magnetic resonance imaging fMRI data: default mode network (DMN) connection strength, task-state activation pattern; Clinical feature data: Y-BOCS score, disease duration, drug response history.

[0037] The preprocessing module is used for data format conversion and quality check, voxel-based morphometry (VBM) preprocessing, surface-based morphometry (SBM) preprocessing, using COMBAT algorithm to eliminate multi-site data batch effect, standardization feature extraction, specifically: Data format conversion and quality check: Use dcm2nii tool to convert raw DICOM format magnetic resonance imaging (MRI) data to NIFTI format, and perform data quality check; Use IQR (interquartile range) image quality index in CAT12 toolbox to control data quality, and select subject data with quality score higher than B to ensure data quality; Voxel-based morphometry (VBM) preprocessing: NIFTI format T1-weighted MRI images were preprocessed using SPM12 software, including bias field correction, noise removal, and skull stripping; DARTEL (Diffeomorphic Anatomical Registration using Exponential Lie Algebra) was used to spatially standardize the structural images of the subjects, generating volume maps of gray matter and white matter; Spatial smoothing was performed on the gray matter volume results, with a smoothing kernel of 8mm to reduce noise and enhance feature continuity; AAL90 atlas was used to extract gray matter volume values (unit: cm³) of each brain region as feature indicators for subsequent analysis; Surface-based morphometry (SBM) preprocessing: The CAT12 toolbox based on SPM12 was used to perform SBM analysis on the segmented data, extracting the gray-white matter interface as the cortical surface, and calculating basic cortical indicators such as cortical thickness (distance from gray-white matter interface to gray matter surface); The surf tool of CAT12 was used to further extract cortical indicators, including sulcal depth (distance between the highest and lowest points of the gray matter outer surface), local gyrus index (number of gyrus folds in gray matter), and fractal dimension (cortical complexity); The aparc2009 cortical atlas was used as the atlas for cortical indicator extraction, providing more abundant cortical structure features for subsequent analysis.

[0038] Batch effect removal: COMBAT (Combining Microarray Data Sets for Batch Effect Removal) algorithm was used to process the data to remove the batch effect of cortical features between different sites; Batch effect refers to systematic bias introduced during data collection or processing, which may affect the interpretation and comparability of results. Through COMBAT processing, the data becomes more comparable between different batches while preserving the biological differences between samples.

[0039] The feature selection module is used to test and screen the different brain region features between the OCD group and the healthy control group, and to standardize the features, including: Based on 5 indicators of the cortex (gray matter volume, cortical thickness, sulcal depth, local gyrus index, and fractal dimension), a total of 682 features; Two-sample T-test was used to compare the differences in each brain structure feature between the OCD group and the healthy control group, and to screen out brain regions with significant differences (p<0.01); MinMaxScaler was used to standardize the data, with a value range of [-1, 1]. The standardization formula is as follows:

[0040] wherein, min ( X ) and max ( X ) are the minimum and maximum values of the feature X , respectively.

[0041] The model construction module comprises: a supervised classification unit, which adopts a support vector machine or a random forest algorithm to optimize parameters by grid search to construct a binary classifier of obsessive-compulsive disorder patients and healthy controls; an unsupervised clustering unit, which adopts a k-means algorithm to perform principal component analysis dimension reduction on the difference features, determines the optimal clustering number through a contour coefficient and an elbow diagram distortion value, and divides obsessive-compulsive disorder neurobiological subtypes (such as “extensive cortical atrophy type” and “local functional abnormality type”); Specifically: The support vector machine SVM in the supervised classification unit is a binary classification model, which finds an optimal hyperplane (or multiple hyperplanes) to separate data points of different classes; a Gaussian radial basis function (RBF) or a linear kernel is selected as a kernel function; the search range of a regularization parameter C is set to (0.00001, 10, 1000), and the search range of a kernel coefficient γ is set to (0.00001, 10, 100); Random Forest is an ensemble learning method composed of multiple decision trees, and each tree is constructed based on different random subsets and features; the search range of the number of decision trees is set to (5, 40, 1), and the maximum depth of the decision tree is set to (5, 40, 1).

[0042] In actual application, the evaluation method of the classification model also includes: 30% of the entire data set (220 subjects, 70 features) is randomly selected as a validation set, and the remaining 70% is used as a training set; the performance of the model is evaluated by five-fold cross-validation in the training set, and the specific steps are as follows: The training set is divided into five folds, four of which are used for training and one of which is used for validation; The parameter combination of SVM and random forest RF is optimized by using a grid search method; The above process is repeated 100 times to calculate the average receiver operating characteristic curve (ROC), average sensitivity, and average specificity on the test set; The average accuracy of the model is evaluated by a permutation test to ensure the reliability and significance of the model performance.

[0043] The unsupervised clustering unit comprises: Feature extraction: Principal Component Analysis (PCA) is performed on the difference features extracted in the previous step to extract principal components that can explain 95% of the variance, reducing the feature dimension and retaining the main information. Determine the optimal clustering: use k-means clustering algorithm to cluster analysis of obsessive-compulsive disorder patients to divide different subtypes.

[0044] Use the silhouette coefficient (Silhouette) and the distortion (Distortion) and slope (Slope) values in the elbow chart to determine the optimal number of clusters.

[0045] Silhouette coefficient: used to evaluate the effect of clustering algorithm, the calculation formula is:

[0046] Among them, a is the average distance of the sample to all other points in the same cluster (intra-cluster distance), b is the average distance to all points in the nearest other cluster (average distance of the nearest other cluster); the value of the silhouette coefficient ranges from -1 to 1, and the closer the value is to 1, the better the clustering effect; Elbow chart: shows the Distortion value corresponding to different cluster numbers (K value), finds the point where the Distortion value sharply decreases and tends to be stable, and determines the optimal cluster number; Slope chart: shows the slope between the Distortion values corresponding to adjacent cluster numbers, helps to determine whether the rate of Distortion decrease is fast enough, and further confirms the optimal cluster number.

[0047] Use the Jaccard similarity coefficient to evaluate the stability and consistency of the clustering results. The calculation formula of Jaccard coefficient is:

[0048] Among them, A and B are two sets of clusters, and A ∩ B ∣ represents the number of common samples in the two clusters, and A ∪ B ∣ represents the number of samples in the union of the two clusters. The value of Jaccard coefficient ranges from 0 to 1, and the closer the value is to 1, the more similar the two clustering results are.

[0049] The prognosis prediction module is used to input the subtype label output by the unsupervised clustering unit and the clinical feature data into a logistic regression model to predict the individual treatment remission probability; The output module is used to generate a subtype classification result.

[0050] Those skilled in the art will appreciate that embodiments of the disclosure can be supplied as a method, a system, or a computer program product. Accordingly, the disclosure can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the disclosure can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer-readable program code. The disclosure is described in reference to the flowchart and / or block diagrams of the method, apparatus (system) and computer program product according to embodiments of the disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagrams and a combination of flows and / or blocks in the flowchart and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus produce the functions specified in the flowchart and / or block diagrams of the method, apparatus (system) and computer program product. Figure 1 one or more flows and / or blocks Figure 1 means for performing the functions specified in the flowchart and / or block diagrams of the method, apparatus (system) and computer program product. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means, which implement the functions specified in the flowchart and / or block diagrams of the method, apparatus (system) and computer program product. Figure 1 one or more flows and / or blocks Figure 1 means for performing the functions specified in the flowchart and / or block diagrams of the method, apparatus (system) and computer program product. These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operational steps are performed on the computer or other programmable data processing apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable data processing apparatus provide a process for implementing the functions specified in the flowchart and / or block diagrams of the method, apparatus (system) and computer program product. Figure 1 one or more flows and / or blocks Figure 1 means for performing the functions specified in the flowchart and / or block diagrams of the method, apparatus (system) and computer program product. Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the disclosure but not to limit the protection scope of the disclosure, and although the disclosure has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical personnel in the art can make various changes, modifications or equivalent replacements to the specific embodiments of the disclosure after reading the disclosure, but these changes, modifications or equivalent replacements are all within the protection scope of the disclosed claims.

Claims

1. A machine learning-based obsessive-compulsive disorder subtype classifier, characterized by: The classifier includes: Data input module, preprocessing module, feature selection module, model building module, prognosis prediction module and output module; The data input module is used to collect multimodal data of patients with obsessive-compulsive disorder, including structural magnetic resonance imaging (sMRI) data, functional magnetic resonance imaging (fMRI) data, and clinical characteristic data; The preprocessing module is in communication with the data input module, and is used to perform voxel-based morphological measurement and surface-based morphological measurement preprocessing on structural magnetic resonance imaging (sMRI) data, and to eliminate batch effects of multi-site data using a COMBAT algorithm; The feature selection module is in communication with the preprocessing module, and is used to detect and screen the different brain region features between the obsessive-compulsive disorder group and the healthy control group, and to perform standardization processing on the features; The model building module is in communication with the feature selection module, and the model building module includes: For supervised classification units, support vector machines or random forest algorithms were used to optimize parameters through grid search to construct binary classifiers for OCD patients and healthy controls; Unsupervised clustering unit, using k-means algorithm to perform principal component analysis and dimensionality reduction on the differential features, and the silhouette coefficient and elbow plot distortion value were used to determine the optimal number of clusters to divide the neurobiological subtypes of obsessive-compulsive disorder; The prognosis prediction module is used to input the subtype labels and clinical feature data output by the unsupervised clustering unit into a logistic regression model to predict the probability of individual treatment remission; Output module, used to generate subtype classification results.

2. The obsessive-compulsive disorder subtype classifier based on machine learning according to claim 1, characterized in that: The voxel-based morphological measurement preprocessing of structural magnetic resonance imaging (sMRI) data comprises: Perform bias field correction, noise removal, and skull stripping on sMRI data; Gray matter volume maps were generated using the DARTEL algorithm; The gray matter volume values ​​of each brain region were extracted based on the AAL90 atlas.

3. The obsessive-compulsive disorder subtype classifier based on machine learning according to claim 1, characterized in that: The surface-based morphological measurement preprocessing of structural magnetic resonance imaging (sMRI) data comprises: The gray-white matter interface in sMRI data was extracted as the cortical surface, and basic cortical indices were calculated based on the cortical surface. Cortical thickness, sulcus depth, local gyrus index and fractal dimension were calculated based on basic cortical indicators.

4. The obsessive-compulsive disorder subtype classifier based on machine learning according to claim 1, characterized in that: The method for determining the optimal number of clusters in the unsupervised clustering unit includes: Calculate the silhouette coefficient: in, a is the average distance from the sample to all other points in the same cluster (intra-cluster distance), b is the average distance to all points in the closest cluster; The elbow plot was drawn based on the silhouette coefficient and the slope of the distortion value of the number of adjacent clusters was calculated.

5. The obsessive-compulsive disorder subtype classifier based on machine learning according to claim 1, characterized in that: The classifier also includes a verification module, which is used to evaluate model performance, including: using the Jaccard similarity coefficient to evaluate clustering stability.

6. The obsessive-compulsive disorder subtype classifier based on machine learning according to claim 1, characterized in that: The feature selection module performs standardization processing on features, including: in, min ( X ) is characterized by X The minimum value of max ( X ) is characterized by X The maximum value of .

7. A method for classifying obsessive-compulsive disorder subtypes, characterized in that: The method is implemented based on the classifier according to any one of claims 1 to 6, and the method comprises: Obtain multimodal data on patients with OCD; The COMBAT algorithm was used to eliminate batch effects and extract voxel-based morphometric features and surface-based morphometric features. Based on the t-test, the different brain region features were screened and standardized; binary classifier for OCD patients and healthy controls; Patient subtypes were divided by PCA dimensionality reduction and k-means clustering; Integrate subtype labels with clinical characteristics to predict treatment response and output subtype classification results.

8. A computer device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the method for classifying subtypes of obsessive-compulsive disorder according to claim 7.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the method for classifying subtypes of obsessive-compulsive disorder as claimed in claim 7.