A microsatellite instability patient classification method and device, computer equipment and storage medium

By performing pathway enrichment and tumor immune enrichment analysis on the expression profile data of microsatellite instability patients, and using a classification model to classify patients, personalized treatment plans can be recommended. This solves the problem of ineffective classification in existing technologies and improves the targeting and effectiveness of treatment.

CN122221055APending Publication Date: 2026-06-16BGI GENOMICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BGI GENOMICS CO LTD
Filing Date
2024-12-06
Publication Date
2026-06-16

Smart Images

  • Figure CN122221055A_ABST
    Figure CN122221055A_ABST
Patent Text Reader

Abstract

The application provides a microsatellite instability patient classification method. The method comprises the following steps: performing pathway enrichment analysis on expression profile data of a target patient, screening the analysis results, obtaining a plurality of first enrichment scores, inputting the first enrichment scores into a first classification model, and determining whether the target patient belongs to a first category; patients in the first category are suitable for a combination drug regimen; if not, performing tumor immune enrichment analysis on the expression profile data to obtain a plurality of second enrichment scores; inputting the plurality of second enrichment scores into a second classification model to determine whether the target patient belongs to a second category or a third category; patients in the second category are suitable for a single antibody drug regimen, and patients in the third category are suitable for both the single antibody drug regimen and the combination drug regimen. The method can provide a reference basis when formulating a treatment plan, and is helpful to improve the prognosis of microsatellite instability patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biomedical technology, and in particular to a method, apparatus, computer device, and storage medium for classifying microsatellite unstable patients. Background Technology

[0002] Immune checkpoint inhibitors (ICIs) are a novel cancer treatment method that primarily works by modulating the body's own immune system to fight tumor cells. Clinically, this therapy can achieve long-term remission or even cure in some cancer patients, but not all patients respond to it. Microsatellite instability (MSI) status has been approved as a predictive biomarker for immunotherapy in various cancers. However, not all MSI patients benefit from immunotherapy. Traditional algorithms primarily focus on predicting the efficacy of ICI treatment without considering patient classification to select the most suitable treatment regimen. Summary of the Invention

[0003] The purpose of this application is to at least address one of the aforementioned technical deficiencies, particularly the problem in conventional techniques that cannot classify patients with microsatellite instability to provide a reference for selecting appropriate treatment options.

[0004] Firstly, this application provides a method for classifying patients with microsatellite instability, including:

[0005] Obtain expression profile data from the target patients;

[0006] Pathway enrichment analysis was performed on the expression profile data, and the results of the pathway enrichment analysis were screened to obtain multiple first enrichment scores.

[0007] Multiple first-enrichment scores are input into the first-classification model to determine whether the target patient belongs to the first category; patients in the first category are suitable for combination therapy.

[0008] If not, tumor immune enrichment analysis is performed on the expression profile data to obtain multiple second enrichment scores;

[0009] Multiple second enrichment scores are input into the second classification model to determine whether the target patient belongs to the second or third category; patients in the second category are suitable for monoclonal antibody treatment, while patients in the third category are suitable for both monoclonal antibody treatment and combination therapy.

[0010] In one embodiment, the process of constructing the first classification model includes:

[0011] Obtain the first training set; the annotations on the training samples in the first training set are used to indicate whether the training samples belong to the first category;

[0012] Pathway enrichment analysis was performed on each training sample in the first training set, and the results of the pathway enrichment analysis for each training sample in the first training set were filtered to obtain multiple first training enrichment scores.

[0013] Multiple first base models are set up, and each first base model is trained according to the annotation of the training samples in each first training set and the multiple first training enrichment scores.

[0014] Determine the model performance of each first base model, and select the first base model with the best model performance as the first classification model.

[0015] In one embodiment, the annotation process for training samples in the first training set includes:

[0016] Select the clustering parameter with the highest clustering stability from multiple sets of candidate clustering parameters as the target clustering parameter;

[0017] The first dataset is clustered according to the target clustering parameters to obtain the target clustering results; the first dataset includes the first training set;

[0018] Clusters that cannot be further classified based on tumor immune enrichment analysis in the target clustering results are identified as belonging to the first category, while clusters that can be further classified based on tumor immune enrichment analysis in the target clustering results are identified as not belonging to the first category.

[0019] Based on the target clustering results, label whether the samples in the first training set belong to the first category.

[0020] In one embodiment, the process of determining cluster stability includes:

[0021] For any set of candidate clustering parameters, the first dataset is clustered according to the set of candidate clustering parameters to obtain the first clustering result;

[0022] Multiple random samplings are performed on the first dataset, and the datasets obtained from each sampling are clustered according to the clustering parameters of the candidate set, resulting in multiple second clustering results;

[0023] The clustering stability of the candidate clustering parameters is determined based on the similarity between each second clustering result and the first clustering result.

[0024] In one embodiment, the first dataset also includes a first validation set for determining the model performance of each first base model.

[0025] In one embodiment, the path enrichment analysis results corresponding to each training sample in the first training set are screened, including:

[0026] Determine the classification contribution of each feature in the pathway enrichment analysis results corresponding to each training sample in the first training set;

[0027] Select the first number of features with the largest classification contribution as the first target features;

[0028] Based on the first target feature, the pathway enrichment analysis results corresponding to each training sample in the first training set are screened.

[0029] In one embodiment, the pathway enrichment analysis results of the expression profile data are screened, including:

[0030] Based on the first target characteristics, the pathway enrichment analysis results of the expression profile data are screened.

[0031] In one embodiment, the process of constructing the second classification model includes:

[0032] The second dataset is obtained from the training samples in the first dataset that do not belong to the first category; the second dataset includes a second training set and a second validation set.

[0033] Tumor immune enrichment analysis was performed on the training samples in the second dataset to obtain multiple corresponding second training enrichment scores.

[0034] Clustering is performed based on the multiple second training enrichment scores corresponding to each training sample in the second dataset to label whether the training samples in the second dataset belong to the second or third category;

[0035] Set up multiple second base models, and train each second base model separately based on the annotation of each training sample in the second training set and the multiple second training enrichment scores;

[0036] The model performance of each second base model is determined by the annotation of the second validation set, and the second base model with the best model performance is selected as the second classification model.

[0037] Secondly, this application provides a microsatellite instability patient classification device, comprising:

[0038] The data acquisition module is used to acquire expression profile data of the target patients;

[0039] The first enrichment analysis module is used to perform pathway enrichment analysis on expression profile data and to filter the pathway enrichment analysis results of expression profile data to obtain multiple first enrichment scores.

[0040] The first classification module is used to input multiple first enrichment scores into the first classification model to determine whether the target patient belongs to the first category; patients in the first category are suitable for combination therapy.

[0041] The second enrichment analysis module is used to perform tumor immune enrichment analysis on expression profile data when the target patient does not belong to the first category, and obtain multiple second enrichment scores.

[0042] The second classification module is used to input multiple second enrichment scores into the second classification model to determine whether the target patient belongs to the second or third category; patients in the second category are suitable for monoclonal antibody treatment regimens, while patients in the third category are suitable for both monoclonal antibody treatment regimens and combination therapy regimens.

[0043] Thirdly, this application provides a computer device including one or more processors and a memory storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, they perform the microsatellite instability patient classification method in any of the above embodiments.

[0044] Fourthly, this application provides a storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the microsatellite instability patient classification method in any of the above embodiments.

[0045] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:

[0046] Based on the microsatellite instability (MSI) patient classification method in this embodiment, pathway enrichment analysis and tumor immune enrichment analysis are performed on patient expression profile data to gain a deeper understanding of the patient's disease status, immune characteristics, and biological features from different perspectives. Pathway enrichment analysis of expression profile data helps identify differences in the degree of enrichment at the biological pathway level among patients, determining whether they belong to the first category. Tumor immune enrichment analysis targeting the tumor immune gene set focuses on tumor immune status, providing a classification basis for patients who do not belong to the first category. The classification results obtained by this method correspond to different treatment plans, improving the targeting and effectiveness of treatment planning and contributing to improving the prognosis of MSI patients. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 A flowchart illustrating a microsatellite instability patient classification method provided in one embodiment of this application;

[0049] Figure 2 This is a schematic diagram of the process of training a first classification model in one embodiment of this application;

[0050] Figure 3 This is a schematic diagram of the process of annotating the first training set in one embodiment of this application;

[0051] Figure 4 This is a flowchart illustrating the process of determining the clustering stability of each group of candidate clustering parameters in one embodiment of this application;

[0052] Figure 5 This is a schematic diagram of the process for filtering the first training set in one embodiment of this application;

[0053] Figure 6 This is a schematic diagram showing the performance evaluation results of each first basic model using the area under the working characteristic curve in one embodiment of this application;

[0054] Figure 7 This is a schematic diagram illustrating the comprehensive results of performance evaluation of each first basic model using multiple metrics in one embodiment of this application;

[0055] Figure 8 This is a PR curve of each first basic model in one embodiment of this application;

[0056] Figure 9 This is a schematic diagram of the process of training a second classification model in one embodiment of this application;

[0057] Figure 10 This is a schematic diagram showing the performance evaluation results of each second basic model using the area under the working characteristic curve in one embodiment of this application;

[0058] Figure 11 This is a schematic diagram illustrating the comprehensive results of performance evaluation of each second basic model using multiple metrics in one embodiment of this application;

[0059] Figure 12 This is a PR curve of each of the second basic models in one embodiment of this application;

[0060] Figure 13 A heatmap of cytokine signals obtained from classifying the U-CAN dataset;

[0061] Figure 14 A heatmap of cytokine signals obtained from classifying the TCGA-COADREAD dataset;

[0062] Figure 15 A heatmap of cytokine signals obtained from classifying the TCGA-UCEC dataset;

[0063] Figure 16 A heatmap of cytokine signals obtained from classifying the TCGA-STAD dataset;

[0064] Figure 17 A heatmap of cytokine signals obtained from classification of the Nipicol dataset;

[0065] Figure 18 A heatmap of cytokine signals obtained from classifying the ImmunoMSI dataset;

[0066] Figure 19 Progression-free survival analysis of patients in three categories after immunotherapy for the Nipicol and ImmunoMSI datasets;

[0067] Figure 20 Progression-free survival analysis of all patients in the Nipicol and ImmunoMSI datasets after immunotherapy;

[0068] Figure 21 This study analyzes the progression-free survival of patients in three categories on the Nipicol and ImmunoMSI datasets after treatment with combination therapy regimens.

[0069] Figure 22 This study analyzes the progression-free survival of patients in three categories after monoclonal antibody therapy using the Nipicol and ImmunoMSI datasets under combination therapy regimens.

[0070] Figure 23 Progression-free survival analysis of patients in the Nipicol and ImmunoMSI datasets in Category I after treatment with monoclonal antibody regimens and combination therapy regimens, respectively;

[0071] Figure 24 Progression-free survival analysis of patients in the second category of the Nipicol and ImmunoMSI datasets after treatment with monoclonal antibody regimens and combination therapy regimens, respectively;

[0072] Figure 25 Progression-free survival analysis of patients in category 3 of the Nipicol and ImmunoMSI datasets after treatment with monoclonal antibody regimens and combination therapy regimens, respectively;

[0073] Figure 26 This is an internal structural diagram of a computer device provided in one embodiment of this application. Detailed Implementation

[0074] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0075] This application provides a method for classifying patients with microsatellite instability, including steps S102 to S110.

[0076] S102, Obtain expression profile data of the target patient.

[0077] It is understandable that the target patients are those with microsatellite instability who need to be classified. Expression profiling data, obtained from patients through RNA sequencing or RNA microarray analysis, reflects the gene expression levels in biological samples. Gene expression levels change under different physiological states and disease conditions. By acquiring patients' expression profiling data, we can understand the activity level of genes in their bodies, providing fundamental data for subsequent analysis. In the classification of microsatellite instability patients, different categories may have specific gene expression patterns; therefore, expression profiling data is an important basis for classification. Specifically, high-throughput sequencing or microarray technology can be used to obtain the expression profiling data of the target patients. After obtaining the expression profiling data, it can be preprocessed, including standardization and normalization. Specifically, standardization can be performed using specific gene expression level standardization metrics such as TPM (Transcripts Per Million), and normalization can be performed using the ScaleData function in the Seurat4 R package.

[0078] S104. Pathway enrichment analysis was performed on the expression profile data, and the results of the pathway enrichment analysis were screened to obtain multiple first enrichment scores.

[0079] Pathway enrichment analysis is a bioinformatics method used to determine the degree of gene enrichment within a specific biological pathway. Specifically, it involves single-sample gene set enrichment analysis (SSGSEA) using databases containing biological pathway information. Commonly used databases include KEGG (Kyoto Encyclopedia of Genes and Genomes) and Reactome. Single-sample gene set enrichment analysis is a method for enriching gene sets within a single sample. The first enrichment score measures the degree of enrichment of each gene in the expression profile within the specific pathway contained in the used database.

[0080] Pathway enrichment analysis is based on the assumption that if a group of genes synergistically function in a specific biological pathway, then these genes may simultaneously undergo expression changes under specific physiological or pathological conditions. By performing pathway enrichment analysis on expression profile data, disease-related biological pathways can be identified, thus providing clues for disease diagnosis and treatment. Since many biological pathways are referenced in pathway enrichment analysis, the results will contain scores corresponding to each pathway. An excessive number of these scores can hinder data analysis. This embodiment selects some important features from the pathway enrichment analysis results to accelerate the analysis process.

[0081] S106. Input multiple first enrichment scores into the first classification model to determine whether the target patient belongs to the first category. Patients in the first category are suitable for combination therapy.

[0082] Understandably, the first classification model is used to classify target patients based on input features (here, multiple first enrichment scores). Commonly used treatment regimens for microsatellite instability patients include monoclonal antibody regimens and combination therapy regimens. Monoclonal antibody regimens refer to treatment using only one monoclonal antibody drug, such as using only pembrolizumab (anti-PD-1). Combination therapy refers to treatment using two or more monoclonal antibody drugs simultaneously, such as using nivolumab (anti-PD-1) and ipilimumab (anti-CTLA-4) simultaneously.

[0083] In the microsatellite instability patient classification, different categories of patients respond differently to two different treatment regimens. Some patients achieve significantly better outcomes with combination therapy than with monoclonal antibody therapy (e.g., a difference in prognosis greater than a certain threshold); these patients are considered suitable for combination therapy. Similarly, some patients achieve significantly better outcomes with monoclonal antibody therapy than with combination therapy (e.g., a difference in prognosis greater than a certain threshold); these patients are considered suitable for both monoclonal antibody therapy and combination therapy.

[0084] Analysis of extensive data on microsatellite instability patients revealed significant differences in first-enrichment scores between patients receiving only combination therapy and those receiving only monoclonal antibody therapy, or those receiving both monoclonal antibody and combination therapy. The first classification model learns from these differences in first-enrichment scores from training samples to identify whether a patient belongs to the first category based on multiple first-enrichment scores. The multiple first-enrichment scores of the target patient are input into the first classification model. The model classifies the patient according to the learned rules to determine if they belong to the first category. If the target patient belongs to the first category, it provides a reference for treatment selection, favoring combination therapy. The specific drug combination for the combination therapy can be determined based on the patient's specific condition and clinical experience. If the target patient does not belong to the first category, further classification analysis is required, leading to step S108.

[0085] S108, if not, then multiple second enrichment scores are obtained based on tumor immune enrichment analysis of the expression profile data.

[0086] It is understandable that tumor immune enrichment analysis is a single-sample gene set enrichment analysis based on a tumor immune gene set. The tumor immune gene set is a collection of multiple genes related to tumor immunity. The second enrichment score is the result of tumor immune enrichment analysis, used to measure the degree of enrichment of gene expression in a specific tumor immune gene set in a target patient.

[0087] If a target patient is determined not to belong to category one in the first classification model, further analysis is needed to determine whether the patient belongs to category two or three. Since patients not belonging to category one show relatively small differences at the pathway level, further classification of these patients is required from other aspects. Tumor immune enrichment analysis can analyze tumor immune characteristics based on the expression profile data of individual patients, providing a more personalized assessment of the enrichment of tumor immune genes in the patient's set. The tumor immune gene set is closely related to the tumor's immune response; by analyzing the degree of enrichment in this gene set, the patient's tumor immune status can be understood. The selection of the tumor immune gene set can be obtained by collecting and summarizing immune cell enrichment pathways from relevant literature in the field of tumor research. Analysis of data related to satellite instability patients belonging to category one revealed that the tumor immune enrichment analysis results of multiple patients in this category reflected no difference in their tumor immune status.

[0088] However, data analysis of microsatellite instability patients not belonging to the first category revealed significant differences in tumor immune responses between patients treated with monoclonal antibody alone and those treated with both monoclonal antibody and combination therapies. Specifically, patients not belonging to the first category showed enrichment in both positive and negative regulation of T cells, an immune cell pathway. K-means clustering of this group also showed two distinct subgroups when k=2. Therefore, when classifying patients not belonging to the first category, tumor immune gene set enrichment analysis can be performed to obtain quantitative parameters of tumor immunology. A second enrichment score can then be extracted based on the analysis results to determine the patient's tumor immune status, providing a reference for further classification of patients not belonging to the first category. This can be achieved using the ssGSEA function provided by the R package IOBR (Immune Oncology Biomarker Research).

[0089] S110 inputs multiple second enrichment scores into the second classification model to determine whether the target patient belongs to the second or third category. Patients in the second category are treated with monoclonal antibody regimens, while patients in the third category are treated with both monoclonal antibody regimens and combination therapy regimens.

[0090] As can be understood, the secondary classification model is used to classify target patients based on input features (here, multiple secondary enrichment scores). As explained above, the tumor immune status of a patient can be determined based on multiple secondary enrichment scores, which in turn determines whether the patient is suitable for either a monoclonal antibody regimen only or a combination therapy regimen. In other words, there are significant differences in the secondary enrichment scores between these two types of patients.

[0091] The secondary classification model learns from the differences in secondary enrichment scores of training samples to identify whether a patient belongs to a second or third category based on multiple secondary enrichment scores. The target patient's multiple secondary enrichment scores are input into the secondary classification model. The model classifies the patient according to the learned rules, determining whether the patient belongs to the second or third category. If the target patient belongs to the second category, a monoclonal antibody treatment regimen is more likely to be chosen. If the target patient belongs to the third category, a monoclonal antibody treatment regimen or a combination therapy regimen will be chosen based on the patient's financial situation.

[0092] Based on the microsatellite instability (MSI) patient classification method in this embodiment, pathway enrichment analysis and tumor immune enrichment analysis are performed on patient expression profile data to gain a deeper understanding of the patient's disease status, immune characteristics, and biological features from different perspectives. Pathway enrichment analysis of expression profile data helps identify differences in the degree of enrichment at the biological pathway level among patients, determining whether they belong to the first category. Tumor immune enrichment analysis targeting the tumor immune gene set focuses on tumor immune status, providing a classification basis for patients who do not belong to the first category. The classification results obtained by this method correspond to different treatment plans, improving the targeting and effectiveness of treatment planning and contributing to improving the prognosis of MSI patients.

[0093] In one embodiment, please refer to Figure 2 The construction process of the first classification model includes steps S202 to S208.

[0094] S202, Obtain the first training set. The labels on the training samples in the first training set are used to indicate whether the training samples belong to the first category.

[0095] It is understandable that the first training set is a dataset specifically prepared for training the first classification model, containing expression profile data from multiple different patients as training samples. The first training set can be obtained from MSI patients in the U-CAN (Uppsala Comprehensive Cancer Consortium) dataset. The annotations in the first training set are used to indicate whether each training sample belongs to the first category. These annotations can be done manually or based on the clustering results of the training samples. For the microsatellite instability patient classification problem, obtaining a large amount of patient data as the first training set and accurately annotating it provides a reliable foundation for subsequent training of the first classification model.

[0096] S204. Pathway enrichment analysis is performed on each training sample in the first training set, and the path enrichment analysis results corresponding to each training sample in the first training set are screened to obtain multiple corresponding first training enrichment scores.

[0097] It's understandable that, since the input to the first classification model is the first enrichment score, during the training phase, it's necessary to first obtain multiple first training enrichment scores corresponding to each training sample in the first training set. The first training enrichment scores and the first enrichment scores are of the same type; the former is for training samples in the first training set, while the latter is for the expression profile data of the target patient. Therefore, it's necessary to first perform pathway enrichment analysis on each training sample in the first training set using the same first database, and then select the most important features to obtain the first training enrichment score. Specifically, the importance of a feature can be determined based on its contribution to classifying the training sample into the first category. After determining which types of features need to be retained, in the actual inference phase, the same types of features from the pathway enrichment analysis results of the target patient's expression profile data can also be retained.

[0098] S206. Set up multiple first base models, and train each first base model according to the annotation of the training samples in each first training set and the multiple first training enrichment scores.

[0099] It is understandable that the first base model is the initial model used to train the first classification model, which can include various different classification models, such as decision tree (DT), random forest (RF), support vector machine (Linear Support Vector Classification, linear SVC), logistic regression (logistic), k-nearest neighbor classifier (kNeighborsClassifier), regularized linear (SVM) model based on stochastic gradient descent (sgdclassifier), meta-estimator (adaboost), etc.

[0100] When training the intrinsic parameters of the first base model, it is necessary to select appropriate hyperparameters. This can be achieved using the hierarchical multiple cross-validation (GridSearchCV) function in sklearn on the first training set, thereby selecting the optimal hyperparameters for each first base model. After setting the optimal hyperparameters, the first base models can be trained using multiple first training enrichment scores corresponding to each training sample as inputs and labels as outputs.

[0101] S208, determine the model performance of each first base model, and select the first base model with the best model performance as the first classification model.

[0102] As we can understand, model performance is a metric that measures how well a model performs in a prediction task. It can include metrics such as the area under the operating feature curve (AUROC), average precision (AP), accuracy (ACC), average weighted accuracy (weighted ACC), average recall (weighted recall), average F1 score (weighted F1), Cohen's Kappa coefficient, and Matthews correlation coefficient (MCC).

[0103] In one embodiment, the annotation process of training samples in the first training set includes steps S302 to S308.

[0104] S302: Select the clustering parameter with the highest clustering stability from multiple sets of candidate clustering parameters as the target clustering parameter.

[0105] Clustering parameters refer to the parameter settings used to control the behavior of the clustering algorithm during cluster analysis. Clustering stability is an indicator of the consistency of clustering results in repeated experiments. Target clustering parameters are the final combination of parameters selected for the actual clustering analysis. Cluster analysis is an unsupervised learning method that aims to divide samples in a dataset into different groups or clusters, such that samples within the same cluster have high similarity, while samples between different clusters have significant differences. Different clustering parameter settings affect the quality and stability of the clustering results. Choosing parameters with the highest clustering stability can improve the reliability and reproducibility of the clustering results and reduce uncertainty caused by inappropriate parameter selection. For example, the clustering in this step can be performed using the FindCluster method, and the available clustering parameters include k.param, PCs, and resolution. The final target clustering parameters can be k.param=20, PCs=20, and resolution=0.9.

[0106] S304, Cluster the first dataset according to the target clustering parameters to obtain the target clustering result; the first dataset includes the first training set.

[0107] It is understandable that the first dataset includes the first training set, since the effectiveness of the clustering algorithm is related to the distribution of the objects to be clustered. In some embodiments, the validation set used to verify model performance is not labeled using a clustering algorithm; in this case, the first dataset may only include the first training set. However, in some embodiments, both the validation set and the training set come from the same group of patients (such as all or part of the MSI patients in a public dataset). In this case, the expression profile data of this group of patients can be selected as the first dataset, a portion of which is selected as the first training set to train each first base model, and the remainder is selected as the first validation set to verify the model performance of each first base model. For example, 80% can be randomly selected as the first training set, and the remaining 20% ​​as the first validation set. Specifically, the first validation set can be used to make predictions on the samples in the first validation set using the first base model, and the aforementioned model performance indicators can be calculated based on the labeling and prediction results to verify the performance of each first base model.

[0108] Based on the selected target clustering parameters, the first dataset is clustered using an appropriate clustering algorithm. This allows similar expression profile data to be grouped together into different clusters. This reveals the underlying structure and patterns in the data, providing useful information for subsequent classification and analysis.

[0109] S306, clusters in the target clustering results that cannot be further classified based on tumor immune enrichment analysis are identified as belonging to the first category, and clusters in the target clustering results that can be further classified based on tumor immune enrichment analysis are identified as not belonging to the first category.

[0110] It is understood that the target clustering results will contain at least two clusters, but the specific category to which each cluster belongs is not yet clear. In this step, each cluster in the target clustering results will undergo tumor immune enrichment analysis based on the tumor immune gene set. Then, based on the results of the tumor immune enrichment analysis, it will be further determined whether these clusters can be further classified. If a cluster shows that it cannot be further classified in the tumor immune enrichment analysis, it means that the samples in that cluster have a high degree of consistency in tumor immune status and do not belong to a category that needs further classification; therefore, this cluster belongs to the first category. Conversely, if a cluster shows that it can be further classified in the tumor immune enrichment analysis, it means that the samples in that cluster need further analysis and do not belong to the first category.

[0111] S308, based on the target clustering results, label whether the samples in the first training set belong to the first category.

[0112] It is understandable that labeling each sample in the first training set based on the classification of different clusters in the target clustering results can provide accurate category information for subsequent model training. If the first dataset also includes a first validation set, then the samples in the first validation set can also be labeled based on the target clustering results.

[0113] In one embodiment, please refer to Figure 4 The process of determining cluster stability includes steps S402 to S406.

[0114] S402, for any set of candidate clustering parameters, cluster the first dataset according to the set of candidate clustering parameters to obtain the first clustering result.

[0115] It is understandable that the first clustering result is the sample clustering situation obtained after performing clustering analysis on the first dataset using a specific set of candidate clustering parameters. The first clustering result includes at least two clusters, and the specific number of clusters depends on the clustering algorithm used and the candidate clustering parameters. By using different candidate clustering parameters to cluster the first dataset, the clustering effect under various parameter combinations can be evaluated, providing a basis for selecting the optimal clustering parameters.

[0116] S404. Perform multiple random samplings on the first dataset, and cluster the datasets obtained from each sampling according to the clustering parameters of the candidate set to obtain multiple second clustering results.

[0117] It is understandable that random sampling here is used to randomly select a portion of the samples from the overall dataset, for example, 80%. The second clustering result is the grouping of samples obtained after clustering the randomly sampled dataset. The second clustering result also includes at least two clusters, the specific number of which depends on the clustering algorithm used and the candidate clustering parameters. For each sampling, a certain proportion of samples are randomly selected from the first dataset to form a new dataset. In this step, the new dataset formed by sampling is clustered using the same candidate clustering parameters as in step S402. This process is repeated N (e.g., 100) times to obtain N second clustering results.

[0118] S406. Based on the similarity between each second clustering result and the first clustering result, determine the clustering stability of the group of candidate clustering parameters.

[0119] Similarity is an metric that measures the closeness between two clustering results. If a set of candidate clustering parameters produces similar clustering results on different subsets of data, it indicates that the set of parameters has high stability. The stability of this set of candidate clustering parameters can be evaluated by comparing the similarity between the second clustering results obtained from each random sampling and the first clustering results obtained from the initial clustering of the entire first dataset. The Jaccard index can be used as the similarity metric. The similarity between the second and first clustering results is determined by calculating the Jaccard index between one cluster in the second clustering result and its corresponding cluster in the first clustering result, and between another cluster in the second clustering result and its corresponding cluster in the first clustering result.

[0120] In one embodiment, please refer to Figure 5 The pathway enrichment analysis results corresponding to each training sample in the first training set are screened, including steps S502 to S506.

[0121] S502, determine the classification contribution of each feature in the pathway enrichment analysis results corresponding to each training sample in the first training set.

[0122] S504: Select the first number of features with the largest classification contribution as the first target feature.

[0123] S506, Based on the first target characteristics, the pathway enrichment analysis results corresponding to each training sample in the first training set are screened.

[0124] It is understandable that classification contribution is an indicator of how important a feature is in distinguishing sample categories. In some embodiments, classification contribution can be obtained based on the ReliefF algorithm. Feature selection is crucial to model performance. Using too many irrelevant or redundant features may lead to overfitting, increased training time, and decreased performance. By determining the classification contribution of features, the features most helpful for the classification task can be selected, improving the model's accuracy and efficiency. After determining the classification contribution of each feature, selecting the features with the largest contribution can reduce the dimensionality of the feature space, reduce computational complexity, and improve the model's generalization ability. By selecting the most representative features, key information in the data can be better captured, improving classification accuracy. After determining which features in the pathway enrichment analysis results belong to the first target feature based on classification contribution, the features corresponding to the first target feature in the pathway enrichment analysis results can be retained. In addition, before screening based on the first target feature, the first target feature can be supplemented based on expert experience, adding some features related to immunotherapy. In addition to being used in the training phase, the first target feature also needs to be used to screen the expression profile data of the target patients in the actual inference phase.

[0125] In one specific embodiment, the first target features include: INTERLEUKIN_27_SIGNALING, METABOLISM_OF_ANGIOTENSINOGEN_TO_ANGIOTENSINS, RUNX3_REGULATES_IMMUNE_RESPONSE_AND_CELL_MIGRATION, G_PROTEIN_BETA_GAMMA_SIGNALLING, INTERLEUKIN_2_FAMILY_SIGNALING, HDL_ASSEMBLY, etc.

[0126] In one specific embodiment, the models trained using the first target features and the first training set extracted from the U-CAN dataset as the base models—Support Vector Machine (Linear SVC), Regularized Linear SGD based on stochastic gradient descent, KNeighbors Vote, Logistic Regression, Decision Tree, Random Forest, and Adaboost—can be evaluated using various model performance metrics. Figure 6 The performance of these models is demonstrated. Figure 7The results show the performance of these models evaluated using metrics such as accuracy (ACC), balanced ACC, weighted precision, weighted recall, weighted F1 score, Kappa coefficient, and Matthews correlation coefficient (mcc). Figure 8 The PR (precision-recall) curves for these models are shown. (Summary) Figures 6 to 8 The evaluation results show that the model trained using support vector machines as the base model has the best performance.

[0127] In one embodiment, please refer to Figure 9 The construction process of the second classification model includes steps S902 to S910.

[0128] S902, the second dataset is obtained based on the training samples in the first dataset that do not belong to the first category. The second dataset includes a second training set and a second validation set.

[0129] The second dataset can be understood as a new training set composed of samples from the first dataset that do not belong to the first category. The second dataset includes a second training set for model training and a second validation set for validating model performance. For example, 80% could be randomly selected as the second training set, and the remaining 20% ​​as the second validation set. The samples in both the second training and validation sets are collectively referred to as training samples. In multi-class classification tasks, it is usually necessary to analyze and model different categories separately. By combining samples from the first dataset that do not belong to the first category into the second dataset, we can specifically classify these samples to determine whether they belong to the second or third category. This allows for more targeted use of the data and improves classification accuracy.

[0130] S904, tumor immune enrichment analysis is performed on the training samples in the second dataset to obtain multiple corresponding second training enrichment scores.

[0131] It's understandable that, since the input to the second classification model is the second enrichment score, multiple second training enrichment scores corresponding to each training sample in the second dataset need to be obtained during the training phase. The second training enrichment scores refer to the same tumor immune gene set as the second enrichment scores, except that the former is for the training samples in the second dataset, while the latter is for the target patients. The second training enrichment scores corresponding to the second training set portion of the second dataset will be used as input during the training of the second base model, while the second training enrichment scores corresponding to the second validation set portion will be used as input during the performance validation of the second base model. The tumor immune gene set contains information on multiple tumor immune gene sets, including the gene set name and genes associated with that set, and may also include webpage information from the references that proposed the gene set. It's worth noting that if a gene set has a clearly defined molecular pathway, then the gene set name is the name of the corresponding molecular pathway. Calculating the second enrichment score and the second training enrichment score involves calculating the enrichment score on the gene set within the tumor immune gene set information. Therefore, the gene set in the tumor immune gene set information can be used as the second target feature. The score of the training sample on the second target feature is the second training enrichment score, and the score of the target patient on the second target feature is the second enrichment score. The tumor immune gene set can be obtained by collecting relevant scientific research literature on tumor immune genes.

[0132] S906, clustering is performed based on the multiple second training enrichment scores corresponding to each training sample in the second dataset to label whether the training samples in the second dataset belong to the second category or the third category.

[0133] Understandably, since data analysis revealed a clear stratification in the tumor immunological responses of target patients not belonging to the first category, their second training enrichment scores can be used to divide the samples into two categories. Then, based on the difference in prognosis between monoclonal antibody regimens and combination therapy regimens, the category with the better prognosis from monoclonal antibody regimens can be designated as the second category, and the category with similar prognoses between monoclonal antibody and combination therapy regimens can be designated as the third category. Based on this, the samples can be labeled according to the clustering results of each training sample in the second dataset. Specifically, this step can use k-means clustering, such as the ATC:skmeans method in the R package cola framework.

[0134] S908 sets up multiple second base models and trains each second base model according to the annotation of each training sample in the second training set and the multiple second training enrichment scores.

[0135] It is understandable that the second base model is the initial model used to train the second classification model, which can include various different classification models, such as Random Forest (RF), Linear Support Vector Classification (linear SVC), Logistic Regression (logistic), Regularized Linear (SVM) models based on stochastic gradient descent (SGD training, sgdclassifier), extreme gradient boosting (xgbc), and multi-layer perceptron classifier (mlpc), etc.

[0136] When training the intrinsic parameters of the second base model, it is necessary to select appropriate hyperparameters. The hierarchical multiple cross-validation (GridSearchCV) function in sklearn can be used to perform validation on the second training set, thereby selecting the optimal model hyperparameters for each second base model. After setting the optimal hyperparameters, supervised training can be performed on each second base model using the multiple second training enrichment scores corresponding to each training sample in the second training set as features and the clustering categories as class labels.

[0137] S910: The model performance of each second base model is determined by the annotation of the second validation set, and the second base model with the best model performance is selected as the second classification model.

[0138] Model performance is understood to be a metric that measures how well a model performs in a prediction task. This can include metrics such as the area under the operating feature curve (AUROC), average precision (AP), accuracy (ACC), weighted average accuracy (weighted ACC), weighted average recall (weighted Recall), weighted average F1 score (weighted F1), Cohen's Kappa coefficient, and Matthews correlation coefficient (MCC). Specifically, the use of a second validation set can involve using a second base model to predict samples in the second validation set, and then calculating the aforementioned model performance metrics based on the labeled and prediction results to validate the performance of each second base model.

[0139] In one specific embodiment, the second target features include: EMT1, detox.iCAF, IFNG.iCAF, Normal.Fibroblast, GPAGs, etc.

[0140] In one specific embodiment, models trained using the aforementioned second target features and the second training set as base models—Support Vector Machine (Linear SVC), Logistic Regression, Regularized Linear SGD (based on stochastic gradient descent), Multilayer Perceptron Classifier (MLPClassifier), Random Forest, and Extreme Gradient Boosting (XGBoost)—can be evaluated using various model performance metrics. Figure 10 The results show the performance of these models evaluated using the area under the operating characteristic curve (AUROC). Figure 11 The results show the performance of these models evaluated using metrics such as accuracy (ACC), balanced accuracy (ACC), weighted precision (weighted precision), weighted recall (weighted recall), weighted F1 score, Kappa coefficient, and Matthews correlation coefficient (mcc). Figure 12 The PR (precision-recall) curves for these models are shown. (Summary) Figures 10 to 12 The evaluation results show that the model trained using support vector machines as the base model has the best performance.

[0141] To validate the first and second classification models obtained from training. Figures 13-18 This paper presents heatmaps of cytokine signals obtained from classifying different datasets. The datasets used are U-CAN, TCGA-COADREAD, TCGA-UCEC, TCGA-STAD, Nipicol, and ImmunoMSI. The class labels for U-CAN were obtained using a clustering algorithm and were used for model training. The class labels for the other datasets were predicted by the trained first and second classification models. The cytokine distribution across these datasets is consistent: the first class is more active in the IL family, the second class shows stronger activity in most cytokines, and the third class shows activity in only some cytokines. The correlation between each class and treatment regimen can be obtained by comparing the progression-free survival of patients in the Nipicol and ImmunoMSI datasets. Figures 19 to 25 .in, Figure 19 It is a progression-free survival analysis of patients in three categories after immunotherapy, based on the training set. Figure 20 It is a progression-free survival analysis of all patients in the training set on monoclonal antibodies (PD1 monoclonal antibody) and combination therapy regimens (PD1 monoclonal antibody and CTLA4 monoclonal antibody). Figure 21This is a progression-free survival analysis of the training set of patients in three categories under the combination therapy regimen (PD1 monoclonal antibody and CTLA4 monoclonal antibody). Figure 22 This is a progression-free survival analysis of the training set under monoclonal antibody (PD1 monoclonal antibody) regimens in three categories of patients. Figure 23 The training set is a progression-free survival analysis of patients in the first category, on monoclonal antibody regimens (PD1 monoclonal antibody) and combination regimens (PD1 monoclonal antibody and CTLA4 monoclonal antibody). Figure 24 The training set is a progression-free survival analysis of patients in the second category, on monoclonal antibody regimens (PD1 monoclonal antibody) and combination regimens (PD1 monoclonal antibody and CTLA4 monoclonal antibody). Figure 25 This is a progression-free survival analysis of patients in category 3 in the training set, specifically for monoclonal antibody therapy (PD-1 monoclonal antibody) and combination therapy (PD-1 monoclonal antibody and CTLA4 monoclonal antibody). Figure 23 It is clear that patients in the first category have a better prognosis with combination therapy. Figure 24 It is clear that patients in the second category have a better prognosis with monoclonal antibody treatment. Figure 25 It is clear that patients in the third category have similar prognoses for monoclonal antibody therapy and combination therapy. Figures 19 to 25 This demonstrates the rationality of the correspondence between the categories identified in this application and the medication regimen.

[0142] This application provides a microsatellite instability patient classification device, comprising: a data acquisition module for acquiring expression profile data of a target patient; a first enrichment analysis module for performing pathway enrichment analysis on the expression profile data and filtering the pathway enrichment analysis results to obtain multiple first enrichment scores; a first classification module for inputting the multiple first enrichment scores into a first classification model to determine whether the target patient belongs to the first category. Patients in the first category are suitable for combination therapy. A second enrichment analysis module for performing tumor immune enrichment analysis on the expression profile data when the target patient does not belong to the first category, to obtain multiple second enrichment scores; and a second classification module for inputting the multiple second enrichment scores into a second classification model to determine whether the target patient belongs to the second or third category. Patients in the second category are suitable for monoclonal antibody therapy, while patients in the third category are suitable for both monoclonal antibody therapy and combination therapy.

[0143] Specific limitations regarding the microsatellite instability patient classification device can be found in the above-described limitations of the microsatellite instability patient classification method, and will not be repeated here. Each module in the aforementioned microsatellite instability patient classification device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module. It should be noted that the module division in this embodiment is illustrative and only represents a logical functional division; other division methods may be used in actual implementation.

[0144] This application provides a computer device including one or more processors and a memory storing computer-readable instructions. When executed by one or more processors, the computer-readable instructions perform the microsatellite instability patient classification method in any of the above embodiments.

[0145] Indicatively, such as Figure 26 As shown, Figure 26 This is a schematic diagram of the internal structure of a computer device provided in an embodiment of this application. (Refer to...) Figure 26 The computer device 2600 includes a processing component 2602, which further includes one or more processors, and memory resources represented by memory 2601 for storing instructions, such as application programs, that can be executed by the processing component 2602. The application programs stored in memory 2601 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 2602 is configured to execute instructions to perform the microsatellite instability patient classification method of any of the above embodiments.

[0146] The computer device 2600 may also include a power supply component 2603 configured to perform power management of the computer device 2600, a wired or wireless network interface 2604 configured to connect the computer device 2600 to a network, and an input / output (I / O) interface 2605.

[0147] This application provides a storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the microsatellite instability patient classification method in any of the above embodiments.

[0148] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0149] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0150] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for classifying patients with microsatellite instability, characterized in that, include: Obtain expression profile data from the target patients; Pathway enrichment analysis was performed on the expression profile data, and the results of the pathway enrichment analysis were filtered to obtain multiple first enrichment scores. Multiple first enrichment scores are input into a first classification model to determine whether the target patient belongs to a first category; patients in the first category are suitable for combination therapy. If not, tumor immune enrichment analysis is performed on the expression profile data to obtain multiple second enrichment scores; Multiple second enrichment scores are input into a second classification model to determine whether the target patient belongs to a second or third category; patients in the second category are treated with monoclonal antibody regimens, while patients in the third category are treated with both monoclonal antibody regimens and combination therapy regimens.

2. The microsatellite instability patient classification method according to claim 1, characterized in that, The construction process of the first classification model includes: Obtain the first training set; the annotations on the training samples in the first training set are used to indicate whether the training samples belong to the first category; Path enrichment analysis is performed on each training sample in the first training set, and the path enrichment analysis results corresponding to each training sample in the first training set are filtered to obtain multiple first training enrichment scores. Multiple first base models are set up, and each first base model is trained according to the annotation of the training samples in each first training set and the multiple first training enrichment scores. Determine the model performance of each of the first base models, and select the first base model with the best model performance as the first classification model.

3. The microsatellite instability patient classification method according to claim 2, characterized in that, The annotation process for the first training set includes: Select the clustering parameter with the highest clustering stability from multiple sets of candidate clustering parameters as the target clustering parameter; The first dataset is clustered according to the target clustering parameters to obtain the target clustering result; the first dataset includes a first training set; Clusters in the target clustering results that cannot be further classified based on tumor immune enrichment analysis are identified as belonging to the first category, while clusters in the target clustering results that can be further classified based on tumor immune enrichment analysis are identified as not belonging to the first category. Based on the target clustering results, label whether the samples in the first training set belong to the first category.

4. The method for classifying microsatellite unstable patients according to claim 3, characterized in that, The process of determining the cluster stability includes: For any set of candidate clustering parameters, the first dataset is clustered according to the set of candidate clustering parameters to obtain the first clustering result; Multiple random samplings are performed on the first dataset, and the datasets obtained from each sampling are clustered according to the candidate clustering parameters described in the group, resulting in multiple second clustering results; The clustering stability of the group of candidate clustering parameters is determined based on the similarity between each of the second clustering results and the first clustering results.

5. The method for classifying microsatellite unstable patients according to claim 3, characterized in that, The first dataset also includes the first validation set used to determine the model performance of each of the first base models.

6. The microsatellite instability patient classification method according to claim 2, characterized in that, The process of filtering the pathway enrichment analysis results corresponding to each training sample in the first training set includes: Determine the classification contribution of each feature in the pathway enrichment analysis results corresponding to each training sample in the first training set; Select the first number of features with the largest classification contribution as the first target features; Based on the first target feature, the pathway enrichment analysis results corresponding to each training sample in the first training set are screened.

7. The method for classifying microsatellite unstable patients according to claim 6, characterized in that, The filtering of pathway enrichment analysis results of the expression profile data includes: Based on the first target feature, the pathway enrichment analysis results of the expression profile data are screened.

8. The method for classifying microsatellite unstable patients according to claim 3, characterized in that, The construction process of the second classification model includes: A second dataset is obtained from the training samples in the first dataset that do not belong to the first category; the second dataset includes a second training set and a second validation set; Tumor immune enrichment analysis was performed on the training samples in the second dataset to obtain multiple corresponding second training enrichment scores. Clustering is performed based on the multiple second training enrichment scores corresponding to each training sample in the second dataset to label the training samples in the second dataset as belonging to the second category or the third category; Multiple second base models are set up, and each second base model is trained according to the annotation of each training sample in the second training set and the multiple second training enrichment scores. The model performance of each second base model is determined using the annotations of the second validation set, and the second base model with the best model performance is selected as the second classification model.

9. A microsatellite instability patient classification device, characterized in that, include: The data acquisition module is used to acquire expression profile data of the target patients; The first enrichment analysis module is used to perform pathway enrichment analysis on the expression profile data and to filter the pathway enrichment analysis results of the expression profile data to obtain multiple first enrichment scores. The first classification module is used to input multiple first enrichment scores into the first classification model to determine whether the target patient belongs to the first category; The first category of patients is suitable for combination therapy; The second enrichment analysis module is used to perform tumor immune enrichment analysis on the expression profile data when the target patient does not belong to the first category, and obtain multiple second enrichment scores. The second classification module is used to input multiple second enrichment scores into the second classification model to determine whether the target patient belongs to the second category or the third category; patients in the second category are subject to monoclonal antibody treatment regimens, while patients in the third category are subject to both monoclonal antibody treatment regimens and combination therapy regimens.

10. A computer device, characterized in that, It includes one or more processors and a memory storing computer-readable instructions that, when executed by the one or more processors, perform the microsatellite instability patient classification method as described in any one of claims 1-8.

11. A storage medium, characterized in that, The storage medium stores computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform the microsatellite instability patient classification method as described in any one of claims 1-8.