Small-sample Classification Method for Cancer Molecular Subtypes Based on Meta-Learning
Through a meta-learning-based method, combining the gradient update cycle of MAML and prototype networks, a task-specific embedded algorithm is designed to solve the problem of small sample classification of cancer molecules subtypes and achieve high accuracy classification effect.
Patent Information
- Application Number
- CN202211214085.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-09-30
AI Technical Summary
The prior art is difficult to effectively solve the problem of small sample classification of cancer molecular subtypes, especially due to the scarcity of sample size and the complexity of the data set.
Using a meta-learning-based method, combining the gradient update cycle of MAML and prototype networks, a meta-learning algorithm with specialized embedded task is designed, and a small sample classification of cancer molecular subtypes is constructed by constructing a neural network model, and a feature extraction and classification is used for the gene expression profile and cancer information of the TCGA database.
A 70.06% accuracy was achieved in the 1-shot test setting and an 82.33% accuracy was achieved in the 5-shot test setting, significantly improving the accuracy of cancer molecular subtype classification.
Smart Images

Figure CN115423043B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of bioinformatics and machine learning, and particularly relates to a small-sample classification method for cancer molecular subtypes based on meta-learning. Background Art
[0002] Cancer is a heterogeneous disease with multiple pathogenesis mechanisms and clinical features. Traditional histopathological classification based on the histological appearance and morphological features of tumors only reflects some of the heterogeneous features of cancer. Molecular analysis techniques have shifted cancer classification from traditional morphology to molecular subtypes based on molecular features. Cancer molecular subtype classification can help people understand the pathogenesis of cancer more deeply and is beneficial to the precise diagnosis and personalized treatment of cancer. Different cancer molecular subtypes often mean different gene mutations. The development of high-throughput sequencing technology has made RNA sequencing the main method for studying gene expression regulation, and the data obtained therefrom has provided unprecedented opportunities for people to study cancer molecular subtypes. However, the samples belonging to certain cancer molecular subtypes are extremely scarce, which poses a huge challenge to cancer molecular subtype classification. Comprehensive analysis is a common method to solve this problem, which increases the sample size by integrating data from different experiments or platforms. However, the complex characteristics of gene expression data will cause the aggregation of data from different studies to be negatively affected by batch effects, heterogeneity, and other sources of bias.
[0003] As an emerging deep learning paradigm, few-shot learning has brought a new way to solve problems. Although Samiei et al. have proposed a dataset named TCGA meta-dataset that can be used for few-shot learning in the paper "The tcga meta-dataset clinical benchmark". However, this dataset cannot provide the most basic N-way K-shot training setting for few-shot learning, and most of the tasks it contains are about clinical classification and are not closely related to cancer molecular subtype classification. At present, almost all methods only focus on the molecular subtype classification of specific cancers and do not regard cancer molecular subtype classification as a class of problems. Using the correlation between cancer molecular subtype classification and cancer classification to solve specific cancer molecular subtype classification or cancer classification problems is a novel idea and concept.
[0004] Meta-learning, also known as learning to learn, is the most common few-shot learning framework in recent years. MAML proposed in the paper "Model-agnostic meta-learning for fast adaptation of deep networks" is a representative method of optimization-based meta-learning, which consists of two nested gradient update loops. The outer loop of gradient update is responsible for finding the initial model parameters suitable for all tasks, while the inner loop of gradient update is responsible for task specialization of the model parameters. The paper "Rapid Learning or Feature Reuse?Towards Understanding the Effectiveness of MAML" believes that feature reuse is the dominant factor in the effectiveness of MAML, which provides the idea of using a more deterministic classification method to replace the last fully connected layer of the network (hereinafter referred to as the network head, and the other parts of the network are simply referred to as the network body). The prototypical network proposed in the paper "Prototypical networks for few-shot learning" is a classic method of metric-based meta-learning. It embeds instances from all tasks into the same low-dimensional space, where features belonging to the same class are relatively close in the embedding space, while features belonging to different classes are relatively far apart in the embedding space. The prototypical network classifies based on the feature centroid closest to the instance features in the task. In the dataset, the relationship between instances changes with different tasks, and the model needs to be able to change according to the specific task to perform different embeddings for instances of different tasks. Summary of the Invention
[0005] In order to alleviate the shortage of few-shot datasets in the field of bioinformatics, the present invention proposes a few-shot classification method for cancer molecular subtypes based on meta-learning, which realizes the few-shot classification of cancer molecular subtypes by using the meta-learning algorithm.
[0006] The present invention is realized by the following technical solutions:
[0007] A few-shot classification method for cancer molecular subtypes based on meta-learning, the method comprising the following steps:
[0008] Step 1, construct a dataset, the data including the gene expression profile of the TCGA database, the cancer information of the sample and the cancer molecular subtype information;
[0009] Set the scenario for few-shot learning;
[0010] Design the task generation method for the small sample dataset, that is, appropriately intervene in the generation method of the training task, and maintain the task of cancer classification by extracting molecular subtypes of different cancers with a probability of 0.5, and the task of cancer classification by extracting molecular subtypes of the same cancer with a probability of 0.5;
[0011] Step 2. Design a meta-learning algorithm based on task-specialized embedding, and this algorithm includes the following processing:
[0012] Start by creating an embedding model g for feature extraction φ ;
[0013] Obtain a task;
[0014] Calculate the centroid of the case features of the support set ;
[0015] Calculate the cross-entropy loss of the support set and the query set ;
[0016] Use the calculated cross-entropy loss to perform gradient update on the model g φ and temporarily update the model several times;
[0017] Calculate the prediction vector z of the query set ; q ;
[0018] Judge whether it is currently in the training stage:
[0019] Branch 1: If it is in the training stage, continue to calculate the query set loss and accumulate it; judge whether to end the accumulation: if the accumulation ends, permanently update the model; further judge whether to continue the model iteration. If no further iteration is required, the process ends; if the loss of the query set ends, then go to the step of obtaining the next task after the start;
[0020] Branch 2: If it is not in the training stage, further judge whether to continue the model iteration. If the model iteration continues, go to the step of obtaining the next task after the start; if the iteration does not continue, the algorithm process ends;
[0021] The optimization objective of the algorithm is:
[0022]
[0023] Among them,
[0024]
[0025] Perform more updates with the support set:
[0026]
[0027] Among them, φ * is the optimization objective of the algorithm, φ is the model parameter, α is the learning rate at the task level, is the loss function, i is the task number, φ i is the updated model parameter of the i-th task, x q is the query set for few-shot learning in the gene expression profile elements, y q is the query set in the cancer information elements, is the gradient function, x s is the support set for few-shot learning in the gene expression profile elements, y s is the support set in the cancer information elements, is the mathematical expectation, is the i-th task satisfies the task distribution is the model parameter φ i when calculating the prediction vector z q of the entire process;
[0028] Step 3, construct the original neural network model;
[0029] Step 4, train and evaluate the original neural network model, specifically including:
[0030] Determine the evaluation metrics as macro-average precision PRE macro macro-average recall REC macro macro-average F1-score F1 macro and macro-average AUC macro .
[0031] Compared with the prior art, the present invention can achieve the following beneficial technical effects:
[0032] For the cancer and cancer molecular subtype classification problems, the accuracy rate ACC reaches up to 70.06% in the 1-shot test setting and up to 82.33% in the 5-shot test setting. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a flowchart of the few-shot classification method for cancer molecular subtypes based on meta-learning of the present invention;
[0034] Figure 2Schematic diagram of the generation process of the small sample data set of the TCGA database for the present invention; (2a) Cancer molecular subtype classification in the test stage, (2b) Cancer molecular subtype classification in the training stage, (2c) Cancer classification task.
[0035] Figure 3 Meta-learning based on task-specialized embedding (TSEBML algorithm) of the present invention.
[0036] Figure 4 Schematic diagram of the structure of the original neural network model constructed by the present invention.
[0037] Figure 5 Cancers and cancer molecular subtypes included in the test set for each fold and the corresponding number of samples.
[0038] Figure 6 Schematic diagram of the t-SNE visualization result of the embeddings of all instances learned by the model through the TSEBML algorithm in the embodiment. Detailed implementation manners
[0039] To further elaborate the technical solutions and features of the present invention, the following takes a pumped storage power station project adopting the technical solutions of the present invention as an example and makes a detailed description in conjunction with the drawings, but the content of the present invention is not limited to the content described in the specific embodiments.
[0040] Construct a data set, design an algorithm, construct a neural network model, train the model and evaluate
[0041] As Figure 1 shown, a small sample classification method for cancer molecular subtypes based on meta-learning of the present invention includes the following steps:
[0042] Step 1: Construct a data set, which specifically includes the following processing:
[0043] Step 1-1: Data preparation. The data includes the gene expression profile X of the TCGA database, the cancer information Y1 of the samples and the cancer molecular subtype information Y2 of the samples. Among them, the gene expression profile X and the cancer information Y1 of the samples are downloaded from https: / / xenabrowser.net / datapages / , and the cancer molecular subtype information Y2 of the samples is obtained from the literature "The cancer genome atlas pan-cancer analysis project" and "Pan-cancer machine learning predictors of primary site of origin and molecular subtype".
[0044] Specifically, ① the gene expression profile X and the cancer information Y1 of the samples are obtained as follows: First, download the IlluminaHiSeq pan-cancer normalized data of all cancers from the TCGAHub of UCSCXena. These data are obtained by performing a log(x + 1) transformation on the counts after RSEM normalization downloaded from the TCGA Data Coordination Center, and then performing mean normalization on each gene across all TCGA cohorts to enable comparison between gene expression profiles from different cancers. There are relevant data for 33 cancers in the data, and the gene expression profile of each cancer is saved in a file with the tsv format. Each sample contains the expression levels of the same 20,530 genes. Since the present invention focuses on cancer classification and cancer molecular subtype classification problems, it is necessary to remove all normal samples that have not undergone carcinogenesis. ② The cancer molecular subtype information Y2 of the samples is obtained as follows: Obtain the corresponding table of sample IDs and cancer molecular subtypes from the literature "The cancer genome atlas pan-cancer analysis project" and "Pan-cancer machine learning predictors of primary site of origin and molecular subtype". If there are at least 20 samples for all molecular subtypes of a certain cancer, then select its molecular subtype classification problem for research; otherwise, do not select it. Among the 33 cancers, the molecular subtypes of 14 cancers, namely breast cancer (BRCA), head and neck cancer (HNSC), kidney clear cell carcinoma (KIRC), kidney papillary cell carcinoma (KIRP), low-grade glioma (LGG), liver cancer (LIHC), lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), ovarian cancer (OV), pheochromocytoma and paraganglioma (PCPG), melanoma (SKCM), gastric cancer (STAD), thyroid cancer (THCA), and endometrioid carcinoma (UCEC), are selected by the present invention, while the molecular subtypes of 19 cancers, namely adrenocortical carcinoma (ACC), bladder cancer (BLCA), cervical cancer (CESC), cholangiocarcinoma (CHOL), colon cancer (COAD), diffuse large B-cell lymphoma (DLBC), esophageal cancer (ESCA), glioblastoma multiforme (GBM), kidney chromophobe (KICH), acute myeloid leukemia (LAML), mesothelioma (MESO), pancreatic cancer (PAAD), prostate cancer (PRAD), rectal cancer (READ), sarcoma (SARC), testicular cancer (TGCT), thymoma (THYM), uterine carcinosarcoma (UCS), and uveal melanoma (UVM), are not selected.
[0045] Step 1-2: Set the scenario for small-sample learning;
[0046] In the scenario of few-shot learning, the dataset is set up differently from traditional methods. Let the dataset be where x k is the gene expression profile of the k-th sample, and y k is the cancer information of the k-th sample. Then its class set Before dividing the training set and the test set, first divide the class set Y into the class set of the training set and the class set of the test set, and Y tr ∪Y te = Y; then divide the dataset D into the training set D tr = {(x,y)|(x,y) ∈ D, y ∈ Y tr} and the test set D te = {(x,y)|(x,y) ∈ D, y ∈ Y te}. A task is the basic unit of training and testing in few-shot learning. Each task consists of a support set and a query set , that is Taking the training task as an example, then The N-way K-shot setting in few-shot learning means that each support set needs to contain N categories, and each category needs to have K samples. This also means that the class set of and |Y i tr | = N. For each there is If the query set is further set to Q-query, similarly, for each there is The above process describes how to generate a training task Following the same principle, the same process applies to the acquisition of the test task ;
[0047] Steps 1-3: Design the task generation method for the few-shot dataset, that is, appropriately intervene in the generation method of the training task, and maintain a 50% probability of extracting different molecular subtypes of cancer to form a cancer classification task, and a 50% probability of extracting the same molecular subtypes of cancer to form a cancer classification task. Each molecular subtype classification of cancer corresponds to a task. For example, classifying molecular subtype A and classification subtype B corresponds to two tasks. A task refers to a small dataset containing a support set and a query set extracted from big data.
[0048] If the number of molecular subtypes at this time is less than the way value for small sample training, then other molecular subtypes from the same cancer are further sampled until the requirement is met. Finally, all the sampled molecular subtypes are used to form the task of cancer molecular subtype classification. In the test phase, tasks are generated using all the classes included in a specific cancer molecular subtype classification problem or a cancer classification problem, without setting a fixed way value;
[0049] The idea of this step is as follows: The most ideal situation for establishing a small sample dataset to solve the cancer molecular subtype classification problem is to use only the tasks of cancer molecular subtype classification for training and testing. However, different cancers often have different numbers of molecular subtypes, which is contrary to the N-way setting of few-shot learning. Moreover, the types of tasks for cancer molecular subtype classification are scarce, and it is difficult to achieve ideal learning effects by only using them. Therefore, initially, it was considered to randomly sample all cancer molecular subtypes during training to form tasks, so as to meet the training setting requirements of few-shot learning and significantly increase the types of tasks. However, the molecular subtypes from different cancers essentially form a cancer classification task. To further increase the types of tasks and obtain more knowledge that can be shared by all cancer classification tasks from the cancer classification tasks, samples of cancers whose molecular subtypes are not selected are added to the dataset. As Figure 2 shown, it is a schematic diagram of the task generation process of the small sample dataset for the TCGA database of the present invention; (2a) Cancer molecular subtype classification in the test phase, (2b) Cancer molecular subtype classification in the training phase, (2c) Tasks of cancer classification.
[0050] Step 2: Combine ① the inner loop of gradient update of MAML and ② the classification method based on the nearest feature centroid of the prototype network to design a model of meta-learning based on task-specialized embedding (TSEBML algorithm), where ① can perform task-specialized embedding on instances, which not only increases the certainty of the method but also enables the model to test tasks with any number of classes. ② Performs task-specialized embedding, increasing the specialization ability of the method, thereby improving the accuracy. Regarding ① the inner loop of gradient update of MAML, the specific content is as follows:
[0051] To solve the classification problem, cross-entropy is selected as the loss function, and the expression is as follows:
[0052]
[0053] where l(y) is the label corresponding to y, z is the output vector of the classifier, z[j] is the j-th element in z, and y is the cancer information element in the class set;
[0054] The optimization objective of MAML, the expression is as follows:
[0055]
[0056] Among them,
[0057]
[0058] using the support set for more updates as an extension, that is:
[0059]
[0060] Among them, f θ is the classification model, θ is the parameter of the classification model, is the distribution of the task, is the task of molecular classification of cancer subtypes, α is the learning rate at the task level, is the loss function, θ * is the ideal optimization objective, i is the task number, θ i is the optimization objective of the i-th task, is the gradient function, x s is the support set for few-shot learning the gene expression profile element in, y s is the support set the cancer information element in, x q is the query set for few-shot learning the gene expression profile element in, y q is the query set the cancer information element in;
[0061] Regarding the classification method based on the nearest feature centroid of the ② prototype network, it specifically includes the following:
[0062] Set the centroid of the instance features of the support set of the j-th class, and the expression is as follows:
[0063]
[0064] Among them, g φ is the embedding model for feature extraction, φ is the model parameter, x is the gene expression profile element in the dataset, y is the cancer information element in the dataset,
[0065] x q The corresponding prediction vector is:
[0066] z q = [-d(g φ (x q ), c i,1 ), -d(g φ (x q ), c i,2 ),..., -d(g φ(x q ),c i,N )] (6)
[0067] where d is a distance function, and for each -d(g φ (x q ),c i,j ) is used as a metric for predicting whether x q belongs to class y i,j ;
[0068] The entire process of calculating the above-mentioned z φ is represented by the function h q , that is:
[0069] h φ (x q ,S i ) = z q (7)
[0070] Then the optimization objective of the prototype network is:
[0071]
[0072] Let Y i be the class set of task T i . For each y i,j ∈Y i (j = 1, 2,..., N), there is
[0073] Out of the consideration of keeping the training process and the testing process consistent, a classification method based on the nearest feature centroid is adopted to completely replace the network head;
[0074] By comparing the optimization objectives of MAML and the prototype network, it is found that the biggest difference between them lies in the different roles played by the support set. In MAML, the support set is used for inner-loop updates of the model parameters, so that the initial parameters shared by all tasks are updated to specialized parameters serving a specific task to complete the classification of query set instances. In the prototype network, the model embeds instances into a low-dimensional space shared by all tasks. The instance features belonging to the same class are relatively closer in the embedding space, while the instance features belonging to different classes are relatively farther apart in the embedding space. The prototype network completes the classification of query set instances by measuring the distances between the instance features of the support set and the query set in the embedding space. The tasks of the TCGA small sample dataset are divided into two types, namely the task of cancer molecular subtype classification and the task of cancer classification. Then, for two instances belonging to the same cancer but different cancer molecular subtypes, they belong to two different classes in the task of cancer molecular subtype classification, but belong to the same class in the task of cancer classification. That is to say, the dataset is still a simple multi-label dataset. Reconfirming the idea of the prototype network, it only considers the single-label situation, that is, instances belonging to the same class in one task will definitely belong to the same class in another task. At this time, it is reasonable to use the same model to embed instances of different tasks. However, when the relationship between instances changes with different tasks, it is hoped that the model can change according to the specific task, so as to perform different embeddings on instances of different tasks to improve the classification effect.
[0075] As Figure 3 shown, it is the flowchart of the meta-learning based on task-specialized embedding (TSEBML algorithm) of the present invention.
[0076] Start, create an embedding model g for feature extraction φ ;
[0077] Obtain a task;
[0078] Calculate the centroid of the case features of the support set ;
[0079] Calculate the cross-entropy loss of the support set and the query set ;
[0080] Use the calculated cross-entropy loss to perform gradient update on the model g φ , temporarily update the model several times, and the number of temporary updates is determined according to the actual situation;
[0081] Calculate the prediction vector z of the query set ; q ;
[0082] Judge whether it is currently in the training stage:
[0083] Branch 1: If it is in the training stage, continue to calculate the query set loss and accumulate it; determine whether to end the accumulation: if the accumulation ends, update the model permanently; further determine whether to continue the model iteration. If no further iteration is required, the process ends; if the loss of the query set ends accumulation, then go to the step of obtaining the next task after starting;
[0084] Branch 2: If it is not in the training stage, further determine whether to continue the model iteration. If the model iteration continues, go to the step of obtaining the next task after starting; if no further iteration is performed, the algorithm process ends;
[0085] The optimization objective of the TSEBML algorithm is:
[0086]
[0087] Among them,
[0088]
[0089] Similar to MAML, perform more updates with the support set:
[0090]
[0091] Among them, φ * is the optimization objective of the TSEBML algorithm, φ is the model parameter, α is the task-level learning rate, is the loss function, i is the task number, φ i is the updated model parameter of the i-th task, x q is the gene expression profile element in the query set of few-shot learning and y q is the cancer information element in the query set ; is the gradient function, x s is the gene expression profile element in the support set of few-shot learning and y s is the cancer information element in the support set ; is the mathematical expectation, is the i-th task satisfies the task distribution and is the entire process of calculating the predicted vector z i when the model parameter is φ q ;
[0092] Step 3, construct a neural network model, which specifically includes the following processing:
[0093] Step 3-1: Select a network architecture. According to the characteristic that genes are linearly arranged on chromosomes, a 1D-CNN, that is, a one-dimensional convolutional neural network (CNN), is adopted as the network architecture.
[0094] Step 3-2: Construct an original neural network model.
[0095] As Figure 4 shown, it is a schematic diagram of the structure of the original neural network model constructed in the present invention. The original neural network model is successively composed of a one-dimensional convolutional layer with a kernel size and a stride both of 130 and 32 filter numbers, a ReLU layer, a max-pooling layer with a kernel size and a stride both of 2, a flattening layer, a hidden layer of size 112, a ReLU layer, and an output layer of size equal to the number of classifications. The end of the RNA sequence data with a length of 20530 is supplemented with 10 zeros to form a single-channel input vector with a length of 20540. The present invention does not use a gene screening method to reduce the dimension of the input vector because the importance of the same gene varies greatly in different classification problems. It is hoped that the model can implicitly select the genes most helpful for completing the task by adjusting the parameters of the convolutional layer when performing task specialization.
[0096] Step 4: Train the model and evaluate:
[0097] Step 4-1: Determine evaluation metrics. In the case of binary classification, the performance metrics of the evaluation method include accuracy (ACC), precision (PRE), recall (REC), F1-score (F1), and the area under the receiver operating characteristic (ROC) curve (AUC). When the number of classifications is extended from 2 to N, the present invention uses the macro-average precision (PRE macro ), macro-average recall (REC macro ), macro-average F1-score (F1 macro ), and macro-average AUC (AUC macro ) to comprehensively evaluate the performance of the classification model from the classification situations of all classes. Since in the query set of each task, each class has the same number of samples, this results in REC macro = ACC at this time.
[0098] Step 4-2: Set up multiple comparison methods to evaluate the capabilities of the TSEBML algorithm.
[0099] The classifier baseline proposed by Chen et al. in the paper "Meta-baseline: exploring simple meta-learning for few-shot learning" is a baseline for solving the few-shot classification problem using traditional deep learning methods. Experiments using it have confirmed the remarkable effectiveness of the TSEBML algorithm as a meta-learning method in solving the few-shot classification problem. The classifier baseline first trains a classifier on all classes of the training set, and then at test time, removes the network head and classifies query set instances using a classification method based on the nearest feature centroid. The distance function it uses is the cosine distance (using the negative of the cosine similarity instead of adding 1 to denote the cosine distance). In the training phase of the classifier baseline, the batch size is set to 10, the learning rate is set to 5×10 -6 . The number of epochs is set to 200, and a test is conducted after each epoch.
[0100] The meta-baseline proposed by Chen et al. in the same paper is a baseline for meta-learning methods. Experiments using it are conducted to evaluate the competitiveness of the TSEBML algorithm in the field of meta-learning. The meta-baseline first pre-trains the model using the classifier baseline, then removes the network head, and uses a classification method based on the nearest feature centroid for training and testing. The distance function it uses is the product of the cosine distance and a learnable scalar τ. Similar to the TSEBML algorithm, the meta-baseline calculates the average loss of a batch of tasks for gradient update. In the training phase of the meta-baseline, the task batch size is set to 10, the learning rate is set to 1×10 -4 , and the learning rate of τ is set to 1×10 -1 . The training is iterated 1000 times in total, and a test is conducted every 10 iterations.
[0101] Experiments are conducted on the original network using MAML to evaluate the advantage of the classification method based on the nearest feature centroid compared to the linear layer. In the training phase of MAML, the task batch size is set to 10, the task-level learning rate is set to 1×10 -2 , and the meta-level learning rate is set to 5×10 -4 . The model is updated 5 times with the support set during training and 10 times with the support set during testing. The training is iterated 10000 times in total, and a test is conducted every 50 iterations.
[0102] Remove the network head of the original model, use the method of the prototype network with the Euclidean distance as the distance function for experiments to evaluate the positive impact caused by the task specialization brought by the inner loop of the gradient update of MAML. In the training phase of the prototype network, the learning rate is set to 5×10 -5 . The training is iterated 5000 times in total, and a test is conducted every 50 iterations.
[0103] Remove the network head and the last ReLU layer of the original model, and use the TSEBML algorithm to conduct experiments with Euclidean distance and cosine distance as the distance functions respectively to evaluate the influence of the choice of distance function on the TSEBML algorithm. In the training stage of the TSEBML algorithm, set the task batch size to 10 and the meta-level learning rate to 1×10 -4 . Update the model 5 times with the support set during training and 10 times with the support set during testing. The training is iterated 1000 times in total, and testing is conducted every 10 iterations. The task-level learning rates of the TSEBML algorithm with Euclidean distance and cosine distance as the distance functions are set to 5×10 -2 and 1×10 -2 .
[0104] Step 4-3, Determine the experimental settings;
[0105] Conduct experiments on the TCGA small sample dataset using ten-fold cross-validation. As Figure 5 shown, it shows the cancers and cancer molecular subtypes included in the test set of each fold and the corresponding number of samples. Except for the method that does not use tasks for training, experiments are conducted on each method using two meta-learning training settings: 5-way 1-shot 15-query and 5-way 5-shot 15-query. During testing, keep the shot value of the task unchanged, set the way value to the number of classes of the specific task, and set the query value equal to the shot value. For each molecular subtype-selected cancer in the test set, test the task of its corresponding molecular subtype classification, and for all cancers in the test set, test the tasks of their corresponding cancer classifications. For each classification problem to be tested, 500 tasks will be randomly generated from the classes it contains, and the average value and standard deviation of the evaluation metrics corresponding to these tasks will be calculated.
[0106] Example:
[0107] Under the two meta-learning settings in Step 4-3, input the training set of each fold in the TCGA small sample dataset obtained in Step 1 into the original model described in Step 3, train the model and evaluate it according to the classifier baseline method in Step 4-2 above, and keep the evaluation results ACC, PRE macro , F1 macro and AUC macro of each cancer classification or cancer molecular subtype classification problem described in Step 4-3 when being tested. Take the average value over all tasks as the final result of the classifier baseline method.
[0108] Under the two meta - learning settings in step 4 - 3, first pre - train the original model described in step 3 in each fold using the classifier baseline method in step 4 - 2. Then, input the training tasks of each fold in the TCGA small - sample dataset obtained in step 1 into the original model described in step 3 with the network head removed, train the model and evaluate it according to the meta - baseline method in step 4 - 2 above. Preserve the evaluation results ACC, PRE macro 、F1 macro and AUC macro for each cancer classification or cancer molecular subtype classification problem described in step 4 - 3 when being tested, and take the average value over all tasks as the final result of the meta - baseline method.
[0109] Under the two meta - learning settings in step 4 - 3, input the training tasks of each fold in the TCGA small - sample dataset obtained in step 1 into the original model described in step 3, train the model and evaluate it according to the MAML method in step 2 - 1 above. Preserve the evaluation results ACC, PRE macro 、F1 macro and AUC macro for each cancer classification or cancer molecular subtype classification problem described in step 4 - 3 when being tested, and take the average value over all tasks as the final result of the MAML method.
[0110] Under the two meta - learning settings in step 4 - 3, input the training tasks of each fold in the TCGA small - sample dataset obtained in step 1 into the original model described in step 3 with the network head removed, train the model and evaluate it according to the prototype network method in step 2 - 2 above. Preserve the evaluation results ACC, PRE macro 、F1 macro and AUC macro for each cancer classification or cancer molecular subtype classification problem described in step 4 - 3 when being tested, and take the average value over all tasks as the final result of the prototype network method.
[0111] Under the two meta - learning settings in step 4 - 3, input the training tasks of each fold in the TCGA small - sample dataset obtained in step 1 into the original model described in step 3 with the network head and the last ReLU layer removed, adopt the Euclidean distance as the distance function, train the model and evaluate it according to the TSEBML algorithm method in step 2 - 3 above. Preserve the evaluation results ACC, PRE macro 、F1 macro and AUC macro for each cancer classification or cancer molecular subtype classification problem described in step 4 - 3 when being tested, and take the average value over all tasks as the final result of the TSEBML algorithm method with the Euclidean distance as the distance function.
[0112] Under the two meta - learning settings in step 4 - 3, the training tasks of each fold in the TCGA small - sample dataset obtained in step 1 are input into the original model described in step 3 with the network head and the last ReLU layer removed. Using the cosine distance as the distance function, the model is trained and evaluated according to the TSEBML algorithm method in step 2 - 3 above. The evaluation results ACC, PRE macro , F1 macro , and AUC macro of each cancer classification or cancer molecular subtype classification problem in step 4 - 3 during testing are retained, and the average value over all tasks is taken as the final result of the TSEBML algorithm method with the cosine distance as the distance function.
[0113] In the binary - classification scenario, by comparing and counting the true values and predicted values of the instances, the true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN) are obtained. ACC, PRE, and REC are calculated as follows:
[0114]
[0115]
[0116]
[0117] F1 is further calculated as follows:
[0118]
[0119] And the ROC curve is plotted, and finally AUC is obtained.
[0120] When the number of classes is extended from 2 to N, the macro - average is used to evaluate the performance of the classification model overall from the classification situations of all classes.
[0121]
[0122]
[0123]
[0124]
[0125] Among them, PRE j , REC j , F1 j , and AUC j respectively represent the PRE, REC, F1, and AUC calculated when only the j - th class is regarded as the positive class. Since in the query set of each task, each class has the same number of samples, this results in REC macro = ACC at this time.
[0126] According to the evaluation method in step 4 above, the performance of the method is evaluated by using ten-fold cross-validation. As Figure 6 shown, it is a schematic diagram of the t-SNE visualization result of the embeddings of all instances learned by the model through the TSEBML algorithm in the embodiment. The t-SNE visualization of the embeddings of all instances learned by the model through the TSEBML algorithm, from which it can be seen that the model already has the ability to better complete different tasks. The experimental data of the five methods show that the index of the TSEBML algorithm described in the present invention is the best. For the cancer and cancer molecular subtype classification problems, the highest ACC reaches 70.06% in the 1-shot test setting and 82.33% in the 5-shot test setting. As shown in Table 1 and Table 2, no matter which distance function is used in which test setting, the present invention has a high classification accuracy.
[0127] Table 1 Evaluation results of methods in the 1-shot test setting
[0128]
[0129] Table 2 Evaluation results of methods in the 5-shot test setting
[0130]
[0131] In summary, the present invention can extract shareable knowledge or experience from the tasks of classifying multiple cancer molecular subtypes, and use a very limited number of instances with known labels to quickly specialize the model, enabling it to classify instances with unknown labels in new tasks. It has a significant advantage in solving the small-sample classification problem of cancer molecular subtypes.
[0132] The above is an exemplary description of the present invention. It should be noted that any simple deformation, modification or equivalent replacement that can be made by those skilled in the art without creative labor falls within the protection scope of the present invention without departing from the core of the present invention.
Claims
1. A small-sample classification method for cancer molecular subtypes based on meta-learning, characterized in that, The method comprises the following steps: Step 1: Construct a data set, where the data includes gene expression profiles in the TCGA database, cancer information of samples, and cancer molecular subtype information; Set the scenario for few-shot learning; Design the task generation method for the few-shot data set, that is, appropriately intervene in the generation method of training tasks, and maintain a probability of half to extract molecular subtypes of different cancers to form cancer classification tasks, and a probability of half to extract molecular subtypes of the same cancer to form cancer classification tasks; Step 2: Design a meta-learning algorithm based on task-specialized embedding, and this algorithm includes the following processing: Start by creating an embedding model g for feature extraction φ ; Obtain a task; Calculation support set centroid of use case features; Calculation of the support set and the query set cross-entropy loss; Update the model g using the calculated cross - entropy loss φ Perform gradient updates and temporarily update the model several times; Compute the query set for the predicted vector z q ; Judge whether it is currently in the training stage: Branch 1: If it is in the training phase, continue to calculate the query set loss and accumulate it; judge whether to end the accumulation: if the accumulation ends, update the model permanently; further judge whether to continue the model iteration, if no further iteration is required, the process ends; if the loss of the query set ends, then go to the step of obtaining the next task after the start; Branch 2: If it is not in the training stage, further judge whether to continue model iteration. If model iteration continues, go to the step of obtaining the next task after the start; if iteration does not continue, the algorithm process ends; The optimization objective of the algorithm is: Among them, Update more times with the support set: Among them, φ * is the optimization objective of the algorithm, φ is the model parameter, α is the learning rate at the task level, is the loss function, i is the task number, φ i is the updated model parameter of the i-th task, x q is the query set for few-shot learning in the gene expression profile elements, y q is the query set in the cancer information elements, is the gradient function, x s is the support set for few-shot learning in the gene expression profile elements, y s is the support set in the cancer information elements, is the mathematical expectation, is the i-th task satisfies the task distribution is the entire process of calculating the prediction vector z i when the model parameter is φ q ; Step 3: Construct an original neural network model; Step 4: Train and evaluate the original neural network model, specifically including: The evaluation metrics are determined to be macro-average precision PRE macro , macro-average recall REC macro , macro-average F1-score F1 macro and macro-average AUC macro .
2. The small-sample classification method for cancer molecular subtypes based on meta-learning according to claim 1, wherein, The original neural network model is successively composed of a one-dimensional convolutional layer with a kernel size and a stride of 130 and 32 filter numbers, a ReLU layer, a max pooling layer with a kernel size and a stride of 2, a flattening layer, a hidden layer of size 112, a ReLU layer, and an output layer of size equal to the number of classifications; 10 zeros are added to the end of the RNA sequence data with a length of 20530 to form a single-channel input vector with a length of 20540.
3. The small-sample classification method for cancer molecular subtypes based on meta-learning according to claim 1, wherein The query set 's prediction vector, specifically including: x q The corresponding predicted vector is: z q = [-d(g φ (x q ), c i,1 ), -d(g φ (x q ), c i,2 ),..., -d(g φ (x q ), c i,N )] where d is the cosine distance function, for each use -d(g φ (x q ), c i,j ) as the measure for predicting whether x q is of class y i,j ; Use the function h φ to represent the entire process of calculating z above, i.e.: q
Citation Information
Patent Citations
Trojan horse communication detection method and system combining meta-learning and spatio-temporal feature fusion
CN112929380A
Breast cancer molecular subtype prediction method based on model driver element learning
CN114187472A