Acute myelogenous leukemia subtype typing method based on prototype contrast learning

Through a method based on prototype comparison learning, combined with multimodal data and adaptive optimization strategies, the problem of insufficient accuracy in AML subtype classification is solved, and more accurate subtype classification is achieved, supporting personalized treatment and precision medicine.

CN120277548AActive Publication Date: 2025-07-08ZHEJIANG LAB
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510775583.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-08
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The existing subtype classification methods for acute myeloid leukemia (AML) mainly rely on single-omic data, lacking comprehensive analysis of multiomic data, resulting in insufficient classification accuracy and traditional methods have problems of subjectivity and low resolution.

Method used

Using a method based on prototype comparison learning, the AML patient data set is constructed, feature extraction and fusion is performed, the prototype is initialized using clustering method, and dynamic optimization is learned through prototype comparison. Combining the encoder and feature alignment loss function of multimodal data, the number of subtypes is adaptively determined to achieve accurate classification of subtypes.

Benefits of technology

It improves the accuracy and reliability of AML subtype classification, reduces the risks of misses and overfitting, provides more accurate and clinically valuable subtype classification results, and supports personalized treatment and precision medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277548A_ABST
    Figure CN120277548A_ABST
Patent Text Reader

Abstract

The invention discloses an acute myelogenous leukemia subtype typing method based on prototype comparative learning, which comprises the following steps: extracting and fusing the characteristics of gene expression data, gene mutation data and clinical information of an AML (acute myelogenous leukemia) patient, and taking the fused characteristics as a sample; clustering the samples by adopting a clustering method to obtain an initial prototype; constructing a prototype comparative learning loss function, and carrying out prototype comparative learning; dynamically updating the prototype, redistributing the sample to the nearest prototype, and carrying out iterative optimization; and according to a clustering quality index, adaptively determining a reasonable AML subtype number. According to the method, different types of biomarkers are comprehensively considered, more comprehensive and accurate AML subtype division is provided, fine differences between different subtypes are effectively identified, and the classification precision is improved. The subtype number of the AML is determined in a self-adaptive mode, the limitation of the fixed subtype number in a traditional method is avoided, the model is automatically adjusted according to the actual structure of data, and the subtype division flexibility and robustness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cancer subtype classification, and particularly to a method for classifying subtypes of acute myeloid leukemia based on prototype contrast learning. Background Art

[0002] Acute myeloid leukemia (AML) is a malignant blood disease originating from the bone marrow, with significant differences in its clinical manifestations, treatment responses, and prognoses. The study of AML subtype classification is not only of great significance for deeply understanding the biological mechanism of AML, but also provides a key basis for clinical personalized treatment and prognosis prediction. The heterogeneity of AML is very high, and differences in molecular characteristics, chromosomal abnormalities, gene mutations, and immunophenotypes of different subtypes make the diagnosis and treatment process of AML extremely challenging. Therefore, accurate subtype classification is crucial for improving the survival rate of patients and formulating more precise treatment plans. In recent years, with the rapid development of molecular biology techniques and high-throughput omics techniques, researchers have been able to identify multiple molecular subtypes of AML. These subtypes are not only closely related to the clinical manifestations of patients, but may also determine the response and drug resistance of patients to different treatments. For example, some AML subtypes exhibit specific gene mutations, such as FLT3, NPM1, IDH1 / 2, etc., and these mutations have important predictive value in clinical practice. In addition, the study of AML subtypes helps to identify new therapeutic targets and lays a foundation for the development of targeted therapies.

[0003] With the rapid development of multi-omics techniques, the multi-modal fusion of gene mutations, gene expression, and clinical data provides a new research direction for AML subtype classification. By integrating this multi-level information, the molecular mechanism of AML can be more comprehensively revealed, and the accuracy and clinical application value of subtype classification can be improved. The fusion of multi-omics data can not only reduce the limitations brought by single-omics data, but also reveal the potential associations between different levels of data, improving the accuracy and reliability of AML subtype classification. With the continuous progress of computational methods, especially the development of machine learning and deep learning techniques, the multi-modal fusion of gene mutations, gene expression, and clinical data will play an increasingly important role in the classification of AML.

[0004] Traditional methods for classifying subtypes of acute myeloid leukemia (AML) mainly rely on clinical manifestations, cell morphology, immunophenotype analysis, and chromosome karyotyping. These methods provide a reference for AML classification to a certain extent, but there are still several limitations. Firstly, cell morphology and immunophenotype analysis rely on doctors' experience and are highly subjective, resulting in inconsistent classification results among different laboratories and doctors. Secondly, although chromosome karyotyping can reveal certain chromosomal abnormalities, its resolution is low and it cannot capture subtle gene mutations and molecular-level variations. Therefore, it is difficult to comprehensively reflect the molecular heterogeneity of AML. Most existing AML subtype classifications are based on single omics data and lack comprehensive analysis of different omics data. The complex relationships among multi-omics information such as gene mutations, gene expression, and clinical data have not been fully explored and utilized, leading to significant limitations in the accuracy of existing classification methods. Summary of the Invention

[0005] The object of the present invention is to provide a method for classifying subtypes of acute myeloid leukemia based on prototype contrast learning in view of the deficiencies of the prior art.

[0006] The object of the present invention is achieved by the following technical solutions: A method for classifying subtypes of acute myeloid leukemia based on prototype contrast learning, comprising:

[0007] Firstly, construct an AML patient dataset, extract features from AML patient data, and each group of AML patient data includes gene expression data, gene mutation data, and clinical information;

[0008] Secondly, fuse the gene mutation features, gene expression features, and clinical features of each group of AML patient data, and use the fused features as samples;

[0009] Subsequently, perform prototype initialization: use a clustering method to roughly cluster the samples to obtain initial prototypes;

[0010] Then, construct a prototype contrast learning loss function and perform prototype contrast learning;

[0011] Next, dynamically update the prototypes, reassign the samples to the nearest prototypes, and perform iterative optimization;

[0012] Finally, perform adaptive subtype discovery, and adaptively determine the number of subtypes of acute myeloid leukemia AML according to the clustering quality index.

[0013] Furthermore, construct encoders for gene expression data, gene mutation data, and clinical information modalities respectively and perform pre-training to extract good representations of these three modality data, including:

[0014] Construct an encoder for gene expression data, perform log2(TPM + 1) normalization processing, and screen out variant genes as core features. This encoder uses a Transformer encoder network with a self-attention mechanism to capture the non-linear interactions between genes;

[0015] Construct an encoder for gene mutation data. First, construct a binary matrix, and focus on incorporating mutations with high clinical relevance such as FLT3-ITD, CEBPA, and IDH1 / 2. The structure of this encoder uses a sparse self-attention mechanism;

[0016] Construct an encoder for clinical data. First, perform Z-score normalization on continuous variables and one-hot encoding on categorical variables. The structure of this encoder includes a 3-layer deep autoencoder, including linear layers, batch normalization layers, and activation functions.

[0017] Furthermore, gene mutations, gene expression features, and clinical features are fused, including:

[0018] Perform feature alignment constraints on the extracted features of different modalities, construct a feature alignment loss function, and minimize the feature semantic differences between feature representations of different modalities based on cosine distance:

[0019]

[0020] Among them, and are gene mutations and gene expression features respectively.

[0021] Furthermore, the clustering method is used to roughly cluster the samples to obtain initial prototypes, including:

[0022] Use the K-means clustering method to roughly cluster the samples to obtain initial pseudo-labels , and calculate the prototype vector of each cluster as the initial prototype :

[0023]

[0024] Among them, is the sample set of the current category k, is the center vector of category k.

[0025] Furthermore, the prototype contrast loss function is:

[0026]

[0027] Among them, sim(·) is the cosine similarity, is the i-th sample, and N is the number of samples; is the temperature parameter that controls the sensitivity of the contrastive loss; is the prototype; is the center vector of class k.

[0028] Furthermore, for the dynamic update of the prototype, reassigning the samples to the nearest prototype and iterative optimization are specifically as follows:

[0029] First, calculate the similarity between each sample and all prototypes:

[0030]

[0031] Then, reassign the pseudo-labels, and the sample belongs to the nearest prototype :

[0032]

[0033] Subsequently, update the prototype:

[0034]

[0035] Among them, is the updated prototype, controls the smooth update and avoids over-reliance on the current samples, so that the prototype remains stable during the training process;

[0036] In addition, to ensure the stability of the pseudo-labels of the samples during the iteration, a pseudo-label consistency loss is introduced, and the cross-entropy loss function is used:

[0037]

[0038] Among them, is the soft pseudo-label of sample i on class k, N is the number of samples, K is the number of classes, is the predicted probability that sample i belongs to class k, calculated by Softmax:

[0039]

[0040] Among them, is the sample and the prototype the cosine similarity between them:

[0041]

[0042] Furthermore, for the adaptive subtype discovery, according to the clustering quality index, the reasonable number of subtypes of acute myeloid leukemia (AML) is adaptively determined, specifically as follows:

[0043] First, evaluate the current clustering quality, calculate the Silhouette Score to evaluate the rationality of the classification, and adaptively determine the reasonable number K of AML subtypes, that is, adjust the number of clusters. If the quality of the prototypes is poor, merging / splitting can be performed: merge similar prototypes; refine unstable prototypes; recalculate new pseudo-labels and continue training. Specifically as follows:

[0044] First, calculate the Silhouette Score to evaluate the quality of the current clustering; then, set a threshold to determine whether to adjust the number of prototypes: if the Silhouette Score < 0.5, it is considered that the clustering quality is poor. As an unstable prototype, a new prototype needs to be added; if the Silhouette Score > 0.8, redundant prototypes need to be reduced, calculate the two closest prototypes and merge them.

[0045] In addition, regularization terms are introduced, including: prototype distribution regularization and prototype aggregation and dispersion regularization;

[0046] Among them, prototype distribution regularization prevents all samples from being assigned to a single category:

[0047]

[0048] Among them, is the number of samples in category k, and N is the total number of samples. This is an entropy loss, which encourages the samples to be evenly distributed among different categories.

[0049] Among them, prototype aggregation and dispersion regularization ensures that the prototypes have sufficient distinctiveness from each other:

[0050]

[0051] Finally, the regularization loss is:

[0052] +

[0053] Among them, and are the weight coefficients of the loss function

[0054] To sum up, the total loss function is:

[0055]

[0056] Among them, is the prototype contrast loss function, , and is a learnable parameter.

[0057] The present invention also provides an acute myeloid leukemia subtype classification system based on prototype contrast learning, including:

[0058] A feature extraction module, which constructs a dataset of AML patients and extracts features from the data of AML patients. Each group of AML patient data includes gene expression data, gene mutation data, and clinical information;

[0059] A feature fusion module, which is used to fuse the gene mutation features, gene expression features, and clinical features of each group of AML patient data, and use the fused features as samples;

[0060] A clustering module, which is used for prototype initialization: using a clustering method to roughly cluster the samples to obtain initial prototypes;

[0061] A contrast learning module, which is used to construct a prototype contrast learning loss function and perform prototype contrast learning;

[0062] An optimization module, which is used to dynamically update the prototypes, reassign the samples to the nearest prototypes, and perform iterative optimization;

[0063] A subtype number determination module, which is used for adaptive subtype discovery: adaptively determining the number of subtypes of acute myeloid leukemia AML according to the clustering quality index.

[0064] The present invention also provides an acute myeloid leukemia subtype classification device based on prototype contrast learning, including one or more processors, which are used to implement the above-mentioned acute myeloid leukemia subtype classification method based on prototype contrast learning.

[0065] The present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it is used to implement the above-mentioned acute myeloid leukemia subtype classification method based on prototype contrast learning.

[0066] Compared with the prior art, the beneficial effects of the present invention are as follows: First, multi-modal data fusion can make full use of three different modal data, namely gene expression data, gene mutation data, and clinical information, to capture the complex biological characteristics of AML from multiple levels. Traditional single-modal analysis often fails to fully reveal the heterogeneity of AML. However, this method can effectively integrate the information from different data sources by constructing a multi-modal fusion module, avoiding data redundancy and information loss, thereby improving the accuracy of classification. Second, prototype contrast learning can roughly cluster samples and initialize prototypes, and then continuously optimize the quality of prototypes through contrast learning. This process provides reasonable inspiration for the initial prototypes with the help of clustering methods, and then dynamically updates the samples through the prototype contrast learning loss function to ensure that each sample can finally be accurately assigned to the most appropriate subtype. The dynamic update and contrast learning mechanism of prototypes can gradually optimize the model and reduce the risks of misclassification and overfitting. Third, adaptive subtype discovery adaptively determines the number of AML subtypes through the evaluation of clustering quality indicators. Compared with the traditional method of artificially setting a fixed number of subtypes, this adaptive strategy can be flexibly adjusted according to the actual distribution of data, thereby avoiding the bias caused by artificially setting the number of subtypes and ensuring the accuracy and reliability of subtype classification. Generally speaking, this subtype classification method based on prototype contrast learning can explore the potential laws of AML from multiple perspectives and multi-modal data, and provide more accurate and clinically valuable subtype classification results through an efficient adaptive optimization strategy. This not only provides strong support for the personalized treatment of AML, but also lays a solid foundation for subsequent clinical research and precision medicine. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0068] Figure 1 Schematic flowchart of a method for classifying subtypes of acute myeloid leukemia based on prototype contrast learning provided by an embodiment of the present invention;

[0069] Figure 2 Structural diagrams of different modal encoders provided by an embodiment of the present invention;

[0070] Figure 3 Overall framework diagram of the model provided by an embodiment of the present invention;

[0071] Figure 4 Prototype classification flowchart provided by an embodiment of the present invention;

[0072] Figure 5 A hardware structure diagram provided by an embodiment of the present invention. Specific implementation manners

[0073] The present invention will be described in detail below with reference to the accompanying drawings. Without conflict, the features in the following embodiments and implementation manners can be combined with each other.

[0074] A method for classifying subtypes of acute myeloid leukemia based on prototype contrast learning according to the present invention, as Figure 1 shown, includes the following steps:

[0075] (1) First, encoders are respectively constructed and pre-trained for three modalities of data, namely gene expression data, gene mutation data, and clinical information, to extract good representations of the three modalities of data, as Figure 2 shown, including:

[0076] Construct an encoder for gene expression data , perform log2(TPM + 1) normalization processing, and screen out highly variable genes as core features , the encoder adopts a Transformer encoder network with a self-attention mechanism to capture the non-linear interaction between genes; this encoder includes multi-head attention, the first layer of normalization, a feed-forward network, and the second layer of normalization, which are connected in sequence; among them, the input of the first layer of normalization is the residual combination of the original input and the output of multi-head attention, and the input of the second layer of normalization is the residual combination of the output of the first normalization and the output of the feed-forward network;

[0077] For gene mutation data construct an encoder , first construct a binary matrix (presence / absence of mutation), and focus on incorporating highly clinically relevant mutations such as FLT3-ITD, CEBPA, IDH1 / 2, etc. The encoder structure mainly adopts a sparse self-attention mechanism; this encoder includes a sparse embedding layer, the first layer of normalization, self-attention, the second layer of normalization, a feed-forward network, and the third layer of normalization, which are connected in sequence; among them, the input of the second layer of normalization is the residual combination of the output of the first normalization and the output of self-attention, and the input of the third layer of normalization is the residual combination of the output of the second normalization and the output of the feed-forward network;

[0078] For clinical data construct an encoder , first perform Z-score normalization on continuous variables and one-hot encoding on categorical variables. The encoder The structure mainly includes 3 layers of deep autoencoders, including a linear layer, a batch normalization layer and an activation function; the encoder includes a first linear layer, a first activation layer (such as a ReLU activation function), a first batch normalization, a second linear layer, a second activation layer (such as a ReLU activation function), a second batch normalization and a third linear layer.

[0079] Considering that the existing AML datasets usually contain fewer AML patient samples, pre-training can help obtain more effective feature representation. Therefore, pre-training is performed on the above encoders respectively. The pre-training task is to build an encoder for each modality data to reconstruct the data. The modality encoder mainly includes multiple linear layers and activation functions. Finally, the features of different omics data are obtained through the pre-trained encoders. , and .

[0080] (2) Secondly, a multimodal fusion module is designed and constructed to effectively fuse three modal information, including:

[0081] The multimodal fusion module mainly adopts the cross-modal attention mechanism to fuse the information of different modalities and enhance the interaction between modalities. Specifically, the multimodal fusion module includes feature projection layer, multi-head cross-attention, residual normalization and weighted splicing. In order to obtain a better fusion effect, the features of different modalities are firstly subjected to feature alignment constraints, and a feature alignment loss function is constructed to minimize the feature semantic differences between the feature representations of different modalities based on the cosine distance:

[0082]

[0083] in, and They are gene mutation, gene expression characteristics and clinical characteristics.

[0084] (3) Then, the prototype is initialized and the samples are roughly clustered using a clustering method to obtain the initial prototype, including:

[0085] Use K-means clustering method to roughly cluster the samples and obtain the initial pseudo labels , calculate the prototype vector of each cluster as the initial prototype :

[0086]

[0087] in, is the sample set of the current category k, is the center vector of category k, is the fused feature, i.e., the sample.

[0088] (4) Then, construct the prototype contrast loss function and conduct prototype contrast learning, as Figure 3 shown, including:

[0089] The purpose of constructing the prototype contrast loss function is to make the samples close to their own prototypes , and far from the prototypes of other classes:

[0090]

[0091] Among them, is the prototype vector corresponding to sample i for the pseudo-labeled class , is the center vector of all samples in class k. The former is the prototype of the class corresponding to the sample, and the latter is the center of the global class; sim(·) is the cosine similarity, is the temperature parameter, which controls the sensitivity of the contrast loss.

[0092] (5) Subsequently, update the prototypes dynamically, reassign the samples to the nearest prototypes, and the iterative optimization is specifically as follows:

[0093] First, calculate the similarity of each sample to all prototypes:

[0094]

[0095] Then, reassign the pseudo-labels, and sample belongs to the nearest prototype :

[0096]

[0097] Subsequently, update the prototypes:

[0098]

[0099] Among them, controls the smooth update, avoids over-reliance on the current samples, so that the prototypes remain stable during training, and prevents the model from being unstable or the prototype vectors from fluctuating excessively.

[0100] In addition, in order to ensure the stability of the pseudo-labels of the samples during the iteration process and avoid drastic changes, a pseudo-label consistency loss is introduced, and the cross-entropy loss function is adopted:

[0101]

[0102] Among them, is the soft pseudo-label of sample i for class k, is the predicted probability that sample i belongs to class k, usually calculated by Softmax:

[0103]

[0104] Among them, is the sample and the prototype The cosine similarity between them is:

[0105]

[0106] (6) Finally, for adaptive subtype discovery, according to the clustering quality index, the number of subtypes of acute myeloid leukemia (AML) is adaptively determined to construct a prototype library, specifically:

[0107] First, evaluate the current clustering quality, calculate the Silhouette Score to evaluate the rationality of the classification, and adaptively determine the reasonable number of AML subtypes K, that is, adjust the number of clusters. If the quality of the prototype is poor, merging / splitting can be performed: merge similar prototypes; refine unstable prototypes; recalculate new pseudo-labels and continue training.

[0108] In addition, a regularization term is introduced to prevent the model from collapsing, such as all samples being assigned to the same category or the features being overly concentrated. Among them, the prototype distribution regularization prevents all samples from being assigned to a single category:

[0109]

[0110] Among them, is the number of samples in category k, and N is the total number of samples. This is an entropy loss that encourages the samples to be evenly distributed among different categories.

[0111] Among them, the prototype dispersion regularization ensures that the prototypes have sufficient distinctiveness from each other:

[0112]

[0113] Finally, the regularization loss is:

[0114] +

[0115] Among them, and are the weight coefficients of the loss function, set as learnable hyperparameters, and are optimized during model training.

[0116] To sum up, the total loss function is:

[0117]

[0118] Among them, , and are learnable parameters.

[0119] Finally, after optimizing the model training through the loss function, the final subtypes are obtained. As Figure 4 shown, when there are new patient samples, the fusion features of the new patient samples are calculated through the trained model, and then the similarity is calculated with the current prototypes (prototypes in the prototype library) respectively, and the subtype prototype with the highest similarity is selected as the subtype of the new patient sample.

[0120] The present invention also provides an acute myeloid leukemia subtype classification system based on prototype contrast learning, including:

[0121] A feature extraction module that constructs a dataset of AML patients and extracts features from the AML patient data. Each group of AML patient data includes gene expression data, gene mutation data, and clinical information;

[0122] A feature fusion module for fusing the gene mutation features, gene expression features, and clinical features of each group of AML patient data, and using the fused features as samples;

[0123] A clustering module for initializing prototypes: roughly clustering the samples using a clustering method to obtain initial prototypes;

[0124] A contrast learning module for constructing a prototype contrast learning loss function and performing prototype contrast learning;

[0125] An optimization module for dynamically updating prototypes, reassigning samples to the nearest prototypes, and iteratively optimizing;

[0126] A subtype quantity determination module for adaptive subtype discovery: adaptively determining the subtype quantity of acute myeloid leukemia AML according to the clustering quality index.

[0127] It should be noted that the system embodiment shown in this embodiment matches the content of the above method embodiment, and the content of the above method embodiment can be referred to and will not be elaborated here.

[0128] Corresponding to the foregoing embodiment of an acute myeloid leukemia subtype classification method based on prototype contrast learning, the present invention also provides an embodiment of an acute myeloid leukemia subtype classification device based on prototype contrast learning.

[0129] See Figure 5 , an acute myeloid leukemia subtype classification device provided by an embodiment of the present invention based on prototype contrast learning includes one or more processors for implementing an acute myeloid leukemia subtype classification method based on prototype contrast learning in the above embodiment.

[0130] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0131] An embodiment of the acute myeloid leukemia subtype classification device based on prototype contrast learning of the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or by hardware, or by a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. At the hardware level, as Figure 5 shown, it is a hardware structure diagram of any device with data processing capabilities where the acute myeloid leukemia subtype classification device based on prototype contrast learning of the present invention is located. Except for Figure 5 the processor, memory, network interface, and non-volatile memory shown, generally, according to the actual functions of the any device with data processing capabilities where the device in the embodiment is located, other hardware may also be included, which will not be elaborated here.

[0132] For the realization process of the functions and roles of each unit in the above device, please refer to the realization process of the corresponding steps in the above method for details, which will not be elaborated here.

[0133] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0134] An embodiment of the present invention further provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a method for classifying subtypes of acute myeloid leukemia based on prototype contrast learning in the above embodiment.

[0135] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.

[0136] The above embodiments are only used to illustrate the design concept and characteristics of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design concepts disclosed by the present invention are within the protection scope of the present invention.

Claims

1. A method for classifying subtypes of acute myeloid leukemia based on prototype contrast learning, characterized in that, Including: First, construct a dataset of AML patients, extract features from the AML patient data, and each group of AML patient data includes gene expression data, gene mutation data, and clinical information; Secondly, fuse the gene mutation features, gene expression features, and clinical features of each group of AML patient data, and use the fused features as samples; Subsequently, perform prototype initialization: use a clustering method to roughly cluster the samples to obtain initial prototypes; Then, construct a prototype contrast learning loss function and perform prototype contrast learning; Next, dynamically update the prototypes, reassign the samples to the nearest prototypes, and iteratively optimize; Finally, perform adaptive subtype discovery, and adaptively determine the number of subtypes of acute myeloid leukemia AML according to the clustering quality index.

2. The acute myeloid leukemia subtype classification method based on prototype contrast learning according to claim 1, characterized in that, Construct encoders for gene expression data, gene mutation data, and clinical information modalities respectively and perform pre-training to extract good representations of these three modalities of data, including: Construct an encoder for gene expression data, perform log2(TPM + 1) normalization processing, and screen out mutant genes as core features. This encoder uses a Transformer encoder network with a self-attention mechanism to capture non-linear interactions between genes; Construct an encoder for gene mutation data. First, construct a binary matrix, and focus on incorporating highly clinically relevant mutations such as FLT3-ITD, CEBPA, and IDH1 / 2. The structure of this encoder uses a sparse self-attention mechanism; Construct an encoder for clinical data. First, perform Z-score normalization on continuous variables and one-hot encoding on categorical variables. The structure of this encoder includes a 3-layer deep autoencoder, including a linear layer, a batch normalization layer, and an activation function.

3. The subtype classification method of acute myeloid leukemia based on prototype contrast learning according to claim 1, wherein The fusion of gene mutations, gene expression features, and clinical features includes: Perform feature alignment constraints on the extracted features of different modalities, construct a feature alignment loss function, and minimize the feature semantic differences between different modality feature representations based on cosine distance: Among them, and are gene mutation and gene expression characteristics respectively.

4. The method for classifying subtypes of acute myeloid leukemia based on prototype contrast learning according to claim 1, characterized in that, The use of the clustering method to roughly cluster the samples to obtain initial prototypes includes: Use the K-means clustering method for the samples to perform rough clustering and obtain initial pseudo-labels , and calculate the prototype vector of each cluster as the initial prototype : Among them, is the sample set of the current category k, is the central vector of category k.

5. The acute myeloid leukemia subtype classification method based on prototype contrast learning according to claim 1, characterized in that The prototype contrast loss function is as follows: where sim(·) is the cosine similarity, is the i-th sample, and N is the number of samples; is the temperature parameter that controls the sensitivity of the contrastive loss; is the prototype; is the center vector of class k.

6. The acute myeloid leukemia subtype classification method based on prototype contrast learning according to claim 3, wherein The dynamic update of the prototypes, reassigning the samples to the nearest prototypes, and iterative optimization is specifically: First, calculate the similarity of each sample to all prototypes: Then, reassign the pseudo-labels, and the samples are assigned to the nearest prototype : Subsequently, update the prototypes: Among them, is the updated prototype, controls the smooth update, avoids over-reliance on the current samples, and thus keeps the prototype stable during the training process; In addition, to ensure the stability of the pseudo-labels of the samples during the iteration process, a pseudo-label consistency loss is introduced, and the cross-entropy loss function is adopted : Among them, is the soft pseudo-label of sample i on class k, N is the number of samples, and K is the number of classes. is the predicted probability that sample i belongs to class k, calculated by Softmax: Among them, is the sample and the prototype the cosine similarity between: 。 7. A method for classifying subtypes of acute myeloid leukemia based on prototype contrast learning according to claim 6, characterized in that, The adaptive subtype discovery, adaptively determining a reasonable number of subtypes of acute myeloid leukemia AML according to the clustering quality index, is specifically: First, evaluate the current clustering quality, calculate the Silhouette Score to evaluate the rationality of the classification, and adaptively determine a reasonable number of AML subtypes K, that is, adjust the number of clusters. If the prototype quality is poor, merging / splitting can be performed: merge similar prototypes; refine unstable prototypes; recalculate new pseudo-labels and continue training; specifically as follows: First, calculate the Silhouette Score to evaluate the quality of the current clustering; then, set a threshold to determine whether to adjust the number of prototypes: if the Silhouette Score < 0.5, it is considered that the clustering quality is poor, and as an unstable prototype, a new prototype needs to be added to the data; if the Silhouette Score > 0.8, redundant prototypes need to be reduced, and the two closest prototypes are calculated and merged; In addition, regularization terms are introduced, including: prototype distribution regularization and prototype dispersion regularization; Among them, prototype distribution regularization prevents all samples from being assigned to a single category: wherein, is the number of samples of class k, and N is the total number of samples; Among them, prototype dispersion regularization ensures that the prototypes have sufficient distinctiveness from each other: Finally, the regularization loss is as follows: + Among them, and are the weight coefficients of the loss function; In summary, the overall loss function is as follows: Among them, is the prototype contrast loss function, , and are learnable parameters.

8. An acute myeloid leukemia subtype classification system based on prototype contrast learning, characterized in that, Including: A feature extraction module that constructs an AML patient dataset and extracts features from AML patient data. Each group of AML patient data includes gene expression data, gene mutation data, and clinical information; A feature fusion module that is used to fuse the gene mutation features, gene expression features, and clinical features of each group of AML patient data, and uses the fused features as samples; A clustering module that is used for prototype initialization: a rough clustering of the samples is performed using a clustering method to obtain initial prototypes; A contrastive learning module that is used to construct a prototype contrastive learning loss function and perform prototype contrastive learning; An optimization module that is used to dynamically update the prototypes, reassign the samples to the closest prototypes, and perform iterative optimization; A subtype number determination module that is used for adaptive subtype discovery: adaptively determine the number of subtypes of acute myeloid leukemia AML according to the clustering quality index.

9. An acute myeloid leukemia subtype classification device based on prototype contrast learning, characterized in that, Including one or more processors for implementing a method for classifying subtypes of acute myeloid leukemia based on prototype contrastive learning according to any one of claims 1-7.

10. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it is used to implement a method for classifying subtypes of acute myeloid leukemia based on prototype contrastive learning according to any one of claims 1-7.

Citation Information

Patent Citations

  • Multi-omics data classification method based on global-local graph Transform

    CN118228115A

  • Cancer category detection method based on multi-omics data clustering remarking

    CN119673295A

  • Multi-modal clustering method and device, equipment, storage medium and product

    CN120123792A

  • Classification, Diagnosis and Prognosis of Acute Myeloid Leukemia by Gene Expression Profiling

    US20080305965A1

  • Classification and risk-assignment of childhood acute myeloid leukaemia (AML) by gene expression signatures

    WO2010143941A1