A method for analyzing cancer multi-omics data based on a multi-head attention mechanism

CN116580848BActive Publication Date: 2026-09-25HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310538812.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-15
Publication Date
2026-09-25
Estimated Expiration
2043-05-15

AI Technical Summary

Technical Problem

然而,高通量测序技术获取的数据量巨大,样本数量较少,数据之间存在着大量的噪声且平台之间存在着较大差异

Benefits of technology

[0054]1.本发明中监督的多头注意力机制模型(SMA)在模拟单细胞和癌症多组学数据集上获得100%准确的亚型分类。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116580848B_ABST
    Figure CN116580848B_ABST
Patent Text Reader

Abstract

The application discloses a method for analyzing cancer multi-omics data based on a multi-head attention mechanism, which comprises the following steps: S1, collecting and preprocessing cancer multi-omics data; S2, using a supervised multi-head attention model to complete a classification task of the cancer multi-omics data; and S3, using a decoupling contrast learning model based on the multi-head attention mechanism to complete a clustering task of the cancer multi-omics data. The application can obtain good effects in the classification task and the clustering task, and can analyze the pathogenesis of cancer together with clinical information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and bioinformatics, and more specifically, to a method for analyzing multi-omics data of cancer based on a multi-head attention mechanism. Background Technology

[0002] With the development of high-throughput sequencing technology, the era of precision medicine has arrived. A large amount of biomedical data is growing explosively and being collected and organized in public databases. For example, the large-scale effort of Cancer Genome Atlas (TCGA) has accumulated genomic, transcriptomic, proteomic and clinical data of more than 20 cancers from thousands of patients [1]. Rich data can help researchers understand the heterogeneity of captured biological processes and phenotypes from different perspectives. However, the amount of data obtained by high-throughput sequencing technology is huge, the number of samples is small, there is a lot of noise between the data and there are large differences between platforms. Therefore, there is a huge challenge in extracting valuable information from high-throughput data.

[0003] Single-omics research theoretically allows for efficient and precise analysis of research subjects. Currently, single-omics has become an important research tool in the life sciences and is widely used in genomics and proteomics. With the development of research, in order to gain a deeper understanding of the interactions and regulatory mechanisms between molecules in organisms, multi-omics analysis integrates genomics, epigenomics, transcriptomics, and proteomics—metabolomics systems—in an unbiased manner to analyze the mechanisms and phenotypes of living systems. Currently, multi-omics data has become a research hotspot in many fields, such as cancer research, drug development, agriculture, and the environment. Software and tools for multi-omics data analysis have also emerged, such as R packages like limma, DESeq2, and edgeR, as well as software like MetaboAnalyst and Proteome Discoverer. Furthermore, some researchers have developed various methods for processing multi-omics data, such as multi-kernel learning methods, Bayesian consensus clustering, and machine learning-based dimensionality reduction.

[0004] Deep learning algorithms have recently been widely applied in research on multi-omics data. Researchers have proposed 16 representative deep learning methods for classifying and clustering multi-omics datasets, including Fully Connected Neural Networks (FCNN), Convolutional Neural Networks (CNN), Graph Neural Networks (GCN), Autoencoders (AE), Capsule Networks (CapsNet), and Generative Adversarial Networks (GAN). Some researchers have proposed an end-to-end multimodal deep learning model (scMDC) to represent different data sources and jointly learn latent features of deep embeddings for clustering analysis. Other researchers have proposed a unified multi-task deep learning framework for multi-omics data (OmiEmbedded) that supports dimensionality reduction, multi-omics integration, tumor type classification, phenotypic feature reconstruction, and survival prediction. Still others have proposed a scalable and interpretable multi-omics deep learning framework (DeepOmix) for cancer survival analysis. It is used to extract relationships between clinical survival time and multi-omics data based on a deep learning framework to predict prognosis. Some researchers have proposed neural network methods based on multi-input multi-output (MIMO) deep adversarial learning to accurately model complex data and use consensus clustering and Gaussian mixture models to identify molecular subtypes of tumor samples. Other researchers have proposed using domain component analysis (NCA) algorithms to select relevant features from multi-omics datasets retrieved from the TCGA and Cancer Drug Sensitivity Genomics (GDSC) databases and to develop survival and prediction models. Furthermore, several deep learning and machine learning methods have been applied to the diagnosis and prognosis of tumor subtypes. Summary of the Invention

[0005] The purpose of this invention is to provide a method for analyzing cancer multi-omics data based on a multi-head attention mechanism, so as to overcome the shortcomings of the existing technology.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0007] A method for analyzing multi-omics data in cancer based on a multi-head attention mechanism includes the following steps:

[0008] S1. Data collection and preprocessing of cancer multi-omics data;

[0009] S2. Use a supervised multi-head attention model to complete the classification task of cancer multi-omics data;

[0010] S3. A decoupled contrastive learning model based on a multi-head attention mechanism is used to learn and complete the clustering task of cancer multi-omics data.

[0011] Further, step S1 specifically includes:

[0012] S11. Normalize the cancer multi-omics data and unify the dimensions of different data.

[0013] S12. Merge and integrate data features from different omics, shuffle the order of sample data, add noise to the samples, and generate training data.

[0014] Furthermore, the step of generating the supervised multi-head attention model in step S2 is as follows:

[0015] S21. Design a multi-head attention encoder;

[0016] S22. Create a symmetrical multi-head attention encoder based on a multi-head attention encoder;

[0017] S23. Create a supervised multi-head attention model based on a symmetric multi-head attention encoder.

[0018] Further, step S21 includes:

[0019] S211. Perform positional encoding on cancer multi-omics data to preserve the relationships between different positions in the sequence;

[0020] S212. Feature extraction is performed using a symmetrical multi-head attention mechanism.

[0021] S213. Employ a multi-head attention mechanism to perform separate computations on the features of multi-omics data;

[0022] S214. Perform multiple self-attention processes on the original input sequence, then concatenate the results of each attention process and perform a linear transformation to obtain the final output result.

[0023] Further, step S22 specifically involves: feature sharing of multi-omics data is achieved by sharing a weight matrix. Feature extraction is performed in a symmetric multi-head self-attention encoder, and the learned weight features are shared in the feature mapping. In backpropagation, since the weight matrix is ​​shared, the symmetric multi-head attention encoder updates the weight gradient using the same value. A symmetric multi-head attention encoder is obtained by connecting two independent multi-head attention encoders in parallel.

[0024] Further, step S23 specifically includes:

[0025] S231. A symmetric multi-head attention mechanism encoder is used to extract features from multi-omics data and generate feature matrices W1 and W2.

[0026] S232. Using the feature fusion method of element-wise multiplication, the features of feature matrices W1 and W2 are multiplied element-wise to obtain a fused feature vector.

[0027] S233. The fused feature vector is fed into a three-layer perceptron for normalization and projected onto a new feature space to generate a new feature matrix. The cross-entropy loss function is used to calculate the error between a single predicted sample and the label;

[0028] S234, Calculating the characteristic matrix The total loss function L is obtained by calculating the distance between the label and the target label.

[0029] Furthermore, the formula for calculating the error between a single predicted sample and the label using the cross-entropy loss function in step S233 is as follows:

[0030]

[0031] The formula for calculating the total loss function L in step S234 is as follows:

[0032]

[0033] Furthermore, step S3 specifically includes:

[0034] S31. Project the image from step S233 onto the new feature space. middle There are a total of n-1 pairs of positive samples used for training. There are n-1 pairs of negative samples used for training. The similarity between pairs of samples is measured by cosine distance:

[0035]

[0036] In the formula, i, j ∈ [1, N], in order to calculate and For each view's error, create a cross-entropy loss function. The loss function between positive and negative samples is then:

[0037]

[0038] In the formula, k∈[1,2], and τ is the temperature parameter that controls the softness in the model;

[0039] S32. Decoupling contrastive learning is achieved by removing positive pairs from the denominator. The process is as follows:

[0040]

[0041] S33, Enhanced by calculating all data To obtain the cross-entropy loss of decoupled contrastive learning, allowing the model to identify all positive samples in the dataset, the process is as follows:

[0042]

[0043] S34, Feature Matrix and Cosine similarity is also used to calculate the error between a pair of samples, and the process is as follows:

[0044]

[0045] In the formula, i, j ∈ [1, M], in order to calculate and The error of each view is used to create a clustering loss function. The loss function between each pair of positive and negative samples can be expressed as:

[0046]

[0047] S35. Through learning from all positive and negative sample pairs, the total loss function is expressed as:

[0048]

[0049]

[0050] In the formula, It is the entropy of the probability assignment for subtype clustering, and outputs most of the label features after each loss calculation.

[0051] Furthermore, the feature space The features are clustered using a decoupled contrastive loss function to output cluster labels. and The clustering loss function is used to calculate the clustering operation on the samples to output the clustering features. In the clustering task, the total loss function is:

[0052] L = L D +L C .

[0053] Compared with the prior art, the advantages of the present invention are as follows:

[0054] 1. The supervised multi-head attention mechanism model (SMA) in this invention achieves 100% accurate subtype classification on simulated single-cell and cancer multi-omics datasets.

[0055] 2. This invention uses a Decoupled Comparative Learning Model (DMACL) to learn multi-omics data features and cluster them to identify cancer subtypes. This unsupervised comparative learning method performs subtype analysis by calculating the similarity between multi-omics data samples. The DMACL model shows significant advantages compared to 16 deep learning models.

[0056] 3. This invention achieves good results in both classification and clustering tasks, and can be used in conjunction with the analysis of clinical information to understand the pathogenesis of cancer. Attached Figure Description

[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0058] Figure 1 This invention relates to a multi-head self-attention encoder framework.

[0059] Figure 2 This represents the performance of seven supervised methods on the cancer benchmark dataset used in the classification task of this invention.

[0060] Figure 3 These are the C-index, silhouette score, and Davies Bouldin score of the 11 unsupervised methods used in this invention on a single-cell multi-omics dataset. Based on cluster analysis of the single-cell dataset, three internal indices, C-index, silhouette score, and Davies Bouldin score (a, b, c), were calculated. The number of clusters was 3, and k-means clustering was run more than 1000 times.

[0061] Figure 4 It is the C-index of 11 unsupervised methods on a cancer benchmark dataset used in this invention for clustering tasks.

[0062] Figure 5 These are the Davies Bouldin scores on the cancer benchmark dataset used in the clustering task by the 11 unsupervised methods in this invention.

[0063] Figure 6 This is the Davies-Bouldin score of 11 unsupervised methods on the cancer benchmark dataset used in the clustering task in this invention. Detailed Implementation

[0064] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby providing a clearer and more explicit definition of the scope of protection of the present invention.

[0065] See Figure 1 and Figure 2As shown, this embodiment discloses a method for analyzing cancer multi-omics data based on a multi-head attention mechanism, including the following steps:

[0066] Step S1: Collect and preprocess cancer multi-omics data.

[0067] Specifically, step S1 includes the following steps:

[0068] Step S11: Normalize the cancer multi-omics data and unify the dimensions of different data to facilitate subsequent feature extraction.

[0069] Step S12: Merge and integrate data features from different omics systems. The purpose of integration is to improve data coverage, increase data information content, and enhance data interpretability. Then, shuffle the order of the sample data, add noise to the samples, and generate training data.

[0070] Step S2: Use a supervised multi-head attention model (SMA) to complete the classification task of cancer multi-omics data.

[0071] Specifically, step S2 includes the following steps:

[0072] Step S21: Design a multi-head attention encoder.

[0073] Step S21 includes:

[0074] Step S211: Since the multi-omics data needs to be fed into the SLUCL framework for feature extraction and data dimensionality reduction, the model does not process the multi-omics data in the order of arrangement. Therefore, it is necessary to perform position encoding on the cancer multi-omics data input to the model to preserve the relationship between each position in the sequence.

[0075] Step S212: Through the linear transformation of positional encoding, the attention mechanism module can better capture the relationship between different positions in the sequence data, thereby improving the model's performance. In the feature extraction part, this embodiment uses a symmetrical multi-head attention mechanism for feature extraction. This embodiment uses a fully connected layer to implement the linear transformation of the input, and its process can be represented as follows:

[0076] y pe =x pe W+b

[0077] In the formula, x pe Let W represent the encoded vector at each position, W be the feature weights of the data, and b be the bias vector for each feature weight. Through a linear transformation of the positional encoding, the attention mechanism module can better capture the relationships between different positions in the sequence data, thereby improving the model's performance.

[0078] In the feature extraction section, this embodiment uses a symmetric multi-head attention mechanism for feature extraction. The tensor matrix after cancating the multi-omics data is W. mn The multi-head attention mechanism will W mn The features are calculated separately. In this paper, W is... mn The feature vector is divided into several smaller feature vectors along the last dimension. Each smaller feature vector is called a head, and the number of heads, h, in the multi-head attention mechanism is set to 80. For each head, a dot product attention mechanism is used to calculate its attention weights with respect to the others. Its self-attention output vector is:

[0079]

[0080] Where Q and K represent the feature matrices output by the same head, and V represents the feature matrix obtained by another head. This refers to the dimension of the matrix, which can be used to reduce the dimension of the feature matrix product. Multi-head attention mechanisms essentially perform multiple self-attention processes on the original input sequence; then, the results of each attention group are concatenated and subjected to a linear transformation to obtain the final output. The computational process can be represented as:

[0081] MultiHead(Q,K,V)=Concat(head1,…head n W o

[0082]

[0083]

[0084] Step S213: Employ a multi-head attention mechanism to perform head-by-head computation on the multi-omics data features. In this embodiment, the features are divided into several small feature vectors along the last dimension, each small feature vector being called a head. The number of heads, h, in the multi-head attention mechanism is set to 80. For each head, a dot product attention mechanism is used to calculate its attention weights for the others. The multi-head attention mechanism constructs attention layers based on the size of h. During the forward pass, the feature matrix is ​​fed into the input layer of the feedforward module. Each neuron in the input layer corresponds to a column of the feature matrix, i.e., a feature. Each neuron weights its input and adds a bias, then calculates the output through an activation function, and passes the output to the next layer of neurons. Finally, the output layer outputs the feature matrix.

[0085] Step S214: The multi-head attention mechanism actually performs multiple self-attention processes on the original input sequence, and then concatenates the results of each attention group together to perform a linear transformation to obtain the final output result.

[0086] Step S22: Create a symmetric multi-head attention encoder based on the multi-head attention encoder. Specifically:

[0087] Feature sharing across multiple omics datasets can typically be achieved by sharing weight matrices. Since the learned weight features are identical when feature extraction is performed in a symmetric multi-head self-attention encoder, weights can be shared in the feature map. Furthermore, during backpropagation, because the weight matrix is ​​shared, the symmetric multi-head attention encoder can update the weight gradients using the same values.

[0088] Step S23: Create a supervised multi-head attention model based on a symmetric multi-head attention encoder. This includes the following steps:

[0089] Step S231: By classifying multi-omics data, interactions and relationships between different types of data can be discovered, and key components and pathways in biological systems can be identified. This comprehensive analysis can provide important clues and insights for studying the pathogenesis of complex diseases and finding new therapeutic targets. The three datasets provided in the experiment already contain labels for all samples. Therefore, this experiment proposes using a supervised multi-head attention mechanism (SMA) model to classify cancer types. This experiment uses a symmetric multi-head attention mechanism encoder to extract features from the multi-omics data, generating feature matrices W1 and W2.

[0090] Step S232: The feature fusion method of element-wise multiplication is adopted. The features of feature matrices W1 and W2 are multiplied element by element to obtain a fused feature vector. This method can highlight the unique features of the encoder to improve the performance and generalization ability of the model.

[0091] Step S233: The fused feature vector is fed into a three-layer perceptron (MLP) for normalization, projected onto a new feature space, and a new feature matrix is ​​generated. The error between a single predicted sample and its label is calculated using the cross-entropy loss function, as follows:

[0092]

[0093] Step S234: Calculate the characteristic matrix The total loss function L is obtained by calculating the distance between the label and the target label.

[0094]

[0095] Step S3: Use a decoupled contrastive learning model (DMACL) based on a multi-head attention mechanism to learn and complete the clustering task of cancer multi-omics data.

[0096] The goal of cancer subtype clustering is to group similar cancer samples into the same subtype and minimize the differences between different subtypes, in order to better understand the biological characteristics and molecular mechanisms of cancer and provide patients with better diagnosis, treatment, and prognosis. Unsupervised, decoupled contrastive learning can greatly improve the similarity of matching. In cancer subtyping tasks, since the experiment does not provide usable labels, both positive and negative samples are composed of pseudo-labels generated by data augmentation.

[0097] S31. Project the image from step S233 onto the new feature space. middle There are a total of n-1 pairs of positive samples used for training. There are n-1 pairs of negative samples used for training. The similarity between pairs of samples is measured by cosine distance:

[0098]

[0099] In the formula, i, j ∈ [1, N], in order to calculate and For each view's error, create a cross-entropy loss function. The loss function between positive and negative samples is then:

[0100]

[0101] In the formula, k∈[1,2], and τ is the temperature parameter controlling the softness in the model. Generally, the negative-positive coupling (NPC) multiplier in cross-entropy loss (InfoNCE) often affects the model training results, leading to the following two situations: First, positive samples near the anchor point are considered important information because they are the only positive samples we have. Simultaneously, the gradient of negative samples gradually decreases. Second, when negative samples are far away and have little information, the model may incorrectly reduce the learning rate from positive samples. This means the model will emphasize negative samples more than considering the information from both positive and negative samples in a balanced way. This may cause the model to make errors when processing positive samples, thus reducing the model's accuracy.

[0102] Step S32: Decoupled contrastive learning is achieved by removing positive pairs from the denominator using a decoupled contrastive learning method. The process is as follows:

[0103]

[0104] Step S33: Calculate the augmented data after processing all data. To obtain the cross-entropy loss of decoupled contrastive learning, allowing the model to identify all positive samples in the dataset, the process is as follows:

[0105]

[0106] The concept of "label as representation" is most common in contrastive clustering. The basic idea of ​​this method is to encode labels as feature vectors and input them, along with the feature vectors of the data points, into the clustering model for training. By embedding labels into the feature space, the clustering problem can be transformed into a contrastive learning problem. That is, data points within the same cluster should be closer in the feature space, while data points in different clusters should be farther apart. This allows the clusters to be determined by comparing the similarity between data points.

[0107] Step S34, Feature Matrix and Cosine similarity is also used to calculate the error between a pair of samples, and the process is as follows:

[0108]

[0109] In the formula, i, j ∈ [1, M], in order to calculate and The error of each view is used to create a clustering loss function. The loss function between each pair of positive and negative samples can be expressed as:

[0110]

[0111] Step S35: Through learning from all positive and negative sample pairs, the total loss function is expressed as:

[0112]

[0113]

[0114] In the formula, It is the entropy of the probability assignment for subtype clustering, and outputs most of the label features after each loss calculation.

[0115] The features use a decoupled contrastive loss function to cluster the samples, thereby outputting cluster labels. and The features are calculated using a clustering loss function to cluster the samples, thus outputting the clustering features. The clustering model is an end-to-end training and prediction process; therefore, during training, it's crucial to decouple and simultaneously optimize the contrastive and clustering loss functions. Ultimately, in the clustering task, the total loss function is:

[0116] L = L D +L c .

[0117] The present invention will be further illustrated by the following embodiments:

[0118] This embodiment compares the classification performance of SMA with six common omics data classification methods: (1) lfNN model: Each omics vector is concatenated into a feature vector as the input of the model. Multiple neural networks extract features and use Softmax as the last layer output classification. (2) efNN model: Each omics vector is used as the input of the model. Multiple neural networks extract features and concatenate the outputs into a vector. Softmax is used as the last layer output classification. (3) lfCNN model: It is similar to efNN, but lfCNN adds convolutional and pooling layers. Multiple omics vectors are concatenated into a feature vector and fed into convolutional and pooling layers. The output features are flattened and fed into a fully connected network for final prediction. (4) efCNN model: It is similar to lfNN. Each omics vector is input into convolutional and pooling layers. The output features are flattened, concatenated, and fed into a fully connected neural network for final prediction. (5) moGCN model: It uses GCN to learn the features of omics data and perform classification tasks. To perform omics-specific classification, a multi-layer GCN needs to be built for each type of omics data. (6) moGAT model: The moGCN model replaces GCN with GAT. In the testing method, lfNN, efNN, lfCNN, efCNN, moGCN, moGAT, and SMA models are trained by direct concatenation of preprocessed multi-omics data as input. All models are trained using the same preprocessed data. The classification performance of all models under the equal and heterogeneous conditions in the simulated dataset can be found in Table 1.

[0119] Table 1. Performance of the 7 supervision methods under the condition that all cluster sizes are the same.

[0120]

[0121] The experiment selected samples of 5 clusters, 10 clusters, and 15 clusters of random sizes. These seven supervised methods are essentially designed for sample classification, classifying samples from true clusters (subtypes). To quantitatively evaluate the seven supervised models in the classification task, we used a simple randomized cross-validation method for training and testing. Simultaneously, all models were evaluated using three metrics: Accuracy, F1 macro, and F1 weighted. Table 2 shows the performance of efNN, moGCN, moGAT, and SMA models in the multi-omics classification task. The efCNN model significantly underperformed the other six models in the 15 clusters of random sizes sample classification task. The lfNN model also significantly underperformed the other six models in the 10 clusters of random sizes sample classification task. This may be because the simulated dataset, after multiple convolutional and pooling layers, led to overfitting, causing misclassifications. The lfNN model only achieved optimal performance in the 5 clusters of random sizes sample classification task. This may be because the lfNN model failed to learn multi-omics features during the feature extraction process, leading to misjudgment.

[0122] Similar to the methods used to evaluate lfNN, efNN, lfCNN, efCNN, moGCN, moGAT, and SMA models on simulated datasets in classification tasks, this embodiment will investigate the performance of these models on single-cell datasets. All models use a simple cross-validation method to classify samples from three cancer cell lines, and classification performance is measured using three evaluation metrics: Accuracy, F1 macro, and F1 weighted.

[0123] Table 2 shows the performance of six supervised methods on single-cell multi-omics datasets.

[0124]

[0125] As described in Table 2, it can be seen that lfNN, efNN, moGCN, moGAT, and SMA models all achieved peak performance in the Accuracy, F1 macro, and F1 weighted evaluations, indicating that these models have achieved optimal performance in classification tasks. lfCNN and efCNN models still did not perform as well as other models in the tests. One main reason may be the relatively small number of convolutional and pooling layers, resulting in insufficient feature extraction. Another reason may be the lack of regularization penalties in the model, or other reasons, including but not limited to inappropriate learning rates or batch sizes.

[0126] In the classification task, similar to the methods used to evaluate lfNN, efNN, lfCNN, efCNN, moGCN, moGAT, and SMA models on simulated and single-cell datasets, experiments were conducted on five datasets containing real-world cancer subtypes. These methods classified samples of real-world cancer subtypes. All models were trained and tested using simple cross-validation, and classification performance was measured using three evaluation metrics: accuracy, F1 macro, and F1 weighted. For each cancer dataset, we selected data samples from three omics datasets, obtaining 59, 272, 206, 144, and 198 samples for BRCA, GBM, SARC, LUAD, and STAD, respectively. BRCA includes five cancer subtypes: LuminalA, LuminalB, Basal-like, Normal-like, and HER2-enriched. GBM includes four cancer subtypes: Proneural, Classical, Mesenchymal, and Neural. SARC includes five cancer subtypes: dediferentiated liposarcoma, leiomyosarcoma, undiferentiated pleomorphic sarcoma, myxofbrosarcoma, malignant peripheral nerve sheath tumor, and synovial sarcoma. LUAD includes four cancer subtypes: formerly bronchioid, formerly squamoid, and formerly magnoid. STAD includes Epstein-Barr virus, microsatellite instability, genomically stable, and chromosomal instability.

[0127] like Figure 3 As shown, the SMA model achieved accuracy (1 in all three evaluation metrics: Accuracy, F1 macro, and F1 weighted) in classifying cancer subtypes in BRCA, GBM, SARC, and STAD, indicating precise classification performance. However, when classifying LUAD subtypes, the accuracy, F1 macro, and F1 weighted metrics only achieved high performance of 0.958, 0.93, and 0.91, respectively. Compared to other models, the SMA model outperforms them in classification tasks. This may be because the SMA model can simultaneously consider the positional information of the input sequence, thus capturing global information. Secondly, the SMA model has a deeper structure and more parameters, which helps it learn more features, thereby improving classification accuracy. Therefore, the SMA model can be considered a standard method for classifying multi-omics cancer data.

[0128] This embodiment compares the clustering performance of the DMACL model with 10 common omics data clustering methods: (1) lfAE: Multi-omics data are first concatenated into a feature vector, and then the AE composed of encoder and decoder performs feature clustering. Among them, the ReLU function is used for the activation function of all layers of the encoder and the middle layer of the decoder, and tanh is used for the last layer of the decoder. (2) efAE: It is similar to the lfAE model, except that when processing multi-omics data, the AE extracts features from multiple omics data simultaneously. (3) lfDAE: lfDAE will process the vector features of each omics data independently. It constructs partially corrupted data by adding noise to the input data and restores it to the original input data through encoding and decoding. (4) efDAE: efDAE will process the vector features after concatenating multi-omics data. The other steps are the same as lfDAE. (5) lfVAE: It is similar to the efAE model, multi-omics data are concatenated into a one-dimensional feature vector, and then the VAE (compared to AE, the latent vector of VAE closely follows the unit Gaussian distribution) performs feature clustering analysis. (6) efVAE: It is similar to the lfVAE model, but at the input end of the model, each omics data is used by a VAE for feature clustering analysis. (7) lfSVAE: Compared with lfVAE, this model only replaces the VAE with SVAE (SVAE is a stacked VAE model. In SVAE, all hidden layers follow a unit Gaussian distribution.) and the rest remains unchanged. (8) efSVAE: Each hidden layer of the encoder is fully connected to two output layers, and the sampling steps are the same as VAE. In the evaluation, a multiplier similar to β-VAE is added to the loss function. (9) lfmmdVAE: It is similar to lfVAE, but the VAE is used to train omics data and finally classifies the features of the multi-omics ensemble. (10) efmmdVAE: A VAE is also used to train omics data. Except for the loss function, the rest is the same as efVAE.

[0129] In the clustering task, the experiment used a model to extract features from simulated multi-omics data, obtaining 5-dimensional, 10-dimensional, and 15-dimensional embeddings. The embedding dimensions were set according to the number of clusters in the simulated multi-omics data. Then, the k-means algorithm was used to cluster the dimensionality reduction results of the multi-omics data. Finally, the sample clusters were obtained to compare the performance of eleven unsupervised methods.

[0130] In the clustering task on the simulated dataset, this embodiment first uses the C-index evaluation metric to measure the consistency between the clustering of multi-omics data fusion and the actual clustering. The lower the C-index, the smaller the distance between clustered samples, and the better the clustering effect of the model. As shown in Table 3, most clustering methods have good clustering performance. However, the DMACL model achieves C-index values ​​of 0.002, 0.022, and 0.023 when the clusters have random sizes. Furthermore, when the clusters have the same size, the DMACL model achieves C-index values ​​of 0.005, 0.021, and 0.014. It outperforms other models in this evaluation metric. This may be because the multi-head attention mechanism focuses more on local information in the data when extracting multi-omics data to extract more significant data features. Furthermore, it can be observed that the DMACL model maintains good clustering performance even as the number of clusters increases.

[0131] Table 3 shows the C-index of eleven unsupervised methods on the simulated dataset.

[0132]

[0133] The silhouette score is obtained by calculating the silhouette coefficient of each sample, measuring the degree to which a sample is assigned to the correct cluster. Table 4 shows that the efVAE model has a higher probability of assigning samples to the correct clusters. The DMACL model, under the condition of clusters with the same size, only achieved rankings of 3rd, 5th, and 7th respectively. The poor clustering performance of the DMACL model may be due to poor sample quality in the simulated data, such as the presence of noise or outliers. Secondly, an uneven distribution of cluster sizes in the dataset may also lead to a lower silhouette score. Furthermore, the silhouette score itself has certain limitations, such as inaccurate evaluation of clustering performance with uneven density.

[0134] Table 4. Silhouette scores of 11 unsupervised methods on a simulated dataset.

[0135]

[0136] Table 5 shows that the efVAE model achieved a low Davies Bouldin score in clustering on the simulated data. This may be because VAE encodes the input data into latent vectors and then generates new data from these latent vectors to learn the data distribution. The DMACL model only achieved rankings of 3rd, 6th, and 6th respectively under the condition of random cluster size. The low Davies Bouldin score may be due to insufficient salient features in the simulated dataset, preventing the multi-head attention mechanism from extracting effective features. Secondly, the number of clusters also appears to be a factor affecting the Davies Bouldin score.

[0137] Table 5. Davies Bouldin scores of 11 unsupervised methods on a simulated dataset.

[0138]

[0139] Evaluation of the DMACL model in single-cell data: For the clustering task of single-cell datasets, all models first fuse features of multi-omics data to obtain fused two-dimensional embeddings. Then, the k-means algorithm is used to reduce the dimensionality of the multi-learning data and cluster them. Finally, the performance of eleven unsupervised methods is compared by obtaining the results of single-class clustering. The experiment uses C-index, silhouette score, and Davies Bouldin score to evaluate the clustering effect of the model. As shown in Figure (3), the DMACL model obtains the lowest C-index value and Davies Bouldin score, and a higher silhouette score when clustering samples. Therefore, the DMACL model becomes the best model for clustering single-cell datasets. This may be because single-cell data has long sequence information, and the DMACL model has a multi-head attention mechanism to handle long sequences, thereby reducing the occurrence of gradient vanishing and gradient explosion during model training. In summary, the DMACL model can better capture the features of single-cell data, thereby improving the accuracy of clustering.

[0140] Cancer multi-omics data are characterized by high dimensionality, diversity, and noise. For the clustering task, eleven unsupervised models were first used to fuse the cancer multi-omics data, resulting in 10-dimensional embeddings. Then, the k-means algorithm was used to cluster the multi-omics data. Since the optimal number of clusters was uncertain, experiments were conducted with 1 to 7 clusters. Finally, an unsupervised model was used to perform cluster analysis on the samples. When evaluating the self-supervised clustering model, the C-index, silhouette score, and Davies Bouldin score were used to measure the model's performance. Figure 4As shown, in the clustering experiments of all models, the C-index of the DMACL model is mainly concentrated in the middle part of the radar chart. According to the coordinates of the radar chart, the closer the data is to the center point, the smaller the value. Therefore, the C-index value of the DMACL model reflects the almost accurate clustering of the samples. This may be because the DMACL model has strong generalization ability, which helps the model capture more features, thereby improving the effect of feature extraction and data dimensionality reduction. The efmmdVAE, efVAE, and lfmmdVAE models also have good clustering results and can be used as reference models for cancer multi-omics datasets.

[0141] from Figure 5 In the study, the DMACL model achieved high silhouette scores on most cancer multi-omics datasets, with these scores primarily distributed in the outer ring of the radar chart. The DMACL model only achieved lower scores on the SKCM and LUSC 2-clustering tasks. This may be because when cancer multi-omics datasets have complex structures and numerous data points, 2-clustering may lead to underfitting, failing to capture the essential features of the dataset and resulting in potentially confusing data points between the two resulting clusters.

[0142] Davies Bouldin scores are also an important evaluation metric for analyzing the clustering performance of the DMACL model. Therefore, we also use Davies Bouldin scores to measure the model's performance. Figure 6 As shown, the Davies Bouldin scores are the lowest in 2- and 3-cluster clustering. In 4-, 5-, and 6-cluster clustering, the DMACL model performs worse on LUCS and LIHC. This may be because when the multi-omics dataset itself has a complex structure and few data points, using 4-, 5-, and 6-cluster clustering may lead to overfitting. Dividing the dataset into three clusters may result in unnecessary subdivisions that do not accurately reflect the essential characteristics of the dataset, leading to poor clustering performance.

[0143] Although embodiments of the present invention have been described in conjunction with the accompanying drawings, the patent owner may make various modifications or alterations within the scope of the appended claims, as long as they do not exceed the protection scope described in the claims of the present invention, they shall be within the protection scope of the present invention.

Claims

1. A method for analyzing multi-omics data in cancer based on a multi-head attention mechanism, characterized in that, Includes the following steps: S1. Data collection and preprocessing of cancer multi-omics data; S2. Use a supervised multi-head attention model to complete the classification task of cancer multi-omics data; S3. A decoupled contrastive learning model based on multi-head attention mechanism is used to learn and complete the clustering task of cancer multi-omics data; The steps for generating the supervised multi-head attention model in step S2 are as follows: S21. Design a multi-head attention encoder; S22. Create a symmetrical multi-head attention encoder based on a multi-head attention encoder; S23. Creating a supervised multi-head attention model based on a symmetric multi-head attention encoder; Step S23 specifically involves: S231. Employ a symmetric multi-head attention mechanism encoder to extract features from multi-omics data and generate a feature matrix. and ; S232. Employ an element-wise multiplication feature fusion method to transform the feature matrix. and The features are multiplied element by element to obtain a fused feature vector; S233. The fused feature vector is fed into a three-layer perceptron for normalization and projected onto a new feature space to generate a new feature matrix. The cross-entropy loss function is used to calculate the error between a single predicted sample and the label; S234, Calculating the characteristic matrix The total loss function is obtained by calculating the distance between the label and the target label. ; Step S3 specifically includes: S31. Project the image from step S233 onto the new feature space. middle There are a total of n-1 pairs of positive samples used for training. There are n-1 pairs of negative samples used for training. The similarity between pairs of samples is measured by cosine distance: In the formula, In order to calculate For each view's error, create a cross-entropy loss function. Therefore, the loss function between positive and negative samples is: In the formula, It is the temperature parameter that controls the softness in the model; S32. Decoupling contrastive learning is achieved by removing positive pairs from the denominator. The process is as follows: S33, Enhanced by calculating all data To obtain the cross-entropy loss of decoupled contrastive learning, allowing the model to identify all positive samples in the dataset, the process is as follows: S34, Feature Matrix Cosine similarity is also used to calculate the error between a pair of samples, and the process is as follows: In the formula, In order to calculate The error of each view is used to create a clustering loss function. The loss function between each pair of positive and negative samples can be expressed as: S35. Through learning from all positive and negative sample pairs, the total loss function is expressed as: , In the formula, It is the entropy of the probability assignment for subtype clustering, and outputs most of the label features after each loss calculation.

2. The method for analyzing cancer multi-omics data based on multi-head attention mechanism according to claim 1, characterized in that, Step S1 specifically includes: S11. Normalize the cancer multi-omics data and unify the dimensions of different data. S12. Merge and integrate data features from different omics, shuffle the order of sample data, add noise to the samples, and generate training data.

3. The method for analyzing cancer multi-omics data based on multi-head attention mechanism according to claim 1, characterized in that, Step S21 includes: S211. Perform positional encoding on cancer multi-omics data to preserve the relationships between different positions in the sequence; S212. Feature extraction is performed using a symmetrical multi-head attention mechanism. S213. Employ a multi-head attention mechanism to perform separate computations on the features of multi-omics data; S214. Perform multiple self-attention processes on the original input sequence, then concatenate the results of each attention process and perform a linear transformation to obtain the final output result.

4. The method for analyzing multi-omics cancer data based on multi-head attention mechanism according to claim 1, characterized in that, The specific steps of step S22 are as follows: feature sharing of multi-omics data is achieved by sharing a weight matrix. Feature extraction is performed in a symmetric multi-head self-attention encoder. The learned weight features share weights in the feature mapping. In backpropagation, since the weight matrix is ​​shared, the symmetric multi-head attention encoder updates the weight gradient with the same value. A symmetric multi-head attention encoder is obtained by connecting two independent multi-head attention encoders in parallel.

5. The method for analyzing cancer multi-omics data based on multi-head attention mechanism according to claim 1, characterized in that, The formula for calculating the error between a single predicted sample and the label using the cross-entropy loss function in step S233 is as follows: The total loss function in step S234 The calculation formula is: 。 6. The method for analyzing multi-omics cancer data based on multi-head attention mechanism according to claim 1, characterized in that, The feature space The features are clustered using a decoupled contrastive loss function to output cluster labels. The clustering loss function is used to calculate the clustering operation on the samples to output the clustering features. In the clustering task, the total loss function is: 。

Citation Information

Patent Citations

  • Small sample drug chemical reaction representation and automatic classification method and device

    CN114743615A

  • Text emotion recognition method and device, computer equipment and readable storage medium

    CN116108836A