Single-cell multi-omics clustering method and system based on cross-omics attention fusion

Through the cross-omics attention fusion method, combined with self-attention and deep learning gating mechanism, the problem of insufficient interaction of omics information in single-cell multiomics data analysis is solved, and more efficient cell feature representation and clustering accuracy is achieved.

CN119807783BActive Publication Date: 2025-07-04QUFU NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510299888.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-04
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

Existing single-cell multiomics data analysis methods fail to make full use of mutual information and dynamic interactions between omics, resulting in poor cell clustering effects.

Method used

A method based on cross-omics attention fusion is adopted, and dynamic interaction of multi-omics data is achieved through self-attention mechanism and cross-omics attention mechanism, feature selection and fusion are performed in combination with deep learning gating mechanism, and clustering is performed using the K-means clustering algorithm.

Benefits of technology

It enhances the complementarity and clustering accuracy of cell feature representation, can flexibly respond to changes in different cell types and states, optimizes the fusion feature representation, and improves clustering performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119807783B_ABST
    Figure CN119807783B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of single-cell multi-omics data analysis, and specifically relates to a single-cell multi-omics clustering method and system based on cross-omics attention fusion. Through the self-attention mechanism and the cross-omics attention mechanism, the dynamic interaction of multi-omics data is realized, and a deep learning gating mechanism is introduced to adaptively adjust feature selection and dynamic feature fusion, which can flexibly cope with the changes of different cell types and states, so as to better reveal the heterogeneity of cells; realize the fine-grained processing of cell feature representation to eliminate the influence of noise and improve the clustering accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of single-cell multi-omics data analysis, and specifically relates to a single-cell multi-omics clustering method and system based on cross-omics attention fusion. Background Art

[0002] High-throughput sequencing technology provides essential tools for molecular biology and biomedical research. In particular, single-cell RNA sequencing (scRNA-seq) and single-cell ATAC sequencing (scATAC-seq) respectively reveal gene expression and chromatin accessibility at the single-cell level. However, the analysis of single-cell omics data often fails to capture the full complexity and heterogeneity of cells, resulting in poor cell clustering performance. Therefore, integrating single-cell multi-omics data is crucial for improving the accuracy and robustness of cell clustering.

[0003] Clustering methods based on single-cell multi-omics have received extensive attention due to their excellent performance. These methods focus on effectively integrating information from multi-omics data to capture differences between cell types, thereby improving clustering accuracy. For example, DCCA uses variational autoencoders to extract feature representations from each omics data respectively. It introduces a cross-omics recurrent attention mechanism that cyclically transfers attention maps between different omics data to coordinate their low-dimensional embeddings. DEMOC uses autoencoders to learn low-dimensional feature representations of each omics data. It adopts a collaborative training mechanism that alternately processes each omics data as a reference view to strengthen consistency and complementary information. scMIC uses autoencoders and graph autoencoders to extract latent feature representations from different omics data. It introduces an information fusion module that effectively integrates local neighborhood information with global structural information to generate more discriminative feature representations. GRMEC-SC generates initial clustering results through multiple single-omics clustering methods, uses weighted non-negative matrix factorization to extract consistent low-dimensional representations, and uses graph regularization to integrate the co-clustering affinity matrix with these representations.

[0004] Although existing methods can extract feature representations of each omics, they usually process each omics data independently and fail to fully utilize the mutual information and dynamic interactions between omics. This lack of information exchange limits the capture of cell heterogeneity and reduces the clustering effect. Summary of the Invention

[0005] The present invention provides a single-cell multi-omics clustering method and system based on cross-omics attention fusion.

[0006] The technical solution of the present invention is as follows:

[0007] The present invention provides a single-cell multi-omics clustering method based on cross-omics attention fusion, comprising the following steps:

[0008] S1: Obtain single-cell gene expression data. After processing by the SCANPY method, perform dimensionality reduction through principal component analysis to obtain a gene expression matrix;

[0009] Obtain single-cell chromatin accessibility data. Use the latent semantic indexing method to perform dimensionality reduction to obtain a peak matrix;

[0010] According to the gene expression matrix and the peak matrix, calculate the Euclidean matrix between cells respectively to obtain the corresponding distance matrix. According to the distance matrix, calculate the corresponding similarity by the reciprocal of the Euclidean distance respectively. For each cell, connect the cells whose similarity satisfies the K-nearest neighbor method to construct a gene expression cell graph and a chromatin accessibility cell graph respectively;

[0011] S2: Extract the latent representations from the gene expression matrix and the gene expression cell graph respectively. The extracted latent representations are linearly weighted to obtain the cell feature representation of transcriptomics, denoted as the first cell feature representation;

[0012] Extract the latent representations from the peak matrix and the chromatin accessibility cell graph respectively. The extracted latent representations are linearly weighted to obtain the cell feature representation of epigenomics, denoted as the second cell feature representation;

[0013] S3: The first cell feature representation and the second cell feature representation are respectively processed by the self-attention mechanism to calculate the allocation weights between cells of each omics, and generate the third cell feature representation of the corresponding omics;

[0014] Encode the allocation weights between cells of other omics into the current omics, and perform weighted summation to generate the fourth cell feature representation of the current omics;

[0015] Fuse the fourth cell feature representation of the current omics with the third cell feature representation to obtain the fused feature representation of the current omics;

[0016] S4: After decoupling the fused feature representation, perform element-wise multiplication and addition in sequence to obtain the gated conversion feature; Perform residual connection processing on the gated conversion feature with the first cell feature representation or the second cell feature representation to obtain the adjusted fused feature of the corresponding omics;

[0017] S5: Use the K-means clustering algorithm for the adjusted fused features of the corresponding omics to obtain the initial clustering centers of each omics; Based on the initial clustering centers, calculate the soft assignment matrix of each omics, and integrate the soft assignment matrices of all omics to obtain the clustering labels of single cells.

[0018] In step S3, the first cell feature representation and the second cell feature representation are respectively processed by the self-attention mechanism to calculate the allocation weights between each omics cell, and the third cell feature representation of the corresponding omics is generated. Specifically:

[0019] The cell feature representation to be processed is mapped to the corresponding query vector, key vector and value vector. After the query vector and the key vector perform a dot product, and the similarity between cells is calculated, it is normalized by the softmax function to obtain the allocation weight of the cell. The allocation weight and the value vector are weighted and summed to obtain the third cell feature representation.

[0020] Encode the allocation weights between other omics cells into the current omics, and perform weighted summation to generate the fourth cell feature representation of the current omics. Specifically:

[0021] According to the formula: , generate the fourth cell feature representation of the current omics;

[0022] In the formula, represents the fourth cell feature representation of the current omics; represents the current omics the value vector of the cell in; represents the allocation weights between other omics cells.

[0023] In step S4, after the fused feature representation is decoupled, it is successively subjected to element-wise multiplication and addition to obtain the gated transformation feature. Specifically:

[0024] The fused feature representation is decoupled by multiple linear processes. For the decoupled feature representation, according to the formula: , the gated transformation feature is obtained;

[0025] Among them, ; ;

[0026] In the formula, represents the gated transformation feature; is the feature representation decoupled by the linear transformation ; is the fused feature representation; is the feature representation decoupled by the linear transformation ; represents the transform gate gating mechanism using the sigmoid function; 、 represent linear transformations; 、 、 represent the learnable parameters of the corresponding linear processes; Indicates element-wise multiplication.

[0027] In step S2, the latent representations are extracted respectively, and the extracted latent representations are linearly weighted to obtain the cell feature representation, specifically:

[0028] For the gene expression matrix or peak matrix, respectively according to the formula: , the latent representations are extracted to obtain the first latent representation of transcriptomics or epigenomics;

[0029] In the formula, is the first latent representation of transcriptomics or epigenomics; represents the encoding function; represents the gene expression matrix or peak matrix;

[0030] For the gene expression cell graph or chromatin accessibility cell graph, the latent representations are extracted respectively through graph neural network processing to obtain the second latent representation of transcriptomics or epigenomics; where the latent representation of the th layer of the graph neural network processing is:

[0031] ;

[0032] In the formula, is the second latent representation of transcriptomics or epigenomics; is the Tanh activation function; is the degree matrix; is the identity matrix; represents the adjacency matrix; represents the th layer of the latent representation; represents the th layer of the learnable weight matrix;

[0033] According to the formula: , the first latent representation of transcriptomics and the second latent representation of transcriptomics are linearly weighted to obtain the cell feature representation of transcriptomics; the first latent representation of epigenomics and the second latent representation of epigenomics are linearly weighted to obtain the cell feature representation of epigenomics;

[0034] In the formula, is the learnable coefficient for automatically adjusting the allocation weights of the first latent representation and the second latent representation; represents the cell feature representation of transcriptomics or epigenomics.

[0035] In step S1, according to the distance matrix, after calculating the corresponding similarity by the reciprocal of the Euclidean distance respectively, for each cell, the cells whose similarity meets the K-nearest neighbor method are connected to construct a gene expression cell graph and a chromatin accessibility cell graph respectively, specifically:

[0036] According to the formula: , calculate the similarity;

[0037] In the formula, represents the similarity between cell and cell ; represents the Euclidean distance between cell and cell .

[0038] For each cell, with the node representing the cell and the edge representing the connection relationship between cells, the cells whose similarity meets the K-nearest neighbor method are connected to construct a gene expression cell graph and a chromatin accessibility cell graph respectively.

[0039] In step S5, based on the initial cluster centers, calculate the soft assignment matrix for each omics, which is achieved according to the formula: , ,

[0040] In the formula, represents the soft assignment matrix of the omics; represents the number of cells; represents the number of initial cluster centers; represents the soft assignment score of the omics cell; is the th adjusted fusion feature of the omics; represents the th cluster center of the omics; represents the th cluster center of the omics.

[0041] Integrate the soft assignment matrices of all omics to obtain the clustering labels of single cells, which is achieved according to the formula: ,

[0042] In the formula, represents the clustering label of the th cell; represents selecting the label that maximizes the weighted soft assignment value among all clustering labels.

[0043] Step S5 further includes calculating the loss of the clustering labels of single cells, which is achieved by ,

[0044] wherein, , , , respectively represent the objective loss function, the reconstruction loss function, the feature alignment loss function, and the collaborative supervised clustering loss function; , are respectively used to balance the weights of , .

[0045] The present invention also provides a single-cell multi-omics clustering system based on cross-omics attention fusion, including:

[0046] Cell graph construction module: Obtain single-cell gene expression data, after processing by the SCANPY method, perform dimensionality reduction through principal component analysis to obtain a gene expression matrix;

[0047] Obtain single-cell chromatin accessibility data, perform dimensionality reduction using the latent semantic indexing method to obtain a peak matrix;

[0048] According to the gene expression matrix and the peak matrix, calculate the Euclidean matrix between cells respectively to obtain the corresponding distance matrix. According to the distance matrix, calculate the corresponding similarity by taking the reciprocal of the Euclidean distance respectively. For each cell, connect the cells whose similarity satisfies the K-nearest neighbor method to construct a gene expression cell graph and a chromatin accessibility cell graph respectively;

[0049] Feature extraction module: Extract the latent representations from the gene expression matrix and the gene expression cell graph respectively, and perform linear weighting on the extracted latent representations to obtain the cell feature representation of transcriptomics, denoted as the first cell feature representation;

[0050] Extract the latent representations from the peak matrix and the chromatin accessibility cell graph respectively, and perform linear weighting on the extracted latent representations to obtain the cell feature representation of epigenomics, denoted as the second cell feature representation;

[0051] Multi-omics fusion module: The first cell feature representation and the second cell feature representation are respectively processed by the self-attention mechanism to calculate the allocation weights between cells of each omics, and generate the third cell feature representation of the corresponding omics;

[0052] Encode the allocation weights between cells of other omics into the current omics, perform weighted summation, and generate the fourth cell feature representation of the current omics;

[0053] Fuse the fourth cell feature representation of the current omics with the third cell feature representation to obtain the fusion feature representation of the current omics;

[0054] Fusion feature adjustment module: After the decoupling of the fusion feature representation, it is successively subjected to element-wise multiplication and addition to obtain the gated conversion feature; the gated conversion feature is subjected to residual connection processing with the first cell feature representation or the second cell feature representation to obtain the adjusted fusion feature of the corresponding omics.

[0055] Clustering module: The adjusted fusion features of the corresponding omics respectively use the K-means clustering algorithm to obtain the initial clustering centers of each omics; based on the initial clustering centers, the soft assignment matrices of each omics are calculated, and the soft assignment matrices of all omics are integrated to obtain the clustering labels of single cells.

[0056] Advantages: First, the present invention realizes the dynamic interaction of multi-omics data through the self-attention mechanism and the cross-omics attention mechanism. When processing single-omics data, it can capture the global dependence within the omics through the self-attention mechanism, and at the same time promote the information exchange between different omics data through the cross-omics attention mechanism, enhancing the complementarity of features.

[0057] Second, the present invention introduces a deep learning gating mechanism to adaptively adjust feature selection and dynamic feature fusion, which can flexibly respond to changes in different cell types and states, thereby better revealing cell heterogeneity; realizing fine-grained processing of cell feature representations to eliminate the influence of noise, and further enhancing the discrimination ability of cell feature representations.

[0058] Third, the present invention can optimize the fusion feature representation and improve the clustering accuracy by integrating the reconstruction loss, the feature alignment loss, and the collaborative supervised clustering loss; due to the joint use of multiple loss functions, even in the face of noisy or missing data, it can still maintain high clustering performance. Description of the Drawings

[0059] Figure 1 Flowchart of the clustering method of the present invention.

[0060] Figure 2 Is the single-cell multi-omics clustering framework diagram of the present invention.

[0061] Figure 3 Is the multi-omics feature fusion schematic diagram of the present invention.

[0062] Figure 4 Is the gated depth processing schematic diagram of the present invention. Detailed Embodiments

[0063] The following embodiments are intended to illustrate the present invention rather than further limit the present invention.

[0064] Referring to Figure 1 , the present invention provides a single-cell multi-omics clustering method based on cross-omics attention fusion, including the following steps:

[0065] S1: Obtain single-cell gene expression data. After processing by the SCANPY method, perform dimensionality reduction through principal component analysis to obtain a gene expression matrix;

[0066] Obtain single-cell chromatin accessibility data, and use the latent semantic indexing method for dimensionality reduction to obtain a peak matrix;

[0067] According to the gene expression matrix and the peak matrix, calculate the Euclidean matrix between cells respectively to obtain the corresponding distance matrix. According to the distance matrix, after calculating the corresponding similarity with the reciprocal of the Euclidean distance respectively, for each cell, connect the cells whose similarity meets the K-nearest neighbor method, and construct a gene expression cell graph and a chromatin accessibility cell graph respectively.

[0068] Specifically, for single-cell gene expression data (scRNA-seq), first filter out cells with a low number of detected genes and invalid cells with a low counting depth. Introduce the SCANPY method to select 3000 highly variable genes and perform normalization processing, and further perform dimensionality reduction through principal component analysis (PCA) to obtain a gene expression matrix.

[0069] For single-cell chromatin accessibility data (scATAC-seq), only use the latent semantic indexing (LSI) method to perform dimensionality reduction on the data to obtain a peak matrix.

[0070] Then, perform the following operations on the gene expression matrix and the peak matrix respectively to construct the corresponding gene expression cell graph and chromatin accessibility cell graph.

[0071] According to the gene expression matrix and the peak matrix, calculate the Euclidean matrix between cells respectively to obtain the corresponding distance matrix, which is used to represent the distance between cells, and the smaller the value, the closer the cells are.

[0072] Next, in step S1, according to the distance matrix, after calculating the corresponding similarity with the reciprocal of the Euclidean distance respectively, for each cell, connect the cells whose similarity meets the K-nearest neighbor method, and construct a gene expression cell graph and a chromatin accessibility cell graph respectively. Specifically:

[0073] According to the formula: , calculate the similarity;

[0074] In the formula, represents the similarity between cell and cell , represents the Euclidean distance between cell and cell ;

[0075] For each cell, using nodes to represent cells and edges to represent the connection relationships between cells, connect the cells whose similarity meets the K-nearest neighbor method, and construct a gene expression cell graph and a chromatin accessibility cell graph respectively.

[0076] S2: Refer to Figure 2 , extract the latent representations from the gene expression matrix and the gene expression cell graph respectively, and perform linear weighting on the extracted latent representations to obtain the cell feature representation of transcriptomics, denoted as the first cell feature representation;

[0077] Extract the latent representations from the peak matrix and the chromatin accessibility cell graph respectively, and perform linear weighting on the extracted latent representations to obtain the cell feature representation of epigenomics, denoted as the second cell feature representation.

[0078] Among them, extract the latent representations respectively (that is, extract the latent representations from each omics data and the corresponding cell graph respectively), and perform linear weighting on the extracted latent representations to obtain the cell feature representation. The operations are as follows:

[0079] (1) For the gene expression matrix or the peak matrix, respectively according to the formula: , extract the latent representation to obtain the first latent representation of transcriptomics or epigenomics;

[0080] In the formula, is the first latent representation of transcriptomics or epigenomics; represents the encoding function; represents the gene expression matrix or the peak matrix;

[0081] For the gene expression cell graph or the chromatin accessibility cell graph, extract the latent representations respectively through graph neural network processing to obtain the second latent representation of transcriptomics or epigenomics; among them, the latent representation of the -th layer of the graph neural network processing is:

[0082] ;

[0083] In the formula, is the second latent representation of transcriptomics or epigenomics; is the Tanh activation function; is the degree matrix; is the identity matrix; represents the adjacency matrix; represents the -th layer of the latent representation; represents the -th layer of the learnable weight matrix;

[0084] (2) According to the formula: , the first latent representation of transcriptomics and the second latent representation of transcriptomics are linearly weighted to obtain the cell feature representation of transcriptomics; the first latent representation of epigenomics and the second latent representation of epigenomics are linearly weighted to obtain the cell feature representation of epigenomics;

[0085] wherein, is a learnable coefficient for automatically adjusting the allocation weights of the first latent representation and the second latent representation; represents the cell feature representation of transcriptomics or epigenomics.

[0086] S3: Refer to Figure 3 , the first cell feature representation and the second cell feature representation are respectively processed by the self-attention mechanism to calculate the allocation weights between cells of each omics, and the third cell feature representation of the corresponding omics is generated;

[0087] The allocation weights between cells of other omics are encoded into the current omics, and weighted summation is performed to generate the fourth cell feature representation of the current omics;

[0088] The fourth cell feature representation of the current omics is fused with the third cell feature representation to obtain the fused feature representation of the current omics.

[0089] Preferably, the first cell feature representation and the second cell feature representation are respectively processed by the self-attention mechanism to calculate the allocation weights between cells of each omics, and the third cell feature representation of the corresponding omics is generated, specifically:

[0090] The cell feature representation to be processed is mapped to the corresponding query vector, key vector and value vector. After the query vector and the key vector are dot-producted to calculate the similarity between cells, normalization is performed by the softmax function to obtain the allocation weights of the cells, and the allocation weights are weighted-summed with the value vector to obtain the third cell feature representation.

[0091] In step S3 of the present invention, the self-attention mechanism is introduced to capture the global relationship within the omics, and the feature representation of cells is enhanced through the global dependence relationship between cells of each omics.

[0092] Furthermore, the allocation weights between cells of other omics are encoded into the current omics, and weighted summation is performed to generate the fourth cell feature representation of the current omics, specifically:

[0093] According to the formula: , the fourth cell feature representation of the current omics is generated;

[0094] wherein, represents the fourth cell feature representation of the current omics; represents the current omics the value vector of the cell in; Indicates the distribution weights among other omics cells.

[0095] The multi-omics data of cells shows different characteristics of cells, and different omics data often present various global dependencies among cells. In step S3 of the present invention, in addition to capturing the global relationships within the omics, a cross-omics attention mechanism (encoding the distribution weights among other omics cells into the current omics and performing weighted summation) is used to learn the complementary information between different omics, so that different omics data can cooperate better.

[0096] In the fusion stage, according to the formula: , the fused feature representation of the current omics is obtained , where is the third cell feature representation.

[0097] In step S3 of the present invention, by fusing the self-attention mechanism cell feature representation (the third cell feature representation) and the cross-omics attention cell feature representation (the fourth cell feature representation of the current omics), the global dependencies among cells within each omics and the complementary information between different omics are captured, thereby effectively enhancing the representation ability of the cell features of the current omics.

[0098] S4: After the fused feature representation is decoupled, it is sequentially subjected to element-wise multiplication and addition to obtain a gated transformation feature; the gated transformation feature is subjected to residual connection processing with the first cell feature representation or the second cell feature representation to obtain an adjusted fused feature of the corresponding omics.

[0099] Preferably, after the fused feature representation is decoupled, it is sequentially subjected to element-wise multiplication and addition to obtain a gated transformation feature, referring to Figure 4 , specifically:

[0100] The fused feature representation is decoupled through multiple linear processes, and for the decoupled feature representation, according to the formula: , the gated transformation feature is obtained;

[0101] Among them, ; ;

[0102] In the formula, represents the gated transformation feature; is the feature representation decoupled by the linear transformation ; is the fused feature representation; is the feature representation decoupled by the linear transformation ; represents the transform gate gating mechanism using the sigmoid function; , represent linear transformations; , , represent learnable parameters for corresponding linear processing; represents element-wise multiplication.

[0103] In the process of obtaining the gated transformation features in step S4 of the present invention, a rectified linear unit is used to extract the significant features of cells, and a transform gate is introduced to control the flow of information.

[0104] In addition, a residual connection is introduced to retain the initial cell feature representation, and the adjusted fusion features of the corresponding omics can be obtained according to the formula: ,

[0105] wherein, is the adjusted fusion feature of the th omics, is the gated transformation feature, is the cell feature representation of transcriptomics or epigenomics.

[0106] Although step S3 integrates different omics data to enrich the cell feature representation, however, in the fused cell feature representation, there may be some redundant or noisy feature values. The present invention introduces a deep learning gating mechanism in step S4 to perform fine-grained processing on the fused cell feature representation, thereby highlighting the features that are more important for the single-cell clustering task. Adaptive adjustment of feature selection and dynamic feature fusion can flexibly cope with the changes of different cell types and states, so as to better reveal the heterogeneity of cells.

[0107] S5: The adjusted fusion features of the corresponding omics respectively use the K-means clustering algorithm to obtain the initial clustering centers of each omics; based on the initial clustering centers, calculate the soft assignment matrix of each omics, and integrate the soft assignment matrices of all omics to obtain the clustering labels of single cells.

[0108] Preferably, calculating the soft assignment matrix of each omics based on the initial clustering centers is achieved according to the formula: , ,

[0109] wherein, represents the soft assignment matrix of the th omics; represents the number of cells; represents the number of initial clustering centers; represents the th soft assignment score of the omics cells; is the th nd adjusted fusion feature of the Indicates the th clustering center of omics; Indicates the th clustering center of omics.

[0110] Furthermore, by integrating the soft assignment matrices of all omics, the clustering labels of single cells are obtained, which is achieved according to the formula: , realized;

[0111] In the formula, Indicates the clustering label of the th cell; Indicates the label that maximizes the weighted soft assignment value among all clustering labels.

[0112] In addition, it also includes calculating the loss of the clustering labels of single cells, which is achieved through , realized;

[0113] In the formula, , , , respectively represent the target loss function, reconstruction loss function, feature alignment loss function, and collaborative supervised clustering loss function; , are respectively used to balance the , weights.

[0114] Among them, the reconstruction loss function is used to minimize the difference between the original information and the reconstructed information to ensure the effectiveness of the cell feature representation. The reconstruction loss function is expressed as follows:

[0115] ;

[0116] In the formula, represents the reconstruction loss function; represents the number of cells; represents the gene expression matrix or peak matrix; represents the decoding function of the gene expression matrix or peak matrix; is the adjusted fusion feature of the th omics; is the gene expression cell graph or chromatin accessibility cell graph; represents the decoding function of the gene expression cell graph or chromatin accessibility cell graph; represents the Frobenius norm, which is used to calculate the difference between the reconstructed data and the original data and measures the magnitude of the reconstruction error.

[0117] The feature alignment loss function maximizes the similarity of the same cell features among different omics to achieve feature alignment and eliminate redundant information. The feature alignment loss function is expressed as follows:

[0118] , ;

[0119] In the formula, represents the feature alignment loss function; represents the similarity matrix calculated using the cosine similarity function between the adjusted fusion features of the th omics and the adjusted fusion features of the th omics; is the adjusted fusion feature of the th omics; is the adjusted fusion feature of the th omics.

[0120] The collaborative supervised clustering loss function aims to enhance the information flow among different omics to improve the clustering performance. Specifically as follows:

[0121] After calculating the soft assignment scores of cells within each omics , calculate the target distribution matrix , where ;

[0122] In the formula, represents the soft assignment score of the th cell in the th omics belonging to the th clustering center; represents the updated soft assignment score of the th cell in the th omics belonging to the th clustering center. The numerator represents the square of the soft assignment score of the cell to the clustering center , and then divided by the sum of the soft assignment scores of all cells to the clustering center . This operation can be understood as weighting the soft assignment scores, making the cells with higher scores have a greater impact on the clustering center. The denominator is the sum of the weighted soft assignment scores of all clustering centers, used for normalization, so that is between 0 and 1.

[0123] This application believes that the multi-omics data of cells should have the same clustering distribution. Therefore, the collaborative supervised clustering loss function enhances the information flow among omics by minimizing the KL divergence loss between the soft assignment matrix and the target distribution matrix . At the same time, minimize the soft assignment matrix within the same omics and the target distribution matrix The KL divergence loss is used to improve the clustering accuracy. Therefore, the collaborative supervised clustering loss function is defined as follows:

[0124] ;

[0125] In the formula, represents the collaborative supervised clustering loss function; represents the number of cells; represents the number of initial clustering centers; represents the index of the cell, ranging from 1 to ; , represents the index of the clustering center, ranging from 1 to ; represents the th omics th cell belongs to the th omics th th clustering center's soft assignment score;

[0126] In step S5 of the present invention, by integrating three loss functions (reconstruction loss, feature alignment loss, and collaborative supervised clustering loss), combined with the backpropagation algorithm, the learnable parameters in step S2 are optimized, and the allocation weights of the first latent representation and the second latent representation are adjusted, thereby optimizing the fused feature representation.

[0127] In summary, first, the present invention realizes the dynamic interaction of multi-omics data through the self-attention mechanism and the cross-omics attention mechanism. When processing single-omics data, it can capture the global dependencies within the omics through the self-attention mechanism, and at the same time promote the information exchange between different omics data through the cross-omics attention mechanism, enhancing the complementarity of features.

[0128] Second, the present invention introduces a deep learning gating mechanism to adaptively adjust feature selection and dynamic feature fusion, which can flexibly respond to changes in different cell types and states, thereby better revealing the heterogeneity of cells; realizing fine-grained processing of cell feature representations to eliminate the influence of noise, and further enhancing the discriminative ability of cell feature representations.

[0129] Third, the present invention can optimize the fused feature representation and improve the clustering accuracy by integrating the reconstruction loss, feature alignment loss, and collaborative supervised clustering loss; due to the joint use of multiple loss functions, even in the face of noisy or missing data, it can still maintain high clustering performance.

[0130] To verify the effectiveness of the present invention, the performance of the present invention is tested from the aspect of cell type clustering below.

[0131] Dataset and evaluation metrics

[0132] The present invention conducts experiments on two single-cell multi-omics datasets, PBMC-10K and PBMC-3K, to verify its performance. The PBMC-10K dataset contains 9,631 cells from humans, covering 29,095 genes and 107,194 peaks, with a total of 19 cell types. The PBMC-3K dataset contains 2,883 cells, covering 21,342 genes and 72,540 peaks, with a total of 8 cell types. The cell types of both datasets are manually labeled with true labels, ensuring the accuracy and reliability of the experiments.

[0133] The present invention uses four metrics, ARI, NMI, AMI, and ACC, to evaluate its clustering performance. The ARI score reflects the similarity between the clustering labels and the true labels of any two cell pairs. NMI uses normalized mutual information to measure the similarity between the single-cell clustering result and the true label. AMI adjusts NMI by considering the distribution of the true cell labels. The ACC score reflects the proportion of the number of correctly clustered cells in the total number of cells.

[0134] Comparative analysis of experimental results

[0135] To better demonstrate the effect of the clustering method of the present invention (denoted as scCAF), it is compared with other comparative methods. Among them, the comparative methods adopt the K-means-based clustering method, DCCA, DEMOC, scMCs, scMIC, scDRMAE, GRMEC-SC.

[0136] The K-means-based clustering method is a classic single-cell clustering method. By randomly initializing K clustering centers, assigning each cell to the nearest center, updating the center as the within-cluster mean, and iteratively optimizing until convergence, the single-cell data is divided into K groups;

[0137] DCCA extracts feature representations from each omics data using variational autoencoders. It introduces a cross-omics cyclic attention mechanism that cyclically transfers attention maps between different omics data to coordinate their low-dimensional embeddings;

[0138] DEMOC uses autoencoders to learn the low-dimensional feature representations of each omics data. It adopts a collaborative training mechanism that alternately processes each omics data as a reference view to strengthen consistency and complementary information;

[0139] scMCs fuses the individuality and commonality of omics through autoencoders, attention mechanisms, and contrastive learning to generate common embedded representations, and optimizes them using multi-head attention and KL divergence to achieve multi-omics clustering;

[0140] scMIC extracts latent feature representations from different omics data using autoencoders and graph autoencoders. It introduces an information fusion module that effectively integrates local neighborhood information and global structural information to produce more discriminative feature representations;

[0141] scDRMAE encodes and reconstructs different omics data through two parallel masked autoencoders, introduces a self-attention mechanism to dynamically allocate weights for concatenated omics data, and re-integrates the low-dimensional representation of scRNA-seq into the data processed by self-attention to prevent information loss. Finally, KL divergence is applied to optimize the clustering and reconstruction interaction;

[0142] GRMEC-SC generates initial clustering results through multiple single-omics clustering methods, uses weighted non-negative matrix factorization to extract consistent low-dimensional representations, and uses graph regularization to integrate the co-clustering affinity matrix with these representations.

[0143] PBMC-10K dataset: Table 1 presents the clustering performance of the scCAF method of the present invention and other comparative methods on the PBMC-10K dataset. The experimental results show that relying only on single-omics data of cells, the K-means-based clustering method achieved an ARI score of 0.608 (single-cell gene expression data) and an ARI score of 0.505 (single-cell chromatin accessibility data) respectively. The scMCs, DCCA, and DEMOC methods did not integrate the information of cell multi-omics well and were lower than the clustering performance based on single-omics data. The scCAF method proposed in the present invention achieved the best performance in all four evaluation metrics. Compared with scMIC, scCAF achieved a 5.9% ARI improvement, a 1.5% NMI improvement, a 1.5% AMI improvement, and a 4.5% ACC improvement respectively.

[0144] Table 1 Performance comparison of scCAF and other comparative methods on the PBMC-10K dataset

[0145]

[0146] PBMC-3K dataset: Table 2 presents the clustering performance of the scCAF method of the present invention and other comparative methods on the PBMC-3K dataset. The scCAF method achieved an ARI score of 0.805, an NMI score of 0.766, an AMI score of 0.765, and an ACC score of 0.820, significantly superior to other comparative methods.

[0147] Table 2 Performance Comparison between scCAF and Other Comparative Methods on the PBMC-3K Dataset

[0148]

[0149] The present invention also provides a single-cell multi-omics clustering system based on cross-omics attention fusion, including:

[0150] Cell graph construction module: Obtain single-cell gene expression data, after being processed by the SCANPY method, perform dimensionality reduction through principal component analysis to obtain a gene expression matrix;

[0151] Obtain single-cell chromatin accessibility data, and perform dimensionality reduction using the latent semantic indexing method to obtain a peak matrix;

[0152] According to the gene expression matrix and the peak matrix, calculate the Euclidean matrix between cells respectively to obtain the corresponding distance matrix. According to the distance matrix, calculate the corresponding similarity by taking the reciprocal of the Euclidean distance respectively. For each cell, connect the cells whose similarity satisfies the K-nearest neighbor method, and construct a gene expression cell graph and a chromatin accessibility cell graph respectively;

[0153] Feature extraction module: Extract the latent representation from the gene expression matrix and the gene expression cell graph respectively, and perform linear weighting on the extracted latent representation to obtain the cell feature representation of transcriptomics, denoted as the first cell feature representation;

[0154] Extract the latent representation from the peak matrix and the chromatin accessibility cell graph respectively, and perform linear weighting on the extracted latent representation to obtain the cell feature representation of epigenomics, denoted as the second cell feature representation;

[0155] Multi-omics fusion module: The first cell feature representation and the second cell feature representation are respectively processed by the self-attention mechanism, calculate the allocation weights between cells of each omics, and generate the third cell feature representation of the corresponding omics;

[0156] Encode the allocation weights between cells of other omics into the current omics, perform weighted summation to generate the fourth cell feature representation of the current omics;

[0157] Fuse the fourth cell feature representation of the current omics with the third cell feature representation to obtain the fusion feature representation of the current omics;

[0158] Fusion feature adjustment module: After decoupling the fusion feature representation, perform element-wise multiplication and addition in sequence to obtain the gated conversion feature; Perform residual connection processing on the gated conversion feature and the first cell feature representation or the second cell feature representation to obtain the adjusted fusion feature of the corresponding omics;

[0159] Clustering module: The adjusted and fused features of the corresponding omics are respectively used with the K-means clustering algorithm to obtain the initial clustering centers of each omics; Based on the initial clustering centers, the soft assignment matrices of each omics are calculated, and the soft assignment matrices of all omics are integrated to obtain the clustering labels of single cells.

Claims

1. A single-cell multi-omics clustering method based on cross-omics attention fusion, characterized in that It includes the following steps: S1: Obtain single-cell gene expression data. After processing by the SCANPY method, dimensionality reduction is performed through principal component analysis to obtain a gene expression matrix; Obtain single-cell chromatin accessibility data, and perform dimensionality reduction using the latent semantic indexing method to obtain a peak matrix; According to the gene expression matrix and the peak matrix, the Euclidean matrix between cells is calculated respectively to obtain the corresponding distance matrix. According to the distance matrix, after calculating the corresponding similarity with the reciprocal of the Euclidean distance respectively, for each cell, the cells whose similarity satisfies the K-nearest neighbor method are connected to construct a gene expression cell graph and a chromatin accessibility cell graph respectively; S2: Extract latent representations from the gene expression matrix and the gene expression cell graph respectively. The extracted latent representations are linearly weighted to obtain a cell feature representation of transcriptomics, denoted as the first cell feature representation; Extract latent representations from the peak matrix and the chromatin accessibility cell graph respectively. The extracted latent representations are linearly weighted to obtain a cell feature representation of epigenomics, denoted as the second cell feature representation; S3: The first cell feature representation and the second cell feature representation are respectively processed by a self-attention mechanism to calculate the assignment weights between cells of each omics, and a third cell feature representation of the corresponding omics is generated; Encode the assignment weights between cells of other omics into the current omics, and perform weighted summation to generate a fourth cell feature representation of the current omics; Fuse the fourth cell feature representation of the current omics with the third cell feature representation to obtain a fused feature representation of the current omics; S4: After the fused feature representation is decoupled, it is successively subjected to element-wise multiplication and addition to obtain a gated conversion feature; the gated conversion feature is subjected to residual connection processing with the first cell feature representation or the second cell feature representation to obtain an adjusted fused feature of the corresponding omics; S5: Use the K-means clustering algorithm for the adjusted fused features of the corresponding omics respectively to obtain the initial clustering centers of each omics; Based on the initial clustering centers, calculate the soft assignment matrices of each omics, and integrate the soft assignment matrices of all omics to obtain the clustering labels of single cells.

2. The single-cell multi-omics clustering method based on cross-omics attention fusion according to claim 1, wherein in the step S3, the first cell feature representation and the second cell feature representation are respectively processed by a self-attention mechanism to calculate the assignment weights between cells of each omics, and a third cell feature representation of the corresponding omics is generated, specifically: The cell feature representation to be processed is mapped into corresponding query vectors, key vectors, and value vectors. The query vectors and key vectors are subjected to dot product, and after calculating the similarity between cells, they are normalized by the softmax function to obtain the assignment weights of cells. The assignment weights are weighted and summed with the value vectors to obtain the third cell feature representation.

3. The single-cell multi-omics clustering method based on cross-omics attention fusion according to claim 2, wherein encoding the assignment weights between cells of other omics into the current omics, and performing weighted summation to generate a fourth cell feature representation of the current omics, specifically: According to the formula: , generate the fourth cell feature representation of the current omics; In the formula, represents the fourth cell feature representation of the current omics; represents the current omics of the cells value vector; represents the allocation weight among cells of other omics.

4. The single-cell multi-omics clustering method based on cross-omics attention fusion according to claim 1, wherein in the step S4, after the fused feature representation is decoupled, it is successively subjected to element-wise multiplication and addition to obtain a gated conversion feature, specifically: The fused feature representation is decoupled through multiple linear processes. For the decoupled feature representation, according to the formula: , the gated conversion features are obtained; Among them, ; ; Wherein, represents the gated conversion feature; is the feature representation after linear transformation and decoupling; is the fused feature representation; is the feature representation after linear transformation and decoupling; represents the transformgate gating mechanism using the sigmoid function; and represent linear transformation; and and represent the learnable parameters for the corresponding linear processing; represents element-wise multiplication.

5. The single-cell multi-omics clustering method based on cross-omics attention fusion according to claim 1, wherein In step S2, potential representations are extracted respectively, and the extracted potential representations are linearly weighted to obtain cell feature representations, specifically as follows: A gene expression matrix or a peak matrix, respectively according to the formula: , extract the latent representation to obtain the first latent representation of transcriptomics or epigenomics; In the formula, is the first potential representation of transcriptomics or epigenomics; represents an encoding function; represents a gene expression matrix or a peak matrix; A gene expression cell graph or a chromatin accessibility cell graph is processed by a graph neural network to extract a latent representation respectively, obtaining a second latent representation of transcriptomics or epigenomics; wherein the latent representation of the layer processed by the graph neural network is: ; In the formula, is the second potential representation of transcriptomics or epigenomics; is the Tanh activation function; is the degree matrix; is the identity matrix; represents the adjacency matrix; represents the th layer of potential representation; represents the learnable weight matrix of the th layer; According to the formula: , the first latent representation of transcriptomics and the second latent representation of transcriptomics are linearly weighted to obtain the cell feature representation of transcriptomics; the first latent representation of epigenomics and the second latent representation of epigenomics are linearly weighted to obtain the cell feature representation of epigenomics; In the formula, is a learnable coefficient used to automatically adjust the allocation weights of the first latent representation and the second latent representation; represents the cell feature representation of transcriptomics or epigenomics.

6. The single-cell multi-omics clustering method based on cross-omics attention fusion according to claim 1, characterized in that In step S1, according to the distance matrix, after calculating the corresponding similarity with the reciprocal of the Euclidean distance respectively, for each cell, the cells whose similarity satisfies the K-nearest neighbor method are connected to construct a gene expression cell graph and a chromatin accessibility cell graph respectively, specifically as follows: According to the formula: , calculate the similarity; In the formula, represents the similarity between cells and cells , and represents the Euclidean distance between cells and cells . For each cell, using nodes to represent cells and edges to represent the connection relationships between cells, the cells whose similarity satisfies the K-nearest neighbor method are connected to construct a gene expression cell graph and a chromatin accessibility cell graph respectively.

7. The single-cell multi-omics clustering method based on cross-omics attention fusion according to claim 1, characterized in that In step S5, based on the initial clustering centers, the soft assignment matrix of each omics is calculated according to the formula: , , achieve; Wherein, represents the soft assignment matrix of the omics; represents the number of cells; represents the number of initial clustering centers; represents the soft assignment score of the omics cells; is the th adjusted fusion feature of the omics; represents the th clustering center of the omics; represents the th clustering center of the omics.

8. The single-cell multi-omics clustering method based on cross-omics attention fusion according to claim 7, characterized in that Integrating the soft assignment matrices of all omics to obtain the clustering labels of single cells according to the formula: , implement; In the formula, represents the clustering label of the th cell; represents the label that maximizes the weighted soft assignment value among all clustering labels.

9. The single-cell multi-omics clustering method based on cross-omics attention fusion according to claim 1, characterized in that The step S5 further includes calculating the loss of the clustering label of single cells, which is achieved by , In the formula, , , , respectively represent the target loss function, the reconstruction loss function, the feature alignment loss function, and the collaborative supervised clustering loss function; , are respectively used to balance the , weights.

10. A single-cell multi-omics clustering system based on cross-omics attention fusion, characterized in that, It includes: Cell graph construction module: Obtain single-cell gene expression data, after processing by the SCANPY method, perform dimensionality reduction through principal component analysis to obtain a gene expression matrix; Obtain single-cell chromatin accessibility data, and perform dimensionality reduction using the latent semantic indexing method to obtain a peak matrix; According to the gene expression matrix and the peak matrix, calculate the Euclidean matrix between cells respectively to obtain the corresponding distance matrix. According to the distance matrix, calculate the corresponding similarity using the reciprocal of the Euclidean distance respectively. For each cell, connect the cells whose similarity satisfies the K-nearest neighbor method, and construct a gene expression cell graph and a chromatin accessibility cell graph respectively; Feature extraction module: The gene expression matrix and the gene expression cell graph extract potential representations respectively, and the extracted potential representations are linearly weighted to obtain the cell feature representation of transcriptomics, denoted as the first cell feature representation; The peak matrix and the chromatin accessibility cell graph extract potential representations respectively, and the extracted potential representations are linearly weighted to obtain the cell feature representation of epigenomics, denoted as the second cell feature representation; Multi-omics fusion module: The first cell feature representation and the second cell feature representation are respectively processed by the self-attention mechanism, calculate the assignment weights between cells of each omics, and generate the third cell feature representation of the corresponding omics; Encode the assignment weights between cells of other omics into the current omics, perform weighted summation to generate the fourth cell feature representation of the current omics; Fuse the fourth cell feature representation of the current omics with the third cell feature representation to obtain the fusion feature representation of the current omics; Fusion feature adjustment module: After the fusion feature representation is decoupled, it is successively subjected to element-wise multiplication and addition to obtain a gated conversion feature; the gated conversion feature is subjected to residual connection processing with the first cell feature representation or the second cell feature representation to obtain the adjusted fusion feature of the corresponding omics; Clustering module: The adjusted fusion features of the corresponding omics respectively use the K-means clustering algorithm to obtain the initial clustering centers of each omics; Based on the initial clustering centers, calculate the soft assignment matrix of each omics, and integrate the soft assignment matrices of all omics to obtain the clustering labels of single cells.

Citation Information

Patent Citations

  • Single-cell multi-omics data integration method and system based on graph contrast learning

    CN118571328A

  • Deep clustering method and system based on cross-modal fusion

    WO2022166361A1