Single-cell multi-omics data clustering method based on diffusion auto-encoder
By combining diffusion autoencoder and bidirectional cross-attention mechanism, the problem of decoupling of cross-omic features in single-cell multiomics data is solved, and efficient cell subtype recognition and topological presentation are achieved, improving the accuracy of clustering.
Patent Information
- Application Number
- CN202510643537.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-26
AI Technical Summary
When processing single-cell multiomics data, it is difficult to effectively decouple cross-omics invariant features and omics-specific noise, resulting in insufficient accuracy in cell subtype recognition and simple feature fusion strategies that are difficult to adaptively capture cell feature differences.
Using a diffusion autoencoder-based method, combined with a two-way cross attention mechanism and conditional diffusion model, the semantic features and random noise of cells are decoupled through dimensionality reduction, feature fusion and clustering algorithms to achieve efficient integration and clustering of cross-omics data.
It improves the accuracy of cell clustering, can accurately present cell topology in low-dimensional space, identify different cell subtypes, and enhances the ability to identify rare cell types.
Smart Images

Figure BDA0005408914180000037 
Figure BDA0005408914180000041 
Figure BDA0005408914180000046
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of biological single-cell data classification, and in particular to a single-cell multi-omics data clustering method based on a diffusion autoencoder. Background Art
[0002] With the development of single-cell sequencing technology, multiple omics data of the same cell can be obtained simultaneously, including transcriptome (scRNA-seq), epigenome (scATAC-seq), proteome, etc., allowing researchers to analyze the biological state and cellular heterogeneity of cells in multiple dimensions and perspectives, thereby revealing cell subpopulations and their differentially expressed genes, complex gene regulatory mechanisms, etc. Compared with single-omics data, cell clustering methods based on multi-omics data can reveal cellular heterogeneity more comprehensively and help discover new cell types. However, due to the significant domain offset problem between different omics data, coupled with the high sparsity and high noise of single-cell sequencing data, the integrated analysis of multi-omics data faces great challenges. How to obtain low-dimensional embeddings that can effectively characterize cell states from multi-omics data has become a difficulty in current research.
[0003] To solve this problem, it is necessary to eliminate the semantic differences between different omics, effectively integrate multi-source heterogeneous features, decouple cross-omics invariant features related to cell biological state and omics-specific random noise, use cross-omics co-semantic embedding to represent cell state, and then accurately identify cell topology. Therefore, it is necessary to design a single-cell multi-omics data clustering method that can adaptively integrate multi-omics features and effectively represent cell state, so as to achieve accurate identification of cell subpopulations, thereby laying the foundation for subsequent analysis of differentially expressed genes and gene regulatory mechanisms. Currently, there are some statistical learning-based methods for clustering single-cell multi-omics data. These methods usually use non-negative matrix factorization (NMF), singular value decomposition and other technical means to map multi-omics data into a low-dimensional space, and then use K-means clustering, spectral clustering, community detection and other methods to identify cell subpopulations. For example, scMNMF uses NMF to integrate data dimensionality reduction and cell clustering, enabling effective cell type identification. scMLC constructs single-modal and cross-modal cell connectivity networks and then clusters cells using robust multi-community detection. GSTRPCA uses an irregular tensor decomposition model based on robust principal component analysis to preserve the original data structure and explore hidden relevant features in omics data. In recent years, deep learning-based clustering of single-cell multi-omics data has become a hot topic. These methods typically use omics-specific encoder networks to map the data into a low-dimensional space and then fuse multi-omics features to generate cell co-embeddings for cell clustering. For example, scMM uses two variational autoencoder networks to integrate scRNA-seq and scATAC-seq data in a low-dimensional space. The learned low-dimensional embeddings can be used for cell clustering and visualization. scMVP uses a multi-view deep generative model to learn cell co-embeddings. MultiVI encodes each omics data set into a modality-independent latent representation and then merges these representations into a joint latent space. In addition, some methods further integrate clustering-guided model optimization strategies to further aggregate cells of the same type in low-dimensional space and enhance the differences between different cell types. For example, scFPN uses a pyramid network structure to gradually fuse the same-level features of different omics, and uses a self-supervised optimization module to continuously optimize the cluster centers; scMIC integrates cell attribute information and potential structural relationships between cells from local and global levels, and uses multi-dimensional collaborative supervised clustering strategy learning to obtain more robust clustering results. However, when learning low-dimensional cell embeddings, the above methods do not decouple cross-omics invariant features related to the biological state of cells and omics technology-specific information, making it easy for random noise in the low-dimensional cell embeddings to interfere with the clustering results and affect the accuracy of cell subtype identification.In addition, when fusing single-cell multi-omics features, simple feature fusion strategies such as splicing and addition are difficult to adaptively capture the feature differences of different cell types, making it difficult for the fused features to fully and accurately present cell heterogeneity, which is not conducive to the identification of rare cell types. Summary of the Invention
[0004] In view of this, the present invention provides a single-cell multi-omics data clustering method based on diffuse autoencoders, which combines a bidirectional cross-attention mechanism with a diffuse autoencoder, aiming to decouple random noise and cell semantic features and improve the accuracy of cell clustering.
[0005] The technical solution adopted by the embodiment of the present invention to solve the technical problem is:
[0006] A single-cell multi-omics data clustering method based on diffuse autoencoders, including:
[0007] Step S1: Dimensionality reduction is performed on paired single-cell omics A and omics B data to obtain omics data with consistent dimensions. and Among them, principal component analysis PCA was used to process RNA type data, and latent semantic index LSI was used to process ATAC type data;
[0008] Step S2: A feature fusion module based on a bidirectional cross-attention mechanism is used to integrate the features of the two types of omics data to obtain the common semantic embedding z of the cells;
[0009] Step S3, based on and Generating random encoder features, and using two conditional diffusion models to reconstruct the original data based on z and the random encoder features, respectively, to obtain reconstructed data; wherein, the parameters of the conditional diffusion models are updated using a loss function to minimize the loss function, thereby obtaining an optimized common semantic embedding z';
[0010] In step S4, a K-means clustering algorithm is used to divide the cells into different cell subpopulations using the optimized common semantic embedding z'.
[0011] Step S2 includes: Step S21, using two semantic encoders to respectively convert the omics data and Coded as Z a and Z r ; The semantic compiler is a deep neural network composed of fully connected layers;
[0012] Step S22, using the feature fusion module based on the bidirectional cross attention mechanism to integrate Z a and Z r , we get the common semantic embedding Z of cells; the feature fusion process is as follows:
[0013] q r =W rq z r
[0014] k r =W rk z a
[0015] v r =W rv z a
[0016]
[0017] z1=A r v r
[0018] q a =W aq z a
[0019] k a =W ak z r
[0020] v a =W av z r
[0021]
[0022] z2=A a v a
[0023] z=z1+z2
[0024] Where q r and q a represents the query in the attention mechanism, k r and k a Represents the key in the attention mechanism, v r and v a Represents the value in the attention mechanism, W rq 、W rk 、W rv 、W aq 、W ak and W av represents the weight matrix of the linear network, C represents the feature dimension, h represents the number of heads in the attention mechanism, T represents the matrix transpose operation, and z is the common semantic embedding of the cell.
[0025] Step S3 includes: Step S31, omics data and Gradually add noise to generate random coding features and
[0026]
[0027] Where, and represents the random noise at time t, α t Control the noise weight at time t, represents the standard normal distribution. When T is large enough, and Approximately obeys the standard normal distribution;
[0028] Step S32, based on common semantic embedding z and random coding features and The generation process of the conditional diffusion model is expressed as:
[0029]
[0030] Where, and Respectively represent the time t Predicted values and Predicted value, and represents the inference distribution; is the conditional probability, indicating that given the current state and the latent variable z, derive the data of the previous time step probability; Indicates that in a given Under the conditions of and z, deduce the data of the previous time step probability;
[0031] Step S33: and Use noise prediction network ∈ θ (·) and ∈ λ (·) is parameterized:
[0032]
[0033] The expression of the loss function L is:
[0034]
[0035] Where, Indicates data Noisy data and the expectation at time step t, represents the addition of the network θ prediction to The noise in represents the addition of the network λ prediction to The noise in the
[0036] From the above technical solution, it can be seen that the single-cell multi-omics data clustering method based on diffuse autoencoder provided by the embodiment of the present invention first performs dimensionality reduction processing on the paired single-cell omics A and omics B data to obtain omics data with consistent dimensions. and Among them, principal component analysis PCA is used to process RNA type data, and latent semantic index LSI is used to process ATAC type data; a feature fusion module based on a bidirectional cross-attention mechanism is used to integrate the features of the two types of omics data to obtain the common semantic embedding z of the cells; based on and Generate random encoder features, use two conditional diffusion models to reconstruct the original data based on z and random encoder features respectively, and obtain reconstructed data; use the loss function to update the parameters of the conditional diffusion model to minimize the loss function and obtain the optimized common semantic embedding z'; use the K-means clustering algorithm to divide the cells into different cell subpopulations using the optimized common semantic embedding z'. The present invention combines the bidirectional cross-attention mechanism with the conditional diffusion model, uses decoupling learning to separate cell semantic features and random noise, and uses cross-omics invariant features related to the cell biological state to cluster cells, so that different cell subtypes can be accurately identified. The proposed bidirectional cross-attention mechanism can effectively fuse the features of different omics and present the cell topology in a low-dimensional space. The proposed conditional diffusion model progressively reconstructs the original data based on random coding features and semantic features, thereby achieving the separation of semantic features and omics-specific information. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 This is a framework diagram of the single-cell multi-omics data clustering method based on diffuse autoencoders of the present invention.
[0038] Figure 2 Schematic diagram of the bidirectional cross-attention feature fusion module of the present invention. DETAILED DESCRIPTION
[0039] The technical solutions and technical effects of the present invention are further described in detail below with reference to the accompanying drawings of the present invention.
[0040] To address the shortcomings of existing technologies, this paper proposes a novel framework for clustering single-cell multi-omics data. This framework combines a bidirectional cross-attention mechanism with a diffuse autoencoder, aiming to decouple random noise from cellular semantic features and improve the accuracy of cell clustering. By introducing a bidirectional cross-attention mechanism, the paper effectively captures the correlations between intra- and inter-omics features, adaptively integrating multi-omics features of cells, mitigating the impact of random noise on clustering results, and accurately presenting cellular topology in a low-dimensional space.
[0041] The method includes the following steps: (1) to address the problems of high computational complexity and limited recognition accuracy caused by the high dimensionality of single-cell data, the original data is first mapped to a low-dimensional space using dimensionality reduction methods such as principal component analysis and latent semantic indexing; (2) for each omics data, a semantic encoder is used to obtain a low-dimensional embedding of the cell, and a bidirectional cross-attention mechanism is used to fuse the features of different omics to obtain a common semantic embedding of each cell; (3) a random encoder is introduced into the conditional diffusion model, and the original data is progressively reconstructed based on the random encoding features and the cell semantic features, and the various modules of the model are jointly trained by minimizing the loss function; (4) after the model converges, the low-dimensional embedding of the cell is obtained through the semantic encoder, and then the K-means clustering algorithm is used to divide the cells into different cell subpopulations. The present invention uses a semantic encoder and a random encoder to decouple the cell biological state-related features and the omics-specific features, and with the help of the data distribution progressive learning ability of the diffusion model, the original data is reconstructed through the decoder, and the cell topology structure is accurately captured in the low-dimensional space, thereby achieving accurate recognition of cell subpopulations.
[0042] refer to Figure 1 As shown, the present invention provides a single-cell multi-omics data clustering method based on diffuse autoencoders, which specifically includes:
[0043] Step S1: Dimensionality reduction is performed on paired single-cell omics A and omics B data to obtain omics data with consistent dimensions. and Among them, principal component analysis (PCA) was used to process RNA type data, and latent semantic index (LSI) was used to process ATAC type data.
[0044] Due to the high dimensionality of single-cell sequencing data, cluster analysis based on the original data faces problems such as high time complexity and insufficient clustering accuracy. Therefore, this technology first uses data dimensionality reduction methods such as PCA and LSI to map omics A and omics B into low-dimensional space respectively, requiring that the dimensions of the two omics data remain consistent after dimensionality reduction.
[0045] In step S2, a feature fusion module based on a bidirectional cross-attention mechanism is used to integrate the features of the two types of omics data to obtain the common semantic embedding z of the cells. The specific implementation process of the bidirectional cross-attention feature fusion module is as follows:
[0046] Step S21: Use two semantic encoders to convert the omics data into and Coded as Z a and Z r ; The semantic compiler is a deep neural network composed of fully connected layers;
[0047] Step S22, using the feature fusion module based on the bidirectional cross attention mechanism to integrate Z a and Z r , get the common semantic embedding Z of cells; reference Figure 2 As shown in Figure 2, the feature fusion process is specifically as follows:
[0048] q r =W rq z r (1)
[0049] k r =W rk z a (2)
[0050] v r =W rv z a (3)
[0051]
[0052] z1=A r v r (5)
[0053] q a =W aq z a (6)
[0054] k a =W ak z r (7)
[0055] v a =W av z r (8)
[0056]
[0057] z2=A a v a (10)
[0058] z=z1+z2 (11)
[0059] Where q r and q a represents the query in the attention mechanism, k r and k a Represents the key in the attention mechanism, v r and v a Represents the value in the attention mechanism, W rq 、W rk 、W rv 、W aq 、W ak and W av (After model initialization, it is continuously optimized during training.) represents the weight matrix of the linear network, C represents the feature dimension, h represents the number of heads in the attention mechanism, T represents the matrix transpose operation, and z represents the common semantic embedding of the cells. Single-cell omics data contains both semantic information representing cell states and omics-specific information unrelated to biological processes. Therefore, decoupling these two types of features and using semantic features to cluster cells is an important foundation for improving the accuracy of cell subtype identification.
[0060] Step S3, based on and Generate random encoder features, use two conditional diffusion models to reconstruct the original data based on z and random encoder features respectively, and obtain reconstructed data; use the loss function to update the parameters of the conditional diffusion model to minimize the loss function, and obtain the optimized common semantic embedding z'; thereby decoupling the cell semantic features and random noise, allowing the model to more accurately identify cell topology. The specific implementation is as follows:
[0061] Step S31, omics data and Gradually add noise to generate random coding features and
[0062]
[0063] Where, and represents the random noise at time t, α t Control the noise weight at time t, Represents the standard normal distribution. When T is large enough, it is considered and Approximately obeys the standard normal distribution;
[0064] Step S32, based on common semantic embedding z and random coding features and The generation process of the conditional diffusion model is expressed as:
[0065]
[0066] Where, and Respectively represent the time t Predicted values and Predicted value, and Represents the inference distribution; the conditional diffusion model is used to decouple the semantic features representing the cell state from random noise;
[0067] Step S33: and Use noise prediction network ∈ θ (·) and ∈ λ (·) is parameterized:
[0068]
[0069] is the conditional probability, indicating that given the current state and the latent variable z, derive the data of the previous time step probability; Same thing.
[0070] The parameter update of the conditional diffusion model is achieved by minimizing the following loss function. The expression of the loss function L is:
[0071]
[0072] Where, Indicates data Noisy data and the expectation at time step t, represents the addition of the network θ prediction to The noise in represents the addition of the network λ prediction to The noise in the
[0073] The conditional diffusion model uses random encoder features and cell semantic features as conditions to progressively reconstruct the original data, enabling the semantic encoder to better learn cell semantic features.
[0074] In step S4, after the model converges, the cells are clustered using the semantic features z' of each cell. Specifically, two semantic encoders are used to derive a common semantic embedding z' from the omics A and omics B data for each cell. The cells are then clustered using the K-means clustering algorithm to divide the cells into distinct subpopulations. For more information on the clustering process, refer to existing technical solutions, such as the "Deep K-means clustering" method in the "Method" section of the article "Clustering of single-cell multi-omics data with a multimodal deeplearning method."
[0075] According to an embodiment disclosed in the present invention, the present invention also discloses an electronic device, a readable storage medium, and a computer program product. The electronic device is intended to represent various forms of digital computers, and the device includes a computing unit that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) or a computer program loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for the operation of the device can also be stored. The computing unit, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.
[0076] Many components in a device are connected to the I / O interface, including: input units, such as a keyboard and mouse; output units, such as various types of displays and speakers; storage units, such as magnetic disks and optical disks; and communication units, such as network cards, modems, and wireless communication transceivers. The communication unit allows the device to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks.
[0077] The computing unit can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of computing units include, but are not limited to, central processing units (CPUs), graphics processing units (GPUs), various dedicated artificial intelligence (AI) computing chips, various computing units for running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit performs the various methods and processes described above, such as a single-cell multi-omics data clustering method based on a diffusion autoencoder. For example, in some embodiments, a single-cell multi-omics data clustering method based on a diffusion autoencoder can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed on the device via a ROM and / or a communication unit. When the computer program is loaded into RAM and executed by the computing unit, one or more steps of the single-cell multi-omics data clustering method based on a diffusion autoencoder described above can be performed. Alternatively, in other embodiments, the computing unit can be configured to perform a single-cell multi-omics data clustering method based on a diffusion autoencoder by any other appropriate means (e.g., by means of firmware).
[0078] The program code for implementing the method disclosed in the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0079] A machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium.
[0080] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0081] The above disclosure is only a preferred embodiment of the present invention, and it is certainly not intended to limit the scope of the present invention. A person skilled in the art can understand that all or part of the processes of the above embodiment and equivalent changes made in accordance with the claims of the present invention are still within the scope of the invention.
Claims
1. A single-cell multi-omics data clustering method based on diffuse autoencoders, characterized in that: include: Step S1: Dimensionality reduction is performed on paired single-cell omics A and omics B data to obtain omics data with consistent dimensions. and Among them, principal component analysis PCA was used to process RNA type data, and latent semantic index LSI was used to process ATAC type data; Step S2: A feature fusion module based on a bidirectional cross-attention mechanism is used to integrate the features of the two types of omics data to obtain the common semantic embedding z of the cells; Step S3, based on and Generating random encoder features, and using two conditional diffusion models to reconstruct the original data based on z and the random encoder features, respectively, to obtain reconstructed data; wherein, the parameters of the conditional diffusion models are updated using a loss function to minimize the loss function, thereby obtaining an optimized common semantic embedding z'; In step S4, a K-means clustering algorithm is used to divide the cells into different cell subpopulations using the optimized common semantic embedding z'.
2. The single-cell multi-omics data clustering method based on diffuse autoencoders according to claim 1, characterized in that: The step S2 comprises: Step S21: Use two semantic encoders to convert the omics data into and Coded as Z a and Z r ; The semantic compiler is a deep neural network composed of fully connected layers; Step S22, using the feature fusion module based on the bidirectional cross attention mechanism to integrate Z a and Z r , we get the common semantic embedding Z of cells; the feature fusion process is as follows: q r =W rq z r k r=W rk z a v r =W rv z a z1=A r in r q a =W aq z a k a =In ak With r v a =W av z r z2=A a v a z=z1+z2 Where q r and q a represents the query in the attention mechanism, k r and k a Represents the key in the attention mechanism, v r and v a Represents the value in the attention mechanism, W rq 、W rk 、W rv 、W aq 、W ak and W av represents the weight matrix of the linear network, C represents the feature dimension, h represents the number of heads in the attention mechanism, T represents the matrix transpose operation, and z is the common semantic embedding of the cell.
3. The single-cell multi-omics data clustering method based on diffuse autoencoders according to claim 2, characterized in that: The step S3 comprises: Step S31, omics data and Gradually add noise to generate random coding features and Where, and represents the random noise at time t, α t Control the noise weight at time t, represents the standard normal distribution. When T is large enough, and Approximately obeys the standard normal distribution; Step S32, based on common semantic embedding z and random coding features and The generation process of the conditional diffusion model is expressed as: Where, and Respectively represent the time t Predicted values and Predicted value, and represents the inference distribution; is the conditional probability, indicating that given the current state and the latent variable z, derive the data of the previous time step probability; Indicates that in a given Under the conditions of and z, deduce the data of the previous time step probability; Step S33: and Use noise prediction network ∈ θ (·) and ∈ λ (·) is parameterized:
4. The single-cell multi-omics data clustering method based on diffuse autoencoders according to claim 3, characterized in that: The expression of the loss function L is: Where, Indicates data Noisy data and the expectation at time step t, represents the addition of the network θ prediction to The noise in represents the addition of the network λ prediction to The noise in the 5. An electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 4.
6. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-4.
7. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 4.
Citation Information
Cited By
Incomplete multi-omics cancer subtype data clustering method
CN121034428A