Single-cell multi-omics clustering method based on multi-subspace contrast learning

By employing multi-subspace contrastive learning and a specific encoder network, the problem of limited feature extraction in single-cell multi-omics data integration is solved, achieving more comprehensive data representation and cluster analysis results.

CN120977402AActive Publication Date: 2025-11-18ANHUI UNIV +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511493626.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2025-11-18
Estimated Expiration
2045-10-20

AI Technical Summary

Technical Problem

Existing single-cell multi-omics data integration methods are difficult to effectively integrate unique and shared features between different omics, and ignore the sparsity and noise of the data, resulting in limited feature extraction.

Method used

We employ a multi-subspace contrastive learning approach, which extracts individual and common features through an autoencoder. We enhance robustness by utilizing multi-subspace contrastive learning and attention mechanisms, and optimize feature representation and clustering by combining a decoder network with zero-inflated negative binomial regression and Bernoulli distribution.

Benefits of technology

It enables more comprehensive and informative representation of single-cell multi-omics data, improves the accuracy and reliability of cluster analysis, and captures complementary information and intrinsic connections between omics data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977402A_ABST
    Figure CN120977402A_ABST
Patent Text Reader

Abstract

The invention discloses a single-cell multi-omics clustering method based on multi-subspace contrastive learning. The method comprises the following steps: extracting important personality characteristics of transcriptomes and chromatin accessibility of different omics based on an automatic encoder; multi-subspace contrastive learning extraction is adopted to fully mine general characteristics of accessibility of transcriptome and chromatin, and deep association among omics data is effectively revealed; an automatic encoder for specific omics chromatin accessibility and transcriptome is adopted to reconstruct data so as to better denoise and reconstruct features conforming to omics information; personality information of a specific omics and common information of a plurality of omics are fused, and co-embedding which is more beneficial to clustering is obtained. According to the method, the defect of single-space comparison can be made up, and a more comprehensive expression with a larger amount of information is constructed for fusion and clustering of single-cell multi-omics data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of single-cell multi-omics data analysis technology, and more specifically to a single-cell multi-omics clustering method based on multi-subspace contrastive learning. Background Technology

[0002] Single-cell multi-omics sequencing technologies have developed rapidly, such as sci-CAR, SNARE-seq, Paired-seq, and 10XGenomics Multiome. These sequencing methods simultaneously sequence gene expression and chromatin accessibility within the same cell, providing valuable data resources for studying single-cell multi-omics clustering methods. However, the integration of single-cell multi-omics data remains a significant challenge. This is mainly due to the inherent high sparsity of the data, the huge heterogeneity caused by measurement noise, the significant dimensionality difference (approximately 10-20 times) between single-cell chromatin accessibility sequencing (scATAC-seq) and single-cell RNA sequencing (scRNA-seq) data, and the imbalance of information between omics disciplines. These factors collectively bring enormous computational and analytical challenges.

[0003] The integration of single-cell multi-omics data can facilitate the study of complex biological information. Researchers have been developing single-cell multi-omics integration methods using machine learning and bioanalysis techniques. Some methods are based on nonnegative matrix factorization or principal component analysis to integrate single-cell multi-omics data and address cellular heterogeneity. Meanwhile, with the development of deep learning and its ability to extract expressive features, deep learning methods have gradually become the mainstream technology for single-cell data analysis. Furthermore, graph-based methods use graphs to represent the relationships between different cells, where cells are considered nodes. Integration is then achieved by manipulating this graph. This approach emphasizes neighborhood structure and sometimes aggregates information between neighborhoods, resulting in additional robustness to noise.

[0004] While the methods described above are effective, they still have some problems. First, most of them focus on features shared between different omics, but ignore the unique features of each omics, and cannot better integrate biological features at different levels to learn more distinctive cell clusters. Second, most of them ignore the high-dimensional sparsity and noise inherent in multi-omics data, which limits feature extraction. Summary of the Invention

[0005] The present invention provides a single-cell multi-omics clustering method based on multi-subspace contrastive learning for constructing a more comprehensive and informative representation for the fusion and clustering of single-cell multi-omics data, which can at least solve one of the above-mentioned technical problems.

[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: A single-cell multi-omics clustering method based on multi-subspace contrastive learning includes the following steps: S1. Based on the autoencoder, important individual features are extracted from the transcriptome and chromatin accessibility of different omics. S2. Multi-subspace contrastive learning is used to extract common features of transcriptome and chromatin accessibility, effectively revealing the deep-seated relationships between omics data. S3: Employ autoencoders targeting specific omics chromatin accessibility and transcriptomes to reconstruct data for better noise reduction and reconstruction of features consistent with omics information; S4: The specific information of the specific omics in S3 and the common information of multiple omics in S2 are fused to obtain a co-embedding that is more conducive to clustering.

[0007] Furthermore, in S1, an autoencoder is used to map single-cell multi-omics data to their respective nonlinear embedding spaces, preserving individual characteristics and enhancing robustness to noise and outliers. Specifically, let's set , These are standardized scRNA-seq and scATAC-seq data, respectively, where N is the number of samples, Dr and Da are the number of features, and two independent encoders are used. and To learn the corresponding d-dimensional feature representation { , } Where d is the dimension of the embedding space, the expressions are as follows: ; In these two formulas, It represents the potential low-dimensional features of cells and genes in scRNA-seq data. It is a potential low-dimensional feature representation between cells and peaks in scATAC-seq data.

[0008] Furthermore, S2 further includes: S2.1 Define two views u and r. From a feature perspective, input views u and r into multiple subspace encoders MSEncoder respectively to generate node embeddings for multiple subspaces. This process is represented as follows:

[0009]

[0010] In these two equations, l represents the l-th subspace, and p represents the number of subspaces. , and Let be the node embeddings of the i-th subspace in views u and r, respectively, where d represents the dimension of the embedding space. Views U and R are feature representations obtained through multiple subspace encoders, and are common features of the same omics in views u and r, respectively. It represents the potential low-dimensional features of cells and genes in scRNA-seq data. It is a potential low-dimensional feature representation between cells and peaks in scATAC-seq data; The subspace encoder MSEncoder maps the input feature matrix X to a keyword embedding K and a value embedding V using two independent linear projections, with the corresponding expressions as follows:

[0011]

[0012] In these two formulas, , , representing different projection matrices of the key and value, respectively; S2.2 To make the features captured by multi-subspace contrastive learning more diverse, we should focus on more important node information, make full use of the degree information obtained by the graph structure, and integrate the graph structure information into the attention mechanism to enhance the discriminativeness of the feature representation. The initialization expression for the degree embedding matrix S is: ; In this formula, Let D represent the projection matrix corresponding to the degree, where D is the degree matrix; The expression for the attention weight matrix is ​​as follows: ; In this formula, Att is the output of the entire expression, representing the "attention weight matrix", the Softmax(˙) function is a function that "compresses" a vector of arbitrary real values ​​into a probability distribution vector, Q is the query matrix, and α is a learnable or settable hyperparameter. To compute the embedding representation of a node within each subspace, the value embedding V is broadcast to each subspace as follows: We obtain the node embeddings of p subspaces, expressed as: ; ; S2.3, through calculation and The mutual information maximization between node embeddings across subspaces and the mutual information minimization within subspaces are used as their respective subspace contrastive loss functions. S2.4 The final objective function of subspace contrastive learning is a weighted sum of the mutual information maximization objective and the mutual information minimization objective, which can be formally expressed as:

[0013] In this formula, L contra Let represent the subspace contrast loss, θ' and φ be the encoder parameters in the view u and view r branches respectively, γ be the weighting coefficient balancing the influence of targets within and between spaces, MI(˙) be the mutual information, and k represent the k-th subspace. and Let represent the node embeddings of the k-th and l-th subspaces in view u, respectively. This represents the node embedding in the k-th subspace of view r; S2.5. Train and concatenate views U and R to obtain the latent feature representation. .

[0014] Furthermore, in S2.3, the feature representations of views u and r are output to p subspaces, i.e. To optimize the subspace encoder MSEncoder and ensure that views U and R both possess diverse feature representations of common features of the same omics, a pair of loss functions is used. These functions focus on maximizing the mutual information between subspace representations and minimizing the mutual information between different subspace representations. For the mutual information between subspace representations, the objective is formulated as follows: ; In this formula, θ' and φ are the encoder parameters in the branches of view u and view r, respectively, k represents the k-th subspace, p represents the number of subspaces, and MI(˙) is the mutual information. This represents the node embedding in the k-th subspace of view u. This represents the node embedding in the k-th subspace of view r.

[0015] Furthermore, due to and The mutual information (MI) is unknown, making it impossible to calculate. To estimate MI, a sampling method is used instead of maximizing mutual information to maximize the lower bound estimate of MI, and a contrast log-ratio upper bound (CLUB) is employed to minimize the upper bound estimate of MI. Further: Regarding the estimation of the lower bound for maximizing mutual information (MI), the Jensen-Shannon estimation is used instead of the Kullback-Leibler KL divergence estimation. The calculation formula is as follows:

[0016] In this formula, It is a discriminator that takes the representations of two views as input and scores the consistency between them. Here, the discriminator is simply instantiated as the dot product between two representations, i.e. SP(·) is the activation function. For the joint distribution expectation of two views within the same subspace k, The product of the independent distribution expectations of two views within the same subspace k; In addition to the objectives within each space, constraints are imposed on the relationships between different subspaces within the same view to increase spatial diversity. To this end, an optimization objective is proposed to minimize mutual information. This objective can increase the key features within view u, and because... Therefore, the expression for the optimization objective is: ; In this formula, l represents the l-th subspace. This represents the node embedding in the l-th subspace of view u; Regarding the estimation of the upper limit for minimizing mutual information (MI), the formula for calculating the upper limit of the logarithmic ratio (CLUB) is as follows: ; In this formula, I CLUB Let E represent the expectation of CLUB, P(x,y) represent the joint distribution of x and y, P(x) and P(y) represent the marginal distributions of x and y respectively, log is the natural logarithm, and P(y|x) is the conditional probability of y given x. The key to calculating CLUB is the distribution... In this case, it is assumed that each dimension of x and y is correlated; give Assuming that each dimension of y depends only on the corresponding dimension of x, then using and Substituting x and y respectively, the relevant expression is:

[0017] ; In this formula, P represents the probability density. Represents the covariance matrix. ˙) represents a multivariate Gaussian distribution, and I represents the identity matrix. This represents the variance shared across all dimensions. and They are and The node embedding in the i-th subspace, iThe symbol for the product of all elements or subspaces i represents the product of a high-dimensional Gaussian distribution into a product of independent one-dimensional Gaussian distributions.

[0018] Furthermore, in S3, a decoder network based on the zero-inflated negative binomial regression model ZINB explores the global probability structure of scRNA-seq data. The ZINB distribution is derived from the mean of the negative binomial distribution. Deviation parameters and coefficients used to describe the probability of interpolated events. The definition and expression are:

[0019]

[0020] In these two formulas, Indicates the parameter is u r The negative binomial probability mass function of x and θ. r It is a vector from the raw scRNA-seq data, representing a random variable taking non-negative integer values, u r Let θ be the mean parameter of the distribution, and θ be the deviation parameter. It is the gamma function, ZINB(˙) represents the zero-inflation negative binomial distribution, and π represents the zero-inflation probability. It is x r Indicator function when =0; Specifically, the ZINB-based decoder estimates parameters based on the co-embedding representation Z through three different fully connected layers. : ; ; ; Of these three formulas, yes In matrix form, The proportion of "extra zeros" is determined by the zero-inflation probability parameter. It is the negative binomial mean parameter, representing the expected count of the non-zero portion, and is usually output by the decoder network. It is the negative binomial dispersion parameter, which controls the variance and is usually output by the decoder network. The activation function is sigmoid(·), so the exit probability is between 0 and 1. Furthermore, since the mean and discrete parameters are non-negative, the exponential function exp is chosen as the activation function. and activation function, It is a decoder with a fully connected layer. , and θ These are three learnable parameter matrices; Unlike traditional autoencoders based on mean squared error loss, the loss function of the ZINB-based decoder network is the negative logarithm of the ZINB likelihood, expressed as: ; In this formula, The zero-inflated negative binomial distribution is used to model sparse, overly discrete single-cell counting data. -log(˙) represents taking the probability density and then taking the negative log to obtain the negative log-likelihood loss. The smaller this value is, the better. Considering the extremely sparse and almost binary nature of scATAC-seq data, a decoder network based on the Bernoulli distribution (Ber) is chosen to model the scATAC-seq data, expressed as: ; In this formula, These are vectors from the original scATAC-seq data. It is Ber's average parameter; The decoder based on the Bernoulli distribution (Ber) estimates the input latent variable Z' through a fully connected layer using sigmoid (·) as the activation function. The expression is: ; In this formula, yes a In matrix form, sigmoid(·) is an activation function that compresses real values ​​to the range between 0 and 1. It is a weight parameter matrix. This represents the mapping of the a-th decoder; Finally, the decoder based on the Bernoulli distribution (Ber) is optimized using cross-entropy loss, expressed as follows: ; In this formula, L Ber This represents the Bernoulli loss, and CE(·) is the cross-entropy function. These are Bernoulli variables that are actually observed.

[0021] Furthermore, in S4, two latent representations are learned. and They encode the characteristics of individuals in different omics data, and also represent them through latent features. To capture the commonalities among these features, and in this process, a scale parameter is introduced. and For these two sets of potential representations and To achieve efficient aggregation between the two, we perform element-wise weighted summation, which can be formalized as follows: ; In this formula, Z represents co-embedding representation.

[0022] The accuracy and reliability of cluster analysis can be improved by generating a more discriminative co-embedding representation Z.

[0023] Furthermore, in S4, in order to pursue a more discriminative and informative co-embedding representation Z, the total loss is... Combining individual and common information from multiple omics datasets, and unifying the input targets for scRNA-seq and scATAC-seq data, the total loss expression is as follows:

[0024] In this formula, This indicates taking the minimum value. These represent network parameters, where α1, α2, and γ all represent weight coefficients. The zero-inflated negative binomial distribution is used to model sparse, overly discrete single-cell counting data. The proportion of "extra zeros" is determined by the zero-inflation probability parameter. It is the negative binomial mean parameter, representing the expected count of the non-zero portion, and is usually output by the decoder network. This is the negative binomial dispersion parameter, which controls the variance. It is usually output by the decoder network. -log(˙) indicates that the probability density is taken first, and then the negative logarithm is taken to obtain the negative log-likelihood loss. The smaller this value, the better. CE(·) is the cross-entropy function. For the actual observed Bernoulli variables, is u a The matrix form, u a The mean parameters of the Bernoulli distribution Ber are θ' and θ'. These are the encoder parameters in the branches of view u and view r, respectively. MI(˙) is the mutual information, used to measure the degree of nonlinear dependence between two random variables. k represents the k-th subspace, and p represents the number of subspaces. and Let represent the node embeddings of the k-th and l-th subspaces in view u, respectively. This represents the node embedding in the k-th subspace of view r; The total loss function is continuously minimized and trained for 500 epochs for convergence optimization. Individual feature representations and shared feature representations are learned from multi-omics data and then merged into the co-embedding representation Z for clustering information.

[0025] The beneficial effects of this invention are reflected in: This invention, through multi-subspace contrastive learning, can capture complementary information between omics data in different feature spaces, as well as the intrinsic connections between omics data within the same feature space, thereby making up for the shortcomings of single-space contrastive learning and constructing a more comprehensive and information-rich representation for the fusion and clustering of single-cell multi-omics data. Attached Figure Description

[0026] The accompanying drawings, which are provided to further illustrate this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application.

[0027] Figure 1 This is a schematic diagram of the overall process of the single-cell multi-omics clustering method according to an embodiment of the present invention.

[0028] Figure 2 This is a detailed flowchart of the single-cell multi-omics clustering method according to an embodiment of the present invention.

[0029] Figure 3 This is an embodiment of the present invention. Figure 2 Enlarged diagram of the working principle of the neutron space encoder MSEncoder.

[0030] Figure 4 This is an embodiment of the present invention. Figure 2 A magnified diagram illustrating the working principle of Node Embedding.

[0031] Figure 5 This is the clustering UMAP diagram of the scMVAE-PoE method on the CellMix dataset.

[0032] Figure 6 This is the clustering UMAP graph of the scMVAE-NN method on the CellMix dataset.

[0033] Figure 7 This is the clustering UMAP graph of the scMVAE-Direct method on the CellMix dataset.

[0034] Figure 8 This is the UMAP clustering diagram of the DCCA method on the CellMix dataset.

[0035] Figure 9 This is the UMAP clustering graph of the scMCs method on the CellMix dataset.

[0036] Figure 10 This is the UMAP clustering diagram of the scEMC method on the CellMix dataset.

[0037] Figure 11 This is the clustering UMAP graph of the scMSCL method of this invention on the CellMix dataset.

[0038] Figure 12 This is a structural block diagram of a computer device according to an embodiment of the present invention. Detailed Implementation

[0039] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0040] It should be noted that the meaning of "and / or" throughout the text includes three parallel solutions. Taking "A and / or B" as an example, it includes solution A, solution B, or a solution that simultaneously satisfies A and B. Furthermore, "multiple" refers to two or more. Additionally, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.

[0041] See Figure 1 This invention provides a single-cell multi-omics clustering method based on multi-subspace contrastive learning, comprising the following steps: S1. Based on the autoencoder, important individual features are extracted from the transcriptome and chromatin accessibility of different omics. S2. Multi-subspace contrastive learning is used to extract common features of transcriptome and chromatin accessibility, effectively revealing the deep-seated relationships between omics data. S3: Employ autoencoders targeting specific omics chromatin accessibility and transcriptomes to reconstruct data for better noise reduction and reconstruction of features consistent with omics information; S4: The specific information of the specific omics in S3 and the common information of multiple omics in S2 are fused to obtain a co-embedding that is more conducive to clustering.

[0042] See Figures 2-4 In this embodiment, in S1, the autoencoder, as a classic neural network, can map high-dimensional data to a low-dimensional representation space, effectively ignoring noise and outliers. Based on this characteristic, the autoencoder is used to map single-cell multi-omics data to their respective nonlinear embedding spaces, preserving individual characteristics and enhancing robustness to noise and outliers. Specifically, let's set , These are standardized scRNA-seq and scATAC-seq data, respectively, where N is the number of samples, Dr and Da are the number of features, and two independent encoders are used. and To learn the corresponding d-dimensional feature representation { , } Where d is the dimension of the embedding space, the expressions are as follows: ; In these two formulas, It represents the potential low-dimensional features of cells and genes in scRNA-seq data. It is a potential low-dimensional feature representation between cells and peaks in scATAC-seq data.

[0043] See Figures 2-4 Attention mechanisms may cause models to overemphasize the individual characteristics or noise of each omics, which is detrimental to data fusion and cell clustering. Furthermore, focusing solely on individual characteristics makes it difficult to comprehensively characterize the complementarity between omics, while information shared between omics can reflect their commonalities, which is crucial for achieving high-quality, consistent clustering. Therefore, in this embodiment, S2 further includes: S2.1 Define two views u and r. Unlike the data transformation in traditional contrastive learning, from a feature perspective, views u and r are input into multiple subspace encoders MSEncoder to generate node embeddings for multiple subspaces. This process is represented as follows:

[0044]

[0045] In these two equations, l represents the l-th subspace, and p represents the number of subspaces. , and Let be the node embeddings of the i-th subspace in views u and r, respectively, where d represents the dimension of the embedding space. Views U and R are feature representations obtained through multiple subspace encoders, and are common features of the same omics in views u and r, respectively. It represents the potential low-dimensional features of cells and genes in scRNA-seq data. It is a potential low-dimensional feature representation between cells and peaks in scATAC-seq data; The subspace encoder MSEncoder, as a self-attention-based multi-feature generator, maps the input feature matrix X to a keyword embedding K and a value embedding V using two independent linear projections, with the corresponding expressions as follows:

[0046]

[0047] In these two formulas, , , representing different projection matrices of the key and value, respectively; S2.2 To make the features captured by multi-subspace contrastive learning more diverse, we should focus on more important node information, make full use of the degree information obtained by the graph structure, and integrate the graph structure information into the attention mechanism to enhance the discriminativeness of the feature representation. The initialization expression for the degree embedding matrix S is: ; In this formula, Let D represent the projection matrix corresponding to the degree, where D is the degree matrix; The expression for the attention weight matrix is ​​as follows: ; In this formula, Att is the output of the entire expression, representing the "attention weight matrix", the Softmax(˙) function is a function that "compresses" a vector of arbitrary real values ​​into a probability distribution vector, Q is the query matrix, and α is a learnable or settable hyperparameter. To compute the embedding representation of a node within each subspace, the value embedding V is broadcast to each subspace as follows: We obtain the node embeddings of p subspaces, expressed as: ; ; S2.3, through calculation and The mutual information maximization between node embeddings across subspaces and the mutual information minimization within subspaces are used as their respective subspace contrastive loss functions. S2.4 The final objective function of subspace contrastive learning is a weighted sum of the mutual information maximization objective and the mutual information minimization objective, which can be formally expressed as:

[0048] In this formula, L contraLet represent the subspace contrast loss, θ' and φ be the encoder parameters in the view u and view r branches respectively, γ be the weighting coefficient balancing the influence of targets within and between spaces, MI(˙) be the mutual information, and k represent the k-th subspace. and Let represent the node embeddings of the k-th and l-th subspaces in view u, respectively. This represents the node embedding in the k-th subspace of view r; S2.5. Train and concatenate views U and R to obtain the latent feature representation. .

[0049] See Figures 2-4 In this embodiment, in step S2.3, the feature representations of views u and r are output to p subspaces, i.e. To optimize the subspace encoder MSEncoder and ensure that views U and R both possess diverse feature representations of common features of the same omics, a pair of loss functions is used. These functions focus on maximizing the mutual information between subspace representations and minimizing the mutual information between different subspace representations. For the mutual information between subspace representations, the objective is formulated as follows: ; In this formula, θ' and φ are the encoder parameters in the branches of view u and view r, respectively, k represents the k-th subspace, p represents the number of subspaces, and MI(˙) is the mutual information. This represents the node embedding in the k-th subspace of view u. This represents the node embedding in the k-th subspace of view r.

[0050] See Figures 2-4 In this embodiment, because and The mutual information (MI) is unknown, making it impossible to calculate. To estimate MI, a sampling method is used instead of maximizing mutual information to maximize the lower bound estimate of MI, and a contrast log-ratio upper bound (CLUB) is employed to minimize the upper bound estimate of MI. Further: Regarding the estimation of the lower bound for maximizing mutual information (MI), the Jensen-Shannon estimation is used instead of the Kullback-Leibler KL divergence estimation. The calculation formula is as follows:

[0051] In this formula, It is a discriminator that takes the representations of two views as input and scores the consistency between them. Here, the discriminator is simply instantiated as the dot product between two representations, i.e. SP(·) is the activation function. For the joint distribution expectation of two views within the same subspace k, The product of the independent distribution expectations of two views within the same subspace k; In addition to the objectives within each space, constraints are imposed on the relationships between different subspaces within the same view to increase spatial diversity. To this end, an optimization objective is proposed to minimize mutual information. This objective can increase the key features within view u, and because... Therefore, the expression for the optimization objective is: ; In this formula, l represents the l-th subspace. This represents the node embedding in the l-th subspace of view u; Regarding the estimation of the upper limit for minimizing mutual information (MI), the formula for calculating the upper limit of the logarithmic ratio (CLUB) is as follows: ; In this formula, I CLUB Let E represent the expectation of CLUB, P(x,y) represent the joint distribution of x and y, P(x) and P(y) represent the marginal distributions of x and y respectively, log is the natural logarithm, and P(y|x) is the conditional probability of y given x. The key to calculating CLUB is the distribution... In this case, it is assumed that each dimension of x and y is correlated; give Assuming that each dimension of y depends only on the corresponding dimension of x, then using and Substituting x and y respectively, the relevant expression is:

[0052] ; In this formula, P represents the probability density. Represents the covariance matrix. ˙) represents a multivariate Gaussian distribution, and I represents the identity matrix. This represents the variance shared across all dimensions. and They are and The node embedding in the i-th subspace, i The symbol for the product of all elements or subspaces i represents the product of a high-dimensional Gaussian distribution into a product of independent one-dimensional Gaussian distributions.

[0053] See Figures 2-4Previous studies have shown that scRNA-seq data typically exhibits characteristics such as dispersion, variance greater than the mean, and high sparsity. Nevertheless, some research reports indicate that the zero-inflated negative binomial (ZINB) distribution can effectively describe these data characteristics. Therefore, in this embodiment, in step S3, a decoder network based on the zero-inflated negative binomial regression model ZINB explores the global probability structure of scRNA-seq data. The ZINB distribution is derived from the mean of the negative binomial distribution. Deviation parameters and coefficients used to describe the probability of interpolated events. The definition and expression are:

[0054]

[0055] In these two formulas, Indicates the parameter is u r The negative binomial probability mass function of x and θ. r It is a vector from the raw scRNA-seq data, representing a random variable taking non-negative integer values, u r Let θ be the mean parameter of the distribution, and θ be the deviation parameter. It is the gamma function, ZINB(˙) represents the zero-inflation negative binomial distribution, and π represents the zero-inflation probability. It is x r Indicator function when =0; Specifically, the ZINB-based decoder estimates parameters based on the co-embedding representation Z through three different fully connected layers. : ; ; ; Of these three formulas, yes In matrix form, The proportion of "extra zeros" is determined by the zero-inflation probability parameter. It is the negative binomial mean parameter, representing the expected count of the non-zero portion, and is usually output by the decoder network. It is the negative binomial dispersion parameter, which controls the variance and is usually output by the decoder network. The activation function is sigmoid(·), so the exit probability is between 0 and 1. Furthermore, since the mean and discrete parameters are non-negative, the exponential function exp is chosen as the activation function. and activation function, It is a decoder with a fully connected layer. , and θ These are three learnable parameter matrices; Unlike traditional autoencoders based on mean squared error loss, the loss function of the ZINB-based decoder network is the negative logarithm of the ZINB likelihood, expressed as: ; In this formula, The zero-inflated negative binomial distribution is used to model sparse, overly discrete single-cell counting data. -log(˙) represents taking the probability density and then taking the negative log to obtain the negative log-likelihood loss. The smaller this value is, the better. Considering the extremely sparse and almost binary nature of scATAC-seq data, a decoder network based on the Bernoulli distribution (Ber) is chosen to model the scATAC-seq data, expressed as: ; In this formula, These are vectors from the original scATAC-seq data. It is Ber's average parameter; The decoder based on the Bernoulli distribution (Ber) estimates the input latent variable Z' through a fully connected layer using sigmoid (·) as the activation function. The expression is: ; In this formula, yes a In matrix form, sigmoid(·) is an activation function that compresses real values ​​to the range between 0 and 1. It is a weight parameter matrix. This represents the mapping of the a-th decoder; Finally, the decoder based on the Bernoulli distribution (Ber) is optimized using cross-entropy loss, expressed as follows: ; In this formula, L Ber This represents the Bernoulli loss, and CE(·) is the cross-entropy function. These are Bernoulli variables that are actually observed.

[0056] See Figures 2-4 In this embodiment, in step S4, two latent representations are learned. and They encode the characteristics of individuals in different omics data, and also represent them through latent features. To capture the commonalities among these features, and in this process, a scale parameter is introduced. and For these two sets of potential representations and To achieve efficient aggregation between the two, we perform element-wise weighted summation, which can be formalized as follows: ; In this formula, Z represents co-embedding representation.

[0057] The accuracy and reliability of cluster analysis can be improved by generating a more discriminative co-embedding representation Z.

[0058] See Figures 2-4 In this embodiment, in step S4, in order to pursue a more discriminative and informative co-embedding representation Z, the total loss is... Combining individual and common information from multiple omics datasets, and unifying the input targets for scRNA-seq and scATAC-seq data, the total loss expression is as follows:

[0059] In this formula, This indicates taking the minimum value. These represent network parameters, where α1, α2, and γ all represent weight coefficients. The zero-inflated negative binomial distribution is used to model sparse, overly discrete single-cell counting data. The proportion of "extra zeros" is determined by the zero-inflation probability parameter. It is the negative binomial mean parameter, representing the expected count of the non-zero portion, and is usually output by the decoder network. This is the negative binomial dispersion parameter, which controls the variance. It is usually output by the decoder network. -log(˙) indicates that the probability density is taken first, and then the negative logarithm is taken to obtain the negative log-likelihood loss. The smaller this value, the better. CE(·) is the cross-entropy function. For the actual observed Bernoulli variables, is u a The matrix form, u a The mean parameters of the Bernoulli distribution Ber are θ' and θ'. These are the encoder parameters in the branches of view u and view r, respectively. MI(˙) is the mutual information, used to measure the degree of nonlinear dependence between two random variables. k represents the k-th subspace, and p represents the number of subspaces. and Let represent the node embeddings of the k-th and l-th subspaces in view u, respectively. This represents the node embedding in the k-th subspace of view r; The total loss function is continuously minimized, and the algorithm is trained for 500 epochs for convergence optimization. Individual feature representations and shared feature representations are learned from multi-omics data, and these representations are then merged into a co-embedding representation Z for clustering information.

[0060] To comprehensively evaluate the clustering performance of the proposed scMSCL method, experiments were conducted on four multimodal single-cell datasets, including six competing methods. Specifically: scEMC proposed a SAN module, combined with a Transformer structure to capture global structural relationships in different feature spaces; scMCs used multiple autoencoders to map omics data from high-dimensional space to low-dimensional space before analysis; scMVAE proposed three strategies—scMVAE-PoE, scMVAE-NN, and scMVAE-Direct—to learn the joint latent features of data fusion and clustering. Among them, scMVAE-Direct directly connects the original features of each omics, scMVAE-NN fuses low-dimensional features extracted from different omics, and scMVAE-PoE estimates the joint posterior distribution through expert model multiplication; in addition, DCCA projects different omics to their corresponding low-dimensional spaces and uses a "teacher-student" mechanism to achieve effective fusion of multi-omics data.

[0061] To verify the superiority of the scMSCL of this invention, this embodiment uses NMI and ARI as evaluation indicators. The higher the value of either, the better the performance. The specific evaluation structure is shown in Table 1 below: Table 1. Evaluation Index Results of Six Comparison Algorithms

[0062] As can be seen from the table above, the NMI and ARI metrics of the scMSCL of this invention exhibit superior performance. See also... Figures 5-11 The UMAP clustering visualization of scMSCL in this invention is also the most compact, with the largest distance between each cluster, demonstrating a significant advantage. Among them, UMAP (Uniform Manifold Approximation and Projection) is a non-linear dimensionality reduction algorithm, mainly used to map high-dimensional data to a low-dimensional space (such as 2D or 3D) while preserving the local and global structure of the data.

[0063] This invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the single-cell multi-omics clustering method based on multi-subspace contrastive learning described above.

[0064] See Figure 12The present invention also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the single-cell multi-omics clustering method based on multi-subspace contrastive learning as described above.

[0065] This invention also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the steps of the single-cell multi-omics clustering method based on multi-subspace contrastive learning described above.

[0066] It is understood that the systems, devices, and storage media provided in the embodiments of the present invention correspond to the methods provided in the embodiments of the present invention, and the explanations, examples, and beneficial effects of the relevant content can be referred to the corresponding parts of the above-described single-cell multi-omics clustering method based on multi-subspace contrastive learning.

[0067] It should be noted that those skilled in the art will understand that all or part of the steps implemented in the embodiments of the present invention can be implemented entirely or partially by software, hardware, firmware, or any combination thereof. When implemented in hardware, it can be implemented entirely or partially by purchasing standard parts or modifications. When implemented in software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid state disks (SSDs)).

[0068] In summary, to overcome the limitations of single-space contrastive analysis, this invention provides a single-cell multi-omics clustering method based on multi-subspace contrastive learning. This method can jointly model single-cell transcriptomics and epigenetic data. First, it can fully explore the specificity of different omics and the potential common representations among them. Then, it integrates them into a co-embedding representation, which can analyze cellular heterogeneity. Specifically, an information extraction and fusion module is designed to refine the individual and common features of heterogeneous omics data. Through multi-subspace contrastive learning, this scMSCL can capture complementary information between omics data in different feature spaces, as well as the intrinsic connections between omics data within the same feature space, constructing a more comprehensive and informative representation for the fusion and clustering of single-cell multi-omics data. Extensive experiments show that this scMSCL outperforms other benchmark methods, demonstrating its superior clustering ability.

[0069] It should be understood that the examples and embodiments described herein are for illustrative purposes only and are not intended to limit the invention. Those skilled in the art can make various modifications or changes based on them. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the invention should be included within the protection scope of the invention.

Claims

1. A single-cell multi-omics clustering method based on multi-subspace contrastive learning, characterized in that, Includes the following steps: S1. Based on the autoencoder, important individual features are extracted from the transcriptome and chromatin accessibility of different omics. S2. Multi-subspace contrastive learning is used to extract common features of transcriptome and chromatin accessibility, effectively revealing the deep-seated relationships between omics data. S3: Employ autoencoders targeting specific omics chromatin accessibility and transcriptomes to reconstruct data for better noise reduction and reconstruction of features consistent with omics information; S4: The specific information of the specific omics in S3 and the common information of multiple omics in S2 are fused to obtain a co-embedding that is more conducive to clustering.

2. The single-cell multi-omics clustering method based on multi-subspace contrastive learning as described in claim 1, characterized in that, In S1, an autoencoder is used to map single-cell multi-omics data to their respective nonlinear embedding spaces, preserving individual characteristics and enhancing robustness to noise and outliers. Specifically, let's set , These are standardized scRNA-seq and scATAC-seq data, respectively, where N is the number of samples, Dr and Da are the number of features, and two independent encoders are used. and To learn the corresponding d-dimensional feature representation { , } Where d is the dimension of the embedding space, the expressions are as follows: ; In these two formulas, It represents the potential low-dimensional features of cells and genes in scRNA-seq data. It is a potential low-dimensional feature representation between cells and peaks in scATAC-seq data.

3. The single-cell multi-omics clustering method based on multi-subspace contrastive learning as described in claim 1, characterized in that, S2 further includes: S2.1 Define two views u and r. From a feature perspective, input views u and r into multiple subspace encoders MSEncoder respectively to generate node embeddings for multiple subspaces. This process is represented as follows: In these two equations, l represents the l-th subspace, and p represents the number of subspaces. , and Let be the node embeddings of the i-th subspace in views u and r, respectively, where d represents the dimension of the embedding space. Views U and R are feature representations obtained through multiple subspace encoders, and are common features of the same omics in views u and r, respectively. It represents the potential low-dimensional features of cells and genes in scRNA-seq data. It is a potential low-dimensional feature representation between cells and peaks in scATAC-seq data; The subspace encoder MSEncoder maps the input feature matrix X to a keyword embedding K and a value embedding V using two independent linear projections, with the corresponding expressions as follows: In these two formulas, , , representing different projection matrices of the key and value, respectively; S2.2 To make the features captured by multi-subspace contrastive learning more diverse, we should focus on more important node information, make full use of the degree information obtained by the graph structure, and integrate the graph structure information into the attention mechanism to enhance the discriminativeness of the feature representation. The initialization expression for the degree embedding matrix S is: ; In this formula, Let D represent the projection matrix corresponding to the degree, where D is the degree matrix; The expression for the attention weight matrix is ​​as follows: ; In this formula, Att is the output of the entire expression, representing the "attention weight matrix", the Softmax(˙) function is a function that "compresses" a vector of arbitrary real values ​​into a probability distribution vector, Q is the query matrix, and α is a learnable or settable hyperparameter. To compute the embedding representation of a node within each subspace, the value embedding V is broadcast to each subspace as follows: We obtain the node embeddings of p subspaces, expressed as: ; ; S2.3, through calculation and The mutual information maximization between node embeddings across subspaces and the mutual information minimization within subspaces are used as their respective subspace contrastive loss functions. S2.4 The final objective function of subspace contrastive learning is a weighted sum of the mutual information maximization objective and the mutual information minimization objective, which can be formally expressed as: In this formula, L contra Let represent the subspace contrast loss, θ' and φ be the encoder parameters in the view u and view r branches respectively, γ be the weighting coefficient balancing the influence of targets within and between spaces, MI(˙) be the mutual information, and k represent the k-th subspace. and Let represent the node embeddings of the k-th and l-th subspaces in view u, respectively. This represents the node embedding in the k-th subspace of view r; S2.

5. Train and concatenate views U and R to obtain the latent feature representation. .

4. The single-cell multi-omics clustering method based on multi-subspace contrastive learning as described in claim 3, characterized in that, In step S2.3, the feature representations of views u and r are output to p subspaces, i.e. To optimize the subspace encoder MSEncoder and ensure that views U and R both possess diverse feature representations of common features of the same omics, a pair of loss functions is used. These functions focus on maximizing the mutual information between subspace representations and minimizing the mutual information between different subspace representations. For the mutual information between subspace representations, the objective is formulated as follows: ; In this formula, θ' and φ are the encoder parameters in the branches of view u and view r, respectively, k represents the k-th subspace, p represents the number of subspaces, and MI(˙) is the mutual information. This represents the node embedding in the k-th subspace of view u. This represents the node embedding in the k-th subspace of view r.

5. The single-cell multi-omics clustering method based on multi-subspace contrastive learning as described in claim 4, characterized in that, because and The mutual information (MI) is unknown, making it impossible to calculate. To estimate MI, a sampling method is used instead of maximizing mutual information to maximize the lower bound estimate of MI, and a contrast log-ratio upper bound (CLUB) is employed to minimize the upper bound estimate of MI. Further: Regarding the estimation of the lower bound for maximizing mutual information (MI), the Jensen-Shannon estimation is used instead of the Kullback-Leibler KL divergence estimation. The calculation formula is as follows: In this formula, It is a discriminator that takes the representations of two views as input and scores the consistency between them. Here, the discriminator is simply instantiated as the dot product between two representations, i.e. SP(·) is the activation function. For the joint distribution expectation of two views within the same subspace k, The product of the independent distribution expectations of two views within the same subspace k; In addition to the objectives within each space, constraints are imposed on the relationships between different subspaces within the same view to increase spatial diversity. To this end, an optimization objective is proposed to minimize mutual information. This objective can increase the key features within view u, and because... Therefore, the expression for the optimization objective is: ; In this formula, l represents the l-th subspace. This represents the node embedding in the l-th subspace of view u; Regarding the estimation of the upper limit for minimizing mutual information (MI), the formula for calculating the upper limit of the logarithmic ratio (CLUB) is as follows: ; In this formula, I CLUB Let E represent the expectation of CLUB, P(x,y) represent the joint distribution of x and y, P(x) and P(y) represent the marginal distributions of x and y respectively, log is the natural logarithm, and P(y|x) is the conditional probability of y given x. The key to calculating CLUB is the distribution... In this case, it is assumed that each dimension of x and y is correlated; give Assuming that each dimension of y depends only on the corresponding dimension of x, then using and Substituting x and y respectively, the relevant expression is: ; In this formula, P represents the probability density. Let N(˙) denote the covariance matrix, N(˙) denote the multivariate Gaussian distribution, and I denote the identity matrix. This represents the variance shared across all dimensions. and They are and The node embedding in the i-th subspace, Π i The symbol for the product of all elements or subspaces i represents the product of a high-dimensional Gaussian distribution into a product of independent one-dimensional Gaussian distributions.

6. The single-cell multi-omics clustering method based on multi-subspace contrastive learning as described in claim 1, characterized in that, In S3, a decoder network based on the zero-inflated negative binomial regression model ZINB explores the global probability structure of scRNA-seq data. The ZINB distribution is derived from the mean of the negative binomial distribution. Deviation parameters and coefficients used to describe the probability of interpolated events. The definition and expression are: In these two formulas, Indicates the parameter is u r The negative binomial probability mass function of x and θ. r It is a vector from the raw scRNA-seq data, representing a random variable taking non-negative integer values, u r Let θ be the mean parameter of the distribution, and θ be the deviation parameter. It is the gamma function, ZINB(˙) represents the zero-inflation negative binomial distribution, and π represents the zero-inflation probability. It is x r Indicator function when =0; Specifically, the ZINB-based decoder estimates parameters based on the co-embedding representation Z through three different fully connected layers. : ; ; ; Of these three formulas, yes In matrix form, The proportion of "extra zeros" is determined by the zero-inflation probability parameter. It is the negative binomial mean parameter, representing the expected count of the non-zero portion, and is usually output by the decoder network. It is the negative binomial dispersion parameter, which controls the variance and is usually output by the decoder network. The activation function is sigmoid(·), so the exit probability is between 0 and 1. Furthermore, since the mean and discrete parameters are non-negative, the exponential function exp is chosen as the activation function. and Activation function, It is a decoder with a fully connected layer. , and W θ These are three learnable parameter matrices; Unlike traditional autoencoders based on mean squared error loss, the loss function of the ZINB-based decoder network is the negative logarithm of the ZINB likelihood, expressed as: ; In this formula, The zero-inflated negative binomial distribution is used to model sparse, overly discrete single-cell counting data. -log(˙) represents taking the probability density and then taking the negative log to obtain the negative log-likelihood loss. The smaller this value is, the better. Considering the extremely sparse and almost binary nature of scATAC-seq data, a decoder network based on the Bernoulli distribution (Ber) is chosen to model the scATAC-seq data, expressed as: ; In this formula, These are vectors from the original scATAC-seq data. It is Ber's average parameter; The Bernoulli distribution-based decoder estimates the input latent variable Z' through a fully connected layer using sigmoid (·) as the activation function. The expression is: ; In this formula, is u a In matrix form, sigmoid(·) is an activation function that compresses real values ​​to the range between 0 and 1. It is a weight parameter matrix. This represents the mapping of the a-th decoder; Finally, the decoder based on the Bernoulli distribution (Ber) is optimized using cross-entropy loss, expressed as follows: ; In this formula, L Ber This represents the Bernoulli loss, and CE(·) is the cross-entropy function. These are Bernoulli variables that are actually observed.

7. The single-cell multi-omics clustering method based on multi-subspace contrastive learning as described in claim 1, characterized in that, In S4, two latent representations are learned. and They encode the characteristics of individuals in different omics data, and also represent them through latent features. To capture the commonalities among these features, and in this process, a scale parameter is introduced. and For these two sets of potential representations and To achieve efficient aggregation between the two, we perform element-wise weighted summation, which can be formalized as follows: ; In this formula, Z represents co-embedding representation; The accuracy and reliability of cluster analysis can be improved by generating a more discriminative co-embedding representation Z.

8. The single-cell multi-omics clustering method based on multi-subspace contrastive learning as described in claim 7, characterized in that, In S4, in order to pursue a more discriminative and informative co-embedding representation Z, the total loss is... Combining individual and common information from multiple omics datasets, and unifying the input targets for scRNA-seq and scATAC-seq data, the total loss expression is as follows: ; In this formula, This indicates taking the minimum value. These represent network parameters, where α1, α2, and γ all represent weight coefficients. The zero-inflated negative binomial distribution is used to model sparse, overly discrete single-cell counting data. The proportion of "extra zeros" is determined by the zero-inflation probability parameter. It is the negative binomial mean parameter, representing the expected count of the non-zero portion, and is usually output by the decoder network. This is the negative binomial dispersion parameter, which controls the variance. It is usually output by the decoder network. -log(˙) indicates that the probability density is taken first, and then the negative logarithm is taken to obtain the negative log-likelihood loss. The smaller this value, the better. CE(·) is the cross-entropy function. For the actual observed Bernoulli variables, is u a The matrix form, u a The mean parameters of the Bernoulli distribution Ber are θ' and θ'. These are the encoder parameters in the branches of view u and view r, respectively. MI(˙) is the mutual information, used to measure the degree of nonlinear dependence between two random variables. k represents the k-th subspace, and p represents the number of subspaces. and Let represent the node embeddings of the k-th and l-th subspaces in view u, respectively. This represents the node embedding in the k-th subspace of view r; The total loss function is continuously minimized and trained for 500 epochs for convergence optimization. Individual feature representations and shared feature representations are learned from multi-omics data and then merged into the co-embedding representation Z for clustering information.

Citation Information

Patent Citations

  • Processing method for high-order tensor data

    CN109993199A

  • Steel structure welding process quality management recommendation method based on big data processing

    CN118132849A

  • Single cell data multi-clustering method and system based on multi-omics data fusion

    CN119741978A

  • Single-cell multi-omics clustering method based on depth information fusion

    CN119763677A

  • Single-cell multi-omics data clustering method based on cross-omics feature fusion

    CN119905150A