Single-cell multi-omics integration method and system based on adaptive probability distribution modeling
By employing an adaptive probability distribution modeling approach and utilizing a multi-encoder and latent space fusion strategy, the problem of single probability distribution assumptions and insufficient fusion capabilities in the integration of single-cell multi-omics data was solved. This approach enabled more efficient data reconstruction and inter-omics alignment, improving the integration effect and the adaptability of the model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIAMEN UNIV
- Filing Date
- 2026-04-16
- Publication Date
- 2026-07-14
AI Technical Summary
In existing technologies, methods for integrating single-cell multi-omics data suffer from problems such as simplistic probability distribution assumptions, insufficient multi-omics fusion capabilities, and poor data distribution adaptability, making it difficult to effectively characterize nonlinear relationships and achieve alignment between omics.
An adaptive probability distribution modeling approach is adopted, which extracts features through an omics-independent multi-encoder structure and a shared feature layer. Combined with latent space fusion and alignment strategies, the adaptive probability distribution decoder outputs the most matching combination of probability distributions, thereby realizing the joint modeling and reconstruction of multi-omics data.
It improved the accuracy of data reconstruction, solved the problem of differences in feature dimensions and distributions among different omics, enhanced the generalization and expressive capabilities of the model, and achieved a more efficient multi-omics integration effect.
Smart Images

Figure CN122392610A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics and artificial intelligence, specifically to a method and system for integrating single-cell multi-omics data based on a deep probabilistic generative model. Background Technology
[0002] With the rapid development of single-cell sequencing technology, multi-omics technologies such as single-cell RNA sequencing, single-cell chromatin accessibility sequencing, and single-cell DNA methylation sequencing can reveal the functional state and regulatory mechanisms of cells at different levels. Different omics data characterize cell state at the transcriptional, epigenetic, and regulatory levels, exhibiting significant complementarity. Therefore, joint modeling and integrated analysis of multi-omics data are of great significance for cell type identification, cell developmental trajectory analysis, and disease mechanism research.
[0003] However, different single-cell omics data exhibit significant differences in data distribution, data dimensionality, and noise characteristics. For example, single-cell RNA sequencing data typically shows an overly discrete count distribution, while single-cell chromatin accessibility sequencing data is highly sparsity and binary. This heterogeneity makes it difficult for traditional data fusion methods to effectively model the complex relationships between multiple omics.
[0004] Existing methods typically employ the following strategies for multi-omics integration:
[0005] Methods based on linear dimensionality reduction, such as principal component analysis (PCA) or canonical correlation analysis (CCA), are difficult to characterize nonlinear relationships. Methods based on deep learning, such as variational autoencoders (VAE), achieve omics fusion through latent space learning, but most methods assume a fixed probability distribution (such as Gaussian or negative binomial distribution), making it difficult to adapt to the real statistical characteristics of different omics data. Methods based on latent space alignment or adversarial learning can achieve modality alignment, but still rely on the assumption of a fixed distribution, which limits the expressive power of the model.
[0006] The above method has the following problems:
[0007] 1. The probability distribution assumption is too simplistic and cannot adapt to the complex and diverse statistical characteristics of different omics data;
[0008] 2. Limited latent space integration capabilities and insufficient alignment effects among omics;
[0009] 3. The model has poor generalization and adaptability, making it difficult to extend to different types of omics data.
[0010] Therefore, there is an urgent need for an integrated method that can adaptively learn the distribution characteristics of data and achieve unified modeling of multiple omics. Summary of the Invention
[0011] To address the problems of single probability distribution assumptions, insufficient multi-omics fusion capabilities, and poor data distribution adaptability in existing technologies, this invention provides a single-cell multi-omics integration method and system based on adaptive probability distribution modeling. By constructing an adaptive probability distribution decoding mechanism, adaptive modeling of multi-omics data distribution is achieved, and the multi-omics integration effect is improved through latent space fusion and alignment strategies.
[0012] The technical solution adopted by this invention to solve its technical problem is:
[0013] On the one hand, a single-cell multi-omics integration method based on adaptive probability distribution modeling includes the following steps:
[0014] The encoding step employs an omics-independent multi-encoder structure combined with a shared feature layer to extract features from different single-cell omics data, outputting the mean and variance of omics-specific latent variables for each omics. It also performs feature fusion on different single-cell omics data, outputting the mean and variance of omics fusion latent variables representing cross-omics shared information. Omics-specific latent variables are obtained based on the mean and variance of the omics-specific latent variables, and omics fusion latent variables are obtained based on the mean and variance of the omics fusion latent variables.
[0015] The latent space fusion step involves joint modeling based on the omics-specific latent variables and omics fusion latent variables. An alternating training strategy is used to reconstruct omics data in different training iterations by alternately utilizing the omics-specific latent variables and omics fusion latent variables to obtain the latent space distribution of different omics. A fusion loss function is introduced to constrain the latent space distribution, so that the omics-specific latent variables are aligned with the omics fusion latent variables.
[0016] The adaptive probability distribution decoding step adopts an omics-independent multi-decoder structure. Based on the statistical characteristics of the latent space distribution of different omics, it outputs parameters of multiple probability distributions and their combined weights. Through adaptive learning of the combined weights, it determines the probability distribution combination that best matches the latent space distribution of the original omics, thereby realizing the joint modeling and reconstruction of single-cell multi-omics data.
[0017] Preferably, the multi-encoder structure includes multiple omics-independent encoders corresponding to different single-cell omics data. Each omics-independent encoder is composed of a multi-layer fully connected neural network. The multiple omics-independent encoders and the shared feature layer constitute an encoder module. The calculation formulas for the mean and variance in the forward process of the encoder module are as follows:
[0018] ;
[0019] ;
[0020] in, and These represent the mean vector and standard deviation vector of the omics-specific latent variables, respectively. This represents the sequencing data of the m-th single-cell omics. This represents the corresponding omics-independent encoder; Indicates a shared feature layer; and These are parametric mapping functions representing the mean and variance of omics-specific latent variables, respectively. and These represent the mean vector and standard deviation vector of the omics fusion latent variables shared across omics, respectively. and These are the parametric mapping functions for the mean and variance of the cross-omics fusion latent variables, respectively; This represents the total number of omics systems involved in the integration.
[0021] Preferably, the omics-specific latent variables and omics fusion latent variables are represented as follows:
[0022] ;
[0023] ;
[0024] in, This represents the omics-specific latent variable corresponding to the m-th omics; This represents the fusion latent variables shared across omics. and ... It represents a multivariate standard normal distribution.
[0025] Preferably, in the latent space fusion step, in different training iterations, the omics fusion latent variables or omics-specific latent variables are alternately used as inputs to the decoder to reconstruct the corresponding latent space variables of each omics data, so as to guide the alignment of different omics latent spaces; wherein the training strategy is expressed as follows:
[0026] ;
[0027]
[0028] in, Indicates the first Latent space variables obtained from round iteration; This is the current training iteration round; This is the preset initial number of training rounds; Represents omics fusion latent variables shared across omics disciplines; Let m represent the omics-specific latent variable corresponding to the m-th omics; M represents the total number of omics involved in the fusion. is the alternation training control function; f represents the frequency of alternation training.
[0029] Preferably, the multi-decoder structure includes multiple omics-independent decoders corresponding to different single-cell omics data. Each omics-independent decoder is composed of a multi-layer fully connected neural network, used to map latent space variables to probability distribution parameters for reconstructing single-cell omics data. The multiple omics-independent decoders form an adaptive probability distribution decoder module, and the forward computation process of the decoder module satisfies the following formula:
[0030] ;
[0031] ;
[0032] in, The latent space variables are input to the decoder, and the latent space variables are omics-specific latent variables or omics fusion latent variables; This represents the omics-independent decoder corresponding to the m-th type of single-cell omics; This represents the set of candidate probability distribution parameters corresponding to the m-th omics. The parameter vector represents the k-th candidate probability distribution corresponding to the m-th omics.
[0033] Preferably, the combined weights of each candidate probability distribution output by the adaptive probability distribution decoder module The calculation formula is as follows:
[0034] ;
[0035] The reconstruction process of the single-cell omics data satisfies the following mixed distribution pattern:
[0036] ;
[0037] in, A learnable probability distribution weight parameter vector; This represents the combined weight of the probability distribution of the k-th candidate in the m-th omics; Denotes the probability distribution of the k-th candidate; This represents the number of candidate probability distributions for the m-th omics predefined; The parameter vector represents the probability distribution of the k-th candidate omics corresponding to the m-th omics. This represents the sequencing data of the m-th single-cell omics. Represents variables in the latent space Under the given conditions, the probability distribution of the m-th omics data; through adaptive learning of the combined weights. This ensures that the mixed probability distribution matches the true distribution of the original single-cell omics data.
[0038] Preferably, a joint loss function is used to optimize the model parameters throughout the training process. This joint loss function includes a fusion loss function for the latent space fusion step and a reconstruction loss based on a weighted mixture probability distribution for the decoding step. The fusion loss function includes latent space distribution constraint loss, latent space alignment loss, and latent space adversarial loss. The joint loss function is expressed as follows:
[0039] ;
[0040] Where M represents the number of single-cell omics; This represents the reconstruction loss of the m-th omics; This represents the latent space distribution constraint loss; This represents the latent space alignment loss; denoted as the latent space alignment adversarial loss based on adversarial learning; α, β, and γ are the weight coefficients of the corresponding loss terms.
[0041] Preferably, the reconstruction loss is defined based on the log-likelihood form of a weighted mixture probability distribution as follows:
[0042] ;
[0043] in, This represents the m-th type of single-cell omics data; This represents the combined weight of the probability distribution of the k-th candidate in the m-th omics; Denotes the probability distribution of the k-th candidate; The parameter vector represents the probability distribution of the k-th candidate omics corresponding to the m-th omics. This represents the number of candidate probability distributions for the m-th omics predefined; Represents the latent space variable Seeking expectations.
[0044] Preferably, the latent space distribution constraint loss adopts the KL divergence form to constrain the distributional consistency between omics-specific latent variables and omics fusion latent variables, as shown below:
[0045] ;
[0046] ;
[0047] in, This represents the omics-specific latent variable corresponding to the m-th omics; This represents the fusion latent variables shared across omics. The variational posterior distribution represents the latent variables of omics fusion shared across omics. ) represents the variational posterior distribution of the omics-specific latent variable corresponding to the m-th omics; Indicates multiple
[0048] A collection of single-cell omics sequencing data; Represents the divergence function;
[0049] The latent space alignment loss includes a distribution alignment term based on the maximum mean difference (MMD), expressed as follows:
[0050] ;
[0051] in, This represents the function representing the maximum mean difference.
[0052] Furthermore, the adversarial loss is used to align the omics-specific latent variables and omics fusion latent variables in an adversarial distribution, as shown below:
[0053] ;
[0054] in, This represents a modal discriminator; y represents the gradient inversion layer; y represents the omics label of the corresponding latent space variable source. The loss function for adversarial tasks includes the cross-entropy function; It indicates a desire for the expected value.
[0055] On the other hand, a single-cell multi-omics integration system based on adaptive probability distribution modeling includes:
[0056] The encoder module is configured to employ an omics-independent multi-encoder structure combined with a shared feature layer to extract features from different single-cell omics data, outputting the mean and variance of omics-specific latent variables for each omics, and to fuse features from different single-cell omics data, outputting the mean and variance of omics fusion latent variables representing cross-omics shared information; omics-specific latent variables are obtained based on the mean and variance of omics-specific latent variables, and omics fusion latent variables are obtained based on the mean and variance of omics fusion latent variables;
[0057] The latent space fusion module is configured to perform joint modeling based on the omics-specific latent variables and omics fusion latent variables. Through an alternating training strategy, it uses omics-specific latent variables and omics fusion latent variables alternately in different training iterations to reconstruct omics data and obtain the latent space distribution of different omics. A fusion loss function is introduced to constrain the latent space distribution, so that the omics-specific latent variables are aligned with the omics fusion latent variables.
[0058] The adaptive probability distribution decoder module is configured to adopt an omics-independent multi-decoder structure. Based on the statistical characteristics of the latent space distribution of different omics, it outputs parameters of multiple probability distributions and their combined weights. By adaptively learning the combined weights, it determines the probability distribution combination that best matches the latent space distribution of the original omics, thereby realizing the joint modeling and reconstruction of single-cell multi-omics data.
[0059] As can be seen from the above description of the present invention, compared with the prior art, the present invention has the following beneficial effects:
[0060] (1) By constructing an adaptive probability distribution decoder mechanism, the present invention enables the entire model to automatically adjust the probability distribution weights according to the statistical characteristics of the latent space distribution of different single-cell omics, thereby avoiding the modeling errors caused by the traditional method's reliance on fixed probability distribution assumptions and thus improving the accuracy of data reconstruction.
[0061] (2) This invention extracts different omics features through a multi-encoder structure and achieves a unified latent space expression through a latent space fusion module, which effectively solves the problem of feature dimension and distribution differences between different omics and improves the integration effect of multi-omics.
[0062] (3) This invention optimizes the reconstruction loss, distribution constraint loss, alignment loss and adversarial loss by weighting and jointly optimizing the probability distribution parameters and network parameters during the training process, thereby improving the model's convergence stability and expressive power. Attached Figure Description
[0063] Figure 1 This is a flowchart of a single-cell multi-omics integration method based on adaptive probability distribution modeling, according to an embodiment of the present invention.
[0064] Figure 2 This is a diagram of a single-cell multi-omics integrated model based on adaptive probability distribution modeling, as described in an embodiment of the present invention.
[0065] Figure 3 This is a flowchart illustrating the training and testing process according to an embodiment of the present invention.
[0066] Figure 4 This is a block diagram of a single-cell multi-omics integrated system based on adaptive probability distribution modeling, according to an embodiment of the present invention. Detailed Implementation
[0067] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0068] In the description of this invention, it should be noted that the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0069] like Figure 1 As shown in the figure, this embodiment presents a single-cell multi-omics integration method based on adaptive probability distribution modeling, which includes the following steps:
[0070] S101, Encoding Steps: Employing an omics-independent multi-encoder structure combined with a shared feature layer, features are extracted from different single-cell omics data, outputting the mean and variance of omics-specific latent variables for each omics. Feature fusion is also performed on different single-cell omics data, outputting the mean and variance of omics fusion latent variables representing cross-omics shared information. Omics-specific latent variables are obtained based on the mean and variance of the omics-specific latent variables, and omics fusion latent variables are obtained based on the mean and variance of the omics fusion latent variables.
[0071] S102, Latent Space Fusion Step: Joint modeling is performed based on the omics-specific latent variables and omics fusion latent variables. An alternating training strategy is used to alternately utilize the omics-specific latent variables and omics fusion latent variables in different training iterations to reconstruct omics data and obtain the latent space distribution of different omics. A fusion loss function is introduced to constrain the latent space distribution, so that the omics-specific latent variables are aligned with the omics fusion latent variables.
[0072] S103, Adaptive probability distribution decoding step: Adopting an omics-independent multi-decoder structure, outputting parameters of multiple probability distributions and their combined weights based on the statistical characteristics of the latent space distribution of different omics, and determining the combination of probability distributions that best matches the latent space distribution of the original omics through adaptive learning of the combined weights, thereby realizing the joint modeling and reconstruction of single-cell multi-omics data.
[0073] This embodiment provides a single-cell multi-omics integration method based on adaptive probability distribution modeling, which aims to solve the problems of large distribution differences, high noise, and strong feature heterogeneity in the integration of single-cell sequencing data (such as single-cell RNA sequencing data, single-cell chromatin accessibility data, etc.).
[0074] like Figure 2As shown in the figure, this embodiment applies a single-cell multi-omics integration method based on adaptive probability distribution modeling to a single-cell multi-omics integration model based on adaptive probability distribution modeling. The single-cell multi-omics integration model includes an input layer, an encoder module, a latent space fusion module, an adaptive decoder module, and an output layer connected in sequence.
[0075] In practice, the first step at the input layer is to acquire the single-cell multi-omics data to be integrated. The following example, using commonly used single-cell transcriptome (RNA) and single-cell chromatin accessibility (ATAC) data integration, includes the following steps:
[0076] RNA data preprocessing: The raw counting matrix (referring to the input single-cell RNA omics data, represented in the form of a data counting matrix) was normalized using Log1p and the top 2000 hypervariable genes were screened.
[0077] ATAC data preprocessing: TF-IDF transformation was used to preserve the discrete features of the signal and to screen for high-accessibility peaks in the top 10% of cells; Normalization: Different omics data were mapped to the same scale space to eliminate systematic errors caused by sequencing depth.
[0078] In step S101, the encoder module includes independent encoding paths for different omics, and includes the following processing procedures:
[0079] For RNA and ATAC, networks consisting of three fully connected layers were constructed, with the number of neurons decreasing layer by layer (e.g., Input -> 1024 -> 512 -> 256).
[0080] After initial extraction, the features of each omics are merged into a shared feature layer, which consists of two fully connected layers to capture common biological signals across omics.
[0081] At the network's end, two parallel fully connected layers output the mean vector and standard deviation vector of omics-specific latent variables / omics fusion latent variables, respectively.
[0082] Introducing random noise vectors and Based on the reparameterization formula, omics-specific latent variables / omics fusion latent variables are obtained to ensure that the latent space can perform gradient backpropagation while maintaining randomness.
[0083] Specifically, the encoder module includes multiple omics-independent encoding sub-modules corresponding to different single-cell omics data and at least one shared feature layer. The calculation formulas for the mean and variance in the forward process of the encoder module are as follows:
[0084] ;
[0085] ;
[0086] in, and These represent the mean vector and standard deviation vector of the omics-specific latent variables, respectively. This represents the sequencing data of the m-th single-cell omics. This represents the corresponding omics-independent encoder; Indicates a shared feature layer; and These are parametric mapping functions representing the mean and variance of omics-specific latent variables, respectively. and These represent the mean vector and standard deviation vector of the omics fusion latent variables shared across omics, respectively. and These are the parametric mapping functions for the mean and variance of the cross-omics fusion latent variables, respectively; This represents the total number of omics systems involved in the integration.
[0087] It should be noted that, , , and These parameter mapping functions are existing technology. They can take the form of linear layers or multilayer perceptrons (MLPs), which are essentially neural network output layers used to map the hidden features extracted by the encoder to the mean and variance of the latent variable distribution. as well as It also uses a neural network.
[0088] Specifically, the reparameterization formula is as follows:
[0089] ;
[0090] ;
[0091] in, This represents the omics-specific latent variable corresponding to the m-th omics; This represents the fusion latent variables shared across omics. and ... It represents a multivariate standard normal distribution.
[0092] Furthermore, in the latent space fusion step S102, the latent space fusion module is used to achieve unified representation and distribution alignment among different omics latent variables. Its core idea is to use an alternating training strategy to make omics-specific latent variables gradually move closer to fusion latent variables shared across omics, thereby constructing a unified latent space.
[0093] Specifically, for the m-th omics, the encoder outputs omics-specific latent variables. and cross-omics shared fusion latent variables To achieve gradual alignment of the latent space, this invention employs an alternating training strategy to control the decoder input, which can be expressed as:
[0094] ;
[0095]
[0096] in, Indicates the first Latent space variables obtained from round iteration; This is the current training iteration round; This is the preset initial number of training rounds; Represents omics fusion latent variables shared across omics disciplines; Let m represent the omics-specific latent variable corresponding to the m-th omics; M represents the total number of omics involved in the fusion. is the alternation training control function; f represents the frequency of alternation training.
[0097] like Figure 3 As shown, in the early stage of training, the omics fusion latent variables are used first for reconstruction, enabling the decoder to learn the cross-omics shared structure; in the later stage of training, omics-specific latent variables are used for reconstruction, so that the omics-specific latent variables gradually approach the distribution of the fusion latent variables, thereby achieving latent space unification.
[0098] Specifically, the fusion loss function includes spatial distribution constraint loss, latent space alignment loss, and latent space adversarial loss.
[0099] The latent space distribution constraint loss adopts the KL divergence form to constrain the distributional consistency between omics-specific latent variables and omics fusion latent variables, and its definition is as follows:
[0100] ;
[0101] ;
[0102] in, This represents the omics-specific latent variable corresponding to the m-th omics; This represents the fusion latent variables shared across omics. The variational posterior distribution represents the latent variables of omics fusion shared across omics. ) represents the variational posterior distribution of the omics-specific latent variable corresponding to the m-th omics; Indicates multiple
[0103] A collection of single-cell omics sequencing data; This represents the divergence function.
[0104] Furthermore, to further enhance the latent space alignment effect, a distribution alignment loss based on maximum mean difference (MMD) is introduced, as follows:
[0105] ;
[0106] in, This represents the function representing the maximum mean difference.
[0107] Furthermore, the adversarial loss is used to align the omics-specific latent variables and omics fusion latent variables in an adversarial distribution, as shown below:
[0108] ;
[0109] in, This represents a modal discriminator; y represents the gradient inversion layer; y represents the omics label of the corresponding latent space variable source. The loss function for adversarial tasks includes the cross-entropy function; This represents the expectation. The encoder is trained using a gradient inversion layer, which prevents the discriminator from distinguishing the source of latent variables, thereby achieving latent space alignment.
[0110] Furthermore, in the adaptive probability distribution decoding step S103, the adaptive probability distribution decoder module is used to reconstruct single-cell omics data based on latent variables, and adaptively matches the real data distribution through a probability distribution combination method. The adaptive probability distribution decoder module includes multiple omics-independent decoders, each composed of a multi-layer fully connected neural network, used to map latent space variables to probability distribution parameters for reconstructing single-cell omics data.
[0111] Specifically, for the m-th omics, the decoder outputs the parameters of the K candidate probability distributions:
[0112] ;
[0113] ;
[0114] ;
[0115] Furthermore, the reconstruction process of the single-cell omics data satisfies the following mixed distribution pattern:
[0116] ;
[0117] in, The latent space variables are input to the decoder, and the latent space variables are omics-specific latent variables or omics fusion latent variables; This represents the omics-independent decoder corresponding to the m-th type of single-cell omics; This represents the set of candidate probability distribution parameters corresponding to the m-th omics. The parameter vector represents the probability distribution of the k-th candidate omics corresponding to the m-th omics. A learnable probability distribution weight parameter vector; This represents the combined weight of the probability distribution of the k-th candidate in the m-th omics; Denotes the probability distribution of the k-th candidate; This represents the number of candidate probability distributions for the m-th omics predefined; The parameter vector represents the probability distribution of the k-th candidate omics corresponding to the m-th omics. This represents the sequencing data of the m-th single-cell omics. Represents variables in the latent space Under the given conditions, the probability distribution of the m-th omics data; through adaptive learning of the combined weights. This ensures that the mixed probability distribution matches the true distribution of the original single-cell omics data.
[0118] For example, for RNA data, the candidate probability distribution includes a negative binomial distribution:
[0119] ;
[0120] Poisson distribution:
[0121] ;
[0122] For ATAC data, candidate distributions include the Bernoulli distribution:
[0123] ;
[0124] Or Poisson distribution:
[0125] ;
[0126] in, Corresponding to the above , representing the m-th single-cell omics sequencing data; , Equivalent to what was mentioned earlier , where represents the parameter vector of the k-th candidate probability distribution corresponding to the m-th omics.
[0127] Through this hybrid probability distribution structure, the model can automatically select the probability distribution form that best suits the current omics data, thereby improving reconstruction accuracy and model generalization ability.
[0128] Furthermore, the method employs a joint loss function to optimize model parameters throughout the model training process. This joint loss function includes at least a reconstruction loss based on a weighted mixture probability distribution, a latent space distribution constraint loss, a latent space alignment loss, and a latent space adversarial loss, and its overall form is expressed as:
[0129] ;
[0130] Where M represents the number of single-cell omics; This represents the reconstruction loss of the m-th omics; This represents the latent space distribution constraint loss; This represents the latent space alignment loss; denoted as the latent space alignment adversarial loss based on adversarial learning; α, β, and γ are the weight coefficients of the corresponding loss terms.
[0131] Specifically, the reconstruction loss is defined based on the log-likelihood form of a weighted mixture probability distribution as follows:
[0132] The reconstruction loss is defined based on the log-likelihood of a weighted mixture probability distribution as follows:
[0133] ;
[0134] in, This represents the m-th type of single-cell omics data; This represents the combined weight of the probability distribution of the k-th candidate in the m-th omics; Denotes the probability distribution of the k-th candidate; The parameter vector represents the probability distribution of the k-th candidate omics corresponding to the m-th omics. This represents the number of candidate probability distributions for the m-th omics predefined; Represents the latent space variable Seeking expectations.
[0135] Furthermore, after the model training is complete, the fused latent variables can be used directly. As a unified low-dimensional representation of single cells, it is used for cell clustering analysis, cell type identification, cross-omics joint analysis, and differential gene expression. Because this latent variable integrates multi-omics information and completes distribution alignment, it has stronger biological interpretability and higher analytical accuracy.
[0136] As shown in Figure 4, this embodiment also discloses a single-cell multi-omics integration system based on adaptive probability distribution modeling, including:
[0137] Encoder module 401 is configured to employ an omics-independent multi-encoder structure combined with a shared feature layer to extract features from different single-cell omics data, outputting the mean and variance of omics-specific latent variables corresponding to each omics, and to perform feature fusion on different single-cell omics data, outputting the mean and variance of omics fusion latent variables representing cross-omics shared information; omics-specific latent variables are obtained based on the mean and variance of omics-specific latent variables, and omics fusion latent variables are obtained based on the mean and variance of omics fusion latent variables;
[0138] The latent space fusion module 402 is configured to perform joint modeling based on the omics-specific latent variables and omics fusion latent variables. It uses an alternating training strategy to alternately utilize the omics-specific latent variables and omics fusion latent variables in different training iterations to reconstruct omics data and obtain the latent space distribution of different omics. A fusion loss function is introduced to constrain the latent space distribution, so that the omics-specific latent variables are aligned with the omics fusion latent variables.
[0139] The adaptive probability distribution decoder module 403 is configured to adopt an omics-independent multi-decoder structure. Based on the statistical characteristics of the latent space distribution of different omics, it outputs parameters of multiple probability distributions and their combined weights. By adaptively learning the combined weights, it determines the probability distribution combination that best matches the latent space distribution of the original omics, thereby realizing the joint modeling and reconstruction of single-cell multi-omics data.
[0140] The specific implementation process of a single-cell multi-omics integration system based on adaptive probability distribution modeling is basically the same as the above-described method embodiments, and will not be repeated here. It should be noted that each functional module in the system corresponds to a specific step in the method embodiments, and those skilled in the art can understand the specific implementation of the system based on the method flow.
[0141] The above embodiments illustrate the basic principles and implementation methods of the present invention, aiming to help understand the core concept and key steps of the invention. It should be understood that these embodiments are merely examples and do not limit the scope of application of the present invention. Those skilled in the art, based on their understanding of the concept of the present invention, can make various equivalent improvements to specific steps, parameter configurations, or system structures. These improvements also fall within the protection scope of the present invention, as defined in the appended claims.
Claims
1. A single-cell multi-omics integration method based on adaptive probability distribution modeling, characterized in that, Includes the following steps: The encoding step employs an omics-independent multi-encoder structure combined with a shared feature layer to extract features from different single-cell omics data, outputting the mean and variance of omics-specific latent variables corresponding to each omics. It also performs feature fusion on different single-cell omics data, outputting the mean and variance of omics fusion latent variables representing cross-omics shared information. Omics-specific latent variables are obtained based on the mean and variance of omics-specific latent variables, and omics-fusion latent variables are obtained based on the mean and variance of omics-fusion latent variables. The latent space fusion step involves joint modeling based on the omics-specific latent variables and omics fusion latent variables. An alternating training strategy is used to reconstruct omics data in different training iterations by alternately utilizing the omics-specific latent variables and omics fusion latent variables to obtain the latent space distribution of different omics. A fusion loss function is introduced to constrain the latent space distribution, so that the omics-specific latent variables are aligned with the omics fusion latent variables. The adaptive probability distribution decoding step adopts an omics-independent multi-decoder structure. Based on the statistical characteristics of the latent space distribution of different omics, it outputs parameters of multiple probability distributions and their combined weights. Through adaptive learning of the combined weights, it determines the probability distribution combination that best matches the latent space distribution of the original omics, thereby realizing the joint modeling and reconstruction of single-cell multi-omics data.
2. The single-cell multi-omics integration method based on adaptive probability distribution modeling according to claim 1, characterized in that, The multi-encoder structure includes multiple omics-independent encoders corresponding to different single-cell omics data. Each omics-independent encoder is composed of a multi-layer fully connected neural network. The multiple omics-independent encoders and the shared feature layer form an encoder module. The calculation formulas for the mean and variance in the forward process of the encoder module are as follows: ; ; in, and These represent the mean vector and standard deviation vector of the omics-specific latent variables, respectively. This represents the sequencing data of the m-th single-cell omics. This represents the corresponding omics-independent encoder; Indicates a shared feature layer; and These are parametric mapping functions representing the mean and variance of omics-specific latent variables, respectively. and These represent the mean vector and standard deviation vector of the omics fusion latent variables shared across omics, respectively. and These are the parametric mapping functions for the mean and variance of the cross-omics fusion latent variables, respectively; This represents the total number of omics systems involved in the integration.
3. The single-cell multi-omics integration method based on adaptive probability distribution modeling according to claim 2, characterized in that, The omics-specific latent variables and omics fusion latent variables are represented as follows: ; ; in, This represents the omics-specific latent variable corresponding to the m-th omics; Represents the fusion latent variables shared across omics; and ... It represents a multivariate standard normal distribution.
4. The single-cell multi-omics integration method based on adaptive probability distribution modeling according to claim 1, characterized in that, In the latent space fusion step, in different training iterations, the omics fusion latent variables or omics-specific latent variables are alternately used as the input of the decoder to reconstruct the latent space variables of each omics data, so as to guide the alignment of different omics latent spaces. The training strategy is represented as follows: ; ; in, Indicates the first Latent space variables obtained from rounds of iteration; This is the current training iteration round; This is the preset initial number of training rounds; This represents the omics fusion latent variables shared across omics disciplines; Let m represent the omics-specific latent variable corresponding to the m-th omics; M represents the total number of omics involved in the fusion. is the alternation training control function; f represents the frequency of alternation training.
5. The single-cell multi-omics integration method based on adaptive probability distribution modeling according to claim 1, characterized in that, The multi-decoder structure includes multiple omics-independent decoders corresponding to different single-cell omics data. Each omics-independent decoder is composed of a multi-layer fully connected neural network, used to map latent space variables to probability distribution parameters for reconstructing single-cell omics data. The multiple omics-independent decoders form an adaptive probability distribution decoder module, and the forward computation process of the decoder module satisfies the following formula: ; ; in, The latent space variables are input to the decoder, and the latent space variables are omics-specific latent variables or omics fusion latent variables; This represents the omics-independent decoder corresponding to the m-th type of single-cell omics; This represents the set of candidate probability distribution parameters corresponding to the m-th omics. The parameter vector represents the k-th candidate probability distribution corresponding to the m-th omics.
6. The single-cell multi-omics integration method based on adaptive probability distribution modeling according to claim 5, characterized in that, The combined weights of each candidate probability distribution output by the adaptive probability distribution decoder module The calculation formula is as follows: ; The reconstruction process of the single-cell omics data satisfies the following mixed distribution pattern: ; in, A learnable probability distribution weight parameter vector; This represents the combined weight of the probability distribution of the k-th candidate in the m-th omics; Denotes the probability distribution of the k-th candidate; This represents the number of candidate probability distributions for the m-th omics predefined; The parameter vector represents the probability distribution of the k-th candidate omics corresponding to the m-th omics. This represents the sequencing data of the m-th single-cell omics. Represents variables in the latent space Under the given conditions, the probability distribution of the m-th omics data; through adaptive learning of the combined weights. This ensures that the mixed probability distribution matches the true distribution of the original single-cell omics data.
7. The single-cell multi-omics integration method based on adaptive probability distribution modeling according to claim 1, characterized in that, Throughout the training process, a joint loss function is used to optimize the model parameters. This joint loss function includes a fusion loss function for the latent space fusion step and a reconstruction loss based on a weighted mixture probability distribution for the decoding step. Specifically, the fusion loss function comprises latent space distribution constraint loss, latent space alignment loss, and latent space adversarial loss. The joint loss function is expressed as follows: ; Where M represents the number of single-cell omics; This represents the reconstruction loss of the m-th omics; This represents the latent space distribution constraint loss; This represents the latent space alignment loss; denoted as the latent space alignment adversarial loss based on adversarial learning; α, β, and γ are the weight coefficients of the corresponding loss terms.
8. The single-cell multi-omics integration method based on adaptive probability distribution modeling according to claim 7, characterized in that, The reconstruction loss is defined based on the log-likelihood of a weighted mixture probability distribution as follows: ; in, This represents the m-th type of single-cell omics data; This represents the combined weight of the probability distribution of the k-th candidate in the m-th omics; Denotes the probability distribution of the k-th candidate; The parameter vector represents the probability distribution of the k-th candidate omics corresponding to the m-th omics. This represents the number of candidate probability distributions for the m-th omics predefined; Represents the latent space variable Seeking expectations.
9. The single-cell multi-omics integration method based on adaptive probability distribution modeling according to claim 7, characterized in that, The latent space distribution constraint loss, in the form of KL divergence, is used to constrain the distributional consistency between omics-specific latent variables and omics fusion latent variables, as shown below: ; ; in, This represents the omics-specific latent variable corresponding to the m-th omics; express Cross-omics shared fusion latent variables; The variational posterior distribution represents the latent variables of omics fusion shared across omics. ) represents the variational posterior distribution of the omics-specific latent variable corresponding to the m-th omics; This represents a collection of various single-cell omics sequencing data. Represents the divergence function; The latent space alignment loss includes a distribution alignment term based on the maximum mean difference (MMD), expressed as follows: ; in, This represents the function representing the maximum mean difference. Furthermore, the adversarial loss is used to align the omics-specific latent variables and omics fusion latent variables in an adversarial distribution, as shown below: ; in, This represents a modal discriminator; y represents the gradient inversion layer; y represents the omics label of the corresponding latent space variable source. The loss function for adversarial tasks includes the cross-entropy function; It indicates a desire for the expected value.
10. A single-cell multi-omics integration system based on adaptive probability distribution modeling, characterized in that, include: The encoder module is configured to use an omics-independent multi-encoder structure combined with a shared feature layer to extract features from different single-cell omics data, output the mean and variance of the omics-specific latent variables corresponding to each omics, and perform feature fusion on different single-cell omics data to output the mean and variance of the omics fusion latent variables representing cross-omics shared information. Omics-specific latent variables are obtained based on the mean and variance of omics-specific latent variables, and omics-fusion latent variables are obtained based on the mean and variance of omics-fusion latent variables. The latent space fusion module is configured to perform joint modeling based on the omics-specific latent variables and omics fusion latent variables. Through an alternating training strategy, it uses omics-specific latent variables and omics fusion latent variables alternately in different training iterations to reconstruct omics data and obtain the latent space distribution of different omics. A fusion loss function is introduced to constrain the latent space distribution, so that the omics-specific latent variables are aligned with the omics fusion latent variables. The adaptive probability distribution decoder module is configured to adopt an omics-independent multi-decoder structure. Based on the statistical characteristics of the latent space distribution of different omics, it outputs parameters of multiple probability distributions and their combined weights. By adaptively learning the combined weights, it determines the probability distribution combination that best matches the latent space distribution of the original omics, thereby realizing the joint modeling and reconstruction of single-cell multi-omics data.