A precise method for generating single-cell multi-omics matching data
By constructing a chromatin accessibility ATAC model and a chromatin accessibility ATAC model for chromatin accessibility ATAC generation transcriptome RNA model, the problem of high cost and noisy single-cell multi-omics sequencing technology is solved, and high-precision single-cell multi-omics matching data generation is achieved, which is suitable for biological development and disease treatment.
Patent Information
- Application Number
- CN202211449668.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-18
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-11-18
AI Technical Summary
The existing single-cell multiomics sequencing technology is costly and noisy, and the existing algorithms have poor results in generating chromatin accessibility data, and the cross-data set generation ability is insufficient and cannot meet actual needs.
The transcriptome RNA generation chromatin accessibility ATAC model and the chromatin accessibility ATAC generation transcriptome RNA model were constructed. Through pre-training and fine training, the encoder, decoder and discriminator cascade architecture was used, and single-cell multiomics matching data were generated through loss functions such as Wasserstein distance and Focal Loss.
It improves generation accuracy and cross-data set generation capabilities, generates more accurate data, reduces noise, and is suitable for reference data for biological development and disease treatment.
Smart Images

Figure CN115862746B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data mining, and particularly relates to a method for generating single-cell multi-omics matching data, providing reference data for exploring biological development and disease treatment. Background Art
[0002] Recently, with the development of single-cell sequencing technology, there have emerged some sequencing technologies that simultaneously measure multiple omics at the single-cell level, namely cell-matching sequencing technologies. It has been found that multi-omics analysis can obtain more knowledge than single-omics, and using single-cell multi-omics can obtain biological information that cannot be obtained by any single omics. However, single-cell multi-omics sequencing technology is usually complex and noisy; in addition to technical defects, the high cost of single-cell multi-omics sequencing methods also limits the wide application of the technology and the data scale to a certain extent. Therefore, researchers need a method for generating multi-omics matching data, that is, when only one omics sequencing data is available, other omics data is generated by computational methods and combined with the sequencing data to form cell-matching multi-omics data for single-cell multi-omics analysis.
[0003] Previous single-cell multi-omics analysis methods were mainly applied to multi-omics joint analysis, and downstream analysis was carried out through the already sequenced single-cell multi-omics data, but the problems of few current single-cell multi-omics datasets, high sequencing costs, and inability to meet actual needs were not solved.
[0004] To solve the above problems, some algorithms have been developed to generate single-cell multi-omics matching data.
[0005] In 2020, Zilu Zhou et al. proposed an algorithm called cTP-net in Nature Communications to generate cell surface protein data from the sequenced single-cell transcriptome data. cTP-net first used a transfer learning algorithm called SAVER-X to denoise the single-cell transcript data to more accurately obtain the relative RNA abundance of each cell, and then adopted a multi-branch deep neural network to achieve the mapping from single-cell transcriptome to surface protein and improve the generalization ability of the model.
[0006] In 2021, Kevin E. Wu et al. proposed a deep learning algorithm called BABEL in Proceedings of the National Academy of Sciences. BABEL realizes the mutual generation of single-cell chromatin accessibility data and transcriptome data through an interoperable autoencoder network, and at the same time uses the insight that most chromatin accessibility interactions occur within chromosomes to prune the fully connected layers in the network, reducing the running time of the algorithm.
[0007] In 2021, Kodai Minoura et al. proposed a new framework scMM based on deep generative models in Cell Reports Methods for extracting interpretable joint representations and cross-omic generation. scMM interprets the complexity of data through a multi-omics variational encoder based on a multi-expert method and compensates for the limited interpretability of deep learning models through a pseudo-cell generation strategy.
[0008] The above algorithms all tend to generate transcriptomic and proteomic data with relatively simple data distributions, but have poor performance in generating chromatin accessibility data. Chromatin accessibility data has the characteristics of ultra-high dimensions from hundreds of thousands to millions of dimensions, extreme sparsity, and extreme binaryization, and existing algorithms cannot fit this type of data well. At the same time, in actual use, researchers often need to jointly analyze multiple datasets, and existing algorithms all perform poorly in cross-dataset generation. Summary of the Invention
[0009] The object of the present invention is to overcome the deficiencies of the above-mentioned existing technologies and propose an accurate single-cell multi-omics matching data generation method to accurately generate single-cell chromatin accessibility data and transcriptomic data and improve the cross-dataset generation ability.
[0010] The technical solution of the present invention is as follows: preprocess the transcriptomic data and chromatin accessibility data simultaneously measured in single cells; construct a transcriptomic RNA to generate chromatin accessibility ATAC model and a chromatin accessibility ATAC to generate transcriptomic RNA model; perform pre-training and fine-tuning on the models to obtain trained models; obtain cell multi-omics matching data through the trained models in the case of only single omics. The implementation steps are as follows:
[0011] (1) Use single-cell multi-omics sequencing to simultaneously measure the gene expression values and chromatin accessibility values of single cells in the required tissue, generating single-cell multi-omics matching data including single-cell transcriptomics and chromatin accessibility.
[0012] (2) Preprocess the single-cell transcriptomic scRNA-seq data:
[0013] (2a) Delete the genes on the sex chromosomes in the single-cell transcriptomic scRNA-seq data, and delete the low-quality cells with less than 200 or more than 7000 gene expressions.
[0014] (2b) Perform numerical normalization on the deleted data so that the sum of the counts of each cell is the median of all cells, then perform logarithmic transformation on the normalized data and standardize it to zero mean and unit variance.
[0015] (2c) Clip the top and bottom 0.5% values in the data distribution of the standardized data;
[0016] (3) Preprocess the single-cell chromatin accessibility scATAC-seq data:
[0017] (3a) Filter out the peaks that appear in less than 5 cells or more than 10% of the cells in the single-cell chromatin accessibility scATAC-seq data;
[0018] (3b) For the filtered data, determine whether to perform coordinate transformation according to its type:
[0019] If it is Hg38 data, no operation is required;
[0020] If it is Hg19 data, convert the Hg19 coordinates to Hg38 coordinates;
[0021] (4) Use the Leiden algorithm to cluster the transcriptome data, and divide the two largest clusters into the validation set and the test set, and the remaining clusters form the training set;
[0022] (5) Construct a transcriptome RNA generating chromatin accessibility ATAC model:
[0023] (5a) Build an RNA encoder composed of a cascade of an input layer with the dimension of the number of transcriptome genes, a first hidden layer with 256 dimensions, a second hidden layer with 64 dimensions, and a third hidden layer with 16 dimensions. All layers use the LeakyRelu function as the activation function;
[0024] (5b) Build an ATAC decoder composed of a cascade of an input dimension of 16 dimensions, a first hidden layer of 64 dimensions, a second hidden layer of 512 dimensions, and an output layer of the chromatin accessibility data feature dimension. Except for the last layer, the other layers use the LeakyRelu function as the activation function, and the activation function of the last layer of this decoder is set to the sigmoid function;
[0025] (5c) Build an ATAC discriminator composed of a cascade of an input dimension of the chromatin accessibility data feature dimension, a first hidden layer of 1024 dimensions, a second hidden layer of 256 dimensions, a third hidden layer of 64 dimensions, a fourth hidden layer of 16 dimensions, and an output layer of 1 dimension. Except for the 1 dimension of the last layer where no activation function is set, the other layers use the LeakyRelu function as the activation function;
[0026] (5d) Cascade the RNA encoder, the ATAC decoder, and the ATAC discriminator to form a transcriptome RNA generating chromatin accessibility ATAC model, and use the Wasserstein distance as the loss function;
[0027] (6) Input the preprocessed RNA data in step (2) and the preprocessed ATAC data in step (3) into the transcriptome RNA generation chromatin accessibility ATAC model, complete the pre-training of the model by the method of WGAN, then remove the ATAC discriminator, and introduce Focal Loss as the loss function to update the pre-trained network parameters to obtain the trained transcriptome RNA generation chromatin accessibility ATAC model;
[0028] (7) Construct a chromatin accessibility ATAC generation transcriptome RNA model:
[0029] (7a) Establish an ATAC encoder composed of a cascade of a latent representation with an input layer dimension of the chromatin accessibility data feature dimension, a first hidden layer of 512 dimensions, a second hidden layer of 64 dimensions, and a third hidden layer of 16 dimensions. All layers use the LeakyRelu function as the activation function;
[0030] (7b) Establish an RNA decoder composed of a cascade with an input dimension of 16 dimensions, a first hidden layer of 64 dimensions, a second hidden layer of 256 dimensions, and an output layer of the number of transcriptome genes. Except for the last layer, the other layers use the LeakyRelu function as the activation function, and set the activation functions of the last layer of this decoder to the exponential function and the softplus function, and the outputs of these two functions are the mean and divergence of the gene expression distribution respectively;
[0031] (7c) Establish an RNA discriminator composed of a cascade with an input dimension of the number of genes, hidden layers of 1024 dimensions, 256 dimensions, 64 dimensions, and 16 dimensions respectively, and finally projected to 1 dimension. Except for the 1 dimension of the last layer without setting the activation function, the other layers use the LeakyRelu function as the activation function;
[0032] (7d) Cascade the ATAC encoder, RNA decoder, and RNA discriminator to form a chromatin accessibility ATAC generation transcriptome RNA model, and use the Wasserstein distance as the loss function;
[0033] (8) Input the preprocessed RNA data in step (2) and the preprocessed ATAC data in step (3) into the chromatin accessibility ATAC generation transcriptome RNA model, complete the pre-training of the model by the method of WGAN, then remove the RNA discriminator, and introduce the negative binomial distribution loss NB Loss as the loss function to update the pre-trained network parameters to obtain the trained chromatin accessibility ATAC generation transcriptome RNA model;
[0034] (9) Generate matching single-cell multi-omics data for different omics data:
[0035] (9a) For single-cell transcriptome scRNA-seq data only, input the preprocessed single-cell transcriptome data in step (2) into the trained transcriptome RNA to generate chromatin accessibility ATAC model in step (6) to obtain the matching single-cell multi-omics data;
[0036] (9b) For single-cell chromatin accessibility scATAC-seq data only, input the preprocessed single-cell chromatin accessibility data in step (3) into the trained chromatin accessibility ATAC to generate transcriptome RNA model in step (8) to obtain the matching single-cell multi-omics data.
[0037] Compared with the prior art, the present invention has the following advantages:
[0038] 1) Since the present invention designs a cascaded model architecture of an encoder, a decoder, and a discriminator, it can better learn the biological information of the data and improve the generation accuracy compared with the prior art.
[0039] 2) Since the present invention uses a pre-training step, stable generation performance can be maintained during cross-dataset generation.
[0040] 3) Since the present invention uses Focal Loss to fit the data characteristics of chromatin accessibility with class imbalance of 0 and 1 during the training stage, chromatin accessibility data can be accurately generated.
[0041] 4) Since the present invention designs different models and loss functions according to the data characteristics of different omics, the generated data will be denoised, rather than the highly noisy data obtained from the original measurement. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a flowchart for the implementation of the present invention;
[0043] Figure 2 is a diagram of the transcriptome RNA to generate chromatin accessibility ATAC model and chromatin accessibility ATAC to generate transcriptome RNA model constructed in the present invention.
[0044] Figure 3 is a visualization diagram of UMAP for gene expression data generated by the present invention and the prior BABEL method respectively.
[0045] Figure 4 is a visualization diagram of UMAP for chromatin accessibility data generated by the present invention and the prior BABEL method respectively. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The following further elaborates on the embodiments and effects of the present invention in conjunction with the accompanying drawings.
[0047] In this embodiment, a dataset of single-cell combined measurement of human peripheral blood mononuclear cell (PBMC) gene expression RNA and chromatin accessibility ATAC from 10x Genomics is taken as an example. There are 11,909 cells in the dataset, and each cell has 36,601 genes and 103,877 peaks.
[0048] Referring to Figure 1 , a precise single-cell multi-omics matching data generation method in this example is implemented as follows:
[0049] Step 1: Preprocess the single-cell transcriptome scRNA-seq data.
[0050] 1.1) Obtain the single-cell transcriptome scRNA-seq data in the dataset of single-cell combined measurement of human peripheral blood mononuclear cell (PBMC) from 10x Genomics, delete the genes on the sex chromosomes in this data, and delete the low-quality cells with less than 200 or more than 7,000 gene expressions to filter the data, obtaining 11,872 remaining cells and 34,861 genes after filtering;
[0051] 1.2) Perform median normalization on the filtered transcriptome data, that is, divide each column of data by the median of that column of data, then perform logarithmic transformation on the median-normalized data, and standardize it to zero mean and unit variance;
[0052] 1.3) Clip the top and bottom 0.5% values in the data distribution of the standardized transcriptome data to obtain the preprocessed single-cell transcriptome scRNA-seq data X.
[0053] Step 2: Preprocess the single-cell chromatin accessibility scATAC-seq data.
[0054] Obtain the single-cell chromatin accessibility scATAC-seq data in the dataset of single-cell combined measurement of human peripheral blood mononuclear cell (PBMC) from 10x Genomics;
[0055] Delete the peaks that appear in less than 5 cells or more than 10% of the cells in the single-cell chromatin accessibility scATAC-seq data to filter the data, retain 86,111 remaining peaks after filtering, and obtain the preprocessed single-cell chromatin accessibility scATAC-seq data Y.
[0056] Step 3: Data partitioning.
[0057] The preprocessed transcriptome data was clustered using the Leiden algorithm at a resolution of 1.5 to find 17 clusters. The two largest clusters were divided into the validation set and the test set, and the remaining clusters formed the training set. In this example, 8,290 cells were divided into the training set, 1,776 cells into the validation set, and 1,776 cells into the test set.
[0058] Step 4: Construct a transcriptome RNA generating chromatin accessibility ATAC model.
[0059] 4.1) An RNA encoder was established, which consisted of a cascade of layers with an input layer dimension of 34,861 (the number of transcriptome genes), a first hidden layer of 256 dimensions, a second hidden layer of 64 dimensions, and a third hidden layer of 16 dimensions. The LeakyRelu function was used as the activation function for all layers.
[0060] 4.2) An ATAC decoder was established, which consisted of a cascade of layers with an input dimension of 16, a first hidden layer of 64 dimensions, a second hidden layer of 512 dimensions, and an output layer of 86,111 dimensions (the chromatin accessibility data feature dimension). The LeakyRelu function was used as the activation function for all layers except the last layer, and the activation function of the last layer was set to the sigmoid function.
[0061] 4.3) An ATAC discriminator was established, which consisted of a cascade of layers with an input dimension of 86,111 (the chromatin accessibility data feature dimension), a first hidden layer of 1,024 dimensions, a second hidden layer of 256 dimensions, a third hidden layer of 64 dimensions, a fourth hidden layer of 16 dimensions, and an output layer of 1 dimension. The LeakyRelu function was used as the activation function for all layers except the last layer, and no activation function was set for the last layer.
[0062] 4.4) The RNA encoder, ATAC decoder, and ATAC discriminator were cascaded to form a transcriptome RNA generating chromatin accessibility ATAC model, as shown in Figure 2 (a), and the Wasserstein distance was used as the loss function. The formula is as follows:
[0063]
[0064] where P1 represents the distribution of the real data, P2 represents the distribution of the generated data, ∏(P1, P2) is the set of all joint distributions obtained by combining the distributions P1 and P2, γ represents any joint distribution in the set, (a, b) represents the samples obtained by sampling through the joint distribution γ, and E (a,b)~γ [||a - b||] represents the expected value of the distance between this pair of samples, and the infimum of these expected values is defined as the Wasserstein distance between the distributions P1 and P2.
[0065] Step 5, train a chromatin accessibility ATAC model for transcriptome RNA generation.
[0066] 5.1) Input the preprocessed RNA data X from Step 1 into the RNA encoder, output the low-dimensional latent representation, and then input the latent representation into the chromatin accessibility ATAC decoder to generate ATAC data Input the generated ATAC data and the actually measured ATAC data Y into the ATAC discriminator for generative adversarial learning to complete the model pre-training;
[0067] 5.2) Remove the ATAC discriminator, input the preprocessed RNA data X from Step 1 into the pre-trained RNA encoder, output the low-dimensional latent representation, and then input the latent representation into the pre-trained chromatin accessibility ATAC decoder to generate ATAC data
[0068] 5.3) Introduce Focal Loss as the loss function to update the parameters of the pre-trained network until the Focal Loss of the loss function drops to the minimum, thereby completing the model training; The Focal Loss formula is as follows:
[0069]
[0070] where y represents whether each chromatin in the preprocessed single-cell chromatin accessibility scATAC-seq data Y is accessible, y = 1 means accessible, and y = 0 means inaccessible. represents the probability of each chromatin being accessible in the generated ATAC data γ≥0 is a regulatory factor used to balance the weights of easy and difficult samples, and ∝∈[0, 1] is a weight factor used to balance the ratio of 0 and 1 samples.
[0071] Step 6, construct a chromatin accessibility ATAC generation transcriptome RNA model:
[0072] 6.1) Establish an ATAC encoder composed of a cascade of latent representations with an input layer dimension of 86,111 dimensions of chromatin accessibility data features, a first hidden layer of 512 dimensions, a second hidden layer of 64 dimensions, and a third hidden layer of 16 dimensions. All layers use the LeakyRelu function as the activation function;
[0073] 6.2) An RNA decoder is established, which consists of a cascade with an input dimension of 16, a first hidden layer of 64, a second hidden layer of 256, and an output layer of 34,861, which is the number of transcriptome genes. The LeakyRelu function is used as the activation function for all layers except the last layer. The activation functions for the last layer of the decoder are the exponential function and the softplus function, and the outputs of these two functions are the mean and divergence of the gene expression distribution, respectively;
[0074] 6.3) An RNA discriminator is established, which consists of a cascade with an input dimension of 34,861 genes, and hidden layers of 1024, 256, 64, 16, and 1 in sequence. The LeakyRelu function is used as the activation function for all layers except the last layer with a dimension of 1, for which no activation function is set;
[0075] 6.4) The ATAC encoder, RNA decoder, and RNA discriminator are cascaded to form a chromatin accessibility ATAC generating transcriptome RNA model, as shown in Figure 2 (b), and the Wasserstein distance is used as its loss function.
[0076] Step 7: Train the chromatin accessibility ATAC generating transcriptome RNA model.
[0077] 7.1) Input the preprocessed ATAC data Y in Step 2 into the ATAC encoder to output a low-dimensional latent representation, and then input the latent representation into the transcriptome RNA decoder to generate RNA data. Then, input the generated RNA data and the actually measured RNA data X into the RNA discriminator for generative adversarial learning to complete the pre-training of the model.
[0078] 7.2) Remove the RNA discriminator. Input the preprocessed ATAC data Y in Step 2 into the pre-trained ATAC encoder to output a low-dimensional latent representation, and then input the latent representation into the pre-trained transcriptome RNA decoder to generate RNA data.
[0079] 7.3) Introduce the NB Loss as the loss function to update the network parameters of the pre-trained model until the value of the loss function NB Loss decreases to the minimum to complete the model training;
[0080]
[0081] where x represents the expression of each gene in the preprocessed single-cell transcriptome scRNA-seq data X, and θ represents the generated RNA data The mean and dispersion parameters of the negative binomial distribution corresponding to each gene in . The negative binomial distribution reconstructs gene expression through these two parameters. τ represents the gamma function, and ∈ is a small constant to maintain numerical stability.
[0082] Step 8: Generate matching single-cell multi-omics data for different omics data.
[0083] 8.1) For single-cell transcriptome scRNA-seq data only, after preprocessing the single-cell transcriptome data in Step 1, input it into the transcriptome RNA to generate chromatin accessibility ATAC model trained in Step 5 to obtain matching single-cell multi-omics data;
[0084] 8.2) For single-cell chromatin accessibility scATAC-seq data only, after preprocessing the single-cell chromatin accessibility data in Step 2, input it into the chromatin accessibility ATAC to generate transcriptome RNA model trained in Step 7 to obtain matching single-cell multi-omics data.
[0085] The technical effects of the present invention are described below in combination with simulation experiments.
[0086] I. Simulation conditions:
[0087] The computer hardware CPU for the simulation experiment is Intel Core(TM)i7-8700, and the computer hardware memory is 32G; computer software: Python 3.7 and Rstudio integrated development software on the WINDOWS 10 system.
[0088] II. Simulation content:
[0089] Simulation 1: Use Pearson correlation as the evaluation index for the ability of each method to generate RNA data from ATAC data, and use AUROC as the evaluation index for the ability of each method to generate ATAC data from RNA data. Compare the evaluation indexes of the present invention with two existing methods, BABEL and scMM, in multiple single-cell multi-omics data sets that simultaneously measure gene expression and chromatin accessibility in humans and mice. The results are shown in Table 1:
[0090] Table 1 Evaluation of the present invention and two existing methods in human and mouse data sets
[0091]
[0092] As can be seen from Table 1, when generating within the dataset, the performance of the present invention and BABEL is significantly higher than that of scMM, and the present invention is slightly better than BABEL. When generating across datasets, the present invention has a greater advantage over the other two methods, especially when generating ATAC data from RNA data. The simulation results show that the present invention maintains a high accuracy both within the dataset and when generating across datasets.
[0093] Simulation 2: To demonstrate the use in practical applications, the single-cell chromatin accessibility scATAC-seq data of PBMC was used to generate RNA data through the present invention and the existing method BABEL respectively. The UMAP algorithm was used to cluster and visualize the generated RNA data, and each cell was colored using the cell type based on scATAC-seq. The results are as Figure 3 shown. Among them Figure 3 (a) represents the UMAP visualization diagram of the present invention, Figure 3 (b) represents the UMAP visualization diagram of BABEL. UMAP1 and UMAP2 respectively represent the horizontal and vertical coordinates obtained through the UMAP algorithm.
[0094] From Figure 3 it can be seen that the RNA data generated by the present invention can well separate the three major types of cells in PBMC, while the existing BABEL cannot separate B lymphocytes and DC dendritic cells. The simulation results show that although there are different noises in different datasets, the present invention can still complete the task of generating RNA from ATAC.
[0095] Simulation 3: The single-cell transcriptome scRNA-seq data of PBMC was used to generate ATAC data through the present invention and the existing method BABEL respectively. The UMAP algorithm was used to cluster and visualize the generated ATAC data, and each cell was colored using the cell type based on scRNA-seq. The results are as Figure 4 shown. Among them Figure 4 (a) represents the UMAP visualization diagram of the present invention, Figure 4 (b) represents the UMAP visualization diagram of BABEL. UMAP_1 and UMAP_2 respectively represent the horizontal and vertical coordinates obtained through the UMAP algorithm.
[0096] From Figure 4 it can be seen that the chromatin accessibility ATAC data generated by the existing BABEL cannot clearly show the differences between each cell cluster, while the chromatin accessibility ATAC data generated by the present invention can clearly separate each cell cluster. The simulation results show that the present invention can accurately realize the generation of single-cell multi-omics matching data.
Claims
1. A precise method for generating single-cell multi-omics matching data, characterized in that, The following are included: (1) Using single-cell multi-omics sequencing to simultaneously measure the gene expression values and chromatin accessibility values of single cells in the required tissue, generating single-cell multi-omics matching data including single-cell transcriptome and chromatin accessibility; (2) Preprocessing the single-cell transcriptome scRNA-seq data: (2a) Deleting the genes on the sex chromosomes in the single-cell transcriptome scRNA-seq data, and deleting the low-quality cells with less than 200 or more than 7000 gene expressions; (2b) Normalizing the numerical values of the data after deletion so that the sum of the counts of each cell is the median of all cells, then performing logarithmic transformation on the normalized data and standardizing it to zero mean and unit variance; (2c) Clipping the values at the top and bottom 0.5% of the data distribution of the standardized data; (3) Preprocessing the single-cell chromatin accessibility scATAC-seq data: (3a) Filtering out the peaks that appear in less than 5 cells or more than 10% of the cells in the single-cell chromatin accessibility scATAC-seq data; (3b) For the filtered data, determine whether to perform coordinate transformation according to its type: If it is Hg38 data, no operation is required; If it is Hg19 data, convert the Hg19 coordinates to Hg38 coordinates; (4) Using the Leiden algorithm to cluster the transcriptome data, and dividing the two largest clusters into a validation set and a test set, and the remaining clusters form the training set; (5) Constructing a transcriptome RNA generating chromatin accessibility ATAC model: (5a) Establishing an RNA encoder composed of a cascade of an input layer with the dimension of the number of transcriptome genes, a first hidden layer with 256 dimensions, a second hidden layer with 64 dimensions, and a third hidden layer with 16 dimensions. All layers use the LeakyRelu function as the activation function; (5b) Establishing an ATAC decoder composed of a cascade of an input dimension of 16 dimensions, a first hidden layer of 64 dimensions, a second hidden layer of 512 dimensions, and an output layer of the chromatin accessibility data feature dimension. Except for the last layer, the other layers use the LeakyRelu function as the activation function, and the activation function of the last layer of this decoder is set to the sigmoid function; (5c) Establishing an ATAC discriminator composed of a cascade of an input dimension of the chromatin accessibility data feature dimension, a first hidden layer of 1024 dimensions, a second hidden layer of 256 dimensions, a third hidden layer of 64 dimensions, a fourth hidden layer of 16 dimensions, and an output layer of 1 dimension. Except for the 1 dimension of the last layer where no activation function is set, the other layers use the LeakyRelu function as the activation function; (5d) Cascade the RNA encoder, ATAC decoder, and ATAC discriminator to form a transcriptome RNA generating chromatin accessibility ATAC model, and use the Wasserstein distance as the loss function; (6) Input the RNA data preprocessed in step (2) and the ATAC data preprocessed in step (3) into the transcriptome RNA generating chromatin accessibility ATAC model, complete the pre-training of the model by the method of WGAN, then remove the ATAC discriminator, and introduce Focal Loss as the loss function to update the network parameters after pre-training, obtaining the trained transcriptome RNA generating chromatin accessibility ATAC model; (7) Construct a chromatin accessibility ATAC generating transcriptome RNA model: (7a) Establish an ATAC encoder composed of a cascade of a latent representation with an input layer dimension of the chromatin accessibility data feature dimension, a first hidden layer of 512 dimensions, a second hidden layer of 64 dimensions, and a third hidden layer of 16 dimensions. All layers use the LeakyRelu function as the activation function; (7b) Establish an RNA decoder composed of a cascade with an input dimension of 16 dimensions, a first hidden layer of 64 dimensions, a second hidden layer of 256 dimensions, and an output layer of the number of transcriptome genes. Except for the last layer, the LeakyRelu function is used as the activation function for the other layers, and the activation functions of the last layer of this decoder are set as the exponential function and the softplus function, and the outputs of these two functions are the mean and divergence of the gene expression distribution respectively; (7c) Establish an RNA discriminator composed of a cascade with an input dimension of the number of genes, hidden layers of 1024 dimensions, 256 dimensions, 64 dimensions, and 16 dimensions respectively, and finally projected to 1 dimension. Except for the 1 dimension of the last layer where no activation function is set, the LeakyRelu function is used as the activation function for the other layers; (7d) Cascade the ATAC encoder, the RNA decoder, and the RNA discriminator to form a chromatin accessibility ATAC generating transcriptome RNA model, and use the Wasserstein distance as the loss function; (8) Input the RNA data preprocessed in step (2) and the ATAC data preprocessed in step (3) into the chromatin accessibility ATAC generating transcriptome RNA model, complete the pre-training of the model by the method of WGAN, then remove the RNA discriminator, and introduce the negative binomial distribution loss NBLoss as the loss function to update the network parameters after pre-training, obtaining the trained chromatin accessibility ATAC generating transcriptome RNA model; (9) Generate matching single-cell multi-omics data for different omics data: (9a) For single-cell transcriptome scRNA-seq data only, input the single-cell transcriptome data preprocessed in step (2) into the transcriptome RNA generating chromatin accessibility ATAC model trained in step (6) to obtain the matching single-cell multi-omics data; (9b) For single-cell chromatin accessibility scATAC-seq data only, input the single-cell chromatin accessibility data preprocessed in step (3) into the chromatin accessibility ATAC generating transcriptome RNA model trained in step (8) to obtain the matching single-cell multi-omics data.
2. The method according to claim 1, characterized in that, The Wasserstein distance selected by the loss function of the transcriptome RNA generating chromatin accessibility ATAC model in step (5d) and the loss function of the chromatin accessibility ATAC generating transcriptome RNA model in step (7d) is as follows: Among them, P1 represents the distribution of real data, P2 represents the distribution of generated data, П(P1, P2) is the set of all joint distributions obtained by combining distributions P1 and P2, γ represents any joint distribution in the set, (x, y) represents a sample obtained by sampling through the joint distribution γ, E (x,y)~γ [||x - y||] represents the expected value of the distance between this pair of samples, and the infimum of these expected values is defined as the Wasserstein distance between distributions P1 and P2.
3. The method according to claim 1, characterized in that In step (6), the pre-training of the transcriptome RNA generating chromatin accessibility ATAC model is completed by the WGAN method as follows: (6a) Input the RNA data preprocessed in step (2) into the RNA encoder to output a low-dimensional latent representation; (6b) Input the latent representation into the chromatin accessibility ATAC decoder to generate ATAC data; (6c) Input the generated ATAC data and the actually measured ATAC data into the ATAC discriminator for generative adversarial learning to complete the model pre-training.
4. The method according to claim 1, wherein In step (6), for the pre-trained transcriptome RNA generating chromatin accessibility ATAC model, the Focal Loss is introduced as the loss function to update the pre-trained network parameters as follows: (6d) Input the RNA data preprocessed in step (2) into the pre-trained RNA encoder to output a low-dimensional latent representation; (6e) Input the latent representation into the pre-trained chromatin accessibility ATAC decoder to generate ATAC data; (6f) Update the network parameters after pre-training using the Focal Loss function until the value of the Focal Loss function: decreases to the minimum to complete the model training. where y = 1 indicates chromatin accessibility in actual measurement, y = 0 indicates chromatin inaccessibility in actual measurement, p ∈ [0, 1] is the probability of generated chromatin accessibility, γ ≥ 0 is a regulatory factor used to balance the weights of easy and difficult samples, and ∝ ∈ [0, 1] is a weight factor used to balance the ratio of 0 and 1 samples.
5. The method according to claim 1, characterized in that In step (8), the pre-training of the chromatin accessibility ATAC generating transcriptome RNA model is completed by the WGAN method as follows: (8a) Input the ATAC data preprocessed in step (3) into the ATAC encoder to output a low-dimensional latent representation; (8b) Input the latent representation into the transcriptome RNA decoder to generate RNA data; (8c) Input the generated RNA data and the actually measured RNA data into the RNA discriminator for generative adversarial learning to complete the model pre-training.
6. The method according to claim 1, wherein In step (8), for the pre-trained chromatin accessibility ATAC generating transcriptome RNA model, the negative binomial distribution loss NB Loss is introduced as the loss function to update the pre-trained network parameters as follows: (8d) Input the ATAC data preprocessed in step (3) into the pre-trained ATAC encoder to output a low-dimensional latent representation; (8e) Input the latent representation into the pre-trained transcriptome RNA decoder to generate RNA data; (8f) Use the loss function NB Loss to update the pre-trained network parameters until the value of the loss function NB Loss, NBL(x; μ, θ), drops to the minimum, i.e.: NBL(x; μ, θ) = -θ(log(θ + ∈) - log(θ + μ)) - μ(log(μ + ∈) - log(θ + μ)) - logτ(x + θ) + logτ(x + 1) + logτ(θ + ∈), to complete the model training. Among them, x represents the gene expression of the transcriptome, μ and θ represent the mean and dispersion parameters of the negative binomial distribution corresponding to each gene respectively. The negative binomial distribution reconstructs the gene expression through these two parameters. τ represents the gamma function, and ∈ is a small constant to maintain numerical stability.
Citation Information
Patent Citations
Method for analyzing active pathway in single cell multi-omics based on graph neural network
CN115240772A
Method for determining PARP inhibitor responsiveness and improving PARP inhibitor therapy
EP3822637A1