A method for generating single-cell sequencing data
By building a deep neural network model, the high cost and data quality problems of single-cell sequencing technology are solved, high-quality single-cell sequencing data is generated, and cell differentiation tracking and heterogeneity analysis are supported.
Patent Information
- Application Number
- CN202411113195.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-14
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-08-14
AI Technical Summary
Existing single-cell sequencing technologies are expensive and difficult to expand existing data. In addition, data integration from different sequencing platforms will introduce batch effects, affecting data quality and the accuracy of subsequent analysis.
A deep neural network model, including an encoder, diffusion module, and decoder, was constructed to generate high-quality single-cell sequencing data by preprocessing and training the comprehensive scRNA-seq dataset.
Generating high-quality scRNA-seq data on multiple sequencing platforms reduces model training costs and inference time, provides higher data quality and diversity, and supports downstream tasks such as cell differentiation tracking and heterogeneity analysis.
Smart Images

Figure CN119132389B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for generating single-cell sequencing data, and belongs to the technical field of bioinformatics. Background Art
[0002] Cells are the fundamental structural and functional units of organisms. To study complex biological systems such as developing embryos, tumors, and microorganisms, it is necessary to understand the behavior, cellular heterogeneity, and intercellular interactions of the individual cells that comprise these systems. With the rapid advancement of technology, a large number of single-cell sequencing technologies, primarily focused on single-cell transcriptomes, have emerged. Traditional sequencing (bulk RNA-seq) generates average transcriptome data from a population of cells and is unable to analyze individual cells. Currently under development, single-cell sequencing (scRNA-seq) analyzes the genome, transcriptome, and epigenome at the level of individual cells. Single-cell multi-omics sequencing simultaneously analyzes the composition and characteristics of multiple biomacromolecules, including the genome, transcriptome, epigenome, and proteome, in a single cell. This advancement in technology has enabled researchers to gain a deeper understanding of the function and diversity of individual cells. Data generated by single-cell sequencing and analyzed through single-cell data can reveal the mechanisms and regulatory networks underlying biological processes. Single-cell data is difficult to analyze due to its high dimensionality, information redundancy, and significant noise. Appropriate dimensionality reduction and feature selection can effectively remove noise and support downstream biological analyses such as visualization, clustering, and trajectory inference. Single-cell data analysis mainly involves data preprocessing (dimensionality reduction, data interpolation, batch correction, etc.) and biological downstream task analysis (clustering, communication between cells, trajectory analysis, etc.).
[0003] Single-cell technologies enable the study of genomics, transcriptomics, and multi-omics in individual cells within tissues. Single-cell RNA sequencing (scRNA-seq) detects gene expression and other information at the single-cell level. However, the single-cell data generated using single-cell sequencing cannot be directly applied to downstream analytical tasks. Instead, these data require dimensionality reduction, denoising, and missing value filling. High-dimensional single-cell data contain a large number of zero values, which can arise from biochemical zeros (i.e., gene expression is not present in the cell) and technical errors (failure to detect the gene due to low expression or other factors). The large number of missing values in the data can affect biological signals. Single-cell data often come from multiple experiments with varying capture times, processing personnel, reagent batches, equipment, and even technical platforms. These differences can lead to data anomalies or batch effects, which can potentially confound cell-specific information during data merging. As scRNA-seq data continue to grow, effective data batch integration is crucial.
[0004] scRNA-seq is a high-throughput sequencing technology that analyzes thousands to millions of cells in a single experiment. It can identify heterogeneity among cell types, states, and populations, and provides gene expression information but not the spatial location of cells. Transcriptome (ST) data is a low-throughput sequencing technology that excels at capturing spatial information of gene expression. It can understand the distribution of gene expression in tissue sections and obtain average expression values for multiple cells, but it cannot obtain expression values for individual cells. ST provides complementary information about gene expression and spatial location to scRNA-seq of dissociated cells. Because a large amount of scRNA-seq data has been generated, it is hoped that the relationship between gene expression and spatial location learned from ST can be used to recover the cell location information of scRNA-seq. Recovering the location information of scRNA-seq data can facilitate downstream analysis tasks, such as cell-to-cell communication.
[0005] The study of scRNA-seq data plays a vital role in many aspects. scRNA-seq can help identify and describe different cell types in heterogeneous cell populations, recognize the heterogeneity of cancer cells and develop personalized treatments, understand the complexity of immune cell populations and their responses to stimuli, analyze the differentiation and regulation of stem cells and discover stem cell markers, and deeply understand the regulatory networks that control gene expression.
[0006] When performing downstream analysis on scRNA-seq data, sufficient, high-quality scRNA-seq data helps researchers draw relatively accurate conclusions. However, factors such as the high cost of sequencing, low measurement accuracy for rare cell types, and data privacy can affect the quantity and quality of sequencing data, potentially impacting the accuracy of subsequent analyses. To address this issue, researchers can conduct multiple experiments using different sequencing technologies to collect more samples. This not only increases experimental costs, but the integration of data from different sequencing platforms can also introduce batch effects, leading to a decrease in data quality. Furthermore, due to privacy concerns and other factors, researchers also find it difficult to expand existing data. Summary of the Invention
[0007] The present invention aims to solve the technical problems of high cost and difficulty in expanding existing data in scRNA-seq data sequencing experiments in the prior art, and further proposes a method for generating single-cell sequencing data;
[0008] The technical solution adopted by the present invention to solve the above problems is: the present invention proposes a method for generating single-cell sequencing data, comprising:
[0009] S1: Build a deep neural network model;
[0010] S2: Obtaining comprehensive scRNA-seq datasets;
[0011] S3. Preprocessing the scRNA-seq comprehensive dataset;
[0012] S4: Training of deep neural network models based on preprocessed scRNA-seq comprehensive datasets;
[0013] S5: Based on the trained deep neural network model, the features of the test data are denoised, de-noised, and reconstructed to generate single-cell sequencing data.
[0014] Optionally, the deep neural network model in S1 consists of an encoder, a diffusion module, and a decoder;
[0015] The encoder and decoder form an autoencoder;
[0016] The encoder is used to extract gene expression data features from input data;
[0017] The diffusion module is used to add noise and denoise the extracted gene expression data features;
[0018] The decoder is used to reconstruct the denoised gene expression data features to generate single-cell sequencing data.
[0019] Optionally, the comprehensive scRNA-seq dataset in S2 includes scRNA-seq datasets under multiple sequencing technologies, scRNA-seq datasets with multiple attributes, and scRNA-seq datasets of cell growth and development at pseudo-time nodes.
[0020] Optionally, the steps for obtaining scRNA-seq datasets using multiple sequencing technologies include:
[0021] S20101: The deep neural network model has a skipping number of 5 in the inference phase and is compared with the existing four generative models: scDiffusion, cscGAN, LSH-GAN, and sciGAN.
[0022] S20102: Use SCC, PCC, Wasserstein, MMD, and ILISI metrics to evaluate the quality of single-cell gene expression data generated by these models and use UMAP to visualize single-cell gene expression data;
[0023] S20103: Use different sequencing technologies and combine single-cell gene expression data to obtain scRNA-seq datasets using multiple sequencing technologies.
[0024] Optionally, the steps for obtaining a scRNA-seq dataset with multiple attributes include:
[0025] S20201: Classify the single-cell data in the muris_mam_spl_T_B dataset into organ types and cell types. Organ types include mammary gland and spleen, and cell types include B cells and T cells.
[0026] S20202: Construct training labels based on the divided organ type and cell type single-cell data, including breast T cells, breast B cells, spleen T cells, and spleen B cells;
[0027] S20203: Train the diffusion module based on the training labels and guide the diffusion module to generate scRNA-seq datasets with multiple attributes.
[0028] Optionally, the steps for obtaining a scRNA-seq dataset of cell growth and development at a pseudo-time node include:
[0029] S20301: Obtain Wadding-OT data, select the cell states at the first 15 time nodes in the Wadding-OT data, and use the time nodes as data labels;
[0030] S20302: Set the number of jump steps of the deep neural network model at the inference node to 5, use the cell state data and data labels of the first 15 time nodes in the Wadding-OT data to train the deep neural network model, and obtain the scRNA-seq dataset of cell growth and development at pseudo-time nodes.
[0031] In the present invention, cfDiffusion can generate high-quality scRNA-seq data on datasets from multiple sequencing platforms, ensuring the generation of higher-quality scRNA-seq data with multiple attributes while significantly reducing the model training cost and inference time.
[0032] Optionally, the steps for preprocessing the scRNA-seq comprehensive dataset in S3 include:
[0033] S301: Convert the scRNA-seq comprehensive dataset into a gene expression matrix X n×m ;
[0034] S302: Gene expression matrix X n×m Perform normalization operation and transform the gene expression matrix X n×m The total technical scale of cells in is 10000. An offset of 1 is added and the logarithm is taken to obtain the normalized gene expression matrix S ori ;
[0035] The expression of the normalization operation is:
[0036]
[0037] In formula (1), the rows of X represent different cells, with a total of n cells, and the columns of X represent genes, with a total of m genes in each cell.
[0038] Optionally, the steps of training the deep neural network model in S4 include:
[0039] S401: The normalized gene expression matrix S ori Input encoder, the encoder uses 2 MLP, S ori After the encoder, a 128-dimensional embedding layer x0 is obtained;
[0040] S402: Input x0 into the fully connected network in the diffusion module for continuous noise addition, perform continuous denoising through reverse diffusion, and output the denoised gene expression matrix S rec ;
[0041] S403: The denoised gene expression matrix S rec Input decoder, the decoder uses 3 MLPs, x0 passes through the decoder to obtain the same matrix as the gene expression matrix S ori The gene expression matrix S of the same dimension rec , and the gene expression matrix S rec Reconstruct and generate single-cell sequencing data;
[0042] The calculation formula of the embedding layer x0 is:
[0043] x0=Encoder(s ori ) (2);
[0044] In formula (2), Encoder is the encoder;
[0045] Gene expression matrix S rec The calculation formula is:
[0046] S rec =Decoder(x0) (3);
[0047] In formula (3), Decoder is the decoder;
[0048] The denoised gene expression matrix S rec The reconstructed expression is:
[0049] Loss=MSE(S rec ,S ori ) (4).
[0050] Optionally, the diffusion module in S402 includes a D1 layer, a D2 layer, a U1 layer, a U2 layer, a FG1 layer, and a FG2 layer;
[0051] Output denoised gene expression matrix S rec The steps include:
[0052] S40201: Get the embedding x at t time steps t , the time information t and label information y are integrated into the Chunk in the fully connected network, and x0 is updated based on each Chunk to obtain the noisy x0;
[0053] The deep neural network model constructed by the present invention can simulate single-cell data under a pseudo-time scale, providing high-quality data support for tracking cell differentiation and development trajectories, analyzing cell-to-cell communication, revealing cell heterogeneity, and other analyses.
[0054] S40202: Train the diffusion module to obtain the reverse diffusion diffusion module, set the number of jumps to 3, and in the Tth time step, change the x in x0 to t Input the trained fully connected network to get the denoised embedding x t-1 ;
[0055] S40203: In the T-1th time step, x t-1 After passing through the D1 layer, it directly passes through the U2 layer to calculate the denoised x t-2 ;
[0056] S40204: Repeat S40202-S40203 to update the noise sequence and jump process until the denoised x0 is obtained;
[0057] The present invention uses a jump calculation method to accelerate the reasoning speed of the deep neural network model, and increases the fusion degree of the data generated by the deep neural network model and the original data.
[0058] After denoising, x t-2 The calculation formula is:
[0059] ∈ t-2 =FC2{FC1[D1(x t-1 ,t-1,y)+U2(x save ,t-1,y)]} (5);
[0060] x t-2 =Denoise(x t-1 ,∈ t-2 ) (6);
[0061] The expressions for updating the noise sequence and the jump process are:
[0062]
[0063] In formula (7), is the high-level feature numbered b-1 at time step t.
[0064] Optionally, the diffusion module for obtaining reverse diffusion in S40202 specifically includes:
[0065] The diffusion module is parameterized as ∈ θ , and add label information y during the training process to obtain the learned noise ∈ θ , and subtract the learned noise ∈ from the noised x0 θ After that, the diffusion module of reverse diffusion is obtained by removing the logarithm from the Bayesian formula and introducing the hyperparameter k;
[0066] After denoising, x t for:
[0067]
[0068] The Bayesian formula and the expressions for taking logarithms are:
[0069]
[0070] Introducing a hyperparameter k:
[0071]
[0072] In formula (10), k is used to adjust the fidelity and diversity of the generated data, and its value range is (1, 2).
[0073] The beneficial effects of the present invention are:
[0074] 1. This paper proposes the cfDiffusion model. Based on the Diffusion model, by introducing a Classifier-Free method and a high-level feature caching mechanism, it significantly reduces the model's training cost and inference time while ensuring the generation of higher-quality scRNA-seq data with multiple attributes.
[0075] 2. The cfDiffusion model in this paper can generate high-quality scRNA-seq data on datasets from multiple sequencing platforms. Compared with other state-of-the-art data generation models, the cfDiffusion model demonstrates superior performance across multiple evaluation metrics.
[0076] 3. The cfDiffusion model constructed in the present invention can also simulate single-cell data under a pseudo-time scale, providing high-quality data support for tracking cell differentiation and development trajectories, analyzing intercellular communication, revealing cell heterogeneity, and other analyses. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1A flowchart of a method for generating single-cell sequencing data provided by the present invention;
[0078] Figure 2 A structural diagram of the deep neural network model provided by the present invention;
[0079] Figure 3 Comparison chart of cfDiffusion and scDiffusion simulation data quality assessment under multiple conditions provided by the present invention
[0080] Figure 4 The quality assessment diagram of cfDiffusion and scDiffusion simulation data under the combined conditions provided by the present invention;
[0081] Figure 5 Distribution diagrams of Alles, Baron_Human, Marshall, and Mizrak data simulated by the five UMAP visualization methods provided by the present invention;
[0082] Figure 6 The UMAP visualization method provided by the present invention simulates the distribution of WOT data and the random forest ROC curve diagram;
[0083] Figure 7 The UMAP visualization provided by the present invention shows the distribution of real data and generated data of muris, pbmc68k, and human_lung in two-dimensional space. DETAILED DESCRIPTION
[0084] Specific implementation method 1: In this specific implementation method, the cfDiffusion model is a deep neural network model, combined with Figure 1-Figure 7 This embodiment is described as follows. Figure 1 As shown, the method for generating single-cell sequencing data described in this embodiment includes the following steps:
[0085] S1. Obtain scRNA-seq dataset.
[0086] This implementation compares cfDiffusion (setting the number of jump steps of cfDiffusion to 5 during the inference phase) with four existing generative models: scDiffusion, cscGAN, LSH-GAN, and sciGAN. The quality of the single-cell gene expression data generated by these models is evaluated using the SCC, PCC, Wasserstein, MMD, and ILISI metrics, and UMAP is used to visualize the single-cell gene expression data. In this experiment, scRNA-seq datasets were selected, including Alles, Baron_Human, Marshall, and Mizrak. Alles and Mizrak use Drop-seq sequencing technology, Baron uses InDrop sequencing technology, and Marshall uses 10X sequencing technology. The cell types of these four datasets are different. The reason for selecting datasets with different characteristics is to evaluate the robustness of each model in generating single-cell gene expression data in different scenarios.
[0087] The scRNA-seq dataset obtained in this embodiment is as follows:
[0088] The single-cell data in the muris_mam_spl_T_B dataset has two attributes: organ type and cell type, and is composed of mammary gland T cells, mammary gland B cells, spleen T cells, and spleen B cells. When training cfDiffusion, this embodiment uses organ type (breast, spleen) and cell type (B cells, T cells) as two labels for pairing single-cell gene expression data, and trains Diffusion together. In the inference stage (setting the number of jump steps of cfDiffusion to 5), the organ type and cell type are paired into four labels: mammary gland T cells, mammary gland B cells, spleen T cells, and spleen B cells, and the label information is used to guide Diffusion to generate specific single-cell gene expression data. According to Figure 3 As shown in A, most of the generated data is overlaid on the real data, and some generated data is scattered around the real data. The model can learn the attribute characteristics of organs and cell types. This embodiment selects the marker gene Cd74 of breast B cells and plots the expression distribution of gene Cd74 in real cells other than breast B cells, real breast B cells, and generated breast B cells into a box plot as shown in Figure 3As shown in .CD, the Wilcoxon rank sum test calculated the p-value of Cd74 gene expression in the breast B cell data simulated by cfDiffusion and the real breast B cell data to be 0.07618, indicating that there is no significant difference between the two distributions. However, the p-value of Cd74 gene expression in the breast B cell data simulated by scDiffusion and the real breast B cell data is 0.00017, indicating that there is a significant difference between the two distributions.
[0089] Wadding-OT involves 18 days of differentiation of pluripotent stem cells (iPSCs), with a recording interval of 0.5 days. A total of 254,197 cell differentiation states were sampled, retaining 19,423 genes. Due to the large dataset, this implementation selects cell states at 15 time points: Day 0, Day 0.5, Day 1, Day 1.5, Day 2, Day 2.5, Day 3, Day 4.5, Day 5, Day 5.5, Day 6, Day 6.5, Day 7, Day 7.5, and Day 8, for a total of 82,920 cells. This implementation uses these time points as labels for the gene expression data. cfDiffusion is trained using the paired gene expression data and labels. During the inference phase, the number of cfDiffusion jumps is set to 5. This implementation generates 5,600 cell states for each time point, as shown in Tables 13 and 14. The data generated by cfDiffusion is evaluated using various metrics. In UMAP's dimensionality reduction and visualization of data, Figure 6 As shown in Figure 2, the generated data and the real data distribution are roughly the same, and the SCC and PCC values are both greater than 0.99, which indicates that the gene expression pattern of the generated single-cell data is extremely similar to the gene expression pattern of the real single-cell data. Figure 6 As shown, the ILISI value is 0.8198, representing the degree of mixing between the generated single-cell data and the real single-cell data. Furthermore, the training, testing, and AUC values of the random forest model indicate that the trained random forest model is not overfitting. Due to the similarity between the generated data and the real data, it is largely impossible to distinguish whether the single-cell data is model-generated. This embodiment uses real data and model-generated data at different time points to train a K-nearest neighbor classifier. The average prediction accuracy at each time point is approximately 0.5, which also reflects the high quality of the single-cell data generated by cfDiffusion at each time point.
[0090] Tabula Muris: Single-cell RNA sequencing was performed on the Illumina NovaSeq 6000 platform on 20 tissues from 3-month-old mice. Single-cell data from 12 organs were selected, including Bladder, Heart and Aorta, Kidney, Limb and Muscle, Liver, Lung, Mammary Gland, Marrow, Spleen, Thymus, Tongue, and Trachea. A total of 57,004 cells were sequenced, each with 18,996 genes.
[0091] PBMC68k: Peripheral blood mononuclear cells (PBMCs) are blood cells that are an important component of the immune system, fighting infection and protecting the body from harmful pathogens. In biomedical research, PBMCs are commonly used to study global immune responses to disease outbreaks and progression, pathogen infection, vaccine development, and a variety of other clinical applications. This dataset contains single-cell data for 11 cell types: CD14+Monocyte, CD19+B, CD34+, CD4+THelper2, CD4+ / CD25 T Reg, CD4+ / CD45RA+ / CD25-Naive T, CD4+ / CD45RO+Memory, CD56+NK, CD8+Cytotoxic T, CD8+ / CD45RA+Naive Cytotoxic, and Dendritic. A total of 68,579 cells are included, with each cell having 17,789 genes.
[0092] Waddington-OT: A mouse embryonic fibroblast (MEF) cell reprogramming dataset used to study the effects of induced pluripotent stem cells (iPSCs) and growth differentiation factor (GDF9) on reprogramming. This dataset contains cells at different time stamps during the 18-day reprogramming process. Data from three datasets were used, and cells with fewer than 10 expression counts and genes expressed in fewer than three cells were filtered out, resulting in a total of 82,920 cells, each containing 19,423 genes.
[0093] Homo Sapiens: Measured on the Illumina NextSeq 500 platform, there are fibroblast, macrophage, memory bcell, There are six types of cells, nkcell, plasmcell, which come from Bladder, Blood, Spleen, Thymus, and Vasculature organs. The number of cells is 46,177, and each cell has 28,231 genes.
[0094] Human_PF_Lung: AT1, AT2, BCells, Basal, Ciliated, DifferentiatingCiliated, EndothelialCells, Fibroblasts, HAS1HighFibroblasts, KRT5- / KRT17+, Ly measured on human lung cells mphaticEndothelialCells,MUC5AC+High,MUC5B+,Macrophages,MastCells,MesothelialCells,Monocytes,Myofibroblasts,NKCells,PLIN2+Fibroblasts,PlasmaCells,ProliferatingEpithelialCel There are 31 types of cells, including ls, Proliferating Macrophages, Proliferating TCells, SCGB3A2+, SCGB3A2+SCGB1A1+, Smooth Muscle Cells, TCells, Transitional AT2, cDCs, and pDCs. The number of cells is 114,396, and each cell has 27,281 genes.
[0095] muris_mam_spl_T_B: Consists of mammalian mammary T cells, mammary B cells, splenic T cells, and splenic B cells. Cells with expression counts less than 10 and genes expressed in less than 3 cells are filtered out, thus retaining 11,330 cells, each with 14,652 genes.
[0096] Alles: The Alles dataset consists of 14 cell types: LVM longitudinal visc. muscle, amnioserosa, developing midgut, epidermis, fat body, germ cells, head mesoderm / hemocyte differentiation, midgut, muscle, neurogenesis, neurons, undifferentiated cells, visceral muscle, and yolk. Cells with expression counts less than 10 and genes expressed in less than 3 cells were filtered out, retaining 4,614 cells, each with 14,850 genes.
[0097] Baron Human: Using InDrop, a droplet-based single-cell RNA-seq method, we obtained the transcriptome of human pancreatic cells, including 14 cell types: acinar, activated_stellate, alpha, beta, delta, ductal, endothelial, epsilon, gamma, macrophage, mast, quiescent_stellate, schwann, and t_cell. Cells with expression counts less than 10 and genes expressed in less than 3 cells were filtered out, retaining 8,569 cells, each with 16,359 genes.
[0098] Marshall: Single-cell transcriptome data of Mus musculus liver were measured on the 10X platform. There were 16 types of cells in total, including B1, B2, B3, B4, B5, B6, B7, CD8 T, Cd4 T, Monocyte, Monocyte derived Macrophage, NK, Redpulp Macrophage, Treg, cDC, and pDC. Cells with expression counts less than 10 and genes expressed in less than 3 cells were filtered out, retaining 6022 cells, each with 16548 genes.
[0099] Mizrak: Drop-seq technology was used to analyze the V-SVZ in male and female mice aged 8-10 weeks. There were 17 types of cells, including Astrocyte, Astrocytes, COP, Doublet, Endothelial, Ependymal, Microglia, MuralCell+Fibroblast, Mural Cell+Fibroblast, Mural+Fibroblast, Neuron, OPC, Oligodendrocyte, Oligodendrocytes, T Cell, aNSC+TAC+NB, T cell, totaling 42,374 cells, each with 48,529 genes.
[0100] S2, build a deep neural network model, such as Figure 2 As shown, the cfDiffusion model framework constructed in this embodiment is used to generate single-cell gene expression data;
[0101] S201: The cfDiffusion model framework consists of an encoder, a diffusion model, and a decoder. The encoder extracts key features from gene expression data, while Diffusion continuously adds and removes noise from these features, ultimately reconstructing them in the decoder to generate the gene expression data.
[0102] S202: Diffusion Backbone network structure. This is designed based on the UNet model, where the D1, D2, U1, U2, FC1, FC2, and Mid modules all use MLP. This implementation integrates time information t and label information y into the D1, D2, U1, U2, and Mid modules. The right half of the figure shows how time information, label information, and feature information are integrated. t is the feature information of a certain time step, and z is the feature information passed in by the previous module. For example, module D2 obtains the feature information passed in by D1, and module Mid obtains the feature information passed in by D2.
[0103] S203: When the number of jump steps is 3 and the module designated for caching information is U2, the model accelerates the process of generating single-cell gene expression data during the inference phase. First, the feature information of U2 at time step t is cached. In the following time steps t-1, t-2, and t-3, the cached feature information is passed to module U1. After U1 completes feature information processing, it is fused with the feature information of module D1. During the entire process of processing feature information, it is only processed by modules D1, U1, FC1, and FC2, and not by the gray modules (D2, Mid, U2). The operation process of time step t-4 is consistent with that of time step t, realizing the update of cached information.
[0104] S3. Deep neural network model training.
[0105] S301: Train the autoencoder. n×m S is obtained by normalization (scaling the total count of each cell to 10,000, adding a small offset of 1, and finally taking the logarithm) ori As the input of VAE, the rows of X represent different cells, there are n cells in total, and the columns of X represent genes, and each cell has m genes in total.
[0106]
[0107] This operation not only helps to reduce the difference in sequencing depth between cells and the impact of extreme values on analysis, but also stabilizes the variance of the data and converts the distribution of the original data into an approximate normal distribution, which is consistent with the Gaussian distribution used in the diffusion process, making it easier for Diffusion to learn the reverse process. The VAE encoder uses two MLPs, S ori After the encoder, we get the 128-dimensional embedding layer (Embedding) x0:
[0108] x0=Encoder(s ori ) (2);
[0109] The decoder of VAE uses 3 MLPs, and x0 is obtained by the decoder with the original gene expression matrix S ori Same dimensions:
[0110] S rec =Decoder(x0) (3);
[0111] The goal of optimization is to maximize S rec Reconstructed into S ori :
[0112] Loss=MSE(S rec ,S ori ) (4);
[0113] S302: Training Classifier-Free Diffusion. In the forward process of Diffusion, adding noise to the embedding x0 in a series of random time steps can obtain the embedding x at the tth time step. t :
[0114]
[0115] In formula (5), ∈ fuses t-1 Gaussian distributions. When t is large enough, x t Approximately obeys Gaussian distribution. Due to the high sparseness and high dimensionality of gene expression matrix, directly using U-Net convolution operation is not able to capture the correlation between two genes that are far apart in the vector, or the captured correlation is weak. MLP can solve this problem. Therefore, using a fully connected network instead of convolution layer can make Diffusion generate single cell data with better performance. Improved U-Net such as Figure 2 B. The time information t and label information y are integrated into each Chunk. Taking the first Chunk as an example, the specific operation is shown in formula (6):
[0116] x t =x t⊙y+t (6);
[0117] In formula (6), ⊙ is the dot product operation of the matrix, after a Chunk t In the improved U-Net, let the number of Chunks be n and n be an odd number, and The information of each Chunk is connected through a jump, directly connecting the low-level information to the next Chunk, as shown in formula (7):
[0118] U i+1 =U i (·)+D i (·) (7);
[0119] In formula (7), U i After Chunk, D i It is before Chunk, A Chunk is equivalent to the middle layer of the U-Net network. t The entire process of , t, and y passing through the entire U-Net is shown in the pseudo code Algorithm 1.
[0120]
[0121]
[0122] S5: Train Diffusion so that it learns from x in the reverse process t Noise reduction is x t-1 , the process can be expressed as: p(x t-1 |x t ). The Diffusion model is parameterized as ∈ θ , and add label information y during the training process, assuming ∈ θ (x t ,t,y) obeys the standard normal distribution, according to formula (8):
[0123]
[0124] In fact, the noise learned by Diffusion ∈ θ It is approximately subject to normal distribution. According to TweedieEstimator, we can further obtain formula (9):
[0125]
[0126] Then by Bayes' formula and taking the logarithm:
[0127]
[0128] Introducing a hyperparameter k:
[0129]
[0130] In formula (11), k is used to adjust the fidelity and diversity of the generated data, and its value range is (1, 2).
[0131] S501: Diffusion inference acceleration;
[0132] S50101: In the reverse diffusion process, x i and its neighboring feature x i-1 、x i-2 Similarly, in consecutive time steps, redundant calculations are eliminated by skipping some U-Net layers. To better understand the entire accelerated reasoning process, this implementation sets the number of jumps to 3, as shown in Figure 2 As shown in C, each time it jumps to branch chunk5, that is, upsampling U1, U i The features above are called high-level features, D i The features above are called low-level features, and a high-level feature will be saved for skip calculation. Figure 3 As shown in C, in this embodiment, it can be observed that the features of the U1 layer at the Tth time step in the inference process are saved as x save , x t After passing through the entire U-Net and denoising, we get the embedding x t-1 In the T-1th time step, x t-1 It does not pass through D2, M, and U1 layers, but passes through U2 directly after D1 calculation. The calculated ∈ t-2 As shown in formula (12):
[0133] ∈ t-2 =FC2{FC1[D1(x t-1 ,t-1,y)+U2(x save ,t-1,y)]} (12);
[0134] Finally, further denoising is obtained to obtain x t-2 :
[0135] x t-2 =Denoise(x t-1 ,∈ t-2 ) (13);
[0136] S50102: In the T-2 and T-3 time steps, this embodiment can obtain x according to the jump calculation method.t-3 、x t-4 , so far it has jumped 3 times. In the T-4th time step, x t-4 Need to go through the entire U-Net to update x save , and then use the same jump calculation method to get x t-5 、x t-6 、x t-7 The same process is repeated until x0 is obtained in this embodiment.
[0137] S502: If Figure 7 As shown in Table 1-5, in the inference phase of the model, the generation of a specific type of single-cell gene expression data using no classifier is slow. We use the jump calculation method to accelerate its inference speed. Since the jump interval is too small, the comparison of experimental results is not very significant, so the jump interval is set to 5. Figure 7 Table 2 shows the change in the quality of the generated single-cell gene expression data as the number of hop intervals increases. The most obvious trend is the gradual reduction in time overhead, as shown in Table 2. Tables 1-4 show that when the number of hop intervals increases from 5 to 20, the quality of the generated single-cell expression data decreases slightly, indicating that the features between adjacent time steps are similar and there is some redundancy. As shown in Table 5, when the number of hop intervals reaches 50, the quality of the generated single-cell expression data decreases significantly, indicating that some important features are deleted. Although data quality decreases between 5 and 20 hop intervals, the overall performance of the metrics shows that the data generated by cfDiffusion is not inferior to that of scDiffusion. For example, in experiments on the WOT dataset, the random forest is more difficult to distinguish the data generated by cfDiffusion than the scDiffusion data. The increase in the ILISI metric also indicates that the fusion of the cfDiffusion data with the original data is increasing.
[0138] Table 1
[0139]
[0140] Table 2
[0141]
[0142] Table 3
[0143]
[0144] Table 4
[0145]
[0146] Table 5
[0147]
[0148] like Figure 7 Figure 2 shows UMAP visualization of the distribution of real and generated data for muris, pbmc68k, and human_lung in two-dimensional space. The quality of the generated data changes when the model jumps to 5, 10, 20, and 50 steps. (A) The quality of the data generated by cfDiffusion for muris changes when the jumps to 5, 10, 20, and 50 steps. (B) The quality of the data generated by cfDiffusion for pbmc68k changes when the jumps to 5, 10, 20, and 50 steps. (C) The quality of the data generated by cfDiffusion for human_lung changes when the jumps to 5, 10, 20, and 50 steps.
[0149] Now generalize all the conditions, first define a time step to update x save The sequence ξ={p1,p2,…,p m}, where p i ∈[0,999],m≤1000. If the jump step size is fixed, then |p i -p i+1 |=c, c is a constant; if the jump step size is dynamically changing, then P(|p i -p i+1 |=|p i+1 -p i+2 |)>0, where P(·) represents the probability of something happening. Specify a jump to any U b , b represents the number. Since the structure of U-Net is symmetrical, U b The corresponding chunk is D b Update x save The jumping process is shown in formula (14):
[0150]
[0151] in represents the high-level feature numbered b-1 at time step t.
[0152] Example
[0153] As described in S1 of the specific implementation method, this example compares cfDiffusion (setting the number of jump steps of cfDiffusion to 5 in the inference phase) with the four existing generative models of scDiffusion, cscGAN, LSH-GAN, and sciGAN. The quality of the single-cell gene expression data generated by these models is evaluated using the SCC, PCC, Wasserstein, MMD, and ILISI indicators, and the single-cell gene expression data is visualized using UMAP. In this experiment, the Alles, Baron_Human, Marshall, and Mizrak datasets were selected. Alles and Mizrak use Drop-seq sequencing technology, Baron uses InDrop sequencing technology, and Marshall uses 10X sequencing technology. The cell types of these four datasets are different. The reason for selecting datasets with different characteristics is to evaluate the robustness of each model in generating single-cell gene expression data in different scenarios. Tables 6, 7, 8, and 9 respectively show the performance of the five models on the quality evaluation indicators of Alles data, Baron_Human data, Marshall data, and Mizrak data. The method cfDiffusion of this embodiment performs best on the Baron_Human, Marshall, and Mizrak datasets, and has the highest ILISI score on the Alles dataset. The scores of the other indicators are very close to those of scDiffusion and are almost negligible. In particular, the MMD and ILISI indicators show that the distribution difference between the single-cell gene expression data generated by cfDiffusion and the real single-cell gene expression data is small, and the degree of fusion is high. At the same time, this embodiment uses UMAP to visualize the highly variable genes in the single-cell gene expression data, such as Figure 5 As shown in Figure 3, the distribution of single-cell gene expression data generated by cfDiffusion and scDiffusion is close to the distribution of true single-cell gene expression data, while the distribution of single-cell data generated by other methods is far from the distribution of true single-cell gene expression data.
[0154] Table 6
[0155]
[0156] Table 7
[0157]
[0158] Table 8
[0159]
[0160] Table 9
[0161]
[0162] As described in S2 of the specific implementation manner, the single-cell data in the muris_mam_spl_T_B dataset has two attributes, organ type and cell type, which are composed of breast T cells, breast B cells, spleen T cells, and spleen B cells. When training cfDiffusion, this embodiment uses organ type (breast, spleen) and cell type (B cells, T cells) as two labels for pairing single-cell gene expression data, and trains Diffusion together. In the inference stage (the number of jump steps of cfDiffusion is set to 5), the organ type and cell type are respectively combined into four labels of breast T cells, breast B cells, spleen T cells, and spleen B cells, and the label information is used to guide Diffusion to generate specific single-cell gene expression data. According to Figure 3 As shown in A, most of the generated data is overlaid on the real data, and some generated data is scattered around the real data. The model can learn the attribute characteristics of organs and cell types. In this example, the marker gene Cd74 of breast B cells is selected, and the expression distribution of gene Cd74 in real cells other than breast B cells, real breast B cells, and generated breast B cells is plotted as a box plot, as shown in Figure 3 As shown in C and D. The Wilcoxon rank sum test calculated that the p-value of Cd74 gene expression in the breast B cell data simulated by cfDiffusion and the real breast B cell data was 0.07618, indicating that there was no significant difference between the two distributions. However, the p-value of Cd74 gene expression in the breast B cell data simulated by scDiffusion and the real breast B cell data was 0.00017, indicating that there was a significant difference between the two distributions. When scDiffusion generates data on breast T cells, breast B cells, spleen T cells, and spleen B cells, it is necessary to train classifiers of two attributes, one for guiding the model to generate data of a specific organ, and the other for guiding the model to generate data of a specific cell type. However, cfDiffusion does not need to train any classifiers, which greatly saves training costs. Under multiple conditions, according to Figure 3 B and Table 10. Table 10 compares the quality of the muris_T_B data simulated by cfDiffusion and scDiffusion. The distribution of the single-cell data generated by scDiffusion is significantly different from that of the real data, and scDiffusion does not perform as well as cfDiffusion.
[0163] Table 10
[0164]
[0165] The single-cell data in the Sapiens dataset also has two attributes. Unlike the muris_mam_spl_T_B dataset, the organ types and cell types in Sapiens are not always in pairs. Therefore, multiple conditions are combined, such as combining the organ type Bladder and the cell type fibroblast into one condition `fibroblast_Bladder`. In this example, cfDiffusion and scDiffusion are trained separately in this way, and the simulated data are visualized by observing Figure 4 .A It can be seen that the distribution of the data simulated by the two methods is roughly similar to that of the real data, such as Figure 4 As shown in Figure B, the five colored lines in the QQ-plot are generally close to the reference line (dashed line), indicating that the expression distribution of the gene AKIRIN2 in the five organs (Bladder, Blood, Spleen, Thymus, and Vasculate) in the simulated data by the two methods is similar to that in the real data. The data quality of cfDiffusion and scDiffusion simulations is shown in Tables 11 and 12. Table 11 compares the quality of Homo sapiens data simulated by cfDiffusion and scDiffusion, while Table 12 compares the realism of the Homo sapiens data simulated by the two methods using random forest analysis. In terms of various indicators, cfDiffusion almost completely outperforms scDiffusion.
[0166] This experiment shows that scDiffusion is only suitable for simulating single-cell gene expression data under a single condition. The data quality simulated under multiple conditions is far inferior to that of cfDiffusion. This is likely due to improper guidance of multiple classifiers. This experiment also shows that cfDiffusion has good performance under all conditions.
[0167] Table 11
[0168]
[0169] Table 12
[0170]
[0171] As described in S3 of the specific implementation method, Wadding-OT includes pluripotent stem cells (iPSCs) undergoing 18 days of differentiation state, with a recording interval of 0.5 days. A total of 254,197 cells were sampled for differentiation state, and 19,423 genes were retained. Because the data set is too large, this example selects the cell states of 15 time nodes, including Day 0, Day 0.5, Day 1, Day 1.5, Day 2, Day 2.5, Day 3, Day 4.5, Day 5, Day 5.5, Day 6, Day 6.5, Day 7, Day 7.5, Day 8, a total of 82,920 cells. In this example, time nodes are used as labels for these gene expression data, and paired gene expression data and labels are used to train cfDiffusion. The number of jump steps of cfDiffusion is set to 5 in the inference stage, as shown in Tables 13 and 14. Table 13 compares the quality of WOT data simulated under pseudo time scales by two methods, and Table 14 compares the authenticity of WOT data simulated under pseudo time scales by two methods using random forest. In this example, 5600 cell states are generated for each time node, and various indicators are used to evaluate the data generated by cfDiffusion. Figure 6 As shown in the UMAP dimensionality reduction and visualization of the data, the generated data and the real data distribution are roughly the same, and the SCC and PCC values are both greater than 0.99, which indicates that the gene expression pattern of the generated single-cell data is extremely similar to that of the real single-cell data. The ILISI value is 0.8198, which represents the degree of mixing between the generated single-cell data and the real single-cell data. Figure 6 As shown, the training, test, and AUC values of the random forest model indicate that the trained random forest model is not overfitting. However, due to the similarity between the generated data and the real data, it is largely impossible to distinguish whether the single-cell data is model-generated. In this example, the real data and model-generated data at different time points were used to train the K-nearest neighbor classifier. The average prediction accuracy at each time point was approximately 0.5, which also reflects the high quality of the single-cell data generated by cfDiffusion at each time point.
[0172] In this example, cfDiffusion and scDiffusion are clearly superior in various metrics, including MMD, ILIST, and random forest evaluation. The numerical differences in other metrics are minimal or comparable. This suggests that cfDiffusion outperforms scDiffusion to a certain extent.
[0173] Table 13
[0174] Method cfDiffusion scDiffusion SCC↑ 0.98569 0.99261 PCC↑ 0.9933 0.9997 Wasserstein↓ 0.01637 0.00457 MMD↓ 2.6222 2.7052 ILISI↑ 0.81984 0.79245 KNN_AUC↓ 0.51 0.51 KNN_ACC↓ 0.49969 0.49969
[0175] Table 14
[0176] Method AUC↓ ACC_Train↓ ACC_Test↓ OOB_Score↓ cfDiffusion 0.679 0.635689018 0.628147612 0.629048078 scDiffusion 0.7444 0.679112397 0.685021708 0.67895964
[0177] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as a preferred embodiment as above, it is not intended to limit the present invention. Any technician familiar with the present profession can make some changes or modifications to equivalent embodiments of equivalent changes using the technical content disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent replacement and improvement of the above embodiments made according to the technical essence of the present invention, within the spirit and principles of the present invention, without departing from the content of the technical solution of the present invention, shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A method for generating single-cell sequencing data, characterized by: It includes the following steps: S1: Build a deep neural network model; S2: Obtaining comprehensive scRNA-seq datasets; S3. Preprocessing the scRNA-seq comprehensive dataset; S4: Training of deep neural network models based on preprocessed scRNA-seq comprehensive datasets; The steps for training a deep neural network model in S4 include: S401: The normalized gene expression matrix Input encoder, the encoder uses 2 MLPs, After the encoder, a 128-dimensional embedding layer is obtained ; S402: The fully connected network in the diffusion module is input for continuous noise addition, and continuous denoising is performed through reverse diffusion, and the denoised gene expression matrix is output. ; The diffusion module in S402 includes D 1st floor, D 2nd floor, U 1st floor, U 2nd floor, FG 1st floor and FG 2nd floor; Output denoised gene expression matrix The steps include: S40201: Get the embedding at t time steps x t , the time information t and label information y are integrated into the Chunk in the fully connected network, and each Chunk is Update and get the noised ; S40202: Train the diffusion module to obtain the reverse diffusion diffusion module, set the number of jumps to 3, and in the Tth time step, in Input the trained fully connected network to get the denoised embedding ; The diffusion module for obtaining reverse diffusion in S40202 specifically includes: The diffusion module is parameterized as , and add label information during training , get the learned noise , and after adding noise Subtract the learned noise Afterwards, the diffusion module of reverse diffusion is obtained by taking the logarithm of the Bayesian formula and introducing the hyperparameter k; After denoising for: The Bayesian formula and the expressions for taking logarithms are: Introducing a hyperparameter have to: In formula (10), Used to adjust the fidelity and diversity of generated data, the value range is ; S40203: In the T-1 time step, After passing through the D1 layer, it passes directly through the U2 layer and the denoised ; S40204: Repeat S40202-S40203, iteratively update the noise sequence and jump process until the denoised ; After denoising The calculation formula is: The expressions for updating the noise sequence and the jump process are: In formula (7), For The time step number is High-level characteristics of S403: The denoised gene expression matrix Input decoder, the decoder uses 3 MLPs, After the decoder, we get the gene expression matrix Gene expression matrices of the same dimensions , and the gene expression matrix Reconstruct and generate single-cell sequencing data; Embedding layer The calculation formula is: In formula (2), For the encoder; Gene expression matrix The calculation formula is: In formula (3), For the decoder; The denoised gene expression matrix The reconstructed expression is: S5: Based on the trained deep neural network model, the features of the test data are denoised, de-noised, and reconstructed to generate single-cell sequencing data.
2. The method for generating single-cell sequencing data according to claim 1, wherein: The deep neural network model in S1 consists of an encoder, a diffusion module, and a decoder; The encoder and decoder constitute an autoencoder; The encoder is used to extract gene expression data features of input data; The diffusion module is used to add noise and remove noise to the extracted gene expression data features; The decoder is used to reconstruct the denoised gene expression data features to generate single-cell sequencing data.
3. The method for generating single-cell sequencing data according to claim 1, wherein: The comprehensive scRNA-seq dataset in S2 includes scRNA-seq datasets under multiple sequencing technologies, scRNA-seq datasets with multiple attributes, and scRNA-seq datasets of cell growth and development at pseudo-time nodes.
4. The method for generating single-cell sequencing data according to claim 3, wherein: The steps of obtaining scRNA-seq datasets using the multiple sequencing technologies include: S20101: The number of jump steps of the deep neural network model in the inference phase is set to 5, and compared with the four existing generative models: scDiffusion, cscGAN, LSH-GAN, and sciGAN; S20102: Use SCC, PCC, Wasserstein, MMD, and ILISI metrics to evaluate the quality of single-cell gene expression data generated by these models and use UMAP to visualize single-cell gene expression data; S20103: Use different sequencing technologies and combine single-cell gene expression data to obtain scRNA-seq datasets using multiple sequencing technologies.
5. The method for generating single-cell sequencing data according to claim 3, wherein: The steps of obtaining the scRNA-seq dataset with multiple attributes include: S20201: Divide the single-cell data in the muris_mam_spl_T_B dataset into organ types and cell types, where the organ types include mammary gland and spleen, and the cell types include B cells and T cells. S20202: constructing training labels based on the divided organ type and cell type single-cell data, wherein the training labels include breast T cells, breast B cells, spleen T cells, and spleen B cells; S20203: Training a diffusion module based on the training labels, and guiding the diffusion module to generate scRNA-seq datasets with multiple attributes.
6. The method for generating single-cell sequencing data according to claim 3, characterized in that: The steps of acquiring the scRNA-seq dataset of cell growth and development at the pseudo time node include: S20301: Obtain Wadding-OT data, select the cell states at the first 15 time nodes in the Wadding-OT data, and use the time nodes as data labels; S20302: Set the number of jump steps of the deep neural network model at the inference node to 5, use the cell state data and data labels of the first 15 time nodes in the Wadding-OT data to train the deep neural network model, and obtain an scRNA-seq dataset of cell growth and development at pseudo time nodes.
7. The method for generating single-cell sequencing data according to claim 1, wherein: The steps for preprocessing the scRNA-seq comprehensive dataset in S3 include: S301: Converting the scRNA-seq comprehensive dataset into a gene expression matrix ; S302: Gene expression matrix Perform normalization operation to transform the gene expression matrix The total technical scale of cells in the normalized gene expression matrix is 10000. An offset of 1 is added and the logarithm is taken to obtain the normalized gene expression matrix. ; The expression of the normalization operation is: (1); In formula (1), The rows represent different cells, with a total of cells, The columns represent genes, and each cell has A gene.