Single-cell RNA-SEQ data processing
By applying a noise regularization process that adds random noise based on gene expression levels to scRNA-seq data, the challenges of dropout events and high noise levels are addressed, resulting in improved accuracy of gene-gene correlation inference and a more reliable gene co-expression network.
Patent Information
- Application Number
- JP2022517965
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-09-25
- Filing Date
- 2020-09-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-09-25
AI Technical Summary
Current methods for processing single-cell RNA sequencing (scRNA-seq) data face challenges such as dropout events and high noise levels, which can lead to false-positive gene-gene correlations and affect the accuracy of gene co-expression network construction.
A noise regularization process is applied to scRNA-seq data by adding random noise to the expression values of genes, determined by their expression levels, to reduce gene-gene correlation artifacts and improve the accuracy of gene-gene correlation inference.
The noise regularization process effectively reduces spurious correlations, enhances the enrichment of top-correlated gene pairs in protein-protein interactions, and improves the consistency among different methods for inferring gene-gene correlations, leading to a more reliable gene co-expression network.
Smart Images

Figure 0007684287000002 
Figure 0007684287000003 
Figure 0007684287000004
Abstract
Description
Technical Field
[0001] The present invention generally relates to methods and systems for processing gene expression data for gene-gene correlation by applying a noise regularization process.
Background Art
[0002] It has been realized to infer gene-gene correlations for constructing gene networks using gene expression data obtained from bulk cell microarrays and RNA sequencing (Ballouz et al., Guidance for RNA-seq co-expression network construction and analysis: safety in numbers. Bioinformatics, 2015. 31(13): p. 2123-2130). However, the analysis results of this expression data are limited to measuring the average gene expression of the entire cell pool. The availability of single-cell RNA sequencing (scRNA-seq) technology has made it possible to profile gene expression at the single-cell resolution level, thereby dissecting the heterogeneity within a seemingly homogeneous cell population and revealing hidden gene-gene correlations masked by bulk expression profiles (Kolodziejczyk et al., The Technology and Biology of Single-Cell RNA Sequencing. Molecular Cell, 2015. 58(4): p. 610-620; Papalexi et al., Single-cell RNA sequencing to explore immune cell heterogeneity. Nature Reviews Immunology, 2018. 18(1): p. 35).
[0003] However, due to technical limitations such as dropout events and high levels of noise, there are challenges in processing scRNA-seq data. To reduce the noise caused by inefficiency and estimate the true expression levels in scRNA-seq data processing, various approaches have been adopted. As the first step in scRNA-seq data analysis, a number of data preprocessing methods have been proposed. These data preprocessing methods may affect gene-gene correlation inference and subsequent gene co-expression network construction, such as the introduction of false-positive gene-gene correlations.
[0004] It will be appreciated that there is a need for methods and systems for processing scRNA-seq data that can efficiently reduce gene-gene correlation artifacts for inferring gene-gene correlations and further constructing gene networks. SUMMARY OF THE INVENTION
[0005] The availability of scRNA-seq data enables the uncovering of hidden gene-gene interactions by dissecting the heterogeneity within a homogeneous cell population and profiling gene expression at the single-cell resolution level. Challenges in processing scRNA-seq data can arise from technical limitations such as dropout (undetected gene expression) and high noise (variability). For the estimation of true expression levels in scRNA-seq data processing, data preprocessing methods for noise mitigation have been adopted. However, these data preprocessing methods may affect gene-gene correlation inference by introducing false-positive gene-gene correlations.
[0006] This application provides a method and system for revealing gene-gene correlations by processing gene expression data and applying a noise regularization process to reduce gene-gene correlation artifacts. The present disclosure also provides a method for improving data processing for gene-gene correlations, the method including processing gene expression data for normalization or complementation, applying a noise regularization process to the normalized or complemented gene expression data, and applying a gene-gene correlation calculation process to obtain correlated gene pairs. In some exemplary embodiments, the gene expression data is single-cell gene expression data. In some exemplary embodiments, the noise regularization process includes adding random noise to the expression values of genes within cells in an expression matrix, and the random noise is determined by the expression level of the gene.
[0007] In some exemplary embodiments, the random noise is determined by: (1) determining the expression distribution of the gene across all cells in the expression matrix; (2) considering the approximately 0.1% to about 20% percentile of the gene expression level as the maximum noise level; (3) generating a random number in the range from 0 to the maximum noise level under a uniform distribution; and (4) adding the random number to the expression value of the gene within the cells in the expression matrix to obtain a noise-regularized expression matrix.
[0008] In some exemplary embodiments, the random noise is determined by: (1) determining the expression distribution of the gene across all cells in the expression matrix; (2) considering the 1% percentile of the gene expression level as the maximum noise level; (3) generating a random number in the range from 0 to the maximum noise level under a uniform distribution; and (4) adding the random number to the expression value of the gene within the cells in the expression matrix to obtain a noise-regularized expression matrix.
[0009] In some exemplary embodiments, the gene-gene correlation calculation process is performed using cell clusters. In some exemplary embodiments, all unique molecular identifier normalization (NormUMI), regularized negative binomial regression (NBR), deep count autoencoder network (DCA), Markov affinity-based graph completion of cells (MAGIC), or single cell analysis via expression recovery (SAVER) are used to process gene expression data for normalization or complementation. In some exemplary embodiments, the method for improving data processing for gene-gene correlation in the present application further includes enriching the gene expression data associated with the correlated gene pairs and / or constructing a gene-gene correlation network based on the correlated gene pairs, and the gene-gene correlation network is cell type specific. In some exemplary embodiments, the method of the present application includes using the gene-gene correlation network to map molecular interactions, guiding experimental design to investigate biological events, discovering biomarkers, guiding comparative network analysis, guiding drug design, identifying changes in gene-gene interactions by comparing the healthy and diseased states of cells, guiding drug development, predicting gene transcriptional regulation, improving drug efficiency, or identifying drug resistance factors.
[0010] The present disclosure provides, at least in part, a gene-gene correlation network, the network being constructed based on correlated gene pairs obtained using a method for improving data processing for gene-gene correlation in the present application, the method including processing gene expression data for normalization or complementation, applying a noise regularization process to the normalized or complemented gene expression data, and applying a gene-gene correlation calculation process to obtain correlated gene pairs.
[0011] The present disclosure provides, at least in part, a computer-implemented method for data processing for gene-gene correlation. The method includes retrieving gene expression data, processing the gene expression data for normalization or complementation, applying a noise regularization process to the normalized or complemented gene expression data, applying a gene-gene correlation calculation process to obtain correlated gene pairs, and constructing a gene-gene correlation network based on the correlated gene pairs. The gene-gene correlation network is cell-type specific. In some exemplary embodiments, the gene expression data is single-cell gene expression data. In some exemplary embodiments, the noise regularization process includes adding random noise to the expression values of genes within cells in the expression matrix, and the random noise is determined by the expression level of the gene.
[0012] In some exemplary embodiments, the random noise is determined by: (1) determining the expression distribution of genes across all cells in the expression matrix; (2) considering the about 0.1 to about 20 percentile of the expression level of the gene as the maximum noise level; (3) generating random numbers in the range from 0 to the maximum noise level under a uniform distribution; and (4) adding the random numbers to the expression values of genes within cells in the expression matrix to obtain a noise-regularized expression matrix.
[0013] In some exemplary embodiments, the random noise is determined by: (1) determining the expression distribution of genes across all cells in the expression matrix; (2) considering the 1 percentile of the expression level of the gene as the maximum noise level; (3) generating random numbers in the range from 0 to the maximum noise level under a uniform distribution; and (4) adding the random numbers to the expression values of genes within cells in the expression matrix to obtain a noise-regularized expression matrix.
[0014] In some exemplary embodiments, the gene-gene correlation calculation process is performed using cell clusters. In some exemplary embodiments, all unique molecular identifier normalization (NormUMI), regularized negative binomial regression (NBR), deep count autoencoder network (DCA), Markov affinity-based graph completion of cells (MAGIC), or single-cell analysis via expression recovery (SAVER) is used to process gene expression data for normalization or complementation.
[0015] In some exemplary embodiments, the computer-implemented method for data processing for gene-gene correlation of the present application further includes performing enrichment on gene expression data associated with correlated gene pairs. In some exemplary embodiments, the computer-implemented method of the present application includes using a gene-gene correlation network to map molecular interactions, guiding experimental design to investigate biological events, discovering biomarkers, guiding comparative network analysis, guiding drug design, identifying changes in gene-gene interactions by comparing the healthy and diseased states of cells, guiding drug development, predicting gene transcriptional regulation, improving drug efficiency, or identifying drug resistance factors.
[0016] The present disclosure provides, at least in part, a computer-based system for data processing for gene-gene correlation. The system includes a database configured to store gene expression data, a memory configured to store instructions, and at least one processor coupled to the memory. The at least one processor is configured to retrieve the gene expression data, process the gene expression data for normalization or complementation, apply a noise regularization process to the normalized or complemented gene expression data, apply a gene-gene correlation calculation process to obtain correlated gene pairs, and construct a gene-gene correlation network based on the correlated gene pairs. The system also includes a user interface that receives queries regarding data processing of gene-gene correlation and can display the results of the correlated gene pairs and the constructed gene-gene correlation network. In some exemplary embodiments, the gene expression data is single-cell gene expression data, and the gene-gene correlation network is cell-type specific. In some exemplary embodiments, the noise regularization process includes adding random noise to the expression values of genes within cells in the expression matrix, and the random noise is determined by the expression level of the genes.
[0017] In some exemplary embodiments, the random noise is determined by: (1) determining the expression distribution of genes across all cells in the expression matrix; (2) considering the about 0.1 to about 20 percentile of the expression level of the gene as the maximum noise level; (3) generating a random number in the range from 0 to the maximum noise level under a uniform distribution; and (4) adding the random number to the expression value of the gene within the cell in the expression matrix to obtain a noise-regularized expression matrix.
[0018] In some exemplary embodiments, the random noise is determined by: (1) determining the gene expression distribution across all cells in the expression matrix; (2) considering the 1st percentile of the gene expression levels as the maximum noise level; (3) generating random numbers in the range from 0 to the maximum noise level under a uniform distribution; and (4) adding the random numbers to the gene expression values in the cells in the expression matrix to obtain a noise-regularized expression matrix.
[0019] In some exemplary embodiments, the gene-gene correlation calculation process is performed using cell clusters. In some exemplary embodiments, total unique molecular identifier normalization (NormUMI), regularized negative binomial regression (NBR), deep count autoencoder network (DCA), Markov affinity-based graph completion of cells (MAGIC), or single cell analysis via expression recovery (SAVER) is used to process gene expression data for normalization or completion. In some exemplary embodiments, at least one processor is further configured to perform enrichment on the gene expression data associated with the correlated gene pairs.
[0020] In some exemplary embodiments, at least one processor is further configured to utilize the gene-gene correlation network to map molecular interactions, to guide experimental design to investigate biological events, to discover biomarkers, to guide comparative network analysis, to guide drug design, to identify changes in gene-gene interactions by comparing the healthy and diseased states of cells, to guide drug development, to predict gene transcriptional regulation, to improve drug efficiency, or to identify drug resistance factors.
[0021] These and other aspects of the invention will be better understood and appreciated in conjunction with the following description and the accompanying drawings. The following description shows various embodiments and numerous specific details by way of illustration and not limitation. Many substitutions, modifications, additions, or rearrangements can be made within the scope of the invention.
Brief Description of the Drawings
[0022]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5A
Figure 5B
Figure 5C
Figure 5D
Figure 6
Figure 7A
Figure 7B
Figure 7C
Figure 8A
Figure 8B
Figure 8C
Figure 9
Figure 10
Figure 11
Mode for Carrying Out the Invention
[0023] Due to the availability of high-throughput gene expression data, it is possible to construct large-scale gene regulatory networks through statistical inference from gene expression data, for example, from a statistical perspective centered on the data. Various statistical network inference methods, such as inference algorithms, are used to estimate interactions. The inferred gene regulatory networks provide information on regulatory interactions between regulators and their potential targets, such as gene-gene interactions or potential protein-protein interactions in complexes. These inferred networks represent statistically significant predictions of molecular interactions obtained from large-scale gene expression data (Emmert-Streib et al., Gene regulatory networks and their applications: understanding biological and medical problems in terms of networks. Frontiers in Cell and Developmental Biology, 2014.2(38)).
[0024] The inferred gene regulatory network can be used to help solve biological and biomedical problems, such as serving as a causal map of molecular interactions, a guide for experimental design, biomarker discovery, a guide for comparative network analysis, or a guide for drug design (Emmert-Streib et al.). Furthermore, the constructed network can be used to provide guidance for further downstream analysis, such as identifying downstream interactions and identifying changes in gene-gene interactions by comparing the healthy and diseased states of cells, which can potentially save time for drug development.
[0025] The inferred gene regulatory network can be used to help solve biological and biomedical problems by functioning as a causal map of molecular interactions, such as for deriving new biological hypotheses regarding molecular interactions or predicting gene transcriptional regulation. Since the predicted links are assumed to correspond to actual physical binding events between molecules, this information can be used to guide laboratory experiments to investigate biological events. In addition, these inferred networks can be used to discover or study biomarkers for diagnostic, prognostic, or predictive purposes. For example, since cancer is a complex disorder related to various pathways rather than individual genes, network-based biomarkers can be used as statistical measures for cancer diagnostic purposes. Furthermore, as more inferred gene regulatory networks become available, it becomes possible to conduct comparative network analysis to understand changes in gene-gene interactions across different physiological conditions or disease states (Emmert-Streib et al.). Therefore, these inferred networks can lead to more efficient design of rational drugs, such as improving drug efficacy or identifying drug resistance factors.
[0026] A gene-gene co-expression network can be regarded as a gene regulatory network constructed from gene-gene correlations inferred from gene expression data, for example, inferred from single-cell RNA sequencing (scRNA-seq) data. Gene-gene co-expression networks can be constructed from different physiological, disease, or treatment conditions. By comparing gene-gene co-expression networks constructed under different conditions, changes in gene interactions across different physiological conditions or disease states can be understood, and such phenotypes under different conditions can be analyzed. For example, the expression of two genes may be highly correlated in one cell type but unrelated in another cell type. scRNA-seq data can unbiasedly capture the entire transcriptome of different cell types in a heterogeneous cell population. This can reveal gene-gene correlations specific to a particular cell type.
[0027] Gene expression is regulated by a network of transcription factors and signaling molecules. Since scRNA-seq data represents each cell as an independent identity representing different types or stages of biological events, it can provide important information for understanding cell and tissue heterogeneity by revealing the dynamics of differentiation and quantifying gene transcription. Correlated expression, especially co-expression between genes, can be beneficial for constructing networks for visualization and interpretation (Stuart et al., A Gene-Coexpression Network for Global Discovery of Conserved Genetic Modules. Science, 2003. 302(5643): p.249-255). Analysis of scRNA-seq data can promote biological discovery because each cell can be classified into different cell types or lineages to improve the understanding of biological processes in different contexts. Therefore, gene-gene correlations revealed from single-cell expression data have the potential to discover cell-type-specific modules and construct more comprehensive networks.
[0028] To analyze scRNA-seq data and infer large-scale regulatory networks under different organs and disease conditions, a correlation metric specifically adapted to single-cell data was developed. The unbiased quantification of the biological relevance of genes was calculated using graph theory tools to identify the major players in organ function and the factors of diseases (Iacono et al., Single-cell transcriptomics unveils gene regulatory network plasticity. Genome Biology, 2019. 20(1): p. 110). A genome-scale gene interaction map was constructed by examining gene-gene pairs for synthetic gene interactions. Based on the network of gene interaction profiles, a functional map was revealed by clustering similar biological processes in coherent subsets (Costanzo, M., et al., The Genetic Landscape of a Cell. Science, 2010. 327(5964): p. 425-431). Here, highly correlated profiles indicate specific pathways that define gene function.
[0029] However, there are challenges in the utilization of scRNA-seq data due to technical limitations such as dropout events (e.g., gene expression not detected by scRNA-seq), high levels of noise (variation), and very large data volumes. In addition, only a small fraction of the transcripts present in each cell are sequenced by scRNA-seq, which leads to unreliable quantification of low- and medium-expressing genes. A large proportion of genes, for example, over 90% of the gene population, have zero or low read counts due to low capture and sequencing efficiency. While many of the observed zero counts reflect true zero expression, a large portion of this count may be due to technical limitations (Huang et al., SAVER: gene expression recovery for single-cell RNA sequencing. Nature Methods, 2018. 15(7): p. 539-542). Furthermore, the observed sequencing depth can vary dramatically between cells. Variations in cell lysis during sequencing, reverse transcription efficiency, and molecular sampling can also contribute to the variation (Hicks et al., Missing data and technical variability in single-cell RNA-sequencing experiments. Biostatistics, 2017. 19(4): p. 562-578).
[0030] To estimate the true expression levels in the processing of scRNA-seq data, which reduces noise caused by low efficiency and includes expression normalization and dropout complementation, various data preprocessing methods have been adopted. Data normalization is often required to remove technical noise while retaining the true biological signal. The high dropout rate of scRNA-seq refers to the large proportion of genes with zero counts due to technical limitations in the detection of transcripts (Svensson et al., Power analysis of single-cell RNA-sequencing experiments. Nature Methods, 2017.14: p. 381; Ziegenhain et al., Comparative Analysis of Single-Cell RNA Sequencing Methods. Molecular Cell, 2017.65(4): p. 631-643.e4). To process dropouts and recover true gene expression, various data complementation methods can be used to preprocess scRNA-seq data such as cell clustering, detection of differentially expressed genes, and trajectory analysis (Tian et al., Benchmarking single cell RNA-sequencing analysis pipelines using mixture control experiments. Nature Methods, 2019.16(6): p. 479-487).
[0031] There are challenges in the application of imputation methods for false gene-gene correlations. This is because these methods are designed to reverse engineer gene networks to measure gene-gene correlations. Andrews et al. tested multiple imputation methods on a small simulated dataset and found that dropout imputation generates false positive gene-gene correlations (Andrews, T. and M. Hemberg, False signals induced by single-cell imputation [version 1; peer review: 4 approved with reservations]. F1000Research, 2018, 7(1740)). Some of the representative scRNA-seq normalization / imputation methods for data preprocessing introduce spurious or inflated correlations due to oversmoothing or overfitting of the data, which affects the inference of gene-gene correlations. With these methods, correlation artifacts can be introduced for gene pairs where co-expression is not expected. Since false signals and correlation artifacts can be introduced into data processing, the obtained gene pairs with the highest correlations from these methods may have weak enrichment in protein-protein interactions.
[0032] In machine learning, adding noise to data under specific conditions can reduce overfitting and improve the robustness of the results (Bishop, Training with noise is equivalent to Tikhonov regularization. Neural computation, 1995. 7(1): p. 108-116; Neelakantan et al., Adding gradient noise improves learning for very deep networks. arXiv preprint arXiv:1511.06807, 2015; Smilkov et al., Smoothgrad: removing noise by adding noise. arXiv preprint arXiv:1706.03825, 2017).
[0033] The present disclosure provides a method and a system for processing scRNA-seq data by utilizing a novel noise regularization method that can efficiently reduce gene-gene correlation artifacts for inferring gene-gene correlations and further constructing gene networks, thereby satisfying the aforementioned requirements. A gene co-expression network can be constructed using the gene-gene correlations derived after applying the noise regularization method of the present application. The resulting network was verified at multiple levels to confirm the reliability of the network construction. The quality of the inferred biological network was evaluated using known interactions in the protein-protein interaction database.
[0034] In some exemplary embodiments, the noise regularization method of the present application is implemented to process the preprocessed scRNA-seq data by adding noise uniformly distributed to the expression level of each gene. By using the gene-gene correlations obtained by adding the noise regularization method of the present application and reducing the artifacts in the gene-gene correlations, a gene co-expression network can be reconstructed. In some exemplary embodiments, multiple known cell modules such as immune cell modules were successfully revealed. This was not visible in the absence of the noise regularization method of the present application. In some exemplary embodiments, when the noise regularization method of the present application is added, cell type marker genes are more highly evaluated in terms of network topology characteristics, for example, evaluated with higher values of degree and PageRank, and their important roles in each cell cluster are identified. The noise regularization method of the present application provides the advantage of increasing the robustness of data processing by reducing over-smoothing or over-fitting of the expression data.
[0035] In some exemplary embodiments, the present application provides a computer-implemented method for improving data processing for gene-gene correlation. The method includes processing gene expression data for normalization or complementation, applying a noise regularization process to the normalized or complemented gene expression data, and applying a gene-gene correlation calculation process to obtain correlated gene pairs. In some exemplary embodiments, the present application provides a computer-based system for data processing for gene-gene correlation. The system includes a database configured to store gene expression data, a memory configured to store instructions, and at least one processor coupled to the memory. The at least one processor is configured to retrieve the gene expression data, process the gene expression data for normalization or complementation, apply a noise regularization process to the normalized or complemented gene expression data, apply a gene-gene correlation calculation process to obtain correlated gene pairs, and construct a gene-gene correlation network based on the correlated gene pairs. The system further includes a user interface configured to receive a query regarding data processing for gene-gene correlation and display the results of the correlated gene pairs and the constructed gene-gene correlation network.
[0036] As shown in FIG. 1, an exemplary computer-based system of the present application for data processing for gene-gene correlation includes one or more databases, a central processing unit (CPU) including one or more processors, a memory coupled to the CPU for storing instructions, and a user interface. In some exemplary embodiments, the computer-based system of the present application further includes algorithms for data normalization or complementation and various reports. In some exemplary embodiments, the database includes gene expression data, genomic data, or protein-protein interaction data. In some exemplary embodiments, the user interface is configured to receive a query for data processing, display the correlated gene pairs, or display the gene-gene correlation network.
[0037] In some exemplary embodiments, the random noise is determined by: (1) determining the gene expression distribution across all cells in the expression matrix; (2) considering the 1st percentile of the gene expression levels as the maximum noise level; (3) generating random numbers in the range from 0 to the maximum noise level under a uniform distribution; and (4) adding the random numbers to the gene expression values in the cells of the expression matrix to obtain a noise-normalized expression matrix.
[0038] In some exemplary embodiments, the expression value of gene i in cell j is represented as V, and the random noise can be determined by: (i) calculating the expression distribution of gene i after applying various data preprocessing methods; (ii) determining the 1st percentile of the expression values of gene i, represented as M, where M is used as the maximum value of the noise level; and (iii) generating uniform distribution random numbers in the range of 0 to M and adding this random number to V.
[0039] In some exemplary embodiments, random noise is generated and added to the expression value of gene i in cell j of the expression matrix processed by a specific method, for example, V. The random noise is determined by: (1) determining the expression distribution of gene i across all cells; (2) considering the 1st percentile of gene i expression, represented as M, as the maximum noise level; (3) using 0.1 as the maximum noise level when M is equal to zero; (4) generating random numbers in the range of 0 to M under a uniform distribution; and (5) adding the random numbers to V to obtain a noise-normalized expression matrix.
[0040] In some exemplary embodiments, the noise regularization process includes obtaining an expression matrix processed by a specific scRNA-seq preprocessing method, where the expression matrix contained the expression of n genes in m cells. Assuming that V is the expression value of gene i in cell j, random noise is generated and added to V. The random noise is determined by the following steps, as shown in the exemplary flowchart of FIG. 2, namely, (1) determining the expression distribution of gene i across all cells; (2) regarding the first percentile from the expression distribution of gene i as the maximum noise level of gene i, denoted as M, and if M is smaller than the minimum value m, using m as the maximum noise level; (3) generating a random number in the range of 0 to M under a uniform distribution; (4) adding this random number to V to obtain a noise-regularized expression value; and (5) repeating this procedure for all entries in the expression matrix.
[0041] The exemplary embodiments disclosed herein satisfy the aforementioned requirements by providing a computer-implemented method for improving the processing of gene expression data for gene-gene correlation by applying a noise regularization process to the normalized or complemented gene expression data.
[0042] In some exemplary embodiments, a computer-implemented method for improving the data processing of gene expression data for gene-gene correlation is provided by applying a noise regularization process to the normalized or complemented gene expression data. These satisfy the long-standing need to efficiently reduce gene-gene correlation artifacts for inferring gene-gene correlations and further constructing gene networks.
[0043] The term "a" should be understood to mean "at least one", and the terms "about" and "approximately" should be understood to allow for standard variations, as understood by those skilled in the art, and endpoints are included when ranges are provided.
[0044] As used herein, the terms "include", "includes", and "including" are meant to be non-limiting and are understood to mean "comprise", "comprises", and "comprising", respectively.
[0045] In some exemplary embodiments, the present disclosure provides a computer-implemented method for improving data processing for gene-gene correlation, including processing gene expression data for normalization or complementation, applying a noise regularization process to the normalized or complemented gene expression data, and applying a gene-gene correlation calculation process to obtain correlated gene pairs. In some exemplary embodiments, the noise regularization process is applied before applying the gene-gene correlation calculation process. In some exemplary embodiments, the gene expression data is single-cell gene expression data.
[0046] As used herein, the term "gene-gene correlation" means pairs of genes that exhibit similar expression patterns across an entire sample. When two genes are co-expressed, the expression levels of these two genes both increase and decrease together. Co-expressed genes are often involved in the same biological pathway, generally regulated by the same transcription factor, or otherwise functionally related.
[0047] As used herein, the term "normalization" refers to a process of organizing a dataset to reduce redundancy and improve data consistency, including aligning adjustment values or adding adjustments to conform to a specific distribution. The normalization process may remove systematic variations (e.g., variability in experimental conditions, machine parameters) and enable bias-free comparison between samples.
[0048] As used herein, the term "completion" means the process of replacing missing data with substituted values. Missing data can cause problems such as introducing a significant amount of bias, for example, by causing a decrease in efficiency that may affect the representativeness of the results. Completion includes the process of replacing missing data with estimated values based on other available information, thereby enabling the analysis of the dataset using standard techniques.
[0049] Exemplary embodiments The embodiments disclosed herein provide a method for improving the processing of gene expression data for gene-gene correlation by applying a noise regularization process to the normalized or completed gene expression data.
[0050] In some exemplary embodiments, the present disclosure provides a method for improving data processing for reducing gene-gene correlation artifacts, which includes processing scRNA-seq data for normalization or completion, applying a noise regularization process to the normalized or completed gene expression data, and applying a gene-gene correlation calculation process to obtain correlated gene pairs, wherein the noise regularization process includes adding random noise to the expression values of genes in cells in the expression matrix.
[0051] In some exemplary embodiments, the random noise is determined by: (1) determining the expression distribution of genes across all cells in the expression matrix; (2) considering the gene expression levels from about 0.1% to about 20% as the maximum noise level; (3) generating random numbers in the range from 0 to the maximum noise level under a uniform distribution; and (4) adding the random numbers to the expression values of genes in cells in the expression matrix to obtain a noise-regularized expression matrix.
[0052] In some specific exemplary embodiments, the random noise is determined by: (1) determining the gene expression distribution across all cells in the expression matrix; (2) considering the gene expression levels at about 0.1 to about 20 percentile, about 0.1 percentile, about 0.5 percentile, about 1 percentile, about 1.5 percentile, about 2 percentile, about 3 percentile, about 4 percentile, about 5 percentile, about 7 percentile, about 10 percentile, about 15 percentile, about 20 percentile, or about 25 percentile as the maximum noise level; (3) generating random numbers in the range from 0 to the maximum noise level under a uniform distribution; and (4) adding the random numbers to the gene expression values in the cells in the expression matrix to obtain a noise-regularized expression matrix. The computer-implemented method of the present application further includes constructing a gene-gene correlation network based on correlated gene pairs.
[0053] In some exemplary embodiments, the computer-implemented method of the present application includes using the gene-gene correlation network to map molecular interactions, guiding experimental design to investigate biological events, discovering biomarkers, guiding comparative network analysis, guiding drug design, identifying changes in gene-gene interactions by comparing the healthy and diseased states of cells, guiding drug development, predicting gene transcriptional regulation, improving drug efficiency, identifying drug resistance factors, providing guidance for further downstream analysis, deriving new biological hypotheses regarding molecular interactions, providing statistical measures for cancer diagnosis purposes, guiding comparative network analysis to understand changes in gene-gene interactions across different physiological or disease states, understanding changes in gene-gene interactions for analyzing specific phenotypes under different conditions, revealing the dynamics of differentiation for quantifying gene transcription, or further including discovering biomarkers for diagnostic, prognostic, or predictive purposes.
[0054] It should be understood that the present method or system is not limited to any of the above methods or systems for improving the processing of gene expression data for gene-gene correlation. The sequential labeling by numbers and / or letters of the method steps provided herein does not mean to limit the method or any of its embodiments to a specific indicated order. Various published documents including patents, patent applications, published patent applications, accession numbers, technical papers, and academic papers are cited herein. Each of these cited documents is incorporated herein by reference in its entirety and for all purposes. Unless otherwise stated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which the present invention belongs.
[0055] The present disclosure will be more fully understood by reference to the following examples provided to explain the present disclosure in more detail. These should not be construed as limiting the scope of the present disclosure.
Example
[0056] Database and Method Obtaining scRNA-seq Datasets Bone marrow scRNA-seq data was retrieved from the Human Cell Atlas Data Portal (https: / / preview.data.humancellatlas.org / ). The retrieved dataset contains profiling data of 378,000 immune cells by the 10× platform. To reduce the computational load, 50,000 cells were randomly sampled from the original dataset. Subsequently, genes expressed in less than 100 cells (0.2%) were further filtered. In the output, 12,600 genes remained in the final benchmark dataset. Single-cell analysis such as clustering or dimensionality reduction was performed using the Seurat R package Version 3.0.
[0057] Normalization or Completion of Data For data normalization or completion, multiple methods are applied in the data preprocessing step, including all unique molecular identifier normalization (NormUMI), regularized negative binomial regression (NBR; Hafemeister et al., Normalization and variance stabilization of single-cell RNA-seq data using regularized negative binomial regression. bioRxiv, 2019: p. 576827), deep count autoencoder (DCA) network (Eraslan et al., Single-cell RNA-seq denoising using a deep count autoencoder. Nature Communications, 2019. 10(1): p. 390), cell Markov affinity-based graph completion (MAGIC; van Dijk, et al., Recovering Gene Interactions from Single-Cell Data Using Data Diffusion. Cell, 2018. 174(3): p. 716 - 729.e27), or single-cell analysis via expression recovery (SAVER; Huang et al.). NBR, SAVER, and DCA were run with default parameters according to the tool's instructions. MAGIC was run using parameters of the number of principal components npca = 30, the power of the Markov affinity matrix t = 6, and the number of nearest neighbors k = 30. NormUMI and NBR are normalization methods. The methods of DCA, MAGIC, and SAVER are completion methods.
[0058] Gene-gene correlation calculation The Spearman correlation for each gene pair was calculated within each cluster of cells, from cluster 0 to cluster 9 and so on. A gene is considered to be expressed in one cluster if it is expressed in either more than 1% of the cells or the larger of 50 cells within that cluster. The correlation of gene pairs within one cluster was considered a valid correlation when both genes were expressed within the cluster. The most effective correlations in 10 clusters (cluster 0 - 9) were recorded as the final correlations for the specific gene pairs.
[0059] Data enrichment by protein - protein interaction Human protein - protein interaction (PPI) data was retrieved from the STRING database (Szklarczyk, et al., STRING v10: protein - protein interaction networks, integrated over the tree of life. Nucleic Acids Research, 2014.43(D1): p.D447 - D452). Gene pairs were ranked by the Spearman correlation coefficient of each method. Then, gene pairs with high ranks (the top n gene pairs) were obtained, and the fraction of pairs that appeared in the protein - protein interaction database was counted.
[0060] Noise regularization Noise regularization was applied to data processing. Random noise determined by gene expression levels was added to the expression matrix before proceeding to correlation calculation. Random noise is generated and added to the expression value of gene i in cell j of the expression matrix processed by a specific method, for example, V. Random noise is generated by: (1) determining the expression distribution of gene i across all cells, (2) considering the 1st percentile of gene i expression, represented as M, as the maximum noise level, (3) using 0.1 as the maximum noise level when M is equal to zero, (4) generating a random number in the range of 0 to M under a uniform distribution, and (5) adding the random number to V to obtain a noise - regularized expression matrix.
[0061] Network construction The Spearman correlation for each gene pair was calculated within the cells of each cluster. Within each cluster, gene pairs were ranked by their Spearman correlation. Since basic cellular functions require housekeeping genes, they are expected to be expressed in all cells regardless of tissue type or cell type. To construct cell-type specific interaction modules, housekeeping genes were removed from the network construct. The list of removed housekeeping genes included the housekeeping gene list obtained from Eisenberg et al. (Eisenberg et al., Human housekeeping genes, revisited. Trends in Genetics, 2013. 29(10): p. 569-574). Furthermore, typical housekeeping genes, such as ACTB, B2M, and cytoskeletal genes from ribosome, TCA, reactome, as well as genes encoding mtDNA, were added to the list of removed housekeeping genes. After removing the housekeeping genes, the top 1,000 ranked gene pairs were obtained from each cluster and combined to construct a draft network. The importance of each node in the network was measured by the degree and PageRank values using the igraph R package by Csardi et al. (Csardi et al., The igraph software package for complex network research. InterJournal, Complex Systems, 2006. 1695(5): p. 1-9). Subsequently, the network was cleaned up by removing links not referring to protein-protein interactions in the STRING database.The final network was visualized using Cytoscape by Shannon et al. (Shannon et al., Cytoscape: A Software Environment for Integrated Models of Biomolecular Interaction Networks. Genome Research, 2003. 13(11): p. 2498-2504) and the R package RCy3 by Ono et al. (Ono et al., CyREST: Turbocharging Cytoscape Access for External Tools via a RESTful API. F1000Research, 2015. 4: p. 478-478). The network layout was generated using the EntOptLayout Cytoscape plugin by Agg et al. (Agg et al., The EntOptLayout Cytoscape plug-in for the efficient visualization of major protein complexes in protein-protein interaction and signaling networks. Bioinformatics, 2019).
[0062] Example 1. Pretreatment of data using representative normalization / completion methods A number of representative normalization / completion methods were benchmarked focusing on their impact on gene-gene correlation inference. The global scaling normalization method was the one with the least data manipulation by normalizing gene expression of each cell by total expression. Usually, this method is followed by logarithmic transformation and z-score scaling. Since logarithmic transformation and z-score scaling do not change rank-based correlations, only total UMI normalization (referred to as NormUMI) was included in the comparison. A framework that utilizes "regularized negative binomial regression" (referred to as NBR) to normalize and stabilize the variance of scRNA-seq data was included. This can remove the influence of technical noise while maintaining biological heterogeneity. Three additional methods representing different complementary methodological categories were also included. For example, (i) MAGIC is a data smoothing approach that uses shared information across similar cells to remove noise and fill in dropout values, (ii) SAVER is a model-based approach that models the expression of each gene under the negative binomial distribution assumption and outputs the posterior distribution of the true expression, and (iii) DCA is a deep learning-based autoencoder for capturing the complexity and non-linearity of scRNA-seq data and reconstructing gene expression.
[0063] These five exemplary normalization / completion methods, e.g., NormUMI, NBR, DCA, MAGIC, and SAVER, were applied to the bone marrow scRNA-seq data of the Human Cell Atlas Project (Regev et al., The Human Cell Atlas. eLife, 2017.6: p.e27041) by comparing gene-gene correlations derived from the preprocessing methods. For the other four methods except NormUMI, it was shown that gene-gene correlations were significantly increased by introducing correlation artifacts for gene pairs where co-expression was not expected. The gene pairs with the highest correlations in these methods had weak enrichment in protein-protein interactions. This suggests that there may be false signals and correlation artifacts introduced in the preprocessing of the data. False signals may be introduced in the preprocessing of the data due to over-smoothing or overfitting.
[0064] Example 2. Calculation of gene-gene correlations within a single cell Actual bone marrow scRNA-seq data from the Human Cell Atlas preview dataset was used as a benchmark dataset (Regev et al.) for various data preprocessing methods. The complete dataset contained 378,000 bone marrow cells that could be grouped into 21 cell clusters, covering all major immune cell types, as shown in Figure 3 and Table 1. 50,000 cells were randomly sampled from the original dataset. Genes expressed in less than 0.2% (100 cells) were excluded from this subset. The final dataset contained 12,600 genes, resulting in over 79 million possible gene pairs.
Table 1
[0065] The overview of the benchmark framework is shown in Figure 4. As shown in Figure 4, five representative data preprocessing methods, e.g., NormUMI, NBR, DCA, MAGIC, and SAVER, were applied to a single-cell expression data matrix, e.g., bone marrow single-cell expression data. Gene-gene correlations were calculated directly from the resulting matrix (shown as Path 1). The enrichment of the derived gene-gene correlations in protein-protein interactions and the consistency among the methods were evaluated. It was discovered that artificial correlation relationships can be introduced by the data preprocessing procedure. (Shown as Path 2) A noise regularization step was introduced, and after applying the random noise determined by the gene expression level (red region) to the expression matrix, the correlation calculation was performed. This noise regularization step effectively reduces spurious correlations and enables the construction of a gene co-expression network using an improved gene-gene correlation metric.
[0066] The expression of two genes can be highly correlated in one cell type but unrelated in another cell type. To capture gene-gene correlations across different cell types, gene-gene Spearman correlations were calculated within the 10 largest clusters, e.g., within more than 500 cells per cluster, in a benchmark dataset containing CD4 T cells, CD8 T cells, natural killer cells, B cells, pre-B cells, CD14+ monocytes, FCGR3A+ monocytes, erythrocytes, granulocyte-macrophage precursors, and hematopoietic stem cells (Figures 3 and 4). For each pair of genes, the highest correlation among the 10 clusters was recorded as the final correlation.
[0067] Example 3. Observation of Artifacts Using Data Preprocessing Methods Five representative data preprocessing methods, e.g., NormUMI, NBR, DCA, MAGIC, and SAVER, were applied to bone marrow scRNA-seq data from the Human Cell Atlas project. The distribution of overall gene-gene correlations in five different data matrices processed by different methods was compared. Since most gene pairs were expected to have no association, the correlation distribution was expected to peak at 0. As shown in Figure 5A, NormUMI generated a correlation distribution that peaked at 0. However, for the other four methods, as shown in Figure 5A, much higher median correlations occurred with respect to the Spearman correlation coefficient (NormUMI ρ = 0.023, NBR ρ = 0.839, MAGIC ρ = 0.789, DCA ρ = 0.770, SAVER ρ = 0.166).
[0068] After accessing the interaction between two genes and applying a specific data preprocessing method, it was clarified whether the higher correlation reflects a higher probability of either a functional or physical interaction between the two genes. Proteins encoded by co-expressed genes interact with each other more frequently than random protein pairs. If the resulting higher correlation is true, co-expressed genes should have relatively higher enrichment in the protein-protein interaction database, while spurious correlations should weaken the enrichment. Using the STRING database (Szklarczyk et al.) containing 5,772,157 interacting gene pairs, protein-protein interaction enrichment in the top-ranked co-expressed gene pairs was evaluated. The top gene pairs from each method (by correlation rank) were selected. The overlapping part of these pairs with the STRING database was calculated as shown in Figure 5B. As a result, it was shown that NormUMI showed 80% and 47% overlap with STRING for the top 100 and 10,000 gene pairs, respectively, and had the highest protein-protein interaction enrichment. In contrast, the top gene pairs from NBR had less overlap with the expected STRING (less than 2%), while MAGIC and DCA had similar protein-protein interaction enrichments in the range of 11% - 22%. SAVER showed relatively good results, but the enrichment was only half of that of NormUMI.
[0069] Gene pairs were randomly sampled and the random pairs were overlaid with PPIs to estimate the background enrichment level (Figure 5D). The estimated background enrichment level was approximately 3.6%, indicating that the PPI enrichment of NBR was even lower than the background. This simple method directly associates physical interactions with gene co-expression, and if the same assumptions are made in all methods, the results also provide a useful comparison between data preprocessing methods.
[0070] Figures 5A-5C show the results of observing artifacts such as pseudo-gene-gene correlations when processing gene expression data using data preprocessing methods. The distribution of correlations differed among these methods, as shown in Figure 5A. NormUMI had a central distribution close to 0, while NBR, DCA, and MAGIC had distinct inflated correlation distributions. The lines indicate the median. Figure 5B shows the enrichment of the top correlated gene pairs in the protein-protein interactions of each method. The x-axis indicates the top n gene pairs. The y-axis indicates the fraction of the n gene pairs that appear in the STRING protein-protein interaction database. The enrichment degree of NormUMI was the highest, followed by SAVER, MAGIC, DCA, and NBR. Figure 5C shows that there was low consistency among the methods for inferring highly correlated gene pairs. The lower triangle shows the overlap of the top 5000 gene pairs among the methods. This was highest between NormUMI and DCA. Only 30 gene pairs were ranked in the top 5,000 by both methods. The upper triangle compares the exact ranks of the shared pairs among the methods, showing a low degree of agreement.
[0071] The consistency of highly correlated gene pairs derived from five data preprocessing procedures was compared. For the top 5,000 gene pairs from each method, a one-to-one comparison was performed. As a result, it was shown that the overlap of gene pairs among the methods was minimal. For example, among the top 5,000 pairs, only one gene pair was shared by NormUMI and NBR. The most overlapping was between NormUMI and DCA, with only 30 gene pairs shared by the two methods (lower triangle in Figure 5C). The ranks of the overlapping pairs in each method were further compared. As a result, it was shown that there was no clearly defined or distinct relationship among these methods (upper triangle in Figure 5C). Although this approach did not yield complete quantitative results, it was shown that the high correlations derived from these data preprocessing methods were likely artifacts.
[0072] Example 4. Unrelated genes as negative control gene pairs Negative control gene pairs were used to investigate potential causes of spurious correlations. Negative control gene pairs were defined by the following criteria: (i) two genes should not appear as interacting pairs within the STRING database, (ii) two genes should not share any gene ontology (GO) terms (Ashburner et al., Gene ontology: tool for the unification of biology. The Gene Ontology Consortium. Nature genetics, 2000. 25(1): p. 25-29; The Gene Ontology Consortium, The Gene Ontology Resource: 20 years and still going strong. Nucleic Acids Research, 2018. 47(D1): p. D330-D338), and (iii) two genes should not be on the same chromosome.
[0073] Scatter plots of the expression values of the gene pair of MB21D1 and OGT, for example, negative gene control pairs, after applying different data preprocessing methods are shown in FIG. 6. There was no existing evidence indicating the correlation between these two genes. Only 3 out of 6534 cells in cluster 2 had non-zero expression values in both genes in the original expression matrix. Five representative data preprocessing methods, such as NormUMI, NBR, DCA, MAGIC, and SAVER, were applied to the analysis. MB21D1 and OGT, one of the negative control gene pairs, had high correlations after applying the NBR (ρ = 0.843), DCA (ρ = 0.828), or MAGIC (ρ = 0.739) processing methods in cell cluster #2. Visualization suggested that these correlation artifacts could be caused by over-smoothing of the data.
[0074] Among the five methods, NormUMI was the only method that maintained zero counts from raw data. In the analysis using NormUMI, out of 6,534 cells, 6,110 cells (93.5%) had zero values in both genes, 3 cells (0.04%) had non - zero values in both genes, and 1.3% and 5.2% of the cells had non - zero values for MB21D1 and OGT respectively. In the other four methods, zeros changed significantly from the original expression matrix. After applying these procedures, some degree of over - smoothing occurred in all of the processed data, especially in the "double - zero regions" within the original data, creating correlation artifacts as shown in Figure 6. NBR is not a complementary method but only minimally shifted zero values, and due to the different adjusted per - cell sizes, an artificial rank correlation was introduced.
[0075] Example 5. Reduction of spurious correlations by applying noise regularization methods Noise regularization methods were applied to reduce spurious correlations. Random noise was added to each entry of the expression matrix processed by pre - treatment methods, such as NormUMI, NBR, DCA, MAGIC, and SAVER. As an example, the expression value of gene i in cell j is denoted as V. The noise was generated by the following steps: (i) calculating the expression distribution of gene i after applying various data pre - treatment methods; (ii) determining the 1st percentile of the expression values of gene i, denoted as M, where M is used as the maximum value of the noise level; and (iii) generating a uniform distribution random number in the range of 0 to M and adding this random number to V.
[0076] After applying this noise regularization method to each preprocessing method, the gene-gene correlation was recalculated. FIG. 7A shows the results of a Spearman correlation analysis, e.g., the correlation distribution, after applying noise regularization to each method according to an exemplary embodiment. Different colors indicate different methods. The results show that, as shown in FIG. 7A with respect to the distribution of correlations, the median correlation in all five methods shifts to 0, indicating that the application of noise regularization reduces the inflation of correlations.
[0077] Figure 7B shows the enrichment of top-correlation gene pairs in protein-protein interactions after applying noise regularization according to an exemplary embodiment. The x-axis shows the top n gene pairs. The y-axis shows the fraction of the n gene pairs that appear in the STRING protein-protein interaction database. Different colors indicate different methods. The error bars of the solid lines indicate the 99% confidence intervals based on 10 repetitions. In all methods, a substantial improvement in protein-protein interaction enrichment in the top-correlation genes was observed. NBR previously had the lowest enrichment degree in protein-protein interactions. However, after applying the noise regularization method, NBR shows the highest enrichment degree in protein-protein interactions. Among the top 100, 1,000, and 10,000 correlation gene pairs in NBR, 99.0%, 96.8%, and 67.7% of the gene pairs can be found in the protein-protein interaction database, corresponding to improvements of 99.0-fold, 50.9-fold, and 31.6-fold, respectively. DCA previously had an average protein-protein interaction enrichment of about 12% in previous results. After noise regularization, DCA has an enrichment of about 97.6% for the top 100 pairs and about 55.8% for the top 10,000 pairs, corresponding to an improvement of about 5-fold. NormUMI, which previously showed the highest enrichment degree, also showed an improvement of about 1.1 - 1.3-fold. To test whether these results of noise regularization are robust and reproducible, the procedure was repeated 10 times with different random seeds to generate random noise. The enrichment performance of protein-protein interactions was stable between each repetition. The standard deviation of NBR at many points was less than 0.1% (the error bars represent the 99% confidence intervals in Figure 7B).
[0078] Figure 7C shows the consistency among methods after applying noise regularization when inferring highly correlated gene pairs. There were more overlapping gene pairs among different methods. Among the top 5,000 gene pairs, there were 2,851 (57%) overlapping pairs between NormUMI and NBR (lower triangle in Figure 7C), and there was a significant correlation among the overlapping gene pairs (Spearman correlation = 0.50, P-value = 1.77e-181, upper triangle in Figure 7C). Among other methods, a certain degree of consistency was shown, especially among highly ranked genes. As shown in Figure 7C, the degree of consistency among different methods was higher compared to the results generated without applying noise regularization as shown in Figure 5C. For example, more than 50% of the gene pairs were shared between NormUMI and NBR after applying noise regularization.
[0079] Example 6. Gene-Gene Correlation Network Inferred from scRNA-seq Data Using the gene-gene correlations revealed by scRNA-seq, a more comprehensive network can be reconstructed to reveal cell type-specific modules. The combination of NBR and noise regularization of the present application described in the previous examples generated the highest protein-protein interaction enrichment among all methods. Therefore, as described in the previous examples, the gene-gene correlation network was reconstructed using the gene-gene correlations derived by applying NBR and noise regularization of the present application to scRNA-seq data.
[0080] Housekeeping genes typically reflect basic and general cellular functions, so housekeeping genes with links were removed from the network construct to focus on cell-type specific interactions. The top 1,000 gene pairs with the highest correlations were obtained from each cluster (cluster #0 to cluster #9), and the network was reconstructed. Two algorithms from graph theory, degree and PageRank, were used to measure the importance of each gene in the network. The value of the degree of a gene in the network is equal to the number of links (interactions) the gene has (Bondy et al., Graph Theory. 2008: Springer Publishing Company, Incorporated. 654). Since important genes tend to connect to more genes, important genes should have relatively high degree values. In addition to the amount of links, PageRank is thought to evaluate the quality of links to a gene by measuring the overall popularity of the gene (Page et al: Bringing order to the web. 1999, Stanford InfoLab).
[0081] Compared with the network constructed without noise regularization, the network constructed with the addition of noise regularization can better represent biological functions in terms of topological structure. Furthermore, genes with high degree or PageRank values also tend to have important functions in the immune system. For example, LYZ, CD79B, and NKG7 are important marker genes for monocytes, B cells, and natural killer cells, respectively. These three genes had high PageRank and degree values within the noise-regularized network. In contrast, as shown in FIGS. 8A and 8B, when noise regularization was not applied, CD79B and NKG7 were not present in the network at all. Furthermore, the network was further improved using known protein-protein interaction information (Cheng et al., Inferring Transcriptional Interactions by the Optimal Integration of ChIP-chip and Knock-out Data. Bioinformatics and biology insights, 2009.3: p. 129-140; Sayyed-Ahmad et al., Transcriptional regulatory network refinement and quantification through kinetic modeling, gene expression microarray data and information theory. BMC Bioinformatics, 2007.8(1): p. 20). Only the gene-gene correlations found in the STRING protein-protein interaction database were retained. Subsequently, EntOptLayout (Agg et al.) was applied. EntOptLayout is a network algorithm that efficiently visualizes various modules within a network.
[0082] The final network revealed a plurality of cell type-related modules that match the cell types within the benchmark dataset, as shown in Figure 8C. This network formed distinct immune cell type-related modules. For example, the upper right corner represents the B cell and pre-B cell modules, and CD78A and CD79B were evaluated to have a higher PageRank (node size in Figure 8C). Similarly, the lower right corner represents the natural killer cell module, and the middle right region represents T cells, as well as the transition from cytotoxic CD8 T cells to natural killer cells. As a result, after performing noise regularization, it was shown that a gene-gene co-expression network that better reflects the network existing in biology can be reconstructed using scRNA-seq data.
[0083] Figures 8A-8C show gene-gene correlation networks inferred from scRNA-seq data. Figures 8A and 8B show comparisons of the degree and PageRank of each gene in the correlation networks constructed before and after applying noise regularization. Genes presented in one network and not present in the other were assigned zero values in the non-present network. Cell-type marker genes such as NKG7, CD79B, or HBB had relatively high degree and PageRank after noise regularization. Figure 8C shows the construction of a network with improved gene-gene correlation. The scRNA-seq data was processed by applying NBR and noise regularization. Further, links that did not exist in the protein-protein interaction were removed. As shown in Figure 8C, the node size is proportional to the PageRank of the gene. Cell-type marker genes such as CD79A, CD79B, NKG7, GNLY, LYZ, or STMN1 have high PageRank, indicating their importance in different cell types. Cell-type related genes also formed cell-type specific modules. Figure 9 shows the enrichment of top-correlated gene pairs in the reactome pathways before and after applying noise regularization. The x-axis shows the top n gene pairs. The y-axis shows the fraction of n gene pairs that appear in the same pathway of the reactome database. The dashed and solid lines represent before and after noise regularization, respectively.
[0084] Example 7. Determination of the Optimal Noise Level The optimal noise level added during noise regularization was determined by comparing it with the expression level of each gene. Different noise levels, such as 0.1, 1, 2, 5, 10, or 20 percentiles of the expression level of each gene, were tested by applying five representative data preprocessing methods, such as NormUMI, NBR, DCA, MAGIC, and SAVER. The results, as shown in Figure 10, indicate that the 1st percentile optimally generated the highest protein-protein interaction enrichment for all five methods. Subsequently, as shown in Figure 11, random noise in the range of approximately 0 to 1 percentile of the gene expression level was generated and added to the expression matrix. This noise regularization process significantly reduced the false correlations between top gene pairs by generating more reliable gene-gene relationships.
[0085] As shown in Figure 11, the noise regularization process involves obtaining an expression matrix processed by a specific scRNA-seq preprocessing method, which contained the expression of n genes in m cells. Assuming that V is the expression value of gene i in cell j, the following steps, namely, (1) determining the expression distribution of gene i across all cells, (2) considering the 1st percentile from the expression distribution of gene i as the maximum noise level of gene i, denoted as M (if M is smaller than the minimum value m, then m is used as the maximum noise level), (3) generating random numbers in the range of 0 to M under a uniform distribution, (4) adding this random number to V to obtain the noise-regularized expression value, and (5) repeating this procedure for all entries in the expression matrix, generate random noise and add it to V. The technical ideas included in the present disclosure are appended below. [Appendix 1] A method for improving data processing for gene-gene correlation, comprising: processing gene expression data for normalization or complementation; applying a noise regularization process to the normalized or complemented gene expression data; and applying a gene-gene correlation calculation process to obtain correlated gene pairs. [Appendix 2] The method according to Appendix 1, wherein the gene expression data is single-cell gene expression data. [Appendix 3] The method according to Appendix 1, wherein the noise regularization process includes adding random noise to the expression values of genes in cells in the expression matrix. [Appendix 4] The method according to Appendix 3, wherein the random noise is determined by the expression level of the gene. [Appendix 5] The random noise is determined by: determining the expression distribution of the gene across all of the cells in the expression matrix; regarding about 0.1% to about 20% of the expression level of the gene as the maximum noise level; generating a random number in the range from 0 to the maximum noise level under a uniform distribution; and adding the random number to the expression values of the gene in the cells in the expression matrix to obtain a noise-regularized expression matrix. [Appendix 6] The random noise is determined by: determining the expression distribution of the gene across all of the cells in the expression matrix; regarding 1% of the expression level of the gene as the maximum noise level; generating a random number in the range from 0 to the maximum noise level under a uniform distribution; and adding the random number to the expression values of the gene in the cells in the expression matrix to obtain a noise-regularized expression matrix. [Appendix 7] The method according to Appendix 1, wherein the gene-gene correlation calculation process is performed within a cell cluster. [Appendix 8] The method according to Appendix 1, further comprising performing enrichment on the gene expression data associated with the correlated gene pairs. [Appendix 9] The method according to Appendix 1 or 3 or 4 or 5 or 6, wherein full-unique molecule identifier normalization (NormUMI), regularized negative binomial regression (NBR), deep count autoencoder network (DCA), Markov affinity-based graph completion of cells (MAGIC), or single-cell analysis via expression recovery (SAVER) is used to process gene expression data for normalization or complementation. [Appendix 10] The method according to Appendix 1 or 3 or 4 or 5 or 6, further comprising constructing a gene-gene correlation network based on the correlated gene pairs. [Appendix 11] The method according to Appendix 10, wherein the gene-gene correlation network is cell type-specific. [Appendix 12] Using the gene-gene correlation network to map molecular interactions, guiding experimental design to investigate biological events, discovering biomarkers, guiding comparative network analysis, guiding drug design, identifying changes in gene-gene interactions by comparing the healthy and diseased states of cells, guiding drug development, predicting gene transcriptional regulation, improving drug efficiency, or identifying drug resistance factors. The method according to Appendix 10. [Appendix 13] A gene-gene correlation network, wherein the network is constructed based on correlated gene pairs, and the correlated gene pairs are obtained using the method according to Appendix 1. [Appendix 14] A computer-implemented method for data processing for gene-gene correlation, comprising: Retrieving gene expression data; Processing the gene expression data for normalization or complementation; Applying a noise regularization process to the normalized or complemented gene expression data; Applying a gene-gene correlation calculation process to obtain correlated gene pairs and constructing a gene-gene correlation network based on the correlated gene pairs. [Appendix 15] The method according to Appendix 14, wherein the gene expression data is single-cell gene expression data. [Appendix 16] The method according to Appendix 14, wherein the noise regularization process includes adding random noise to the expression values of genes within cells in the expression matrix. [Appendix 17] The method according to Appendix 16, wherein the random noise is determined by the expression level of the gene. [Appendix 18] The random noise is Determining the expression distribution of the gene across all of the cells in the expression matrix; Regarding about 0.1 to about 20 percentiles of the expression level of the gene as the maximum noise level; Generating random numbers in the range from 0 to the maximum noise level under a uniform distribution; The method according to Supplementary Note 16, which is determined by adding the random numbers to the expression values of the gene in the cells in the expression matrix to obtain a noise-normalized expression matrix. [Supplementary Note 19] The random noise is Determining the expression distribution of the gene across all of the cells in the expression matrix; Regarding the 1 percentile of the expression level of the gene as the maximum noise level; Generating random numbers in the range from 0 to the maximum noise level under a uniform distribution; The method according to Supplementary Note 16, which is determined by adding the random numbers to the expression values of the gene in the cells in the expression matrix to obtain a noise-normalized expression matrix. [Supplementary Note 20] The method according to Supplementary Note 14, wherein the gene-gene correlation calculation process is performed within a cell cluster. [Supplementary Note 21] The method according to Supplementary Note 14, further comprising performing enrichment on the gene expression data associated with the correlated gene pairs. [Supplementary Note 22] The method according to Supplementary Note 14 or 16 or 17 or 18 or 19, wherein full unique molecular identifier normalization (NormUMI), regularized negative binomial regression (NBR), deep count autoencoder network (DCA), Markov affinity-based graph completion of cells (MAGIC), or single-cell analysis via expression recovery (SAVER) is used to process gene expression data for normalization or complementation. [Supplementary Note 23] The method according to Supplementary Note 14, wherein the gene-gene correlation network is cell type-specific. The method according to appended note 14 or 16 or 17 or 18 or 19, further comprising using the gene-gene correlation network to map molecular interactions, guiding experimental design to investigate biological events, discovering biomarkers, guiding comparative network analysis, guiding drug design, identifying changes in gene-gene interactions by comparing the healthy and diseased states of cells, guiding drug development, predicting transcriptional regulation of genes, improving drug efficiency, or identifying drug resistance factors. [Appended note 25] A system for generating a gene-gene network, A database configured to store gene expression data, A memory configured to store instructions, At least one processor coupled to the memory, wherein the at least one processor Retrieves the gene expression data, Processes the gene expression data for normalization or complementation, Applies a noise regularization process to the normalized or complemented gene expression data, Applies a gene-gene correlation calculation process to obtain correlated gene pairs, Constructs a gene-gene correlation network based on the correlated gene pairs, and at least one processor configured to execute instructions for performing the above, A user interface coupled to the processor, configured to receive queries for gene-gene correlation and display the results of the correlated gene pairs and the constructed gene-gene correlation network. [Appended note 26] The system according to appended note 25, wherein the gene expression data is single-cell gene expression data. [Appended note 27] The system according to appended note 25, wherein the noise regularization process includes adding random noise to the expression values of genes within cells in the expression matrix. [Appended note 28] The system according to appended note 27, wherein the random noise is determined by the expression level of the gene. [Appended note 29] The random noise Determines the expression distribution of the gene across all of the cells in the expression matrix, Considers the approximately 0.1% to approximately 20% percentile of the expression level of the gene as the maximum noise level, Generates a random number in the range from 0 to the maximum noise level under a uniform distribution, The system according to appended note 27, which is determined by adding the random number to the expression value of the gene in the cell in the expression matrix to obtain a noise-regularized expression matrix. [Appended note 30] The random noise is determining the expression distribution of the gene across all of the cells in the expression matrix, regarding the 1st percentile of the expression level of the gene as the maximum noise level, generating a random number in the range from 0 to the maximum noise level under a uniform distribution, The system according to appended note 27, which is determined by adding the random number to the expression value of the gene in the cell in the expression matrix to obtain a noise-regularized expression matrix. [Appended note 31] The system according to appended note 25, wherein the gene-gene correlation calculation process is performed using cell clusters. [Appended note 32] The system according to appended note 25, wherein the at least one processor is further configured to perform enrichment on the gene expression data associated with the correlated gene pairs. [Appended note 33] Total unique molecular identifier normalization (NormUMI), regularized negative binomial regression (NBR), deep count autoencoder network (DCA), Markov affinity-based graph completion of cells (MAGIC), or single-cell analysis via expression recovery (SAVER) is used to process gene expression data for normalization or complementation, the system according to appended note 25 or 27 or 28 or 29 or 30. [Appended note 34] The system according to appended note 25, wherein the gene-gene correlation network is cell-type specific. [Appended note 35] The at least one processor is further configured to utilize the gene-gene correlation network to map molecular interactions, to guide experimental design to investigate biological events, to discover biomarkers, to guide comparative network analysis, to guide drug design, to identify changes in gene-gene interactions by comparing the healthy and diseased states of cells, to guide drug development, to predict transcriptional regulation of genes, to improve drug efficiency, or to identify drug resistance factors, the system according to appended note 25 or 27 or 28 or 29 or 30.
Claims
**Claim 1** A method for improving data processing for gene-gene correlation, comprising: processing gene expression data for normalization or complementation; applying a noise regularization process to the normalized or complemented gene expression data, the noise regularization process including adding random noise to the expression values of genes in cells in an expression matrix, the random noise determining the expression distribution of the genes across all of the cells in the expression matrix, considering about 0.1 to about 20 percentiles of the expression levels of the genes as the maximum noise level, generating random numbers in a range from 0 to the maximum noise level under a uniform distribution, and adding the random numbers to the expression values of the genes in the cells in the expression matrix to obtain a noise-regularized expression matrix; applying a gene-gene correlation calculation process to the gene expression data to which the noise regularization process has been applied to obtain correlated gene pairs. **Claim 2** The method according to claim 1, wherein the gene expression data is single-cell gene expression data. **Claim 3** The random noise is determined by: determining the expression distribution of the genes across all of the cells in the expression matrix; considering 1 percentile of the expression levels of the genes as the maximum noise level; generating random numbers in a range from 0 to the maximum noise level under a uniform distribution; and adding the random numbers to the expression values of the genes in the cells in the expression matrix to obtain a noise-regularized expression matrix. The method according to claim 1. **Claim 4** The method according to claim 1, wherein the gene-gene correlation calculation process is performed within a cell cluster. **Claim 5** The method according to claim 1, further comprising performing enrichment on the gene expression data associated with the correlated gene pairs. **Claim 6** The method according to claim 1 or 3, wherein full-identity molecule identifier normalization (NormUMI), regularized negative binomial regression (NBR), deep count autoencoder network (DCA), cell Markov affinity-based graph completion (MAGIC), or single-cell analysis via expression recovery (SAVER) is used to process gene expression data for normalization or complementation.
7. The method according to claim 1 or 3, further comprising constructing a gene-gene correlation network based on the correlated gene pairs.
8. The method according to claim 7, wherein the gene-gene correlation network is cell type-specific.
9. Using the gene-gene correlation network to map molecular interactions, guiding experimental design to investigate biological events, discovering biomarkers, guiding comparative network analysis, guiding drug design, identifying changes in gene-gene interactions by comparing the healthy and diseased states of cells, guiding drug development, predicting transcriptional regulation of genes, improving drug efficiency, or identifying drug resistance factors, the method according to claim 7.
10. A computer-implemented method for data processing for gene-gene correlation, comprising: retrieving gene expression data; processing the gene expression data for normalization or complementation; applying a noise regularization process to the normalized or complemented gene expression data, the noise regularization process including adding random noise to the expression values of genes within cells in an expression matrix, the random noise being determined by determining the expression distribution of the genes across all of the cells in the expression matrix, considering the about 0.1 to about 20 percentile of the expression levels of the genes as the maximum noise level, generating random numbers in the range from 0 to the maximum noise level under a uniform distribution, and adding the random numbers to the expression values of the genes within the cells in the expression matrix to obtain a noise-regularized expression matrix, the applying. Applying a gene-gene correlation calculation process to the gene expression data to which the noise regularization process has been applied to obtain correlated gene pairs, and constructing a gene-gene correlation network based on the correlated gene pairs, the method comprising.
11. The method according to claim 10, wherein the gene expression data is single-cell gene expression data.
12. The random noise is Determining the expression distribution of the gene across all of the cells in the expression matrix; Regarding the 1st percentile of the expression level of the gene as the maximum noise level; Generating a random number in the range from 0 to the maximum noise level under a uniform distribution; Adding the random number to the expression value of the gene in the cell in the expression matrix to obtain a noise-regularized expression matrix, the method according to claim 10, determined by.
13. The method according to claim 10, wherein the gene-gene correlation calculation process is performed within a cell cluster.
14. The method according to claim 10, further comprising enriching the gene expression data associated with the correlated gene pairs.
15. Total unique molecular identifier normalization (NormUMI), regularized negative binomial regression (NBR), deep count autoencoder network (DCA), cell Markov affinity-based graph completion (MAGIC), or single-cell analysis via expression recovery (SAVER) is used to process gene expression data for normalization or complementation, the method according to claim 10 or 12.
16. The method according to claim 10, wherein the gene-gene correlation network is cell-type specific.
17. Using the gene-gene correlation network to map molecular interactions, guiding experimental design to investigate biological events, discovering biomarkers, guiding comparative network analysis, guiding drug design, comparing the healthy and diseased states of cells to identify changes in gene-gene interactions, guiding drug development, predicting gene transcriptional regulation, improving drug efficiency, or identifying drug resistance factors, the method according to claim 10 or 12.
18. A system for generating a gene-gene network, comprising A database configured to store gene expression data, a memory configured to store instructions, at least one processor coupled to the memory, the at least one processor being configured to: retrieve the gene expression data, process the gene expression data for normalization or complementation, apply a noise regularization process to the normalized or complemented gene expression data, the noise regularization process including adding random noise to the expression values of genes within cells in an expression matrix, the random noise determining the expression distribution of the genes across all of the cells in the expression matrix, considering about 0.1 to about 20 percentiles of the expression levels of the genes as the maximum noise level, generating random numbers in a range from 0 to the maximum noise level under a uniform distribution, and adding the random numbers to the expression values of the genes within the cells in the expression matrix to obtain a noise-regularized expression matrix, apply a gene-gene correlation calculation process to the gene expression data to which the noise regularization process has been applied to obtain correlated gene pairs, construct a gene-gene correlation network based on the correlated gene pairs, and execute instructions for doing so; and a user interface coupled to the processor, configured to receive queries for gene-gene correlation and display the results of the correlated gene pairs and the constructed gene-gene correlation network.
19. The system according to claim 18, wherein the gene expression data is single-cell gene expression data.
20. The random noise is determined by: determining the expression distribution of the genes across all of the cells in the expression matrix, considering 1 percentile of the expression levels of the genes as the maximum noise level, generating random numbers in a range from 0 to the maximum noise level under a uniform distribution, and adding the random numbers to the expression values of the genes within the cells in the expression matrix to obtain a noise-regularized expression matrix. The system according to claim 18.
21. The system according to claim 18, wherein the gene-gene correlation calculation process is performed using cell clusters.
22. The system according to claim 18, wherein the at least one processor is further configured to perform enrichment on the gene expression data associated with the correlated gene pairs.
23. Total unique molecular identifier normalization (NormUMI), regularized negative binomial regression (NBR), deep count autoencoder network (DCA), Markov affinity-based graph completion of cells (MAGIC), or single-cell analysis via expression recovery (SAVER) is used to process gene expression data for normalization or complementation, according to the system of claim 18 or 20.
24. The system according to claim 18, wherein the gene-gene correlation network is cell type specific.
25. The at least one processor is further configured to utilize the gene-gene correlation network to map molecular interactions, to guide experimental design to investigate biological events, to discover biomarkers, to guide comparative network analysis, to guide drug design, to identify changes in gene-gene interactions by comparing the healthy and diseased states of cells, to guide drug development, to predict transcriptional regulation of genes, to improve drug efficiency, or to identify drug resistance factors, according to the system of claim 18 or 20.
Citation Information
Patent Citations
Method for identifying expression distinguishers in biological samples
US20180251849A1