Correction method for count data set of single-cell RNA-Seq analysis, single-cell RNA-Seq analysis method, analysis method for cell type composition ratio, and apparatus and computer program for executing these methods

By weighting single-cell RNA-Seq data based on total RNA content and using signature gene sets, the method addresses challenges in analyzing diverse tissues and improves the accuracy of cell type composition ratios.

JP7689737B2Active Publication Date: 2025-06-09KARYDO THERAPEUTIX INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2021576209
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-02-06
Filing Date
2021-02-06
Publication Date
2025-06-09
Estimated Expiration
2041-02-06

AI Technical Summary

Technical Problem

Existing methods for single-cell RNA-Seq analysis face challenges such as the requirement for fresh tissues, difficulties in analyzing rare pathological samples, and experimental artifacts in gene expression, particularly in handling diverse tissues and organs.

Method used

A deconvolution method that weights single-cell RNA-Seq count datasets based on total RNA content for each cell type, using signature gene sets to characterize cell types and correct for RNA content variations, thereby improving the accuracy of cell type composition ratios in various tissues.

Benefits of technology

The method effectively estimates the proportion of each cell type closer to actual tissue ratios and can handle a greater variety of tissues, reducing deviations in cell type composition ratios compared to previous methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007689737000031
    Figure 0007689737000031
  • Figure 0007689737000032
    Figure 0007689737000032
  • Figure 0007689737000033
    Figure 0007689737000033
Patent Text Reader

Abstract

The present application discloses a correction method for a single-cell RNA-seq analysis count data set, said method comprising weighting the single-cell RNA-seq analysis count data set, which was acquired from analysis target cells or estimated with respect to the analysis target cells, on the basis of the total RNA content of each cell type corresponding to the analysis target cells.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification discloses a method for correcting a count dataset of single-cell RNA-Seq analysis, a method for analyzing single-cell RNA-Seq, a method for analyzing the composition ratio of cell types, and an apparatus and a computer program for executing these methods.

Background Art

[0002] Human organs are composed of about 1×10 8 -3×10 12 cells. Changes in the cell composition and / or cell phenotype of an organ are closely related to its dysfunction, remodeling, and regeneration. Each organ is a mixed population of cells. Therefore, single-cell RNA-Seq (Single-cell RNA-Seq or scRNA-Seq) analyzes the comprehensive gene expression profile for the cell population of each organ in order to capture changes in the cell composition and / or cell phenotype of the organ, decomposes the analysis data to the expression level of single cells, and derives information on single-cell changes (Non-Patent Documents 1 to 5). Therefore, scRNA-Seq is said to be a powerful method for generating a detailed molecular cell atlas of normal and diseased organs.

[0003] However, scRNA-Seq has limitations. First, for scRNA-Seq, tissue collected from an organ must be used to recover individual cells using digestive enzymes or physical disruption. On this premise, the tissue needs to be fresh to recover such cells. In other words, tissues collected by surgery or the like are generally cryopreserved for several months to several years, and such preserved tissues cannot be used for scRNA-Seq. In pathological tissue diagnosis after surgery, even if it is found to be a rare disease, it is difficult to newly obtain such rare pathological samples, and samples that can be used for RNA expression analysis are generally cryopreserved. In addition, the collection of human tissue is generally a biopsy, and there is a problem that the sample amount is small. Even if the whole organ can be collected by autopsy or the like, in the case of large organs such as the heart and brain, it is not impossible but not practical to separate individual cells from the whole organ for the purpose of scRNA-Seq.

[0004] Also, in many cases, in the study of drug effects and / or etiology, it is necessary to analyze the effects of drugs and / or pathological conditions in different organs of the same subject. However, in the case of humans, there is also a problem that it is difficult to conduct an analysis by collecting multiple types of organs from the same person.

[0005] Furthermore, scRNA-Seq also has problems with experimental artifacts in gene expression. As such an example, it has been reported that abnormal gene expression is induced in cells during the cell separation process.

[0006] To solve the above problems, computational deconvolution of whole-organ RNA datasets has been proposed. Whole-organ RNA database deconvolution extracts RNA from the collected test tissue without separating cells for each cell type, obtains the sequence information of the expressed RNA by RNA-Seq, and then estimates the expression level of each RNA for each cell type based on the proportion of cell types contained in the test tissue calculated by a computer. This method enables RNA expression analysis using not only fresh tissues but also cryopreserved tissues. In addition, this method also enables simultaneous purification of RNA from multiple organs.

[0007] Several computer analysis methods for deconvolving whole-organ RNA-Seq data have been proposed so far (Non-Patent Documents 6-19). These methods use almost all of the RNA-Seq data of the corresponding organ to calculate the cell type composition of the organ to be analyzed.

[0008] Recently, Multi-Subject Single Cell deconvolution (MuSiC) (Non-Patent Document 17), Dampened Weighted Least Squares (DWLS) (Non-Patent Document 18), and Complete Deconvolution for Sequencing data (CDSeq) method (Non-Patent Document 19) have been reported. These three methods are said to be superior to the methods described in Non-Patent Documents 6 to 16 reported previously.

Prior Art Documents

Non-Patent Documents

[0009]

Non-Patent Document 1

Non-Patent Document 9

Non-Patent Document 10

Non-Patent Document 11

Non-Patent Document 12

Non-Patent Document 13

Non-Patent Document 14

Non-Patent Document 15

Non-Patent Document 16

Non-Patent Document 17

[0010] However, the methods, synthetic datasets, cultured cells, mixtures of several tissues, and / or RNA-Seq data derived from one to four actual organs described in Non-Patent Documents 17 to 19 have only been verified to be useful. In other words, the applicability to more diverse actual organs has not been examined. The present inventors evaluated the performance of MuSiC (Non-Patent Document 17) and the DWLS method (Non-Patent Document 19). These are the latest two methods that perform deconvolution on one to four actual organs and are shown to be superior to other previous methods. However, as shown in the verification of the effects described later, the ratios of cell types calculated by the computer using the MuSiC or DWLS method deviated from those experimentally estimated by actual scRNA-Seq studies, and the degrees of deviation were also various. In particular, the deviation was significant in skeletal muscle and the heart.

[0011] Therefore, in order to eliminate such a deviation, an object of the present invention is to provide a deconvolution method for RNA-Seq data for estimating the proportion of each cell type, which is closer to the proportion of various cells in an actual tissue. Another object of the present invention is to provide a deconvolution method for RNA-Seq data that can handle a greater variety of tissues.

Means for Solving the Problems

[0012] One embodiment of the present invention relates to a method for correcting a count data set of single-cell RNA-Seq analysis, which includes weighting a count data set of single-cell RNA-Seq analysis obtained from an analysis target cell or predicted for the analysis target cell, based on the total RNA content for each cell type corresponding to the analysis target cell.

[0013] Preferably, the weighting is performed based on the expression of a signature gene set that characterizes each cell type, and the signature gene set includes a predetermined number of genes.

[0014] One embodiment of the present invention relates to a method for analyzing single-cell RNA-Seq, which includes weighting a count data set of single-cell RNA-Seq analysis obtained from an analysis target cell or predicted for the analysis target cell, based on the total RNA content for each cell type corresponding to the analysis target cell, and analyzing the RNA expression pattern in each cell type that constitutes an analysis target organ including the analysis target cell, based on the weighted count data set of single-cell RNA-Seq analysis.

[0015] One embodiment of the present invention is to weight a count data set of single-cell RNA-Seq analysis obtained from cells to be analyzed or predicted for cells to be analyzed based on the total RNA content for each cell type corresponding to the cells to be analyzed, and to analyze the composition ratio of cell types constituting an organ to be analyzed based on the weighted count data set of single-cell RNA-Seq analysis. The present invention relates to a method for analyzing the composition ratio of cell types constituting an organ to be analyzed.

[0016] One embodiment of the present invention relates to a correction device (10) for a count data set of single-cell RNA-Seq analysis. The correction device (10) includes a control unit (101). The control unit (101) weights a count data set of single-cell RNA-Seq analysis obtained from cells to be analyzed based on the total RNA content for each cell type corresponding to the cells to be analyzed.

[0017] One embodiment of the present invention relates to an analysis device for single-cell RNA-Seq. The analysis device (20) includes a control unit (201). The control unit (201) weights a count data set of single-cell RNA-Seq analysis obtained from cells to be analyzed or predicted for cells to be analyzed based on the total RNA content for each cell type corresponding to the cells to be analyzed, and analyzes the RNA expression pattern in each cell type constituting an organ to be analyzed including the cells to be analyzed based on the weighted count data set of single-cell RNA-Seq analysis.

[0018] One embodiment of the present invention relates to an analysis device for the composition ratio of cell types constituting an organ to be analyzed. The analysis device (20) includes a control unit (201). The control unit (201) weights a count data set of single-cell RNA-Seq analysis obtained from cells to be analyzed or predicted for cells to be analyzed based on the total RNA content for each cell type corresponding to the cells to be analyzed, and analyzes the composition ratio of cell types constituting an organ to be analyzed including the cells to be analyzed based on the weighted count data set of single-cell RNA-Seq analysis.

[0019] One embodiment of the present invention relates to a correction program for a count data set of single-cell RNA-Seq analysis, which, when executed by a computer, causes the computer to execute a process including a step of weighting a count data set of single-cell RNA-Seq analysis obtained from or predicted for analysis target cells based on the total RNA content for each cell type corresponding to the analysis target cells.

[0020] One embodiment of the present invention relates to an analysis program for single-cell RNA-Seq, which, when executed by a computer, causes the computer to execute a process including a step of weighting a count data set of single-cell RNA-Seq analysis obtained from or predicted for analysis target cells based on the total RNA content for each cell type corresponding to the analysis target cells, and a step of analyzing an RNA expression pattern in each cell type constituting an analysis target organ including the analysis target cells based on the weighted count data set of single-cell RNA-Seq analysis.

[0021] One embodiment of the present invention relates to an analysis program for the composition ratio of cell types constituting an analysis target organ, which, when executed by a computer, causes the computer to execute a process including a step of weighting a count data set of single-cell RNA-Seq analysis obtained from or predicted for analysis target cells based on the total RNA content for each cell type corresponding to the analysis target cells, and a step of analyzing the composition ratio of cell types constituting the analysis target organ including the analysis target cells based on the weighted count data set of single-cell RNA-Seq analysis.

Advantages of the Invention

[0022] The present invention can estimate the ratio of each cell type closer to the ratio of various cells in an actual tissue from an RNA sequence database. Further, according to the present invention, the ratio of each cell type can be estimated in more diverse tissues.

Brief Description of the Drawings

[0023]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Mode for Carrying Out the Invention

[0024] 1. Correction method, analysis method, and analysis method for the composition ratios of cell types constituting the analysis target organ One embodiment of the present invention relates to a method, an apparatus, and a program for correcting a count data set of single-cell RNA-Seq analysis. 1-1. Method for correcting a count data set of single-cell RNA-Seq analysis

[0025] A method for correcting a count data set of single-cell RNA-Seq (scRNA-Seq) analysis (hereinafter, also simply referred to as "correction method") includes weighting a count data set of single-cell RNA-Seq analysis obtained from or predicted for cells to be analyzed, based on the total RNA content for each cell type. 1-1-1. RNA-Seq analysis

[0026] In the present specification, RNA is not limited as long as it can be analyzed by RNA-Seq analysis. RNA may include mRNA, non-coding RNA, microRNA, and the like.

[0027] RNA is not limited as long as it exists in a living organism. The living organism is not limited as long as it is a multicellular organism having organs. The living organism may be an animal or a plant, but is preferably an animal. Preferred animals include mammals such as humans, mice, rats, dogs, cats, rabbits, cows, horses, goats, sheep, pigs, etc., and birds such as chickens, etc. More preferably, they are mammals such as humans, mice, dogs, cats, cows, horses, pigs, etc., still more preferably humans, mice, dogs, or cats, etc., even more preferably humans or mice, and most preferably humans. Also, the living organisms include both those having a disease and those not having a disease. The cells to be analyzed are not limited as long as they exist in the organs of the living organism. The organs are preferably cells whose cell composition within the organ is known.

[0028] An organ refers to a collection of several tissues existing in an organism, which has a certain independent form and specific functions. For example, when the organism is a mammal, the organs may include circulatory system organs (such as the heart, arteries, veins, lymphatic vessels, etc.), respiratory system organs (such as the nasal cavity, paranasal sinuses, larynx, trachea, bronchi, lungs, etc.), digestive system organs (such as the lips, cheeks, palate, teeth, gums, tongue, salivary glands, pharynx, esophagus, stomach, duodenum, jejunum, ileum, cecum, appendix, ascending colon, transverse colon, sigmoid colon, rectum, anus, liver, gallbladder, bile ducts, biliary tract, pancreas, pancreatic duct, etc.), urinary system organs (such as the urethra, bladder, ureters, kidneys), nervous system organs (such as the cerebrum, cerebellum, midbrain, brainstem, spinal cord, peripheral nerves, autonomic nerves, etc.), female genital system organs (such as ovaries, fallopian tubes, uterus, vagina, etc.), breasts, male genital system organs (such as the penis, prostate gland, testes, epididymis, vas deferens), endocrine system organs (such as the hypothalamus, pituitary gland, pineal gland, thyroid gland, parathyroid glands, adrenal glands, etc.), integumentary system organs (such as the skin, hair, nails, etc.), hematopoietic system organs (such as blood, bone marrow, spleen, etc.), immune system organs (such as lymph nodes, tonsils, thymus, etc.), osteochondral organs (such as bones, cartilage, skeletal muscles, connective tissues, ligaments, tendons, diaphragm, peritoneum, pleura, adipose tissues (brown adipose, white adipose), etc.), sensory system organs (such as the eyeball, eyelids, lacrimal glands, external ear, middle ear, inner ear, cochlea, etc.). In the present invention, the target tissues preferably include the heart, cerebrum, lungs, kidneys, adipose tissue, liver, skeletal muscle, testes, spleen, thymus, bone marrow, pancreas, skin (for example, including the epidermis, papillary layer, and reticular layer above the subcutaneous tissue). As organs, the aorta, brain, fat, heart, kidneys, large intestine, liver, lungs, bone marrow, pancreas, skin, skeletal muscle, spleen, and thymus are preferred.

[0029] RNA-Seq analysis is an analysis called so-called transcriptome analysis. It is a method of comprehensively obtaining reads containing sequence information from the RNA present in the target sample, mapping the reads to a reference sequence, and analyzing the expressed genes and their count numbers (also called read numbers). The count number corresponds to the gene expression level. The count data of RNA-Seq analysis may include the gene names of the expressed genes and / or their registration numbers in the gene database, and the count numbers of the reads of each gene.

[0030] RNA-Seq analysis can be performed using a DNA sequencer called a next-generation sequencer or a third-generation sequencer. Examples of next-generation sequencers include, for example, MiSeq9 (registered trademark), HiSeq (registered trademark), NextSeq (registered trademark), MiSeq (registered trademark) of Illumina, Inc. (San Diego, CA); Ion Proton (registered trademark), Ion PGM (registered trademark) of Thermo Fisher Scientific (Waltham, MA); GS FLX + (registered trademark), GS Junior (registered trademark) of Roche (Basel, Switzerland), etc. Examples of third-generation sequencers include PacBio Sequel (trademark), etc.

[0031] The count data set of scRNA-Seq analysis is a set of count data generated based on gene expression analysis of individual cells of an organism and / or gene expression predicted by a computer analysis method. For example, the count data set of scRNA-Seq analysis can be count data obtained from actual individual cells by RNA-Seq analysis. Also, the count data set of scRNA-Seq analysis may be, for example, a count data set predicted by deconvolving, by a computer analysis method, the count data obtained by RNA-Seq analysis from an entire organ based on the reference cell composition ratio by the method described in Non-Patent Documents 6-19. As a method for predicting the count data set of scRNA-Seq analysis, for example, the Complete Deconvolution for Sequencing data (CDSeq) method (Non-Patent Document 19) is preferable.

[0032] 1-1-2. Calculation of Weight Coefficient Based on RNA Content in Different Cell Types Present in Each Organ A method for calculating a weight coefficient for weighting a count data set of single-cell RNA-Seq analysis obtained from an analysis target cell or predicted for an analysis target cell based on the total RNA content for each cell type will be described.

[0033] First, in order to perform weighting, it is necessary to obtain information on what cell types each organ is composed of. The cell composition of each organ can be obtained from scRNA-Seq data described in Non-Patent Document 5 or Non-Patent Document 2, or from databases registered with NIH or the like. The composition of these cell types is information obtained by actually analyzing the cell type composition of the tissue of each organ. Such a cell composition of each organ is also referred to as a "reference cell type". The reference cell type usually includes a scRNA-Seq count dataset of genes expressed in each cell type. In addition, the reference cell type includes the composition ratio (also referred to as a reference) of the reference cell type in each organ, which is associated with a label indicating the name or abbreviation of each cell type.

[0034] In the calculation of the weight coefficient, for the aorta, fat, kidney, large intestine, liver, lung, bone marrow, pancreas, skin, spleen, and thymus, it is preferable to use the cell type composition in each organ described in Non-Patent Document 5 as the reference cell type and its composition ratio. For skeletal muscle, it is preferable to use the ratio of cell types described in Non-Patent Document 2 as the reference cell type and its composition ratio.

[0035] For the heart, in accordance with the separation analysis of cardiomyocytes and non-muscle cells, it is preferable to correct the composition ratio of the reference cell type described in Non-Patent Document 5 and use it as the composition ratio of the reference cell type. Specifically, the composition ratio of cardiomyocytes (3.1%) adopted in Non-Patent Document 5 is extremely low compared to the ratio (30% to 40%) generally agreed upon in the field of tissue anatomy based on various past studies. Therefore, in this example, it is preferable to set the composition ratio of cardiomyocytes to 30%, divide the remaining 70% by the composition ratio of non-muscle cell types, and use it as the composition ratio of the reference cell type.

[0036] For the brain, it is preferable to determine the reference cell types and their composition ratios based on the report in NIH (http: / / www.nervenet.org / papers / BrainRev99.html#Numbers). As labels for each cell type in the brain, the corresponding cell type labels of the scRNA-Seq data described in Non-Patent Document 5 were used. First, the classes of cell types in the brain were classified into four classes: "neurons", "glial cells", "endothelial cells", and "others", and the ratios of each class were set to 75:23:7:4, respectively. This ratio follows the estimated ratio in the mouse brain (http: / / www.nervenet.org / papers / BrainRev99.html#Numbers). Next, according to Non-Patent Document 5, "neurons", "glial cells", and "others" were further classified into more detailed cell type classes. Specifically, the "neuron" class was further classified into "neuronal cells-excitatory neurons and some neural stem cells" and "neuronal cells-inhibitory neurons". The "others" class was classified into "pericytes-NA" and "oligodendrocyte progenitor cells-NA". Regarding the "glial cell" class, it was classified based on the following three premises. i) "Glial cells" were classified into four cell types: "microglial cells-NA", "astrocytes-NA", "Bergmann glial cells-NA", and "oligodendrocytes NA" according to Non-Patent Document 5. ii) The composition ratios of these four glial cell types follow the description in Non-Patent Document 5. iii) Since it has been reported that "microglial cells" account for 10-15% of the cells in the whole brain, the ratio of "microglial cells-NA" in the whole brain was set to 0.1.Based on these premises, the ratios of each brain cell type are set as "macrophage - NA" (about 0.2%), "microglial cell - NA" (10.0%), "astrocyte - NA" (about 2.2%), "Bergmann glial cell - NA" (about 2.1%), "pericytes - NA" (about 1.5%), "endothelial cells - NA" (about 6.4%), "neurons - excitatory neurons, and some neural stem cells" (about 47.5%), "neurons - inhibitory neurons" (about 21.3%), "oligodendrocytes - NA" (about 8.7%), "oligodendrocyte progenitor cells - NA" (about 1.9%), and these can be used as the reference cell types and their composition ratios in the brain. The reference cell types used in this specification and the composition ratios of those cell types are shown in the composition ratio list of reference cell types at the end of this book.

[0037] Furthermore, when calculating the weight coefficients, in addition to the composition ratios of the above - mentioned reference cell types, gene expression in each cell type, that is, the count data of scRNA - Seq analysis in each cell type, is required. However, generally, the number of genes targeted by scRNA - Seq analysis ranges from 20,000 to 30,000.

[0038] It may be possible to use the count data of all these genes, but it is efficient to select genes (signature genes) that can characterize each cell type and calculate the weight coefficients using the count data of these gene sets. Such a signature gene set that characterizes each cell type can be calculated, for example, by the method described below.

[0039] First, when selecting signature genes, in the count data of scRNA - Seq analysis, it is preferable to delete the counts derived from the three genes Rn45s, Akap5, and Lrrc17, which are spike - in genes labeled with ERCC and reported as non - mRNA artifacts although they greatly affect the total count, as well as the counts derived from each gene. Also, the RNA counts derived from each gene are such that the total count of each cell in the scRNA - Seq dataset is 100, 10 3 、10 4 、105 、10 6 In cases such as when it is 10, it is preferable to convert and normalize.

[0040] For the selection of signature genes, for example, a classifier generated by training an artificial intelligence such as a random forest can be used. Using the composition ratio of the reference cell types of each organ and the count data set of scRNA-Seq analysis reported for each reference cell type, an artificial intelligence is trained to generate a classifier. For example, when using a random forest as the artificial intelligence, the important features of the classifier are extracted as the signature gene names of each cell type, and the "Mean Decrease Gini" value is used as the importance index of each gene. Genes with a high "Mean Decrease Gini" value are extracted as signature genes. The signature genes can be extracted from about 100 genes to 2000 genes from the ones with a higher "Mean Decrease Gini" value and used as a signature gene set. Next, for each cell type present in each organ, a weight coefficient for correcting the count data set of scRNA-Seq analysis with the RNA content is calculated.

[0041] For calculating the weight coefficient, for each organ, the count data of scRNA-Seq analysis of the signature gene set in each cell type of the reference cell type (also referred to as signature gene scRNA-Seq data) and the count data obtained by RNA-Seq analysis of all the RNA contained in the whole organ (also referred to as whole-organ RNA-Seq data) can be used. Both the signature gene scRNA-Seq count data and the whole-organ RNA-Seq count data are used after being normalized.

[0042] As the whole-organ RNA-Seq data, publicly available count datasets of RNA-Seq analysis can be used. The whole-organ RNA-Seq data of mice can be obtained from "i-organs.atr.jp". The whole-organ RNA-Seq data of humans can be obtained from "The Human Protein Atlas" (https: / / www.proteinatlas.org / ; heart (ERR315328) and kidney (ERR315494)). The weight coefficients can be calculated according to the following method.

[0043]

Number

[0044]

Number

[0045]

Number

Number

[0046] 1-1-3. Correction of the count data set for scRNA-Seq analysis Using the calculated weight coefficients, the scRNA-Seq count data sets obtained from the cells to be analyzed or predicted for the cells to be analyzed are weighted based on the total RNA content for each cell type corresponding to the cells to be analyzed.

[0047] Assuming that the distribution of the weight coefficients follows a Gaussian distribution, the mean and variance of the weighted counts of genes for each cell to be analyzed are calculated according to the following formula (3).

Number

Number

Number

[0048] Based on the mean and variance of the weight coefficients of each analyzed cell, assuming that the mean and variance of the weight coefficients in the corresponding cell type follow a Gaussian distribution, the mean and variance of the weight coefficients in the corresponding cell type are calculated according to the following formula (4).

Equation

[0049] In Equation (4), k, C k , and N k represent the cell type, the group of analyzed cells labeled with cell type k, and the number of analyzed cells in C k , respectively.

[0050] According to the above formula (2), the mean value, variance, and quartiles of the weight coefficients for each cell type present in each organ calculated are shown in the weight coefficient list described later.

[0051] The count dataset of scRNA-Seq analysis weighted by this method is also referred to as the estimated scRAN-Seq count dataset.

[0052] 2. Analysis of the composition ratio of cell types in each organ and analysis of the total RNA expression pattern in the cell types Using the weight coefficients calculated in the above 1-1-2., the composition ratio of the cell types constituting the organ to be analyzed can be analyzed. The analysis of the composition ratio of the cell types constituting the organ to be analyzed includes calculating the composition ratio of the cell types constituting the organ to be analyzed including the cells to be analyzed based on the count dataset of scRNA-Seq analysis weighted in the above 1-1-3. In other words, the composition ratio of the cell types obtained by this method is an estimated composition ratio.

[0053] In addition, by using the weight coefficients calculated in the above 1-1-2., it is possible to analyze the total RNA expression patterns in the cell types that make up the organ to be analyzed. The analysis of the total RNA expression pattern is to obtain the estimated count data of scRNA-Seq analysis. Here, the total RNA is intended to include the RNA expressed from the signature gene set and other genes.

[0054] For example, the analysis of the composition ratio of cell types and the total RNA expression patterns in each cell type can be designed and calculated simultaneously using an algorithm based on Bayes' theorem. The calculation can follow the following formula (5).

Number

Number

Number

Number

Number

Number

Number

[0055] As the initial r and X, use the composition ratio of cell types and the count of the reference dataset weighted by the calculation formula of Equation (4). The hyperparameters α and β are set to 10 -3 、10 -2 、…、10 3 for convenience. The results of combinations of the number of hyperparameters (α and β) of the signature gene set (100 to 2000 genes) that produced a high similarity (showing high Pearson and Spearman correlation coefficients and determining similarity based on a low mean squared error) with the actual whole-organ RNA-Seq can be selected as the best estimation results.

[0056] 3. Correction device for the count dataset of scRNA-Seq analysis, analysis device for scRNA-Seq, and analysis device for the composition ratio of cell types constituting the organ to be analyzed 3-1. Correction device for the count dataset of scRNA-Seq analysis Figure 1 shows the hardware configuration of the correction device 10 for the count dataset of scRNA-Seq analysis.

[0057] The correction device 10 can be a general-purpose computer. The correction device 10 is communicably connected to an input device 111, an output device 112, and a media drive 113. The correction device 10 includes a CPU 101, a memory 102, a ROM (read only memory) 103, a storage device 104, a communication interface (I / F) 105, an input interface (I / F) 106, an output interface (I / F) 107, and a media interface (I / F) 108. Each component within the correction device 10 is connected to each other via a bus 109 so as to be capable of data communication.

[0058] The storage device 104 is composed of a hard disk, a semiconductor memory element such as a flash memory, an optical disk, or the like. The storage device 104 stores an operating system (OS) 1041, a correction program 1042 described later, an algorithm database (DB) DB1, a reference cell type database (DB) DB2, and an organ whole RNA-Seq database (DB) DB3. The correction program 1042 cooperates with the operating system 1041 to cause the computer to function as the correction device 10.

[0059] The CPU 101 is also referred to as the control unit 101 in the present embodiment. The algorithm database DB1 stores mathematical formulas for performing the correction described in 1-1-3 above. In the reference cell type database DB2, labels indicating cell types included in each organ, their composition ratios, and data counts of scRNA-Seq analysis of each cell type are stored in an associated manner. Further, in the reference cell type database DB2, data counts of scRNA-Seq analysis of each corrected cell type are stored in association with a label indicating the organ name and a label indicating the cell type name. In the organ whole RNA-Seq database DB3, count data of RNA-Seq analysis of each whole organ of a mouse or a human is registered for each organ. These data are generated and stored from the known data described in 1-1-2.

[0060] The input device 111 is composed of a touch panel, a keyboard, a mouse, a tablet, a microphone, etc., and performs character input or voice input to the correction device 10. The input device 111 may be connected from outside the control unit 101 or may be integrated with the correction device 10.

[0061] The output device 112 is composed of, for example, a display device such as a display, a printer, etc., and outputs various operation windows, analysis results, etc.

[0062] The media drive 113 can be a USB drive, a floppy disk drive, a CD-ROM drive, or a DVD-ROM drive, etc.

[0063] The communication I / F 105 communicates with an external database or other computers. The output I / F 107 transmits information to the output device 112.

[0064] 3-2. Processing of the correction program Figure 2 shows the flow of the processing of the correction program 1042. The control unit 101 of the correction device 10 first receives a processing start command input by the operator from the input device 111 and starts the processing. In step S1, the control unit 101 selects signature genes that characterize each cell type of the organ to be analyzed according to the method described in 1-1-2. above.

[0065] Next, in step S2, the control unit 101 acquires the scRNA-Seq count data of the signature gene set acquired in step S1 from the reference cell type database DB2.

[0066] Next, in step S3, the control unit 101 acquires the whole-organ RNA-Seq count data from the whole-organ RNA-Seq database DB3. Note that step S3 may be before step S2.

[0067] Next, in step S4, the control unit 101 reads out the expressions (1) to (4) described in the above 1-1-2. from the algorithm database DB1. Using the scRNA-Seq count data of the signature gene set obtained in step S2 and the whole-organ RNA-Seq count data obtained in step S4 for each of the read expressions, the control unit 101 calculates the weight coefficients for each cell type present in each organ based on the expressions described in the above 1-1-2. The control unit 101 stores the calculated weight coefficients in the algorithm database DB1.

[0068] Finally, in step S5, the control unit 101 obtains the count data set of the scRNA-Seq analysis weighted for each cell type according to 1-1-3. and stores it in the reference cell type database DB2.

[0069] Furthermore, the control unit 101 may receive an output process start command input by the operator from the input device 111 and output the weighted scRNA-Seq analysis count data set from the output device 112.

[0070] In this embodiment, an example is shown in which steps S1 to S5 are performed by a single computer. However, for example, step S1, steps S2 to S4, and step S5 may be performed by different computers. That is, the first computer selects signature genes according to step S1, the second computer obtains information on the signature gene sets of each cell type present in each organ from the first computer, performs the processes of steps S2 to S4, and calculates the weight coefficients. Furthermore, a third computer may obtain the weighted scRNA-Seq analysis count data set.

[0071] Furthermore, the first computer may perform steps S1 to S4, and the second computer may perform step S5.

[0072] Also, the first computer may perform step S1, and the second computer may perform steps S2 to S5.

[0073] 3-3. Analytical apparatus for scRNA-Seq and analytical apparatus for the composition ratio of cell types constituting the organ to be analyzed As described in 2. above, the analysis of scRNA-Seq and the analysis of the composition ratio of cell types constituting the organ to be analyzed can be performed simultaneously. Therefore, the analytical apparatus 20 performs both processes.

[0074] Fig. 3 shows the hardware configuration of the analytical apparatus 20. The configuration of the analytical apparatus 20 is basically the same as that of the correction apparatus 10 except for the storage device 204. The storage device 204 stores an analysis program 2042 to be described later instead of the correction program 1042. Further, the storage device 204 stores an algorithm database (DB) DB1, a reference cell type database (DB) DB2, and an organ-wide RNA-Seq database (DB) DB3, similar to the storage device 104.

[0075] 3-4. Analysis program processing Fig. 4 shows the flow of the process of the analysis program 2042. The control unit 201 of the analytical apparatus 20 first receives a process start command input by the operator from the input device 211 and starts the process. In step S11, the control unit 201 reads out the algorithm described in 2. above from the algorithm database DB1.

[0076] Next, in step S13, the control unit 201 acquires organ-wide RNA-Seq count data from the organ-wide RNA-Seq database DB3.

[0077] Subsequently, in step S13, the control unit 201 reads out the weighted scRNA-Seq analysis count data set obtained in 3-2. above from the reference cell type database DB2 and applies it to the algorithm.

[0078] Next, the control unit 201 records, as an estimation result, the composition ratio of cell types constituting the organ to be analyzed estimated by the algorithm and the estimated count data of the scRNA-Seq analysis in the storage device 204.

[0079] For the estimation result, the control unit 201 may output only the composition ratio of cell types constituting the organ to be analyzed from the output device 212, or may output only the estimated count data of the scRNA-Seq analysis from the output device 212. Further, the control unit 201 may output both results from the output device 212.

[0080] 4. Recording Medium with Computer Program The correction program 1042 and the analysis program 2042 may be recorded on a recording medium.

[0081] That is, each program is stored in a recording medium such as a hard disk, a semiconductor memory element such as a flash memory, or an optical disk. Each program may also be stored in a recording medium connectable via a network such as a cloud server. Each program may be provided in a downloadable format or as a program product recorded on a recording medium.

[0082] The storage format of the program on the recording medium is not limited as long as each device can read the program. It is preferable that the storage on the recording medium is non-volatile.

Example

[0083] Examples are shown below to explain the present invention in more detail. However, the present invention is not construed as being limited to the examples.

[0084] I. Method 1. Calculation of Composition Ratio of Reference Cell Types For 14 organs including the aorta, brain, fat, heart, kidney, large intestine, liver, lung, bone marrow, pancreas, skin, skeletal muscle, spleen, and thymus, the composition ratios of reference cell types were calculated based on the scRNA-Seq data described in Non-Patent Document 5 and the databases registered with NIH and the like. For the aorta, fat, kidney, large intestine, liver, lung, bone marrow, pancreas, skin, spleen, and thymus, the ratios of cell types in each organ described in Non-Patent Document 5 were used as the composition ratios of reference cell types.

[0085] For skeletal muscle, the ratios of cell types described in Non-Patent Document 2 were used as the composition ratios of reference cell types.

[0086] For the heart, in accordance with the separation analysis of cardiomyocytes and non-muscle cells, the composition ratios of cell types described in Non-Patent Document 5 were corrected and used as the composition ratios of reference cell types (Referene). Specifically, the ratio of cardiomyocytes (3.1%) adopted in Non-Patent Document 5 is extremely low compared to the ratio (30% to 40%) that has generally been agreed upon in the field of histoanatomy based on various past studies. Therefore, in this example, the ratio of cardiomyocytes was set to 30%, and the remaining 70% was divided by the proportion of non-muscle cell types to obtain the composition ratios of reference cell types.

[0087] Regarding the brain, based on the report in NIH (http: / / www.nervenet.org / papers / BrainRev99.html#Numbers), the composition ratio of the reference cell types was determined. As labels for each cell type in the brain, the corresponding cell type labels of the scRNA-Seq data described in Non-Patent Document 5 were used. First, the cell types in the brain were classified into four classes: "neurons", "glial cells", "endothelial cells", and "others", and the ratios of each class were set to 75:23:7:4, respectively. This ratio follows the estimated ratio in the mouse brain (http: / / www.nervenet.org / papers / BrainRev99.html#Numbers). Next, according to Non-Patent Document 5, "neurons", "glial cells", and "others" were further classified into more detailed cell type classes. Specifically, the "neurons" class was further classified into "nerve cells - excitatory neurons, and some neural stem cells" and "nerve cells - inhibitory neurons". The "others" class was classified into "pericytes - NA" and "oligodendrocyte progenitor cells - NA". Regarding the "glial cells" class, it was classified based on the following three premises. i) "Glial cells" can be classified into four cell types: "microglial cells - NA", "astrocytes - NA", "Bergmann glial cells - NA", and "oligodendrocytes NA" according to Non-Patent Document 5. ii) The ratios of these four glial cell types follow the description in Non-Patent Document 5. iii) Since it is reported that "microglial cells" account for 10 - 15% of the cells in the whole brain, the ratio of "microglial cells - NA" in the whole brain is set to 0.1.Based on these premises, the ratios of each brain cell type were set as "macrophage - NA" (about 0.2%), "microglial cell - NA" (10.0%), "astrocyte - NA" (about 2.2%), "Bergmann glial cell - NA" (about 2.1%), "pericytes - NA" (about 1.5%), "endothelial cell - NA" (about 6.4%), "neuron - excitatory neuron, and some neural stem cells" (about 47.5%), "neuron - inhibitory neuron" (about 21.3%), "oligodendrocyte - NA" (about 8.7%), "oligodendrocyte progenitor cell - NA" (about 1.9%), and these were used as the composition ratios of the reference cell types in the brain. For the human heart and kidney, the composition ratios of the cell types in the mouse heart and the composition ratios of the cell types in the mouse kidney were used as the composition ratios of the reference cell types.

[0088] The composition ratios of the reference cell types for each organ are shown in the list of composition ratios of the reference cell types described below. In addition, for each cell type shown in the list of composition ratios of each reference cell type, scRNA - Seq count data for each cell type are registered in a publicly known database.

[0089] 2. Pretreatment of Data and Normalization of RNA Counts All data processing and analysis were performed using the software "R" version 3.6.1. All cell type labels were the same as those assigned in previously reported scRNA-Seq studies. The gene symbols attached to the scRNA-Seq data were converted and associated with the whole-organ RNA-Seq data by entrez gene IDs derived from "org.Mm.egALIAS2EG" within the "org.Mm.eg.db" R package. Genes with ERCC labels were removed because they are spike-in genes. Furthermore, the RNA counts derived from three genes, Rn45s, Akap5, and Lrrc17, were also removed because they are non-mRNA artifacts that significantly affect the total count. Next, the RNA count derived from each gene was converted and normalized when the total count of each cell in the scRNA-Seq dataset was 100. This normalization process was also performed for each RNA included in the whole-organ RNA-Seq dataset.

[0090] 3. Selection of signature gene sets for cell type identification Using random forest (RF), signature genes for each cell type were computationally selected using the reference cell type composition ratio dataset and scRNA-Seq data described in the previous session. For this selection, the "randomForest" package in R was used for tuning and creating the classifier by RF. The scRNA-Seq data was first split into two parts, with one used as training data for creating the classifier by RF and the other used as test data for calculating the F1 score to verify the accuracy of the classifier. RF analysis was performed on the dataset maintaining the cell type composition ratio described in the previous session. Following the generation of the classifier, the important features of the classifier were extracted as the signature gene names for each cell type, and the "Mean Decrease Gini" value was used as the importance index for each gene.

[0091] 4. Database In this example, all publicly available datasets were used. Whole-organ RNA-Seq data of mice and whole-organ RNA-Seq of myocardial infarction model mice were obtained from "i-organs.atr.jp". Whole-organ RNA-Seq data of humans were obtained from "The Human Protein Atlas" (https: / / www.proteinatlas.org / ; heart (ERR315328) and kidney (ERR315494)). The scRNA-Seq data were obtained from non-patent document 5 (aorta, brain, fat, heart, kidney, large intestine, liver, lung, bone marrow, pancreas, skin, spleen, thymus) and skeletal muscle of the "Mouse Cell Atlas".

[0092] 5. Estimation of gene expression variation analysis The total RNA counts of all genes calculated by prediction were normalized to 1 million copies. The normalized counts of each gene were rounded to an integer and analyzed using the R package "DESeq2 (version 1.24.0)".

[0093] II. Performance verification of the deconvolution method for previously reported whole-organ RNA-Seq data First, for the previously reported methods MuSiC (non-patent document 17) and DWLS method (non-patent document 19), a side-by-side comparison was made between the cell type composition ratios calculated by each method and the cell type composition ratios of the reference obtained from scRNA-Seq data and past reports to verify the performance of each deconvolution method.

[0094] 1. Calculation of cell type composition ratios by previously reported methods MuSiC (non-patent document 17) and the DWLS method (non-patent document 19) were followed according to their respective documents. When performing the quadratic problem solver in the DWLS method, solve.QP (R package: quadprog) was replaced with solve_osqp (R package: osqp).

[0095] 2. Results The comparison results between the estimated composition ratios of cell types in each organ calculated by computer using the MuSiC or DWLS method and the composition ratios of reference cell types in each organ prepared in 2. above are shown in FIGS. 5 and 6. The composition ratios of the estimated cell types in each organ estimated by the MuSiC or DWLS method and the composition ratios of the reference cell types were divergent, and the degrees of divergence were also various. In particular, the divergence was remarkable in skeletal muscle, heart, pancreas, and liver.

[0096] The heart is composed of cardiomyocytes and non-muscle cells. Cardiomyocytes occupy the largest volume of the heart. However, when comparing the number of cells, there are more non-muscle cells than cardiomyocytes. Contrary to this fact, the calculated cell type composition of the heart by the MuSiC or DWLS method is calculated to have cardiomyocytes accounting for 90%. The same tendency was also seen in skeletal muscle.

[0097] Thus, as one of the reasons for the divergence between the composition ratio of the reference cell type and the composition ratio of the estimated cell type, a difference in the total RNA content between different cell types was considered. The total RNA content has been reported to vary from cell to cell in the range of 50,000 transcripts / cell to 300,000 transcripts / cell. In the heart, the volume of cardiomyocytes is said to be 20 to 25 times that of non-muscle cells such as endothelial cells and fibroblasts. Therefore, the total RNA content per cell may vary greatly between muscle cells and non-muscle cells. In fact, this possibility is not considered in the MuSiC and DWLS methods. Such a point is considered to have led to the divergence between the composition ratios of the reference and estimated cell types.

[0098] III. Verification of the Cause of Divergence between the Estimated Whole-Organ RNA-Seq Dataset and the Actual Whole-Organ RNA-Seq Dataset The deviation between the composition ratio of the reference cell types and the estimated composition ratio of the cell types was hypothesized to be due to differences in the total RNA content among the cell types contained in the tissue collected from the organ when extracting the total RNA sample of the organ. This hypothesis was verified by comparing the actual gene expression profile with the estimated gene expression profile. The estimated whole-organ RNA-Seq data is the result of multiplying the composition ratio of the reference cell types obtained in I.1 by the count data obtained in I.2.

[0099] The estimated whole-organ RNA-Seq data is shown in Fig. 7. The estimated whole-organ RNA-Seq data was calculated as the sum of the transcript counts for each gene normalized by weighting based on the composition ratio of known reference cell types for tissues composed of multiple cell types.

[0100] The results shown in Fig. 7 are the indicated number (number of genes) of signature genes corresponding to the top rank numbers in each cell type of each organ used to identify the cell types of each organ calculated by RF. The top rank numbers were set to 100 genes, 300 genes, and 2000 genes among the signature genes. However, in the aorta and kidney, since the total number of signature genes is less than 2000, the aorta was compared with 1577 genes instead of 2000 genes, and the kidney was compared with 1461 genes. The similarity / dissimilarity between the actual gene expression profiles and the estimated gene expression profiles of 14 organs is indicated by the Pearson correlation coefficient.

[0101] As shown in Fig. 7, the Pearson correlation coefficient was less than 0.75 in 10 organs (aorta, brain, heart, large intestine, liver, lung, pancreas, skin, skeletal muscle, thymus). This indicates that it is insufficient to simply multiply the composition ratio of the reference cell types obtained in I.1 by the count data obtained in I.2 to reconstruct the whole-organ RNA-Seq data in these organs.

[0102] IV. Setting and verification of cell type-specific weight coefficients to eliminate the deviation between datasets We calculated the weight coefficients for correcting the RNA content in different cell types existing in each tissue and verified their accuracy.

[0103] 1. Calculation of cell type-specific coefficients The weight coefficients for each cell type existing in each organ were calculated according to the following method.

Equation

[0104] Next, the following formula (1) will be explained.

Equation

Equation

Number

Number

Number

Number

[0105] Based on the calculated mean and variance of the weight coefficients of each cell to be analyzed, assuming that the mean and variance of the weight coefficients in the corresponding cell type follow a Gaussian distribution, the mean and variance of the weight coefficients in the corresponding cell type were calculated according to the following Equation (4).

Number

[0106] 2. Results Using the above calculation formula (2), the weight coefficients of each cell type and their ranges were created (Figure 8). The weight coefficients of muscle cells were actually higher than those of non-muscle cell types in both the heart and skeletal muscle (Figure 8). Using these cell type-specific weight coefficients, the transcript counts of each cell type were weighted. Next, according to the composition ratio of the reference cell type of each cell type contained in each organ, the composition ratio of the reference cell type of each cell type in each organ was further applied to the transcript count weighted by the weight coefficient to generate an RNA-Seq dataset. This calculation method is called estimated whole-organ RNA-Seq (v-RNA-Seq), and the RNA-Seq dataset obtained by estimated whole-organ RNA-Seq is called the estimated whole-organ RNA-Seq dataset.

[0107] Next, the estimated whole-organ RNA-Seq dataset was compared with the corresponding actual whole-organ RNA-Seq dataset. The results are shown in Figure 9. Compared with Figure 7, the divergence of the gene expression profiles shown by each dataset was improved in most organs (Pearson correlation coefficient = 0.8 - 1.0).

[0108] V. Calculation of the composition ratio of cell types in each organ and estimation of the total RNA expression pattern of all cell types contained in each organ Using the specific weight coefficients based on the RNA content for each cell type calculated in the above IV., an algorithm based on Bayes' theorem was designed to simultaneously calculate both the ratio of each cell type contained in each organ and the gene expression pattern in each of those cell types.

[0109] 1. Calculation of the composition ratio of cell types and gene expression pattern The calculation of the composition ratio of cell types and the gene expression pattern followed the following formula (5). The mean and variance of the transcript counts weighted by the weight coefficients in each cell type were calculated according to the above formula (4).

Number

Number

Number

Number

Number

Number

[0110] As the initial r and X, the composition ratio of the cell types and the counts of the reference dataset weighted by the calculation formula of formula (4) were used. The hyperparameters α and β are 10 -3 , 10 -2 , …, 10 3It was set to. The signature gene set (100, 300, 2,000 / 1,577 / 1,461) hyperparameter (α and β) combinations that produced a high similarity (showing high Pearson and Spearman correlation coefficients and determining similarity based on low mean squared error) with the actual whole-organ RNA-Seq were selected as the best estimation results. The outline of this calculation is shown in Fig. 10. The comparison results between the composition ratios of cell types estimated by the method of the present invention and the reference cell type ratios are shown in Figs. 11 and 12. Also, the comparison results between the actual scRNA-Seq and the scRNA-Seq count data estimated by the method of the present invention and the analysis results of t-Distributed Stochastic Neighbor Embedding (t-SNE) are shown in Figs. 13 and 14.

[0111] 2. Verification of cell type identification Whether the estimated scRNA-Seq count data calculated in the above V.1. can identify the cell types present in each organ was verified using t-Distributed Stochastic Neighbor Embedding (t-SNE). In the cells belonging to each cell type present in each organ, the total sampling size was set to 3,000, and the number of sampled cells for each cell type and the estimated scRNA-Seq count data of cell type k were

Number

[0112] 3. Results In the present invention, two hyperparameters α and β were defined to consider the influence of combinations of cell type ratios. Gene expression patterns at different organ levels, such as those of normal and pathological organs, may be different. However, this difference may be due to i) although the gene expression patterns in each cell type appear to be the same, the ratios of each cell type are different, and ii) although the ratios of cell types are the same, differences in gene expression patterns occur among the same cell types. There may also be a combination of i) and ii). Therefore, a wide range of comprehensive combinations of α and β were evaluated, and the optimal combination of cell type composition and weighted transcriptome counts of each cell type was calculated to explain the behavior of the transcriptome at the organ level.

[0113] Using this method, the compositional ratios of cell types in 10 organs (aorta, adipose, heart, kidney, liver, lung, large intestine, bone marrow, skeletal muscle, spleen) were calculated. The results are shown in Fig. 11. Brain, pancreas, skin, and thymus were excluded from the study among the 14 organs used in Figs. 5 and 6 for the following reasons. 1) The actual ratios of cell types were not available. 2) The pancreas is actually derived from pancreatic islets. Although the actual ratios of cell types in pancreatic islets are available, they do not represent the ratios of the entire actual pancreas. 3) In the case of skin or thymus, the Pearson correlation coefficient did not exceed 0.8 even when using cell types and specific weight coefficients.

[0114] The calculated compositional ratios of cell types in the above 10 organs were similar to the reference cell type compositional ratios experimentally determined by actual scRNA-Seq studies (Fig. 11). In particular, the abnormally large ratios of cardiomyocytes and skeletal muscle cells estimated by the MuSiC and DWLS methods were improved by V-scRNA-Seq, respectively. The results are shown in Fig. 11. Also, as shown in Fig. 12, the mean squared error (MSE) with respect to the reference cell type compositional ratios was superior for V-scRNA-Seq in 5 actual organs (adipose, heart, large intestine, liver, skeletal muscle) compared to other methods.

[0115] Also, for 23,131 genes expressed in any of the examined organs excluding skeletal muscle and 14,323 (skeletal muscle) genes in skeletal muscle included in the estimated whole-organ RNA-Seq dataset, the estimated transcript counts corrected with cell type-specific weight coefficients and the composition ratios of reference cell types were calculated according to the method of the present invention, and the corrected estimated transcript counts were compared with gene expression in each cell type in 10 actual organs.

[0116] The Pearson correlation coefficient showed that the estimated transcript counts were comparable to the actual counts for all cell types and organs (Figure 13). Also, the similarity and relevance of annotations of the same or related cell types among different organs were shown (Figure 13).

[0117] t-SNE analysis using all V-scRNASeq data of 10 organs showed that each cell type in each organ could be classified according to its gene expression profile (Figure 14).

[0118] VI. Calculation of changes in cell type ratio and gene expression in diseases Next, we evaluated whether our method could detect changes in cell type ratio and gene expression of each cell type associated with the disease process. Cardiovascular disease is the leading cause of death worldwide (https: / / www.who.int / news-room / fact-sheets / detail / cardiovascular-diseases-(cvds)). There are reports showing that the cell type composition among heart diseases changes over time. Furthermore, as described above, the heart is an organ in which the method according to the present invention can effectively calculate both the composition ratio of cell types and its gene expression pattern, rather than the previously published deconvolution method. Therefore, the method of the present invention was applied to a mouse model of myocardial infarction (MI), and it was examined whether the method according to the present invention could detect both the composition ratio of cell types in the heart and the time-dependent changes in cell type-dependent gene expression known during MI. For the disease model calculation, first, using the composition ratios of the same reference cell types in normal mice at each stage (E, M, L), the weight coefficients were calculated using the whole-organ RNA-Seq data of the sham heart. Next, using the whole-heart RNA-Seq data of the sham / MI model, the composition ratios of cell types and gene expression profiles at each stage were calculated as described above.

[0119] The results are shown in Fig. 15. Methods for creating animal models of myocardial infarction are well-known. The three stages of myocardial infarction are as follows: 1) 1 day after coronary artery ligation (E-MI, early myocardial infarction stage), 2) 7 days after coronary artery ligation (M-MI, early fibrosis stage), 3) 8 weeks after coronary artery ligation (L-MI, cardiac remodeling stage). In this analysis, using the RNA-Seq data of the sham controls (E-sham, M-sham, L-sham) and the composition ratios of reference cell types in normal mouse hearts, the weight coefficients for each cell type were calculated.

[0120] By calculation, two major changes in cell type composition related to the predicted MI were detected, specifically a decrease in cardiomyocytes and an increase in fibroblasts (Fig. 15a). By this method, an increase in myofibroblasts characteristic of the M-MI stage was detected (Fig. 15a). This was also consistent with the reported experimental results.

[0121] According to the present invention, gene expression changes in each cell type during myocardial infarction calculated also detected multiple characteristics predicted from previous experimental studies (Fig. 15b).

[0122] In cardiomyocytes, statistically significant upregulation of the Nppb gene, Sparc gene, and Col4a1 gene (log 2 fold change > 0.7), and downregulation of the Myh6 gene (log 2 fold change < 0.7) were detected (Fig. 15b). In fibroblasts, statistically significant upregulation of the Col4a1 gene, Col1a1 gene, and Sparc gene (log 2A two-fold change > 0.7 was detected (Figure 15b). Here, statistically significant means an adjusted p-value < 0.001. In addition to these landmark genes in known MI pathologies, this method found many other genes whose expression varied depending on the pathology in each cell type (Figure 15b).

[0123] VII. Validation of inferred human scRNA-Seq We verified that the mouse weight coefficients and V-scRNA-Seq could be applied to the deconvolution of the human whole-organ RNA-Seq dataset. Using publicly available human whole-organ RNA-Seq data for the heart and kidney, we calculated the composition ratios and transcriptome profiles of their cell types.

[0124] For each gene stored in the human whole-organ RNA-Seq dataset, the total count of the RNA expressed from it was first normalized to 100. Next, by extracting genes with common names between mouse and human, the mouse gene symbols were made to match those of human. Using these common gene sets between mouse and human, the calculation method described for the mouse dataset was applied. The human whole-organ RNA-Seq data for the heart and kidney were obtained from "The Human Protein Atlas" (https: / / www.proteinatlas.org / ).

[0125] The results are shown in Figure 8. It was shown that the composition ratios of cell types calculated for the human heart and kidney were similar to those of the corresponding normal mouse organs (Figure 16a). Furthermore, the t-SNE analysis results of the inferred scRNA-Seq data for the human heart and kidney showed that classification based on the gene expression profiles of known cell types in each organ was possible (Figure 16b). These results indicate the applicability of cell type-specific weight coefficients and the V-scRNA-Seq framework across different species.

[0126] <List of composition ratios of reference cell types> The following list is arranged in the order of Organ:Cell type:Abbreviation:Reference. The semicolon ";" is intended to separate the data of each cell type. The cell composition ratio is normalized so that the whole organ is "1". Since representative cell types are shown here, the sum of the composition ratios of each cell type in each organ does not necessarily equal 1. Aorta:Aorta-endothelial cell-NA:EC :0.40 ; Aorta:Aorta-erythrocyte-NA:ERC:0.21 ; Aorta:Aorta-fibroblast-NA:FC :0.22 ; Aorta:Aorta-professional antigen presenting cell-NA:PAP:0.16 ; Brain:Brain_Myeloid-macrophage-NA:MAC:0.00 ; Brain:Brain_Myeloid-microglial cell-NA:MI :0.10 ; Brain:Brain_Non-Myeloid-astrocyte-NA:AS :0.02 ; Brain:Brain_Non-Myeloid-Bergmann glial cell-NA:BGC:0.00 ; Brain:Brain_Non-Myeloid-brain pericyte-NA:BP :0.02 ; Brain:Brain_Non-Myeloid-endothelial cell-NA:EC :0.06 ; Brain:Brain_Non-Myeloid-neuron-excitatory neurons and some neuronal stem cells:NEUR2 :0.47 ; Brain:Brain_Non-Myeloid-neuron-inhibitory neurons:NEUR1 :0.21 ; Brain:Brain_Non-Myeloid-oligodendrocyte-NA:OLC:0.09 ; Brain:Brain_Non-Myeloid-oligodendrocyte precursor cell-NA:OPC:0.02 ; Fat:Fat-B cell-NA:B:0.10 ; Fat:Fat-endothelial cell-NA:EC :0.16 ; Fat:Fat-mesenchymal stem cell of adipose-mesenchymal progenitor:MSA:0.43 ; Fat:Fat-myeloid cell-NA:MYE:0.20 ; Fat:Fat-NA-NA:NA :0.01 ; Fat:Fat-natural killer cell-NA:NK :0.01 ; Fat:Fat-T cell-NA:T:0.08 ; Heart:Heart-cardiac muscle cell-NA:CM :0.30 ; Heart:Heart-endocardial cell-NA:ECC:0.02 ; Heart:Heart-endothelial cell-NA:EC :0.20 ; Heart:Heart-fibroblast-NA:FC :0.29 ; Heart:Heart-leukocyte-NA:LEU:0.13 ; Heart:Heart-myofibroblast cell-NA:MYF:0.05 ; Heart:Heart-NA-conduction cells:CC :0.01 ; Heart:Heart-smooth muscle cell-NA:SM :0.01 ; Kidney:Kidney-endothelial cell-NA:EC :0.19 ; Kidney: Kidney - Epithelial cell of proximal tubule - NA: PT: 0.48; Kidney: Kidney - Kidney collecting duct epithelial cell - NA: CD: 0.22; Kidney: Kidney - Leukocyte - NA: LEU: 0.02; Kidney: Kidney - Macrophage - NA: MAC: 0.09; Large Intestine: Large Intestine - Brush cell of epithelium proper of large intestine - Tuft cell - TUF: 0.01; Large Intestine: Large Intestine - Enterocyte of epithelium of large intestine - Enterocyte (Distal) - EN - D: 0.06; Large Intestine: Large Intestine - Enterocyte of epithelium of large intestine - Enterocyte (Proximal) - EN - P: 0.21; Large Intestine: Large Intestine - Enteroendocrine cell - Chromaffin Cell - CHR: 0.01; Large Intestine: Large Intestine - Epithelial cell of large intestine - Lgr5 - amplifying undifferentiated cell - EP1: 0.16; Large Intestine: Large Intestine - Epithelial cell of large intestine - Lgr5 - undifferentiated cell - EP2: 0.10; Large Intestine: Large Intestine - Epithelial cell of large intestine - Lgr5+ amplifying undifferentiated cell (Distal): EP3 - D: 0.03; Large Intestine: Large Intestine - Epithelial cell of large intestine - Lgr5+ amplifying undifferentiated cell (Proximal): EP3 - P: 0.05; Large Intestine: Large Intestine - Epithelial cell of large intestine - Lgr5+ undifferentiated cell (Distal): EP4 - D: 0.08; Large Intestine: Large Intestine - Epithelial cell of large intestine - Lgr5+ undifferentiated cell (Proximal): EP4 - P: 0.12; Large Intestine: Large Intestine - Large intestine goblet cell - Goblet cell (Distal): GB1 - D: 0.09; Large Intestine: Large Intestine - Large intestine goblet cell - Goblet cell (Proximal): GB1 - P: 0.05; Large Intestine: Large Intestine - Large intestine goblet cell - Goblet cell, top of crypt (Distal): GB2 - D: 0.02; Liver: Liver - B cell - NA: B: 0.07; Liver: Liver - Endothelial cell of hepatic sinusoid - NA: EC: 0.33; Liver:Liver-hepatocyte-NA:HE :0.42 ; Liver:Liver-Kupffer cell-NA:KUP:0.11 ; Liver:Liver-natural killer cell-NK / NKT cells:NK2:0.07 ; Lung:Lung-B cell-NA:B:0.02 ; Lung:Lung-ciliated columnar cell of tracheobronchial tree-multiciliated cells:CCC:0.01 ; Lung:Lung-classical monocyte-invading monocytes:CMN:0.07 ; Lung:Lung-epithelial cell of lung-alveolar epithelial type 1 cells, alveolar epithelial type 2 cells, club cells, and basal cells:EP5:0.06 ; Lung:Lung-leukocyte-mast cells and unknown immune cells:LEU2 :0.02 ; Lung:Lung-lung endothelial cell-NA:EC :0.34 ; Lung:Lung-monocyte-circulating monocytes:MN2:0.07 ; Lung:Lung-myeloid cell-dendritic cells, alveolar macrophages, and interstital macrophages:MYE2 :0.01 ; Lung:Lung-NA-lung neuroendocrine cells and unknown cells:NC :0.03 ; Lung:Lung-natural killer cell-NA:NK :0.02 ; Lung: Lung-stromal cell-NA: SC: 0.33; Lung: Lung-T cell-NA: T: 0.03; Marrow: Marrow-B cell-Cd3e+ Klrb1+ B cell: B2: 0.01; Marrow: Marrow-basophil-NA: BAS: 0.00; Marrow: Marrow-common lymphoid progenitor-NA: CLP: 0.04; Marrow: Marrow-granulocyte-NA: GRA: 0.16; Marrow: Marrow-granulocyte monocyte progenitor cell-NA: GMP: 0.02; Marrow: Marrow-granulocytopoietic cell-NA: GC: 0.05; Marrow: Marrow-hematopoietic precursor cell-NA: HPC: 0.08; Marrow: Marrow-immature B cell-NA: IB: 0.06; Marrow: Marrow-immature natural killer cell-NA: INK: 0.01; Marrow: Marrow-immature NK T cell-NA: INKT: 0.01; Marrow: Marrow-immature T cell-NA: IT: 0.02; Marrow: Marrow-late pro-B cell-Dntt- late pro-B cell: LPB1: 0.04; Marrow: Marrow-late pro-B cell-Dntt+ late pro-B cell: LPB2: 0.03; Marrow: Marrow-macrophage-NA: MAC: 0.03; Marrow: Marrow-mature natural killer cell-NA: MNT: 0.01; Marrow: Marrow-megakaryocyte-erythroid progenitor cell-NA: EPC: 0.01; Marrow: Marrow-monocyte-NA: MN: 0.04; Marrow: Marrow-naive B cell-NA: NBC: 0.12; Marrow: Marrow-pre-natural killer cell-NA: PNK: 0.00; Marrow: Marrow-precursor B cell-pre-B cell (Philadelphia nomenclature): PB: 0.11; Marrow: Marrow-regulatory T cell-NA: RT: 0.00; Marrow: Marrow-Slamf1-negative multipotent progenitor cell-NA: MPC1: 0.10; Marrow: Marrow-Slamf1-positive multipotent progenitor cell-NA: MPC2: 0.04; SkMuscle: B cell_Jchain high(Muscle): B3: 0.02; SkMuscle: B cell_Vpreb3 high(Muscle): B4: 0.09; SkMuscle: Dendritic cell(Muscle): DEN: 0.01; SkMuscle: Endothelial cell(Muscle): EC: 0.02; SkMuscle: Erythroblast_Car1 high(Muscle): ERB1: 0.03; SkMuscle: Erythroblast_Car2 high(Muscle): ERB2: 0.16; SkMuscle: Granulocyte monocyte progenitor cell(Muscle): GMP: 0.08; SkMuscle: Macrophage_Ms4a6c high(Muscle): MAC2: 0.13; SkMuscle: Macrophage_Retnla high(Muscle): MAC3: 0.02; SkMuscle: Muscle cell_Tnnc1 high(Muscle): MC1: 0.01; SkMuscle: Muscle cell_Tnnc2 high(Muscle): MC2: 0.03; SkMuscle: Muscle progenitor cell(Muscle): MPC: 0.08; SkMuscle: Neutrophil_Camp high(Muscle): NEUT1: 0.16; SkMuscle: Neutrophil_Prg2 high(Muscle): NEUT2: 0.01; SkMuscle: Neutrophil_Retnlg high(Muscle): NEUT3: 0.12; SkMuscle: Stromal cell(Muscle): SC: 0.02; SkMuscle: T cell(Muscle): T: 0.01; Pancreas: Pancreas - endothelial cell - NA: EC: 0.06; Pancreas: Pancreas - leukocyte - NA: LEU: 0.04; Pancreas: Pancreas - pancreatic A cell - pancreatic A cell: A: 0.24; Pancreas: Pancreas - pancreatic acinar cell - acinar cell: ACI: 0.10; Pancreas:Pancreas - pancreatic D cell - pancreatic D cell:D:0.11; Pancreas:Pancreas - pancreatic ductal cell - ductal cell:DUC:0.12; Pancreas:Pancreas - pancreatic PP cell - pancreatic PP cell:PP:0.05; Pancreas:Pancreas - pancreatic stellate cell - stellate cell:PSC:0.04; Pancreas:Pancreas - type B pancreatic cell - beta cell:BC:0.22; Skin:Skin - basal cell of epidermis - Basal IFE:BE:0.22; Skin:Skin - epidermal cell - Intermediate IFE:EPI:0.12; Skin:Skin - keratinocyte stem cell - Inner Bulge:KSC:0.26; Skin:Skin - keratinocyte stem cell - Outer Bulge:KSC2:0.37; Skin:Skin - leukocyte - NA:LEU:0.01; Skin:Skin - stem cell of epidermis - Replicating Basal IFE:SCE:0.02; Spleen:Spleen - B cell - NA:B:0.77; Spleen:Spleen - macrophage - NA:MAC:0.03; Spleen:Spleen - T cell - NA:T:0.20; Thymus:Thymus - DN1 thymic pro - T cell - DN1 thymocytes:TPT:0.01; Thymus: Thymus - immature T cell - DN4 - DP in transition Cd69 negative rapidly dividing thymocytes: IT3: 0.15; Thymus: Thymus - immature T cell - DN4 - DP in transition Cd69 negative thymocytes: IT2: 0.44; Thymus: Thymus - immature T cell - DN4 - DP in transition Cd69 positive thymocytes: IT4: 0.37; Thymus: Thymus - leukocyte - antigen presenting cell: LEU3: 0.02

[0127] <List of the composition ratios of reference cell types> The following list is arranged in the order of Organ: Singnature.gene.set.number: Cell.type: mean: var: min: first_quantile: Median: third_quantile: max. The ";" is intended to separate the data of each cell type. Aorta: 100: EC : 0.151788089: 0.320824524: 0.01: 0.01: 0.01: 0.020643025: 4.333799883; Aorta: 100: ERC: 24.67386955: 27569.49096: 0.01: 0.268057658: 0.947630617: 4.647035248: 1361.854647; Aorta: 100: FC : 1.120387302: 7.507603394: 0.01: 0.014564004: 0.061908957: 0.516121627: 12.41886869; Aorta: 100: PAP: 0.124841086: 0.130858196: 0.01: 0.01: 0.01211706: 0.032624313: 1.661400716; Aorta:300:EC :0.335653916:0.888982181:0.01:0.01:0.01:0.10214728:5.784560681; Aorta:300:ERC:20.7856725:15487.87723:0.01:0.132647672:0.943529224:3.121387266:1008.048707; Aorta:300:FC :1.122992247:3.637888239:0.01:0.103255397:0.318573613:1.121704441:9.938671176; Aorta:300:PAP:0.16699144:0.283733571:0.01:0.01:0.014927526:0.085629113:3.587288911; Aorta:1577:EC :0.328052831:2.980101375:0.01:0.01:0.01:0.01:12.65695793; Aorta:1577:ERC:0.861381942:10.22004111:0.01:0.01:0.01:0.189752138:24.51986069; Aorta:1577:FC :1.157779979:11.62919843:0.01:0.01:0.011244537:0.227616723:17.14421124; Aorta:1577:PAP:0.236296791:0.806728782:0.01:0.01:0.01:0.01:4.644598908; Brain:100:AS :0.434856565:2.615451954:0.01:0.022455987:0.046046996:0.12257169:18.38001032; Brain:100:BGC:1.072299836:4.247525505:0.052380374:0.193697173:0.441824763:0.815423538:9.813487019; Brain:100:BP :1.286448516:7.416001166:0.01:0.149224444:0.442108329:1.349882731:20.55492602; Brain:100:EC :0.155829289:2.482545762:0.01:0.011392245:0.018610335:0.058316551:33.77084044; Brain:100:MAC:0.048124031:0.004108694:0.01:0.013496606:0.023131426:0.063089991:0.377400612; Brain:100:MI :0.012961869:2.3118E-05:0.01:0.010033301:0.011173433:0.013679577:0.064623845; Brain:100:NEUR1 :2.653825355:114.4922959:0.01:0.014365065:0.065192324:0.488780763:74.1873715; Brain:100:NEUR2 :1.516349011:26.86180756:0.01:0.017777754:0.088547977:0.515962167:38.7956663; Brain:100:OLC:1.47033239:789.3444746:0.01:0.043020985:0.137535878:0.575570849:1014.927739; Brain:100:OPC:1.384613588:14.07919044:0.01:0.012646413:0.040728358:0.702904737:25.50241672; Brain:300:AS :1.74038249:13.24520375:0.01:0.015261047:0.195577225:1.884738702:41.41426285; Brain:300:BGC:1.341778577:3.944690072:0.045824361:0.224444365:0.549709956:1.458466949:8.815660456; Brain:300:BP :1.749492174:12.04296246:0.01:0.014499893:0.199956059:1.774169825:17.37539007; Brain:300:EC :0.205803506:0.364383999:0.01:0.01:0.010304075:0.054289482:7.530690763; Brain:300:MAC:0.010347881:4.91595E-06:0.01:0.01:0.01:0.01:0.024372891; Brain:300:MI :0.010016051:1.25293E-07:0.01:0.01:0.01:0.01:0.02439228; Brain:300:NEUR1 :0.091026428:0.378856142:0.01:0.01:0.01:0.01:5.354404503; Brain:300:NEUR2 :1.881598235:117.9200715:0.01:0.01:0.01:0.010376092:122.6759504; Brain:300:OLC:0.683957014:4.665791152:0.01:0.01:0.031833092:0.255883967:30.43686688; Brain:300:OPC:0.371699481:4.009370499:0.01:0.01:0.01:0.022522107:22.58268291; Brain:2000:AS :1.591406611:8.974491561:0.01:0.01:0.070425016:1.748463704:15.85464919; Brain:2000:BGC:1.125268038:5.864155062:0.01:0.01:0.013973665:0.729619368:9.858942038; Brain:2000:BP :0.010231754:6.39702E-06:0.01:0.01:0.01:0.01:0.03860368; Brain:2000:EC :0.046223928:0.054173766:0.01:0.01:0.01:0.01:2.487765951; Brain:2000:MAC:0.010208662:4.75753E-07:0.01:0.01:0.01:0.01:0.013557299; Brain:2000:MI :0.01008294:2.25461E-06:0.01:0.01:0.01:0.01:0.062541367; Brain:2000:NEUR1 :0.044018483:0.043085508:0.01:0.01:0.01:0.01:1.347998842; Brain:2000:NEUR2 :1.008760807:29.81632686:0.01:0.01:0.01:0.01:51.96076429; Brain:2000:OLC:0.55101755:2.370980823:0.01:0.01:0.01:0.153605575:17.776641; Brain:2000:OPC:0.014554191:0.001657313:0.01:0.01:0.01:0.01:0.498116609; Fat:100:B:0.046754003:0.011173744:0.01:0.010216413:0.015136787:0.037396414:1.302102477; Fat:100:EC :2.414604433:19.41475516:0.01:0.010548113:0.280908337:2.789098145:41.78589451; Fat:100:MSA:0.320323969:1.112225282:0.01:0.01:0.01:0.04529986:10.49111527; Fat:100:MYE:0.071228559:0.045200484:0.01:0.01:0.014329697:0.043635857:3.146909307; Fat:100:NA :2.710777335:28.70862922:0.01:0.01:0.285071051:3.167316468:27.15404948; Fat:100:NK :0.318498521:0.146633229:0.013489934:0.060943058:0.201014136:0.368195143:1.823855366; Fat:100:T:0.45552476:2.077988331:0.01:0.01:0.029935117:0.292900664:16.31685769; Fat:300:B:0.085593432:0.063367131:0.01:0.01:0.01:0.010472349:1.660266129; Fat:300:EC :1.951699824:32.94181255:0.01:0.01:0.01:0.449707234:46.60358358; Fat:300:MSA:0.301040647:2.273597383:0.01:0.01:0.01:0.01:20.45002588; Fat:300:MYE:0.145061513:0.432465824:0.01:0.01:0.01:0.011508338:8.06361426; Fat:300:NA :0.097562096:0.200890568:0.01:0.01:0.01:0.01:2.807266835; Fat:300:NK :0.01:8.01339E-29:0.01:0.01:0.01:0.01:0.01; Fat:300:T:0.010988287:0.000224172:0.01:0.01:0.01:0.01:0.237562315; Fat:2000:B:0.267000961:1.17913974:0.01:0.01:0.01:0.012349464:10.30491319; Fat:2000:EC :1.52498714:15.89721155:0.01:0.01:0.01:0.367139836:29.44172018; Fat:2000:MSA:0.352538559:3.267249585:0.01:0.01:0.01:0.01:26.45242703; Fat:2000:MYE:0.154984661:0.568213675:0.01:0.01:0.01:0.010339534:12.65610984; Fat:2000:NA :0.038760219:0.01129055:0.01:0.01:0.01:0.01:0.520072065; Fat:2000:NK :0.011812948:8.05709E-05:0.01:0.01:0.01:0.01:0.062738677; Fat:2000:T:0.050755542:0.090293712:0.01:0.01:0.01:0.01:4.357269635; Heart:100:CC :0.01:4.19568E-32:0.01:0.01:0.01:0.01:0.01; Heart:100:CM :2.823889258:54.45296231:0.01:0.01:0.077634074:0.644519204:37.63426786; Heart:100:EC :0.383886505:8.95708179:0.01:0.01:0.0145107:0.127136156:60.17501651; Heart:100:ECC:0.142329405:0.121205364:0.01:0.012507964:0.028868296:0.164640401:2.344871042; Heart:100:FC :0.089502111:0.035613316:0.01:0.01:0.013279019:0.057356411:1.287831061; Heart:100:LEU:0.049459847:0.008453581:0.01:0.011171809:0.016984044:0.053998062:1.156060563; Heart:100:MYF:0.197874897:0.106739298:0.01:0.01282829:0.041756576:0.227034214:1.680778788; Heart:100:SM :1.348022055:1.282516297:0.297989562:0.449770835:0.871973921:1.854237253:4.078581375; Heart:300:CC :0.01:2.41086E-32:0.01:0.01:0.01:0.01:0.01; Heart:300:CM :1.706492592:22.04625552:0.01:0.01:0.014028892:0.385791139:22.6212958; Heart:300:EC :0.308624254:0.996382718:0.01:0.01:0.01441405:0.228046496:14.13284018; Heart:300:ECC:0.075306832:0.026880368:0.01:0.01:0.018232343:0.066596459:1.01598398; Heart:300:FC :0.080447546:0.034630326:0.01:0.01:0.012546049:0.04781712:1.673015079; Heart:300:LEU:0.024387404:0.006626406:0.01:0.01:0.011895671:0.021707587:1.457680288; Heart:300:MYF:0.292706873:0.433907162:0.01:0.011944381:0.075827223:0.295975505:6.11236449; Heart:300:SM :0.123010953:0.023413091:0.01:0.01:0.056597312:0.185566209:0.509664676; Heart:2000:CC :0.010313811:7.39796E-07:0.01:0.01:0.01:0.01:0.01312443; Heart:2000:CM :3.112064232:77.80962808:0.01:0.01:0.01:0.390638194:51.1228028; Heart:2000:EC :0.13368653:0.143577296:0.01:0.01:0.01:0.040965266:2.829282942; Heart:2000:ECC:0.310990008:4.392293003:0.01:0.01:0.01:0.01:14.68197108; Heart:2000:FC :0.053694616:0.018926396:0.01:0.01:0.01:0.02231088:1.648764347; Heart:2000:LEU:0.039521164:0.287158212:0.01:0.01:0.01:0.01:9.759389068; Heart:2000:MYF:0.017612063:0.002417966:0.01:0.01:0.01:0.01:0.422266391; Heart:2000:SM :0.077480204:0.044825118:0.01:0.01:0.01:0.01:0.817152076; Kidney:100:CD :0.861873721:5.325131329:0.01:0.01:0.027677075:0.429648536:15.29691009; Kidney:100:EC :0.087685384:0.053268255:0.01:0.01:0.010467171:0.079612624:1.530825402; Kidney:100:LEU:0.019678869:0.000210734:0.01:0.01:0.01:0.029438172:0.048074249; Kidney:100:MAC:0.014345637:0.00027765:0.01:0.01:0.01:0.01:0.096159876; Kidney:100:PT :1.066404822:6.559136387:0.01:0.01:0.042416515:0.656508148:14.31773311; Kidney:300:CD :1.02990839:31.96503195:0.01:0.01:0.01:0.01:46.00176013; Kidney:300:EC :0.148567471:0.401009826:0.01:0.01:0.01:0.01:3.505286789; Kidney:300:LEU:0.01:3.0112E-29:0.01:0.01:0.01:0.01:0.01; Kidney:300:MAC:0.010736962:1.52072E-05:0.01:0.01:0.01:0.01:0.030634931; Kidney:300:PT :1.078855015:17.30970861:0.01:0.01:0.01:0.01:34.86889454; Kidney:1461:CD :0.571606592:3.043433891:0.01:0.01:0.01:0.082063639:10.68411549; Kidney:1461:EC :0.262442826:0.506765362:0.01:0.01:0.01:0.040553741:4.187360011; Kidney:1461:LEU:0.448563929:1.140902526:0.01:0.01:0.025451239:0.131769898:3.072526545; Kidney:1461:MAC:0.015812234:0.000494913:0.01:0.01:0.01:0.01:0.126188915; Kidney:1461:PT :1.009683719:9.1880931:0.01:0.01:0.010200921:0.349203618:18.39498163; Large Intestine:100:CHR:0.012506713:9.45398E-05:0.01:0.01:0.01:0.01:0.059484444; Large Intestine:100:EN-D :0.02890134:0.011490302:0.01:0.01:0.01:0.01:0.914454533; Large Intestine:100:EN-P :0.245917773:0.409541469:0.01:0.01:0.01:0.099277919:5.442989504; Large Intestine:100:EP1:0.010344637:3.46044E-05:0.01:0.01:0.01:0.01:0.113368671; Large Intestine:100:EP2:0.108748193:1.93259855:0.01:0.01:0.01:0.01:19.76858819; Large Intestine:100:EP3-D:0.589057488:2.322023876:0.01:0.01:0.036088665:0.212778459:7.753674897; Large Intestine:100:EP3-P:1.359012809:13.83295307:0.01:0.01:0.038106769:0.79029694:30.18008581; Large Intestine:100:EP4-D:1.334742196:8.776057103:0.01:0.01:0.051754177:0.765595964:15.7202124; Large Intestine:100:EP4-P:1.936669446:21.3101594:0.01:0.01:0.048946915:0.996926491:29.43809417; Large Intestine:100:GB1-D:0.567448195:1.714325103:0.01:0.01:0.040470534:0.529872099:10.85628586; Large Intestine:100:GB1-P:0.206993092:0.398629305:0.01:0.01:0.01:0.035904714:4.782087384; Large Intestine:100:GB2-D:0.020528901:0.001429011:0.01:0.01:0.01:0.01:0.171902125; Large Intestine:100:TUF:0.01:4.6781E-30:0.01:0.01:0.01:0.01:0.01; Large Intestine:300:CHR:0.01:8.55585E-28:0.01:0.01:0.01:0.01:0.01; Large Intestine:300:EN-D :0.028139446:0.0207549:0.01:0.01:0.01:0.01:1.423039804; Large Intestine:300:EN-P :0.225660499:0.725122737:0.01:0.01:0.01:0.01:9.328144914; Large Intestine:300:EP1:0.124803964:1.666430783:0.01:0.01:0.01:0.01:21.98275274; Large Intestine:300:EP2:0.136387448:1.286467877:0.01:0.01:0.01:0.01:12.06235413; Large Intestine:300:EP3-D:0.100450079:0.080406616:0.01:0.01:0.01:0.031250293:1.703547206; Large Intestine:300:EP3-P:2.013779797:104.0491511:0.01:0.01:0.01:0.010203473:93.19098724; Large Intestine:300:EP4-D:0.618244402:4.172513336:0.01:0.01:0.01:0.069821296:14.62455624; Large Intestine:300:EP4-P:2.216779474:81.90841528:0.01:0.01:0.01:0.089417422:72.73878781; Large Intestine:300:GB1-D:0.625078047:4.444453549:0.01:0.01:0.01:0.143932698:20.8490123; Large Intestine:300:GB1-P:0.136413664:0.845071847:0.01:0.01:0.01:0.01:9.557555803; Large Intestine:300:GB2-D:0.584599971:1.070237117:0.01:0.01:0.089996554:0.651617884:4.481089012; Large Intestine:300:TUF:0.01:3.04504E-29:0.01:0.01:0.01:0.01:0.01; Large Intestine:2000:CHR:0.015394585:0.000704757:0.01:0.01:0.01:0.01:0.153127464; Large Intestine:2000:EN-D :0.058661679:0.087600473:0.01:0.01:0.01:0.01:2.68276365; Large Intestine:2000:EN-P :0.186182828:0.342939218:0.01:0.01:0.01:0.018775138:4.579226957; Large Intestine:2000:EP1:0.242031163:1.52035394:0.01:0.01:0.01:0.01:13.54761958; Large Intestine:2000:EP2:0.137616207:0.684872242:0.01:0.01:0.01:0.01:7.342157204; Large Intestine:2000:EP3-D:0.248573596:1.254007981:0.01:0.01:0.01:0.043345374:8.244720325; Large Intestine:2000:EP3-P:1.160239961:8.681215788:0.01:0.01:0.01:0.196549515:14.92215724; Large Intestine:2000:EP4-D:1.024156717:12.94504348:0.01:0.01:0.01:0.153132844:28.16052677; Large Intestine:2000:EP4-P:1.870495516:32.95764143:0.01:0.01:0.01:0.241904836:37.58578581; Large Intestine:2000:GB1-D:0.703346962:4.343873838:0.01:0.01:0.017259145:0.246563149:15.28585118; Large Intestine:2000:GB1-P:0.267045238:3.849090953:0.01:0.01:0.01:0.01:20.24231902; Large Intestine:2000:GB2-D:0.602863068:0.645603881:0.01:0.02646894:0.232151648:0.877725989:3.034948802; Large Intestine:2000:TUF:0.01:6.86261E-32:0.01:0.01:0.01:0.01:0.01; Liver:100:B:0.213797267:0.325687856:0.01:0.01:0.01:0.056571357:2.613370426; Liver:100:EC :0.037007003:0.043483818:0.01:0.01:0.01:0.01:2.711247341; Liver:100:HE :1.577528039:35.04314183:0.01:0.01:0.01:0.183535291:43.73524356; Liver:100:KUP:0.509042737:6.621034858:0.01:0.01:0.01:0.012229472:18.82181524; Liver:100:NK2:0.539076723:10.69009305:0.01:0.01:0.01:0.01:20.43361731; Liver:300:B:0.011314986:7.08967E-05:0.01:0.01:0.01:0.01:0.063914422; Liver:300:EC :0.113993526:0.733401923:0.01:0.01:0.01:0.01:10.81729271; Liver:300:HE :1.561492365:98.12025078:0.01:0.01:0.01:0.01:105.9140195; Liver:300:KUP:0.260603997:2.328280291:0.01:0.01:0.01:0.01:11.82954434; Liver:300:NK2:0.125984933:0.524647684:0.01:0.01:0.01:0.01:4.533412393; Liver:2000:B:0.046524022:0.014063618:0.01:0.01:0.01:0.01:0.629735622; Liver:2000:EC :0.155958403:0.528364338:0.01:0.01:0.01:0.01:8.591024824; Liver:2000:HE :1.498511944:62.89746473:0.01:0.01:0.01:0.01:83.09827976; Liver:2000:KUP:0.255600958:1.055424386:0.01:0.01:0.01:0.04115651:7.789532314; Liver:2000:NK2:0.367333689:2.537527345:0.01:0.01:0.01:0.01:9.737383429; Lung:100:B:0.01857141:0.001028567:0.01:0.01:0.01:0.01:0.129999746; Lung:100:CCC:1.456821314:8.825647444:0.01:0.01:0.01:1.547008977:9.746616082; Lung:100:CMN:0.028888155:0.013039084:0.01:0.01:0.01:0.010769856:0.835026957; Lung:100:EC :0.592266493:3.130564936:0.01:0.01:0.01:0.14637682:17.3946568; Lung:100:EP5:2.184258795:31.36444499:0.01:0.142788:0.468145799:1.199237162:33.16037781; Lung:100:LEU2 :0.01:2.14937E-25:0.01:0.01:0.01:0.01:0.01; Lung:100:MN2:0.015465605:0.000828191:0.01:0.01:0.01:0.01:0.194112768; Lung:100:MYE2 :0.01115847:5.79673E-06:0.01:0.01:0.01:0.010703344:0.016013027; Lung:100:NC :5.849657994:83.74798893:0.01:0.01:0.391509729:8.663540696:29.4471241; Lung:100:NK :0.01:6.00811E-27:0.01:0.01:0.01:0.01:0.01; Lung:100:SC :0.596743132:2.135538046:0.01:0.01:0.025279551:0.325975557:10.78671528; Lung:100:T:0.01:5.256E-27:0.01:0.01:0.01:0.01:0.01; Lung:300:B:0.044904601:0.017056636:0.01:0.01:0.01:0.01:0.498664407; Lung:300:CCC:1.891546254:5.51406788:0.01:0.056520696:0.298394181:3.445486709:6.056462974; Lung:300:CMN:0.01072913:1.36084E-05:0.01:0.01:0.01:0.01:0.032120838; Lung:300:EC :0.716551065:10.93183795:0.01:0.01:0.01:0.025238049:41.15891979; Lung:300:EP5:2.367899563:40.49001777:0.01:0.01:0.075846939:1.72436971:34.37901999; Lung:300:LEU2 :0.01:3.0477E-29:0.01:0.01:0.01:0.01:0.01; Lung:300:MN2:0.011887933:9.99892E-05:0.01:0.01:0.01:0.01:0.080417127; Lung:300:MYE2 :0.01:4.19293E-28:0.01:0.01:0.01:0.01:0.01; Lung:300:NC :3.512526692:52.14415724:0.01:0.01:0.166310027:3.261063433:26.50611444; Lung:300:NK :0.01:8.33236E-28:0.01:0.01:0.01:0.01:0.01; Lung:300:SC :0.453653037:1.848672295:0.01:0.01:0.01:0.044123625:8.947055684; Lung:300:T:0.011154861:3.06752E-05:0.01:0.01:0.01:0.01:0.036561792; Lung:2000:B:0.257379098:0.134236035:0.01:0.011041603:0.089877832:0.287607349:1.094854368; Lung:2000:CCC:3.063528155:12.0082984:0.01:0.061665312:1.300096759:5.627807116:9.698705171; Lung:2000:CMN:0.019441605:0.000835036:0.01:0.01:0.01:0.011049543:0.185526719; Lung:2000:EC :0.690009542:4.988718328:0.01:0.01:0.01:0.165162893:19.51617638; Lung:2000:EP5:1.706230301:7.732375989:0.01:0.065375787:0.461915327:1.751011363:10.1566163; Lung:2000:LEU2 :0.010031388:1.37926E-08:0.01:0.01:0.01:0.01:0.010439427; Lung:2000:MN2:0.010778167:1.48816E-05:0.01:0.01:0.01:0.01:0.033806873; Lung:2000:MYE2 :0.154800639:0.063766588:0.01:0.014829632:0.056271633:0.122672891:0.660438295; Lung:2000:NC :2.838071299:26.0408083:0.01:0.01:0.011684946:2.709452458:15.53664066; Lung:2000:NK :0.011379011:1.8491E-05:0.01:0.01:0.01:0.01:0.026888432; Lung:2000:SC :0.325228195:1.505962612:0.01:0.01:0.01:0.040614985:13.08108162; Lung:2000:T:0.435842093:0.882885258:0.01:0.01:0.037300932:0.313265777:4.183441938; Marrow:100:B2 :0.185902275:0.434241737:0.01:0.01:0.01:0.01:2.712418872; Marrow:100:BAS:0.014519176:0.000265498:0.01:0.01:0.01:0.01:0.068749284; Marrow:100:CLP:2.420824383:24.86621029:0.01:0.01:0.214101778:2.430900549:26.01299376; Marrow:100:EPC:1.904845804:16.12612784:0.01:0.01:0.210311004:1.707937823:19.04839729; Marrow:100:GC :1.621135952:4.828924043:0.01:0.117158592:0.707247043:2.393001514:11.24602065; Marrow:100:GMP:0.181084318:0.236147535:0.01:0.01:0.016960511:0.092667507:3.474723251; Marrow:100:GRA:0.374090145:0.702833462:0.01:0.012786779:0.047776796:0.259564272:7.914217848; Marrow:100:HPC:1.13030154:12.97866174:0.01:0.01:0.01:0.195455929:27.27162771; Marrow:100:IB :0.31131546:5.521495558:0.01:0.01:0.01:0.01:32.72827775; Marrow:100:INK:0.01:3.00023E-30:0.01:0.01:0.01:0.01:0.01; Marrow:100:INKT :0.047342165:0.025679816:0.01:0.01:0.01:0.01:0.709047616; Marrow:100:IT :0.053571303:0.039218934:0.01:0.01:0.01:0.01:1.336584538; Marrow:100:LPB1 :1.608961936:8.767079358:0.01:0.01:0.177612782:1.660130475:16.04030105; Marrow:100:LPB2 :2.376398871:16.3857241:0.01:0.01:0.239978477:2.936407898:16.65895823; Marrow:100:MAC:0.0259025:0.006520404:0.01:0.01:0.01:0.01:0.564218133; Marrow:100:MN :0.038996541:0.026703762:0.01:0.01:0.01091448:0.019318394:1.81463237; Marrow:100:MNT:0.056297223:0.082362769:0.01:0.01:0.01:0.01:1.825894631; Marrow:100:MPC1 :0.367024786:1.315133774:0.01:0.01:0.01:0.090052996:9.312120992; Marrow:100:MPC2 :0.175320535:0.418543065:0.01:0.01:0.01:0.023679329:5.238568571; Marrow:100:NBC:0.095766964:0.126490199:0.01:0.01:0.01:0.036544526:5.940112975; Marrow:100:PB :0.010675228:0.000122283:0.01:0.01:0.01:0.01:0.218063081; Marrow:100:PNK:2.482467108:17.15464909:0.01:0.022599142:0.036680956:3.944061666:14.74284385; Marrow:100:RT :0.012112594:6.19298E-05:0.01:0.01:0.01:0.01:0.040539169; Marrow:300:B2 :0.072463781:0.080445354:0.01:0.01:0.01:0.01:1.517781713; Marrow:300:BAS:0.518609643:3.362888996:0.01:0.01:0.01:0.01:6.621925359; Marrow:300:CLP:1.284579894:25.89444494:0.01:0.01:0.01:0.047828455:31.57063454; Marrow:300:EPC:0.975661564:17.3377232:0.01:0.01:0.01:0.017342857:20.88881381; Marrow:300:GC :0.95647576:6.392217988:0.01:0.01:0.021089981:0.538024117:20.6035697; Marrow:300:GMP:1.437037515:20.63563736:0.01:0.01:0.01:0.149559748:24.18271454; Marrow:300:GRA:0.626599281:4.597668957:0.01:0.01:0.01:0.147438321:23.3880025; Marrow:300:HPC:1.304453698:29.11832644:0.01:0.01:0.01:0.031238635:46.69346883; Marrow:300:IB :0.581646271:13.28705952:0.01:0.01:0.01:0.01:40.52106243; Marrow:300:INK:0.010677684:7.80735E-06:0.01:0.01:0.01:0.01:0.021520634; Marrow:300:INKT :0.011091919:2.26534E-05:0.01:0.01:0.01:0.01:0.030746458; Marrow:300:IT :0.187626221:0.516269114:0.01:0.01:0.01:0.01:4.414912871; Marrow:300:LPB1 :2.327177218:39.45478944:0.01:0.01:0.01:0.669150502:47.664881; Marrow:300:LPB2 :3.027541169:44.75653921:0.01:0.01:0.028413594:1.591414395:32.7034205; Marrow:300:MAC:0.533265095:3.571316901:0.01:0.01:0.01:0.038816232:13.60439309; Marrow:300:MN :0.113166634:1.302606568:0.01:0.01:0.01:0.01:14.07338651; Marrow:300:MNT:0.187233959:0.61635283:0.01:0.01:0.01:0.01:4.786642481; Marrow:300:MPC1 :0.201187132:1.513216898:0.01:0.01:0.01:0.01:14.12861891; Marrow:300:MPC2 :0.049832787:0.104629922:0.01:0.01:0.01:0.01:3.562672271; Marrow:300:NBC:0.186056236:4.138595569:0.01:0.01:0.01:0.01:37.74765959; Marrow:300:PB :0.011186099:0.00027564:0.01:0.01:0.01:0.01:0.317098022; Marrow:300:PNK:2.032067255:22.65497581:0.01:0.01:0.063921568:1.501403207:18.24728518; Marrow:300:RT :0.01:1.13792E-28:0.01:0.01:0.01:0.01:0.01; Marrow:2000:B2 :0.167172108:0.259600952:0.01:0.01:0.01:0.01:2.487743147; Marrow:2000:BAS:2.130613564:8.466073636:0.01:0.056819654:0.825334925:3.54974934:9.299718758; Marrow:2000:CLP:1.098629192:10.50141526:0.01:0.01:0.01:0.265548246:28.26777702; Marrow:2000:EPC:0.420817747:2.493197877:0.01:0.01:0.01:0.093986726:10.33754185; Marrow:2000:GC :0.969074716:3.773959377:0.01:0.01:0.069629699:1.107661712:13.42674231; Marrow:2000:GMP:1.71162988:8.833804552:0.01:0.051021256:0.494409338:1.397984403:12.66009452; Marrow:2000:GRA:0.675219767:3.340806905:0.01:0.01:0.014277322:0.347261219:14.96173219; Marrow:2000:HPC:0.894207305:5.34592313:0.01:0.01:0.01:0.400270601:20.58178485; Marrow:2000:IB :0.369499113:2.05090549:0.01:0.01:0.01:0.010907813:14.24513112; Marrow:2000:INK:0.056216304:0.027923716:0.01:0.01:0.01:0.01:0.699996817; Marrow:2000:INKT :0.111411572:0.167818618:0.01:0.01:0.01:0.01:1.797976351; Marrow:2000:IT :0.11687255:0.091009648:0.01:0.01:0.01:0.01347243:1.739893257; Marrow:2000:LPB1 :1.893858092:16.46027996:0.01:0.01:0.050292558:1.496103716:27.37277965; Marrow:2000:LPB2 :3.302615802:36.83218789:0.01:0.01:0.210685719:3.784145256:27.26089787; Marrow:2000:MAC:0.423312503:0.893001499:0.01:0.01:0.01:0.220012709:4.183752057; Marrow:2000:MN :0.263946748:2.930026775:0.01:0.01:0.01:0.015794453:16.7846469; Marrow:2000:MNT:0.181549588:0.183314907:0.01:0.01:0.01:0.029673303:1.778545776; Marrow:2000:MPC1 :0.362290945:1.790559251:0.01:0.01:0.01:0.035890566:14.05147694; Marrow:2000:MPC2 :0.147777286:0.464731193:0.01:0.01:0.01:0.01:6.22391615; Marrow:2000:NBC:0.137362883:0.348696493:0.01:0.01:0.01:0.01:7.142698962; Marrow:2000:PB :0.058905676:0.294446787:0.01:0.01:0.01:0.01:9.928819957; Marrow:2000:PNK:1.833011369:8.638914818:0.01:0.011389807:0.226801309:2.007042567:9.85192995; Marrow:2000:RT :0.010730026:7.99407E-06:0.01:0.01:0.01:0.01:0.02095039; Pancreas:100:A:0.01:2.95436E-27:0.01:0.01:0.01:0.01:0.01; Pancreas:100:ACI:11.89919455:1239.078518:0.01:0.027755043:1.992745424:10.39200207:299.3981303; Pancreas:100:BC :0.014382556:0.001847759:0.01:0.01:0.01:0.01:0.531578068; Pancreas:100:D:0.01:3.72485E-28:0.01:0.01:0.01:0.01:0.01; Pancreas:100:DUC:0.085962055:0.033775105:0.01:0.01:0.012791323:0.057646561:1.321106184; Pancreas:100:EC :9.397344192:1396.05103:0.01:0.01:0.188111294:0.979741716:253.3114757; Pancreas:100:LEU:0.961108391:10.44635577:0.01:0.014391258:0.041267131:0.085491737:15.68452118; Pancreas:100:PP :0.025692694:0.002933414:0.01:0.01:0.01:0.011266917:0.308624983; Pancreas:100:PSC:2.221217553:63.14401388:0.01:0.01:0.054436934:0.626533357:46.16280282; Pancreas:300:A:0.015787667:0.006866904:0.01:0.01:0.01:0.01:1.19647179; Pancreas:300:ACI:13.64002226:1753.848566:0.01:0.01:0.01:2.740798527:312.48915; Pancreas:300:BC :0.023946869:0.035699395:0.01:0.01:0.01:0.01:2.607600067; Pancreas:300:D:0.01:1.69592E-25:0.01:0.01:0.01:0.01:0.01; Pancreas:300:DUC:0.14580081:0.480585724:0.01:0.01:0.01:0.01:6.811039251; Pancreas:300:EC :1.476887113:58.74553799:0.01:0.01:0.01:0.01:53.07849045; Pancreas:300:LEU:2.617639833:230.4585141:0.01:0.01:0.01:0.01:93.58914858; Pancreas:300:PP :0.01:3.8384E-25:0.01:0.01:0.01:0.01:0.01; Pancreas:300:PSC:0.158792338:0.445999812:0.01:0.01:0.01:0.01:4.047027631; Pancreas:2000:A:0.013012481:0.001619796:0.01:0.01:0.01:0.01:0.584870878; Pancreas:2000:ACI:10.86705529:601.7602352:0.01:0.01:0.815365852:8.845002453:155.7998631; Pancreas:2000:BC :0.012869098:0.000777185:0.01:0.01:0.01:0.01:0.381544425; Pancreas:2000:D:0.010021663:1.6756E-08:0.01:0.01:0.01:0.01:0.01106179; Pancreas:2000:DUC:0.206309149:0.698594703:0.01:0.01:0.01:0.020448076:6.846125984; Pancreas:2000:EC :0.624484527:6.177609328:0.01:0.01:0.01:0.022861064:16.26364263; Pancreas:2000:LEU:0.224386478:1.390589728:0.01:0.01:0.01:0.016633662:7.289777506; Pancreas:2000:PP :0.01169833:5.85563E-05:0.01:0.01:0.01:0.01:0.049072198; Pancreas:2000:PSC:0.22808577:0.454643165:0.01:0.01:0.01:0.024362231:2.982198696; Skin:100:BE:0.088525095:0.433917818:0.01:0.01:0.01:0.01:9.105400656; Skin:100:EPI :1.497255989:8.07664797:0.01:0.023857561:0.156802628:1.353358422:15.12733217; Skin:100:KSC:0.025219498:0.016150493:0.01:0.01:0.01:0.01:1.945234359; Skin:100:KSC2:0.317196037:3.681240333:0.01:0.01:0.01:0.03712216:40.56973381; Skin:100:LEU:6.537198302:190.7648029:0.01:0.01:0.076820057:6.008499214:46.96736574; Skin:100:SCE:0.098152807:0.13388837:0.01:0.01:0.01:0.036533711:2.313070393; Skin:300:BE:0.108357985:1.13393988:0.01:0.01:0.01:0.01:18.55862333; Skin:300:EPI :1.305163535:21.44147441:0.01:0.01:0.01:0.098746599:36.14180497; Skin:300:KSC:0.019268223:0.015211094:0.01:0.01:0.01:0.01:2.485425355; Skin:300:KSC2:0.419843734:9.493485969:0.01:0.01:0.01:0.01:63.04223322; Skin:300:LEU:7.856353791:227.1964209:0.01:0.01:0.01:10.01127642:50.34195125; Skin:300:SCE:0.041529126:0.025377722:0.01:0.01:0.01:0.01:1.015417564; Skin:2000:BE:0.054235631:0.107612828:0.01:0.01:0.01:0.01:4.894140689; Skin:2000:EPI :1.463389585:13.46542866:0.01:0.01:0.041508019:0.580429537:24.10158418; Skin:2000:KSC:0.057082183:0.283160697:0.01:0.01:0.01:0.01:8.947880368; Skin:2000:KSC2:0.358296776:2.180414803:0.01:0.01:0.01:0.01:17.56111049; Skin:2000:LEU:3.062859377:42.92355399:0.01:0.01:0.039353867:2.372801805:22.81755758; Skin:2000:SCE:0.418997541:4.3441421:0.01:0.01:0.01:0.01:13.10047803; SkMuscle:100:B3 :0.576226174:6.536569269:0.01:0.01:0.01:0.010556533:12.0156782; SkMuscle:100:B4 :0.126395942:0.228262628:0.01:0.01:0.01:0.01:3.576026664; SkMuscle:100:DEN:0.01:7.18298E-31:0.01:0.01:0.01:0.01:0.01; SkMuscle:100:EC :0.675764223:2.800705161:0.01:0.01:0.030371209:0.285263553:7.310227407; SkMuscle:100:ERB1 :0.01:4.75706E-24:0.01:0.01:0.01:0.01:0.01; SkMuscle:100:ERB2 :0.010000162:4.76075E-12:0.01:0.01:0.01:0.01:0.010029355; SkMuscle:100:GMP:0.087430963:0.486566491:0.01:0.01:0.01:0.01:6.365643955; SkMuscle:100:MAC2 :0.029195195:0.03837845:0.01:0.01:0.01:0.01:2.308267274; SkMuscle:100:MAC3 :0.013829597:0.000133662:0.01:0.01:0.01:0.01:0.06064265; SkMuscle:100:MC1:28.93245718:212.9191713:0.101008:28.7416311:35.43885636:36.5236684:44.08999214; SkMuscle:100:MC2:15.04398917:182.2259111:0.01:5.74097169:11.36666489:21.01674324:54.58031214; SkMuscle:100:MPC:0.02286293:0.004120317:0.01:0.01:0.01:0.01:0.523549078; SkMuscle:100:NEUT1 :0.010017958:4.95878E-08:0.01:0.01:0.01:0.01:0.012908385; SkMuscle:100:NEUT2 :0.011664165:9.56584E-06:0.01:0.01:0.01:0.011539886:0.017153774; SkMuscle:100:NEUT3 :0.018413406:0.009697599:0.01:0.01:0.01:0.01:1.162636555; SkMuscle:100:SC :0.087895601:0.009532824:0.01:0.01:0.05575928:0.105801958:0.345386647; SkMuscle:100:T:0.109161343:0.059970569:0.01:0.01:0.01:0.028898199:0.752215042; SkMuscle:300:B3 :0.010361993:2.88286E-06:0.01:0.01:0.01:0.01:0.017963856; SkMuscle:300:B4 :0.029963223:0.03376483:0.01:0.01:0.01:0.01:1.858624152; SkMuscle:300:DEN:0.16913301:0.253233149:0.01:0.01:0.01:0.01:1.601330102; SkMuscle:300:EC :0.895315862:11.40577034:0.01:0.01:0.01:0.01:16.86925061; SkMuscle:300:ERB1 :0.01:9.0575E-24:0.01:0.01:0.01:0.01:0.01; SkMuscle:300:ERB2 :0.010088532:1.41868E-06:0.01:0.01:0.01:0.01:0.026024377; SkMuscle:300:GMP:0.037725117:0.063800617:0.01:0.01:0.01:0.01:2.311184743; SkMuscle:300:MAC2 :0.014416578:0.002691851:0.01:0.01:0.01:0.01:0.619487802; SkMuscle:300:MAC3 :0.010123343:3.80338E-07:0.01:0.01:0.01:0.01:0.013083577; SkMuscle:300:MC1:7.592685729:87.58666453:0.01:1.765507418:4.031314363:10.95550861:30.46760561; SkMuscle:300:MC2:15.5601171:833.7239616:0.01:0.01:1.157277635:12.19090123:106.3308553; SkMuscle:300:MPC:0.030102624:0.035158047:0.01:0.01:0.01:0.01:1.75892828; SkMuscle:300:NEUT1 :0.01:5.49289E-23:0.01:0.01:0.01:0.01:0.01; SkMuscle:300:NEUT2 :0.032082425:0.003901068:0.01:0.01:0.01:0.01:0.186659401; SkMuscle:300:NEUT3 :0.01087057:0.000103831:0.01:0.01:0.01:0.01:0.129268097; SkMuscle:300:SC :0.033081414:0.009127034:0.01:0.01:0.01:0.01:0.465886089; SkMuscle:300:T:0.010433177:1.68878E-06:0.01:0.01:0.01:0.01:0.013898597; SkMuscle:2000:B3 :2.27902329:55.61744578:0.01:0.01:0.01:0.01:29.79690193; SkMuscle:2000:B4 :0.081475145:0.454831686:0.01:0.01:0.01:0.01:6.820171116; SkMuscle:2000:DEN:0.023446865:0.001476213:0.01:0.01:0.01:0.01:0.132248841; SkMuscle:2000:EC :1.148049772:26.0796183:0.01:0.01:0.01:0.01:26.51870209; SkMuscle:2000:ERB1 :0.01:2.62163E-29:0.01:0.01:0.01:0.01:0.01; SkMuscle:2000:ERB2 :0.012219673:0.000496529:0.01:0.01:0.01:0.01:0.28302578; SkMuscle:2000:GMP:0.010542119:2.43931E-05:0.01:0.01:0.01:0.01:0.054995847; SkMuscle:2000:MAC2 :0.09516298:0.982113219:0.01:0.01:0.01:0.01:11.65243136; SkMuscle:2000:MAC3 :0.01:2.2775E-30:0.01:0.01:0.01:0.01:0.01; SkMuscle:2000:MC1:7.164084475:71.86606507:0.01:0.787027995:4.186646381:11.04237596:27.32437377; SkMuscle:2000:MC2:14.32686526:568.351699:0.01:0.156707198:2.07884161:13.74566834:82.41595513; SkMuscle:2000:MPC:0.274615674:3.129079571:0.01:0.01:0.01:0.01:14.27447261; SkMuscle:2000:NEUT1 :0.012227059:0.00053056:0.01:0.01:0.01:0.01:0.306938325; SkMuscle:2000:NEUT2 :0.01:2.46619E-30:0.01:0.01:0.01:0.01:0.01; SkMuscle:2000:NEUT3 :0.215387785:5.34538476:0.01:0.01:0.01:0.01:27.06598056; SkMuscle:2000:SC :0.384236497:3.218447212:0.01:0.01:0.01:0.01:8.613897029; SkMuscle:2000:T:0.01:4.29784E-30:0.01:0.01:0.01:0.01:0.01; Spleen:100:B:0.789739043:13.27629372:0.009999999:0.01:0.014974386:0.175114572:76.01353696; Spleen:100:MAC:0.033677171:0.000928003:0.01:0.01052609:0.019320469:0.044604965:0.119540695; Spleen:100:T:0.889478679:2.626844949:0.01:0.028523727:0.229516685:1.134817519:11.21089027; Spleen:300:B:0.6007074:5.718042953:0.009999997:0.01:0.01:0.014069245:24.35418936; Spleen:300:MAC:0.110306986:0.104400051:0.01:0.01:0.01:0.017623009:1.695440132; Spleen:300:T:1.193135349:9.943389873:0.01:0.01:0.01:0.306849232:26.93225583; Spleen:2000:B:0.527834996:2.847863632:0.01:0.01:0.01:0.026519544:14.2891133; Spleen:2000:MAC:0.028947532:0.008489124:0.01:0.01:0.01:0.01:0.526640415; Spleen:2000:T:0.902514553:4.96652535:0.01:0.01:0.01:0.263963255:13.63243794; Thymus:100:IT2:0.892838408:8.146903544:0.01:0.01:0.01:0.187870672:27.72681703; Thymus:100:IT3:1.517353985:8.980022847:0.01:0.01:0.087409456:1.088970761:12.81068485; Thymus:100:IT4:0.509593999:2.415552664:0.01:0.01:0.019557589:0.222231944:12.92113673; Thymus:100:LEU3:0.075703947:0.050678587:0.01:0.01:0.01:0.010156311:0.855605674; Thymus:100:TPT:0.603227667:0.873439087:0.01:0.02145553:0.26503743:0.596217908:2.432418181; Thymus:300:IT2:0.569771808:9.930095561:0.01:0.01:0.01:0.01:31.00492796; Thymus:300:IT3:1.539814872:17.40773119:0.01:0.01:0.01:0.178454647:21.665993; Thymus:300:IT4:0.293142261:3.327556195:0.01:0.01:0.01:0.01:19.65626539; Thymus:300:LEU3:0.100077789:0.113494376:0.01:0.01:0.01:0.01:1.270564448; Thymus:300:TPT:0.01:5.05421E-28:0.01:0.01:0.01:0.01:0.01; Thymus:2000:IT2:0.461222048:2.278633079:0.01:0.01:0.01:0.015257748:11.52396779; Thymus:2000:IT3:1.632787561:11.73465067:0.01:0.01:0.032282811:1.271299967:19.54368881; Thymus:2000:IT4:0.249174221:1.215448065:0.01:0.01:0.01:0.01:7.965207557; Thymus:2000:LEU3:1.037253207:2.514034235:0.01:0.04187717:0.257041468:0.893381107:4.991732332; Thymus:2000:TPT:0.011241477:9.2476E-06:0.01:0.01:0.01:0.01:0.017448864;

Explanation of symbols

[0128] 10 Correction device 101 Control unit 20 Analysis device 201 Control unit

Claims

1. A method for correcting a count data set of single-cell RNA-Seq analysis, which is obtained from the cells to be analyzed or predicted for the cells to be analyzed, and weighted by a weight coefficient based on the total RNA content for each cell type corresponding to the cells to be analyzed, comprising: The weight coefficient is calculated from the count data of single-cell RNA-Seq analysis of a signature gene set characterizing each cell type in each cell type of the reference cell type and the count data obtained by performing RNA-Seq analysis on all the RNA contained in each organ to be analyzed; Both the count data of single-cell RNA-Seq analysis of the signature gene set and the count data obtained by performing RNA-Seq analysis on all the RNA contained in each organ to be analyzed are normalized; The correction method.

2. The correction method according to Claim 1, wherein the signature gene set contains a predetermined number of genes.

3. A method for analyzing single-cell RNA-Seq, which comprises weighting a count data set of single-cell RNA-Seq analysis obtained from the cells to be analyzed or predicted for the cells to be analyzed by a weight coefficient based on the total RNA content for each cell type corresponding to the cells to be analyzed, and analyzing the RNA expression pattern in each cell type constituting the organ to be analyzed including the cells to be analyzed based on the weighted count data set of single-cell RNA-Seq analysis; wherein The weight coefficient is calculated from the count data of single-cell RNA-Seq analysis of a signature gene set characterizing each cell type in each cell type of the reference cell type and the count data obtained by performing RNA-Seq analysis on all the RNA contained in each organ to be analyzed; Both the count data of single-cell RNA-Seq analysis of the signature gene set and the count data obtained by performing RNA-Seq analysis on all the RNA contained in each organ to be analyzed are normalized; The analysis method.

4. To a count data set of single-cell RNA-Seq analysis obtained from the cells to be analyzed or predicted for the cells to be analyzed, weight it with a weight coefficient based on the total RNA content for each cell type corresponding to the cells to be analyzed, Based on the weighted count data set of single-cell RNA-Seq analysis, analyze the composition ratio of cell types constituting the organ to be analyzed including the cells to be analyzed, including A method for analyzing the composition ratio of cell types constituting an organ to be analyzed, The weight coefficient is calculated from the count data of single-cell RNA-Seq analysis of a signature gene set characterizing each cell type in each cell type of the reference cell type and the count data obtained by performing RNA-Seq analysis on all the RNA contained in each organ to be analyzed, Both the count data of single-cell RNA-Seq analysis of the signature gene set and the count data obtained by performing RNA-Seq analysis on all the RNA contained in each organ to be analyzed are normalized, the analysis method.

5. An apparatus for correcting a count data set of single-cell RNA-Seq analysis, The correction apparatus includes a control unit, The control unit Weight the count data set of single-cell RNA-Seq analysis obtained from the cells to be analyzed with a weight coefficient based on the total RNA content for each cell type corresponding to the cells to be analyzed, The weight coefficient is calculated from the count data of single-cell RNA-Seq analysis of a signature gene set characterizing each cell type in each cell type of the reference cell type and the count data obtained by performing RNA-Seq analysis on all the RNA contained in each organ to be analyzed, Both the count data of single-cell RNA-Seq analysis of the signature gene set and the count data obtained by performing RNA-Seq analysis on all the RNA contained in each organ to be analyzed are normalized, the correction apparatus.

6. An analysis apparatus for single-cell RNA-Seq, The analysis apparatus includes a control unit, The control unit Weight the count data set of single-cell RNA-Seq analysis obtained from the cells to be analyzed or predicted for the cells to be analyzed with a weight coefficient based on the total RNA content for each cell type corresponding to the cells to be analyzed, Based on the weighted single-cell RNA-Seq analysis count dataset, analyze the RNA expression patterns in each cell type that constitutes the organ to be analyzed and includes the cells to be analyzed, The weight coefficient is calculated from the count data of the single-cell RNA-Seq analysis of the signature gene set that characterizes each cell type in each cell type of the reference cell type and the count data obtained by performing RNA-Seq analysis on all the RNA contained in each organ to be analyzed, Both the count data of the single-cell RNA-Seq analysis of the signature gene set and the count data obtained by performing RNA-Seq analysis on all the RNA contained in each organ to be analyzed are normalized, The analysis device.

7. An analysis device for the composition ratio of cell types that constitute an organ to be analyzed, The analysis device includes a control unit, The control unit, Based on the total RNA content for each cell type corresponding to the cells to be analyzed, weight the count dataset of the single-cell RNA-Seq analysis obtained from the cells to be analyzed or predicted for the cells to be analyzed with a weight coefficient, Based on the weighted count dataset of the single-cell RNA-Seq analysis, analyze the composition ratio of the cell types that constitute the organ to be analyzed and include the cells to be analyzed, The weight coefficient is calculated from the count data of the single-cell RNA-Seq analysis of the signature gene set that characterizes each cell type in each cell type of the reference cell type and the count data obtained by performing RNA-Seq analysis on all the RNA contained in each organ to be analyzed, Both the count data of the single-cell RNA-Seq analysis of the signature gene set and the count data obtained by performing RNA-Seq analysis on all the RNA contained in each organ to be analyzed are normalized, The analysis device.

8. When executed on a computer, On the computer, A correction program for the count dataset of the single-cell RNA-Seq analysis, which includes a process of weighting the count dataset of the single-cell RNA-Seq analysis obtained from the cells to be analyzed or predicted for the cells to be analyzed with a weight coefficient based on the total RNA content for each cell type corresponding to the cells to be analyzed, The weight coefficient is calculated from the count data of single-cell RNA-Seq analysis of the signature gene set that characterizes each cell type in each cell type of the reference cell types and the count data obtained by performing RNA-Seq analysis on all the RNA contained in each organ to be analyzed, Both the count data of the single-cell RNA-Seq analysis of the signature gene set and the count data obtained by performing RNA-Seq analysis on all the RNA contained in each organ to be analyzed are normalized, The correction program.

9. When executed on a computer, On a computer, A step of weighting a count data set of single-cell RNA-Seq analysis obtained from or predicted for the analysis target cells with a weight coefficient based on the total RNA content for each cell type corresponding to the analysis target cells, A step of analyzing the RNA expression pattern in each cell type constituting the analysis target organ including the analysis target cells based on the weighted count data set of single-cell RNA-Seq analysis, A single-cell RNA-Seq analysis program for executing a process comprising: The weight coefficient is calculated from the count data of single-cell RNA-Seq analysis of the signature gene set that characterizes each cell type in each cell type of the reference cell types and the count data obtained by performing RNA-Seq analysis on all the RNA contained in each organ to be analyzed, Both the count data of the single-cell RNA-Seq analysis of the signature gene set and the count data obtained by performing RNA-Seq analysis on all the RNA contained in each organ to be analyzed are normalized, The analysis program.

10. When executed on a computer, On a computer, A step of weighting a count data set of single-cell RNA-Seq analysis obtained from or predicted for the analysis target cells with a weight coefficient based on the total RNA content for each cell type corresponding to the analysis target cells, A step of analyzing the composition ratio of the cell types constituting the analysis target organ including the analysis target cells based on the weighted count data set of single-cell RNA-Seq analysis, An analysis program for analyzing the composition ratio of cell types constituting an organ to be analyzed, which causes a process including the following to be executed: The weight coefficient is calculated from the count data of single-cell RNA-Seq analysis of a signature gene set that characterizes each cell type in each cell type of the reference cell type and the count data obtained by performing RNA-Seq analysis on all RNA contained in each organ to be analyzed. Both the count data of single-cell RNA-Seq analysis of the signature gene set and the count data obtained by performing RNA-Seq analysis on all RNA contained in each organ to be analyzed are normalized. The analysis program.

Citation Information

Patent Citations

  • Systems and methods for analyzing mixed cell populations

    WO2019018684A1