Information processing device, method for operating information processing device, and program for operating information processing device

The information processing device performs cluster allocation and negative binomial distribution search on the cell population, which solves the accuracy of DEG detection in the multi-subtype cell population, and achieves a more reasonable and reliable DEG detection.

CN120283285APending Publication Date: 2025-07-08FUJIFILM CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380082325.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-02
Filing Date
2023-11-30
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

现有技术在混合存在多种亚型的细胞群中难以有效检测差异表达基因(DEG),因为其假设基因表达量分布为单峰性,而实际可能为多峰性。

Method used

Cluster allocation of cell populations is performed through an information processing device, and a suitable probability distribution is searched using a negative binomial distribution, significant differences between groups are determined, representative index values and degree of difference are calculated, and cluster allocation and distribution search are adjusted to improve accuracy.

Benefits of technology

Even in a multi-subtype cell population, significant differences can be accurately detected, reducing misjudgment, and improving the rationality and credibility of the detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120283285A_ABST
    Figure CN120283285A_ABST
Patent Text Reader

Abstract

An information processing device for detecting expression-varying genes exhibiting specific expressions with respect to cell characteristics of interest on the basis of gene expression level data of a cell population in which a plurality of subtypes are mixed, said information processing device being provided with a processor for processing the expression-varying genes in which the expression-varying genes exhibit specific expressions with respect to the cell characteristics of interest. The processor performs a process for assigning a cluster to which a cell population is estimated to belong in a gene expression amount distribution for each of a plurality of candidate genes, which are candidates for expression-varying genes, to each sample in which the cell population is divided into two groups in accordance with cell characteristics of interest, and for each of the plurality of candidate genes on the basis of the result of the assignment of the clusters, performing a process for determining the expression amount of the candidate genes. A first probability distribution suitable for the distribution of the gene expression levels of the two groups is searched for.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present invention relates to an information processing device, a working method of the information processing device, and a working program of the information processing device. Background Art

[0002] As a technology that is the core of gene analysis, a technology for detecting differentially expressed genes (hereinafter, referred to as DEGs (Differential Expressed Genes)) that show specific expression with respect to a cell property of interest based on gene expression level data of a certain cell population is known.

[0003] As an example, consider the following case: Based on gene expression level data of cancer tissue-derived cells collected from breast cancer (hereinafter, referred to as BRCA (Breast Invasive Carcinoma)) patients, genes (referred to as ER marker genes) that are important in determining the state of estrogen receptor (hereinafter, referred to as ER (Estrogen Receptor)), which is a cell property of interest, are detected. In this case, BRCA patients are divided into two groups according to the size of the proportion of cancer tissue-derived cells having ER, and the gene expression levels of cancer tissue-derived cells are compared between the groups, thereby detecting ER marker genes.

[0004] And, as another example, consider the following case: Based on gene expression level data of Chinese hamster ovary (hereinafter, referred to as CHO (Chinese Hamster Ovary)) cells widely used in the production of antibody drugs, genes that contribute highly to antibody production, which is a cell property of interest, are detected. In this case, the population of CHO cells is divided into two groups according to the amount of antibody produced, and the gene expression levels are compared between the groups, thereby detecting genes that contribute highly to antibody production.

[0005] In <M.D. Robinson, et al., “edgeR: a Bioconductor package for differential expression analysis of digital gene expression data.” Bioinformatics 26, 1392010.> (hereinafter, referred to as Non-Patent Document 1), a method for detecting DEGs called edgeR is described. The method for detecting DEGs described in Non-Patent Document 1 is an algorithm constructed based on the understanding that the distribution of gene expression levels of a cell population (a histogram with the gene expression level on the horizontal axis and the number of samples on the vertical axis) well conforms to the negative binomial distribution.

[0006] In the method for detecting DEGs described in Non-Patent Document 1, negative binomial distributions suitable for the distributions of gene expression levels of two groups of cell populations are searched for separately. Then, based on parameters such as the means and variances of the two searched negative binomial distributions, it is determined by statistical methods whether the difference between the groups is large, and genes determined to have a large difference between the groups are detected as DEGs. In addition, the process of searching for a negative binomial distribution suitable for the distribution of gene expression levels, etc., the process of searching for a probability distribution suitable for a certain distribution is called statistical modeling. Summary of the Invention

[0007] Technical Problem to be Solved by the Invention

[0008] In a cell population, multiple subtypes of cell populations are mixedly present. In such a cell population with multiple subtypes mixedly present, it is expected that multiple clusters (peaks) caused by multiple subtypes will appear in the distribution of gene expression levels. On the other hand, the method for detecting DEGs described in Non-Patent Document 1 assumes a so-called unimodal gene expression level distribution with one cluster, and does not assume a so-called multimodal gene expression level distribution with multiple clusters. Therefore, in the method for detecting DEGs described in Non-Patent Document 1, it may not be possible to detect effective DEGs.

[0009] One embodiment of the technology of the present invention provides an information processing device, a working method of the information processing device, and a working program of the information processing device that can detect effective DEGs even when a cell population with multiple subtypes mixedly present is the analysis object.

[0010] Means for Solving the Technical Problem

[0011] The information processing device of the present invention performs the following processing: Based on gene expression level data of a cell population with multiple subtypes mixedly present, it detects expression-variable genes that show specific expression with respect to a cell characteristic of interest. The information processing device includes a processor, and the processor performs the following processing: For each of a plurality of candidate genes that are candidates for expression-variable genes, it assigns to each sample that divides the cell population into two groups according to the cell characteristic of interest, the cluster that is presumed to belong in the gene expression level distribution; Based on the cluster assignment results, for each of the plurality of candidate genes, it separately searches for a first probability distribution suitable for the distributions of gene expression levels of the two groups.

[0012] Preferably, the processor performs the following processing: Based on the search results of the first probability distribution, for each of the plurality of candidate genes, it determines whether there is a statistically significant difference in the distributions of gene expression levels of the two groups; It detects candidate genes determined to have a statistically significant difference in the distributions of gene expression levels of the two groups as expression-variable genes.

[0013] The processor preferably performs the following processing: calculate the representative index value of each of the two groups according to the distribution of gene expression levels and the first probability distribution searched; calculate a score representing the degree of difference between the two groups according to the representative index value; and make a determination based on the score.

[0014] The processor preferably performs at least one of the following processes: receive the search result of the first probability distribution and re-perform the cluster assignment, and re-search the first probability distribution based on the result of the re-performed cluster assignment.

[0015] The processor preferably performs the following processing: change the number of assigned clusters, search the first probability distribution for each changed number; and determine, based on the search result of the first probability distribution, the number among the changed numbers that can search the first probability distribution most suitable for the distribution of gene expression levels.

[0016] The processor preferably performs the following processing: perform cluster assignment using the second probability distribution.

[0017] The first probability distribution is preferably a negative binomial distribution.

[0018] The processor preferably performs the following processing: generate simulated gene expression level data by simulation based on the searched first probability distribution.

[0019] The processor preferably performs the following processing: generate only simulated gene expression level data of expressed variable genes by simulation.

[0020] A working method of an information processing device, the information processing device performs the following processing: based on gene expression level data of a cell population in which multiple subtypes are mixed, detect expressed variable genes that show specific expression regarding the cell characteristics of interest,

[0021] The working method of the information processing device includes the following steps:

[0022] For each of a plurality of candidate genes that are candidates for expressed variable genes, assign to each sample that divides the cell population into two groups according to the cell characteristics of interest, the cluster presumed to belong in the gene expression level distribution; and

[0023] Based on the cluster assignment result, for each of the plurality of candidate genes, respectively search the first probability distribution suitable for the gene expression levels of the two groups.

[0024] A working program of an information processing device, the information processing device performing the following processing: based on gene expression data of a cell population in which multiple subtypes coexist, detecting expression-altered genes that are specifically expressed with respect to a cell characteristic of interest, the working program of the information processing device causing a computer to execute a process including the following steps: for each of a plurality of candidate genes that are candidates for expression-altered genes, assigning, to each sample that divides the cell population into two groups according to the cell characteristic of interest, a cluster presumed to belong in the gene expression distribution; and based on the cluster assignment result, for each of the plurality of candidate genes, respectively searching for a first probability distribution suitable for the gene expression distributions of the two groups.

[0025] Advantages of the Invention

[0026] According to the technology of the present invention, it is possible to provide an information processing device, a working method of the information processing device, and a working program of the information processing device that can detect effective DEGs even when analyzing a cell population in which multiple subtypes coexist. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a diagram showing ER status information of BRCA patients, gene expression data of cancer tissue-derived cells collected from BRCA patients, and an information processing device.

[0028] Figure 2 It is a diagram showing the situation of dividing BRCA patients into a high-expression group and a low-expression group based on ER status information and the gene expression distribution.

[0029] Figure 3 It is a block diagram of a computer constituting the information processing device.

[0030] Figure 4 It is a block diagram of a processing unit of the CPU of the information processing device.

[0031] Figure 5 It is a diagram showing the processing of the allocation unit.

[0032] Figure 6 It is a diagram showing a negative binomial distribution suitable for the gene expression distribution.

[0033] Figure 7 It is a diagram showing the search result.

[0034] Figure 8 It is a flowchart showing the processing sequence of the detection unit.

[0035] Figure 9 It is a diagram showing the detailed structure of the detection unit.

[0036] Figure 10 It is a diagram showing the function calculation result.

[0037] Figure 11 It is a figure showing the result of fraction calculation.

[0038] Figure 12 It is a flowchart showing the processing sequence of the determination unit.

[0039] Figure 13 It is a figure showing a gene analysis screen displaying the detection result of DEG.

[0040] Figure 14 It is a figure showing the situation where the gene expression level of DEG among the gene expression levels of all genes is input into the ER status prediction model, and the ER status prediction result is output from the ER status prediction model.

[0041] Figure 15 It is a flowchart showing the processing sequence of the information processing device.

[0042] Figure 16 It is a figure showing the second embodiment of allocating clusters using class distribution.

[0043] Figure 17 It is a figure showing the third embodiment of feeding back the search result based on the negative binomial distribution of the search unit to the allocation unit and causing the allocation unit to re-allocate clusters.

[0044] Figure 18 It is a flowchart showing the processing sequence of the information processing device in the third embodiment.

[0045] Figure 19 It is a figure showing the fourth embodiment of changing the number of allocated clusters and performing a search of the negative binomial distribution for each changed number.

[0046] Figure 20 It is a figure showing the situation of determining the number that can search for the negative binomial distribution most suitable for the gene expression level distribution from the changed numbers.

[0047] Figure 21 It is a figure showing the fifth - 1 embodiment of generating simulated gene expression level data based on the searched negative binomial distribution through simulation.

[0048] Figure 22 It is a figure showing the fifth - 2 embodiment of generating simulated gene expression level data of only DEG through simulation.

[0049] Figure 23 It is a table showing the classification details of the groups and subtypes of the samples in the examples.

[0050] Figure 24It is a graph showing the gene expression level distribution of genes detected as DEGs in both the prior art and the technology of the present invention.

[0051] Figure 25 It is a graph showing the gene expression level distribution of genes detected as DEGs in the prior art but not detected as DEGs in the technology of the present invention.

[0052] Figure 26 It is a graph showing the gene expression level distribution of genes detected as DEGs in the technology of the present invention but not detected as DEGs in the prior art. Detailed implementation mode

[0053] [First Embodiment]

[0054] As an example, as Figure 1 shown, the information processing apparatus 10 related to the technology of the present invention performs processing for detecting DEGs (ER marker genes) based on the sample data group 11. The information processing apparatus 10 is, for example, a desktop personal computer, and includes a display 12 for displaying various screens and input devices 13 such as a keyboard, a mouse, a touch panel, and / or a microphone for voice input. The information processing apparatus 10 is, for example, installed in a gene analysis facility and is operated by a user 14 participating in gene analysis in the gene analysis facility.

[0055] The sample data group 11 is input to the information processing apparatus 10. The sample data group 11 includes a plurality of sample data 15. Moreover, the sample data 15 is composed of a group of ER status information 16 and gene expression level data 17. One sample data 15, that is, one group of ER status information 16 and gene expression level data 17, is derived from one BRCA patient 18.

[0056] The ER status information 16 is information indicating the expression status of ER in cancer tissue-derived cells 20 collected from the breast 19 of the BRCA patient 18. When the proportion of cancer tissue-derived cells 20 having ER is greater than a preset threshold, "high expression" is registered in the ER status information 16. On the other hand, when the proportion of cancer tissue-derived cells 20 having ER is less than the threshold, "low expression" is registered in the ER status information 16. The ER status information 16 includes a sample ID (Identification Data) for uniquely identifying each BRCA patient 18. The expression status of ER in the cancer tissue-derived cells 20 registered in the ER status information 16 is an example of the "cell characteristics of interest" related to the technology of the present invention.

[0057] The gene expression level data 17 can be obtained by performing gene expression analysis 21 such as RNA (Ribonucleic acid)-Seq (sequencing) on cancer tissue-derived cells 20. As shown in Table 22, the gene expression level data 17 registers the gene expression level of each gene. The figure that visualizes the gene expression level of each gene by the type and shade of color is a heat Figure 23 .

[0058] Similar to the ER status information 16, the gene expression level data 17 also contains a sample ID. That is, the ER status information 16 and the gene expression level data 17 are associated with each other through the sample ID.

[0059] Multiple cancer tissue-derived cells 20 collected from multiple BRCA patients 18 constitute a cancer tissue-derived cell population 24. The cancer tissue-derived cells 20 have 4 subtypes: "Basal-like", "HER2 (Human Epidermal Growth Factor Receptor Type 2)-enriched", "Luminal A", and "Luminal B". Also, according to multiple documents, it is suggested that there is a subtype called "Normal-like" in the cancer tissue-derived cells 20. That is, the cancer tissue-derived cell population 24 is an example of the "cell population in which multiple subtypes coexist" involved in the technology of the present invention. In addition, when the number of multiple cancer tissue-derived cells 20 (sample number) constituting the cancer tissue-derived cell population 24 is set to N and the number of genes is set to G, the gene expression level data 17 can be expressed as a matrix Y∈NN in which each element has a natural number NN value G×N .

[0060] As an example, as Figure 2 shown, the BRCA patients 18 are divided into any one of two groups, a high-expression group HG and a low-expression group LG, based on the ER status information 16. Specifically, the BRCA patients 18 registered as "high expression" in the ER status information 16 are assigned to the high-expression group HG. In contrast, the BRCA patients 18 registered as "low expression" in the ER status information 16 are assigned to the low-expression group LG. For each of these high-expression group HG and low-expression group LG, a gene expression level distribution (hereinafter, referred to as gene expression distribution) 30 can be generated for each of the candidate genes of multiple DEGs based on the gene expression level data 17. The gene expression distribution 30 is a histogram with the gene expression level of the candidate gene on the horizontal axis and the number of sample i on the vertical axis. In Figure 2In it, the gene expression level distribution 30 of the candidate gene ESR1 is illustrated. Hereinafter, the set of gene expression level distributions 30 generated for each of a plurality of candidate genes is labeled as the gene expression level distribution group 31. In addition, the candidate genes are genes pre-selected by the user 14 from all genes. The number of candidate genes is, for example, about 20,000.

[0061] As an example, as Figure 3 shown, in addition to the aforementioned display 12 and input device 13, the computer constituting the information processing device 10 further includes a storage device 35, a memory 36, a CPU (Central Processing Unit) 37, and a communication unit 38. These are connected to each other via a bus 39.

[0062] The storage device 35 is a mechanical hard disk drive (Hard Disk Drive) built into the computer constituting the information processing device 10 or connected by a cable or network. Alternatively, the storage device 35 is a disk array formed by combining and arranging multiple mechanical hard disks. The storage device 35 stores control programs such as an operating system, various application programs, and various data attached to these programs. In addition, a solid state drive (Solid State Drive) can be used instead of the mechanical hard disk.

[0063] The memory 36 is a working memory for causing the CPU 37 to execute processing. The CPU 37 loads the programs stored in the storage device 35 into the memory 36 and executes processing according to the programs. Thus, the CPU 37 centrally controls each part of the computer. The CPU 37 is an example of the "processor" related to the technology of the present invention. In addition, the memory 36 can also be built into the CPU 37. The communication unit 38 controls the transmission of various information with an external device, such as a device for performing gene expression analysis 21.

[0064] As an example, as Figure 4 shown, a working program 45 is stored in the storage device 35 of the information processing device 10. The working program 45 is an application program for causing the computer to function as the information processing device 10. That is, the working program 45 is an example of the "working program of the information processing device" related to the technology of the present invention.

[0065] When the working program 45 is started, the CPU 37 of the computer constituting the information processing device 10 works in cooperation with the memory 36, etc., and functions as an acquisition unit 50, a distribution generation unit 51, an allocation unit 52, a search unit 53, a detection unit 54, and a display control unit 55.

[0066] The acquisition unit 50 acquires the sample data group 11. The sample data group 11 is output from the acquisition unit 50 to the distribution generation unit 51, the allocation unit 52, and the search unit 53. More specifically, after the sample data group 11 is temporarily stored in the storage device 35, it is appropriately read from the storage device 35 and output to the distribution generation unit 51, the allocation unit 52, and the search unit 53.

[0067] As Figure 2 shown, the distribution generation unit 51 generates the gene expression level distribution group 31 based on the sample data group 11. The gene expression level distribution group 31 is output from the distribution generation unit 51 to the search unit 53, the detection unit 54, and the display control unit 55. Similar to the sample data group 11, after the gene expression level distribution group 31 is also temporarily stored in the storage device 35, it is appropriately read from the storage device 35 and output to the search unit 53, the detection unit 54, and the display control unit 55.

[0068] For each of the multiple candidate genes, the allocation unit 52 allocates the cluster to which each sample i of the high expression group HG and the low expression group LG is presumed to belong in the gene expression level distribution 30. The allocation unit 52 generates the allocation result 60 of the cluster and outputs the generated allocation result 60 to the search unit 53.

[0069] Based on the allocation result 60, for each of the multiple candidate genes, the search unit 53 respectively searches for the probability distribution P that fits the gene expression level distribution 30 of the high expression group HG and the low expression group LG. The search unit 53 generates the search result 61 of the probability distribution P and outputs the generated search result 61 to the detection unit 54.

[0070] Based on the search result 61, the detection unit 54 detects DEG from the multiple candidate genes. The detection unit 54 generates the detection result 62 of DEG and outputs the generated detection result 62 to the display control unit 55. In addition, similar to the sample data group 11, etc., the allocation result 60, the search result 61, and the detection result 62 are also temporarily stored in the storage device 35 and then appropriately read from the storage device 35 and output to the search unit 53, the detection unit 54, and the display control unit 55.

[0071] The display control unit 55 controls the display of various screens on the display 12. The various screens include the gene analysis screen 80 on which the detection result 62 is displayed (refer to Figure 13 ).

[0072] As an example, as Figure 5As shown, the number-of-clusters setting information 65 is input to the allocation unit 52. The number-of-clusters setting information 65 includes a number-of-clusters setting value K. The number-of-clusters setting value K is input by the user 14 via the input device 13. The user 14 inputs, as the number-of-clusters setting value K, the number of clusters presumably latent in the gene expression level distribution 30. In the gene expression level distribution 30 related to the cancer tissue-derived cells 20 of the BRCA patient 18 in this example, based on the understanding, it is presumed that there are two clusters according to the gene HER2. Therefore, in this example, the number-of-clusters setting value K = 2 is set.

[0073] The allocation unit 52 allocates a cluster z to each sample i based on the gene expression level of the gene HER2 i ∈ {1, 2}, and generates an allocation result 60. In the allocation result 60, for each sample ID, the group (high-expression group HG or low-expression group LG) to which the sample i belongs and the allocated cluster (cluster 1 or cluster 2) are registered.

[0074] As an example, as Figure 6 shown, for the sample i in the high-expression group HG and the count data Y of the gene ESR1 iESR1 (HG) is represented by the following formula (1).

[0075] Y iESR1 (HG) =Y iESR1C1 (HG) +Y iESR1C2 (HG) ··· (1)

[0076] where Y iESR1C1 (HG) is the count data of the sample i presumably belonging to cluster 1 in the high-expression group HG, and Yi ESR1C2 (HG) is the count data of the sample i presumably belonging to cluster 2 in the high-expression group HG.

[0077] Similarly, for the sample i in the low-expression group LG and the count data Y of the gene ESR1 iESR1 (LG) is represented by the following formula (2).

[0078] Y iESR1 (LG) =Y iESR1C1 (LG) +Y iESR1C2 (LG) ··· (2)

[0079] where Y iESR1C1 (LG) is the count data of the sample i presumably belonging to cluster 1 in the low-expression group LG, and Yi ESR1C2 (LG)is the count data of sample i estimated to belong to cluster 2 in the low expression group LG.

[0080] If equations (1) and (2) are further generalized, the count data Y of sample i and gene g in group j is i g (j) It is represented by the following formula (3).

[0081] [Formula 1]

[0082]

[0083] In addition, δ zik=1 is the Kronecker delta function.

[0084] The search unit 53 searches for a probability distribution P suitable for the gene expression level distribution 30 based on the assumption that the gene expression level distribution 30 follows the probability distribution P with parameter λ. The search unit 53 searches for the probability distribution P for each cluster. That is, it is assumed that the gene expression level distribution 30 of the gene g related to the group of samples i assigned to the cluster k follows the probability distribution P with parameter λ. k =(p gk 、r gk ) is the probability distribution P(p gk 、r gk ), the search unit 53 performs probability distribution P(p gk 、r gk ) search. In addition, p gk 、r gk A numerical value that determines the shape of the probability distribution P, such as the mean, variance, etc.

[0085] Here, it is known that each count data in the gene expression level distribution 30 well follows the negative binomial distribution P which is a type of probability distribution P. NB (p, r). Therefore, the count data Y of sample i and gene ESR1 in the high expression group HG is iESR1 (HG) It is further represented by the following formula (4).

[0086] Y iESR1 (HG) =Y iESR1C1 (HG) +Y iESR1C2 (HG) ~P NB (p ESR1C1 (HG) 、r ESR1C1 (HG) )+P NB (p ESR1C2 (HG) 、r ESR1C2 (HG) )···(4)

[0087] Similarly, for sample i of the low-expression group LG, the count data Y of gene ESR1 iESR1 (LG) is further represented by the following formula (5).

[0088] Y iESR1 (HG) = Y iESR1C1 (LG) + Y iESR1C2 (LG) ~ P NB (p ESR1C1 (LG) , r ESR1C1 (LG) ) + P NB (p ESR1C2 (LG) , r ESR1C2 (LG) ) ··· (5)

[0089] If formulas (4) and (5) are further generalized, the count data Y of gene g for sample i in group j ig (j) is further represented by the following formula (6).

[0090] [Mathematical formula 2]

[0091]

[0092] In addition, the negative binomial distribution P NB (p, r) is an example of the "first probability distribution" involved in the technology of the present invention.

[0093] The search unit 53, for example, uses the variational Bayesian method or the Markov chain Monte Carlo method, etc., to optimize the parameters p and r of the negative binomial distribution P NB (p, r) of each cluster in such a way that it better fits the gene expression level distribution 30. As an example, as Figure 7 shown, the search result 61 becomes the data of the negative binomial distribution P NB (p, r) of each cluster in which the parameters p and r are optimized and registered in the high-expression group HG and the low-expression group LG respectively for each of the multiple candidate genes.

[0094] As an example, for the flowchart shown in Figure 8 , the detection unit 54 detects DEG in the following order. First, the detection unit 54 determines whether there is a statistically significant difference between the gene expression level distribution 30 of the high-expression group HG and the gene expression level distribution 30 of the low-expression group LG for one candidate gene based on the search result 61 of the negative binomial distribution P NB (p, r) (step ST10).

[0095] When it is determined that there is a statistically significant difference (Yes in step ST20), the detection unit 54 detects the candidate gene as a DEG (step ST30). On the other hand, when it is determined that there is no statistically significant difference (No in step ST20), the detection unit 54 does not detect the candidate gene as a DEG. The detection unit 54 repeats the processes of step ST10, step ST20, and step ST30 until the determination of all candidate genes is completed (No in step ST40).

[0096] As an example, as Figure 9 shown, the detection unit 54 includes a function calculation unit 70, a score calculation unit 71, and a determination unit 72.

[0097] The gene expression level distribution group 31 and the search result 61 are input to the function calculation unit 70. Based on the gene expression level distribution group 31 and the search result 61, the function calculation unit 70 calculates the log-likelihood function l g (j) for the high-expression group HG and the low-expression group LG, respectively. And, the function calculation unit 70 calculates the log-likelihood function l g for all samples i in total of the two combinations of the high-expression group HG and the low-expression group LG. g (j) The function calculation unit 70 outputs the calculation results of the log-likelihood function l g and the log-likelihood function l

[0098] As an example, as Figure 10 shown, the function calculation result 75 is data in which the calculated log-likelihood function l g (HG) of the high-expression group HG, the log-likelihood function l g (LG) of the low-expression group LG, and the overall log-likelihood function l g are registered for each of multiple candidate genes. The log-likelihood functions l g (j) of the high-expression group HG and the low-expression group LG are an example of the "representative index value" related to the technology of the present invention. In addition, the representative index value may be the average value of the gene expression level distributions 30 of the high-expression group HG and the low-expression group LG, etc.

[0099] The score calculation unit 71 calculates a score Λ g indicating the degree of difference between the high-expression group HG and the low-expression group LG based on the function calculation result 75. The score calculation unit 71 outputs the calculation result of the score Λ g (hereinafter, marked as the score calculation result) 76 to the determination unit 72.

[0100] As an example, as Figure 11 shown, the fractional calculation result 76 is the calculated score Λ registered for each of a plurality of candidate genes g data. In addition, the score Λ g is obtained by calculating the following formula (7).

[0101] Λ g = -2(l g - l g (HG) - l g (LG) ) ··· (7)

[0102] The determination unit 72 performs Figure 8 the process of step ST10 of the flowchart shown. More specifically, as Figure 12 shown in the flowchart, the determination unit 72 uses the score Λ g to perform a likelihood ratio test of the null hypothesis that there is no statistically significant difference between the gene expression level distribution 30 of the high-expression group HG and the gene expression level distribution 30 of the low-expression group LG, and calculates the p-value (step ST101). This likelihood ratio test is used to test the degree of compliance of the score Λ g in the χ 2 distribution with a degree of freedom θ = 2K + 1.

[0103] Next, the determination unit 72 corrects the p-value by, for example, Bonferroni correction or the like (step ST102). This correction is performed to solve the problem of multiple testing that can occur when a likelihood ratio test is applied to an object with multiple degrees of freedom such as a gene.

[0104] When the corrected p-value is less than a pre-set significance level (yes in step ST103), the determination unit 72 rejects the null hypothesis established in step ST101 and determines that there is a statistically significant difference between the gene expression level distribution 30 of the high-expression group HG and the gene expression level distribution 30 of the low-expression group LG (step ST104). In this case, the determination unit 72 outputs the detection result 62 that sets this gene as a DEG.

[0105] On the other hand, when the corrected p-value is greater than or equal to the pre-set significance level (no in step ST103), the determination unit 72 adopts the null hypothesis established in step ST101 and determines that there is no statistically significant difference between the gene expression level distribution 30 of the high-expression group HG and the gene expression level distribution 30 of the low-expression group LG (step ST105). In this case, the determination unit 72 does not detect this gene as a DEG. In addition, the significance level is, for example, 0.05.

[0106] As an example, as Figure 13 shown, the gene analysis screen 80 displayed on the display 12 under the control of the display control unit 55 has a first display area 81, a second display area 82, and a third display area 83. Basic information of gene analysis objects such as cell groups, sample numbers, and groupings is displayed in the first display area 81. A DEG detection button 84 is arranged in the second display area 82. By selecting the DEG detection button 84, each processing unit of the CPU 37 is activated to detect DEGs.

[0107] Before the DEG detection button 84 is selected, nothing is displayed in the third display area 83. On the other hand, in the case of detecting DEGs by selecting the DEG detection button 84, as shown in the figure, the display control unit 55 displays a list of groups of the identifiers 85 of the genes detected as DEGs and the gene expression amount distribution 30 of the genes detected as DEGs in the third display area 83. A scroll bar 86 is provided in the third display area 83, and the third display area 83 can be scrolled up and down for display. The user 14 can confirm the DEGs through the display in the third display area 83.

[0108] As an example, as Figure 14 shown, the DEGs detected as described above are used, for example, when selecting input data for the ER status prediction model 90. The ER status prediction model 90 is a machine learning model that takes the gene expression amounts of DEGs of cancer tissue-derived cells 20 collected from the breast 19 of a BRCA patient 18 whose ER high expression or low expression is unknown as input data and the ER status prediction result 91 as output data. The ER status prediction result 91 is either "high expression" when predicting that ER is highly expressed or "low expression" when predicting that ER is lowly expressed. By thus reducing the input data from the gene expression amounts of all genes to the gene expression amounts of DEGs, the prediction accuracy of the ER status prediction result 91 can be further improved.

[0109] Next, regarding the operation based on the above structure, as an example, reference is made to Figure 15 the flowchart shown. If the work program 45 is started in the information processing device 10, then as Figure 4 shown, the CPU 37 of the information processing device 10 functions as an acquisition unit 50, a distribution generation unit 51, an allocation unit 52, a search unit 53, a detection unit 54, and a display control unit 55.

[0110] First, in the acquisition unit 50, sample data groups 11 are acquired (step ST1000). The sample data groups 11 are output from the acquisition unit 50 to the distribution generation unit 51, the allocation unit 52, and the search unit 53.

[0111] On the display 12, a gene analysis screen 80 is displayed under the control of the display control unit 55. In the gene analysis screen 80, when the user 14 selects the DEG detection button 84 (Yes in step ST1100), as Figure 2 shown, in the distribution generation unit 51, a gene expression level distribution 30 of each candidate gene is generated based on the sample data set 11 (step ST1200). A set of gene expression level distributions 30 of each candidate gene, that is, a gene expression level distribution group 31, is output from the distribution generation unit 51 to the search unit 53, the detection unit 54, and the display control unit 55.

[0112] As Figure 5 shown, in the allocation unit 52, a cluster is allocated to each sample i (step ST1300). The allocation result 60 of the cluster is output from the allocation unit 52 to the search unit 53.

[0113] Next, as Figure 6 and Figure 7 shown, in the search unit 53, for each cluster, a negative binomial distribution P NB (p, r) suitable for the gene expression level distribution 30 is searched (step ST1400). The search result 61 of the negative binomial distribution P NB (p, r) is output from the search unit 53 to the detection unit 54.

[0114] Moreover, as Figures 8 - 12 shown, in the detection unit 54, DEGs are detected (step ST1500). The detection result 62 of the DEGs is output from the detection unit 54 to the display control unit 55.

[0115] Finally, as Figure 13 shown, under the control of the display control unit 55, in the third display area 83 of the gene analysis screen 80, an identifier 85 indicating the gene detected as a DEG and a group of gene expression level distributions 30 of the gene detected as a DEG are displayed in a list (step ST1600). Thus, the detection result 62 of the DEGs is provided for the user 14 to view.

[0116] As described above, the CPU 37 of the information processing apparatus 10 includes an allocation unit 52 and a search unit 53. The allocation unit 52 allocates, for each of a plurality of candidate genes that are candidates for DEGs, to each sample i that divides the cancer tissue-derived cell population 24 into a high expression group HG and a low expression group LG according to the ER status information 16, the cluster presumed to belong in the gene expression level distribution 30. The search unit 53, based on the allocation result 60 of the cluster, searches, for each of the plurality of candidate genes, a negative binomial distribution P NB (p, r) suitable for the gene expression level distribution 30 of the high expression group HG and the low expression group LG.

[0117] Therefore, it is possible to search for the negative binomial distribution P of the gene expression level distribution 30 of the cancer tissue-derived cell population 24 that is very suitable for a mixture of multiple subtypes and is expected to have multiple clusters due to multiple subtypes. NB (p, r). Thus, it is possible to reduce the probability of misjudgment when determining whether there is a statistically significant difference between the gene expression level distribution 30 of the high-expression group HG and the gene expression level distribution 30 of the low-expression group LG after determination. Therefore, even when a cell population with a mixture of multiple subtypes such as the cancer tissue-derived cell population 24 is the analysis object, it is possible to detect effective DEGs.

[0118] The detection unit 54 is based on the negative binomial distribution P NB (p, r) search result 61, for each of the multiple candidate genes, determines whether there is a statistically significant difference between the gene expression level distribution 30 of the high-expression group HG and the gene expression level distribution 30 of the low-expression group LG. The detection unit 54 detects a candidate gene determined to have a statistically significant difference between the gene expression level distribution 30 of the high-expression group HG and the gene expression level distribution 30 of the low-expression group LG as a DEG. Therefore, it is possible to detect the most reasonable and convincing gene as a DEG.

[0119] The determination unit 72 of the detection unit 54 calculates the log-likelihood function l for the high-expression group HG and the low-expression group LG respectively according to the gene expression level distribution 30 and the searched negative binomial distribution P NB (p, r). g (j) . The determination unit 72 calculates a score Λ representing the degree of difference between the high-expression group HG and the low-expression group LG according to the log-likelihood function l g (j) g and makes a determination based on the score Λ g NB . Therefore, it is possible to further detect the most reasonable and convincing gene as a DEG.

[0120] The first probability distribution searched by the search unit 53 is the negative binomial distribution P NB (p, r). Therefore, it is possible to perform the detection of DEGs consistent with the view that the gene expression level distribution 30 follows the negative binomial distribution P NB (p, r). In addition, the first probability distribution is not limited to the exemplified negative binomial distribution P NB (p, r). It can be a binomial distribution, a Poisson distribution, etc.

[0121] [Second Embodiment]

[0122] In the above first embodiment, clusters were assigned to each sample i based on the gene expression level of the gene HER2, but it is not limited to this. As an example, as Figure 16As in the distribution unit 95 shown, clusters can be assigned to each sample i according to the one-hot encoded vector z sampled from the categorical distribution P cat (π (j) ). The categorical distribution P i is a distribution that follows the Dirichlet distribution P cat (α) with parameters α (j) , 2, ……, (j) and is a K-dimensional vector π k=1 . As indicated by the superscript in (j), the categorical distributions P K (π Dir ) are used for the high-expression group HG and the low-expression group LG, respectively. The categorical distribution P cat (π (j) ) is an example of the "second probability distribution" involved in the technology of the present invention. cat (π (j) ) is an example of the "second probability distribution" involved in the technology of the present invention.

[0123] Thus, in the second embodiment, the distribution unit 95 uses the categorical distribution P cat (π (j) ) to perform cluster assignment. Therefore, as in the gene expression level distribution 30 related to the cancer tissue-derived cells 20 collected from the BRCA patient 18 in the first embodiment described above, cluster assignment can be performed even without prior knowledge of the clusters. In addition, the second probability distribution is not limited to the categorical distribution P cat (π (j) ), and can also be, for example, a multinomial distribution or the like. Also, in the absence of prior knowledge of the clusters, clusters can be randomly assigned instead of using the categorical distribution P cat (π (j) ) or other second probability distributions. And non-parametric methods such as the k-nearest neighbor method can also be used to assign clusters.

[0124] [Third Embodiment]

[0125] As an example, as shown in Figure 17 and Figure 18 , in the third embodiment, the search result 61 from the search unit 101 is fed back to the distribution unit 100. The distribution unit 100 changes the cluster assignment method based on the search result 61 fed back from the search unit 101 (step ST1420), and performs cluster assignment again with the changed assignment method (step ST1300). The search unit 101 performs a search for the negative binomial distribution P NB (p, r) suitable for the gene expression level distribution 30 again based on the cluster assignment result 60 obtained by the re-cluster assignment (step ST1400). When the distribution unit 100 and the search unit 101 do not satisfy the end condition (No in step ST1410), the following processing is repeated: performing cluster assignment and searching for the negative binomial distribution P NB(Search for (p, r)).

[0126] Regarding the change in the clustering assignment method of step ST1420, for example, if it is the assignment unit 95 of the above-described second embodiment, by changing the class distribution P based on the search result 61 cat (π (j) ) The parameter π (j) is performed. And in step ST1410, for example, in the negative binomial distribution P NB (p, r) searched last time and the negative binomial distribution P NB (p, r) searched this time, if the difference in the parameters p and r is less than a preset threshold, it is considered that the end condition is satisfied. Or, when the above processing is repeated a preset number of times, the end condition may also be satisfied.

[0127] Thus, in the third embodiment, the assignment unit 100 and the search unit 101 perform at least once the following processing: receiving the search result 61 of the negative binomial distribution P NB (p, r) to re-perform the clustering assignment, and based on the re-performed clustering assignment result 60, re-search for the negative binomial distribution P NB (p, r). If such processing is performed at least once, the clustering assignment method and the search for the negative binomial distribution P NB (p, r) will be improved, and it is possible to search for a negative binomial distribution P NB (p, r) that is more suitable for the gene expression level distribution 30. Therefore, it is possible to further reduce the probability of misjudgment when determining whether there is a statistically significant difference between the gene expression level distribution 30 of the high-expression group HG and the gene expression level distribution 30 of the low-expression group LG after the determination.

[0128] [Fourth Embodiment]

[0129] In the above-described first embodiment, the clustering number setting value K is set to one type, but it is not limited thereto. As an example, as Figure 19 shown by the assignment unit 105, it is possible to receive the input of the clustering number setting information 106A, 106B, and 106C that respectively include three types of clustering number setting values (K = 2, K = 3, and K = 4). The three types of clustering number setting values can be input by the user 14 via the input device 13, or can be automatically input by the CPU 37.

[0130] The assignment unit 105 assigns clusters to each sample i based on the three types of clustering number setting values, and outputs the assignment results 107A, 107B, and 107C. The assignment result 107A corresponds to the clustering number setting value K = 2, the assignment result 107B corresponds to the clustering number setting value K = 3, and the assignment result 107C corresponds to the clustering number setting value K = 4.

[0131] The search unit 108 searches for the negative binomial distribution P NB (p, r) based on the allocation result 107A and outputs its search result 109A. Also, the search unit 108 searches for the negative binomial distribution P NB (p, r) based on the allocation result 107B and outputs its search result 109B. In addition, the search unit 108 searches for the negative binomial distribution P N B (p, r) based on the allocation result 107C and outputs its search result 109C.

[0132] As an example, as Figure 20 shown, the CPU of the information processing apparatus according to the fourth embodiment has an evaluation unit 115. The search results 109A, 109B, and 109C from the search unit 108 are input to the evaluation unit 115. Also, the gene expression level distribution group 31 from the distribution generation unit 51 is input to the evaluation unit 115. The evaluation unit 115 evaluates the fitness of the negative binomial distribution P NB (p, r) of the search results 109A, 109B, and 109C to the gene expression level distribution 30 using, for example, an information amount criterion such as the widely applicable information criterion (WAIC: Widely Applicable Information Criterion), and outputs the evaluation result 116 to the allocation unit 105.

[0133] The allocation unit 105 determines the cluster number setting value K that can search for the negative binomial distribution P NB (p, r) that is most suitable for the gene expression level distribution 30 based on the evaluation result 116. In Figure 20 it, an example is shown in which the evaluation result of the search result 109A when the cluster number setting value K = 2 is "no", the evaluation result of the search result 109B when the cluster number setting value K = 3 is "good", and the evaluation result of the search result 109C when the cluster number setting value K = 4 is "yes", and thus it is determined that the cluster number setting value K = 3.

[0134] In this way, in the fourth embodiment, the allocation unit 105 changes the number of allocated clusters. The search unit 108 searches for the negative binomial distribution P NB (p, r) for each changed number. The allocation unit 105 can determine, from the changed numbers, the negative binomial distribution P NB (p, r) that can search for the one most suitable for the gene expression level distribution 30 based on the search results 109A, 109B, and 109C of the negative binomial distribution P NBThe number of (p, r). Therefore, even when it is different from the above-described first embodiment and there is no opinion on the number of clusters, an appropriate number of clusters can be set. Therefore, the probability of misjudgment when determining whether there is a statistically significant difference in the gene expression level distribution 30 of the high-expression group HG and the gene expression level distribution 30 of the low-expression group LG after determination can be further reduced.

[0135] In addition, the number of clusters to be changed is not limited to the three exemplified. For example, it can be nine kinds of K = 2, 3, 4,..., 8, 9, 10, or it can be two kinds of K = 2, 3. When the evaluation results of two or more numbers of clusters are "good", considering the simplification of processing, it is preferable to determine the smaller value. Also, this fourth embodiment can be implemented in combination with the above-described second embodiment and / or the above-described third embodiment.

[0136] [Fifth_1 Embodiment]

[0137] As an example, as Figure 21 shown, the CPU of the information processing device of the fifth_1 embodiment has a generation unit 120. The search result 61 from the search unit 53 is input to the generation unit 120. The generation unit 120 generates simulated gene expression data (hereinafter, referred to as simulated gene expression data) 121 by simulation based on the search result 61. In theory, the generation unit 120 can generate simulated gene expression data 121 without limit. Here, simulation means a process of tracing the following series of processes: generating a gene expression level distribution 30 from the gene expression data 17, assigning clusters to each sample i, and searching for a negative binomial distribution P suitable for the gene expression level distribution 30 based on the assignment result 60 NB (p, r).

[0138] Thus, in the fifth_1 embodiment, the generation unit 120 generates simulated gene expression data 121 by simulation based on the searched negative binomial distribution P NB (p, r). By analyzing the simulated gene expression data 121, the characteristics of the cancer tissue-derived cell population 24 can be clarified. Also, the average value of the gene expression levels of the simulated gene expression data 121 is calculated, and the calculated average value is compared with the gene expression levels of the gene expression data 17, whereby the abnormal values of the gene expression levels caused by measurement noise or the like hidden in the gene expression data 17 can be extracted. Thus, on the basis of removing the extracted abnormal values from the gene expression data 17, the search for the negative binomial distribution P NB (p, r) can be performed, and a more suitable negative binomial distribution P NB (p, r) can be searched.

[0139] [Fifth_2 Embodiment]

[0140] As an example, as Figure 22 shown, in the 5_2nd embodiment, the generation unit 120 only generates the simulated gene expression amount data 121 for the DEGs. In this case, in addition to the search result 61, the detection result 62 from the detection unit 54 is also input to the generation unit 120.

[0141] Thus, in the 5_2nd embodiment, the generation unit 120 only generates the simulated gene expression amount data 121 for the DEGs through simulation. By analyzing the simulated gene expression amount data 121 of the DEGs, the characteristic differences between other genes and the DEGs can be clarified. And, compared with the above 5_1st embodiment that generates the simulated gene expression amount data 121 for all candidate genes, the processing load of the simulation can be reduced.

[0142] It can be configured such that the user 14 can select the above 5_1st embodiment that generates the simulated gene expression amount data 121 for all candidate genes and the present 5_2nd embodiment that only generates the simulated gene expression amount data 121 for the DEGs. And, it can also be configured such that the user 14 can specify the genes for which the simulated gene expression amount data 121 is generated.

[0143] [Examples]

[0144] The following describes the examples of the technology of the present invention. In this example, the ER status information 16 of 18 BRCA patients was obtained from the Legacy Archive of the GDC (Genomic Data Commons) managed by the National Cancer Institute (NCI) in the United States. And, the gene expression amount data 17 of 18 BRCA patients was obtained from The Cancer Genome Atlas (TCGA) operated by the NCI. The gene expression amount data 17 is data calculated by the following method: using the HT (High-throughput) Seq module to align the RNA-Seq data obtained from the cancer tissue-derived cells 20 of 18 BRCA patients to the reference genome.

[0145] Figure 23 Table 125 shows the details of the groups and subtypes of sample i in this example. The number of samples is 1222. Moreover, for 61 of the sample is, the ER status information 16 was not given and there is no label (NaN (Not a Number)), and for the remaining 1161 sample is, the ER status information 16 was given. And, for 255 of the sample is, the subtype information was not given and there is no label, and for the remaining 967 sample is, the subtype information was given. This subtype information utilized the information provided by the TCGA Pan-Cancer Atlas project.

[0146] Among the genes registered with gene expression data 17, 19,140 genes encoding proteins were selected as candidate genes. Similarly to the first embodiment described above, the number of cluster setting value K was set to 2. Clusters were also assigned to sample i with the subtype "Normal-like" and unlabeled sample i. In the search for the negative binomial distribution P NB (p, r), the variational Bayesian method was used. And, as a comparative example, the DEGs were detected using the existing detection method (hereinafter, referred to as the prior art) described in Non-Patent Document 1.

[0147] Figure 24 shows the gene expression level distributions 30 of the genes ESR1, GATA3, and TBC1D9, each of which was detected as a DEG in both the prior art and the technology of the present invention. There are significant differences in the gene expression level distributions 30 of these genes ESR1, GATA3, and TBC1D9 between the high-expression group HG and the low-expression group LG. And the genes ESR1, GATA3, and TBC1D9 have been reported as DEGs in many literatures. In particular, the gene ESR1 is indeed a gene related to the production of ERα. The gene GATA3 is recognized as a breast cancer-specific marker gene. And recently, for example, in the following Document A, etc., the gene TBC1D9 has been reported as an important gene in the subtype "Basal-like".

[0148] Document A <C. Kothari, et al., "Tbc1d9: An important modulator of tumorigenesis in breast cancer." Cancers, 13(14): 3557, 2021.>

[0149] Thus, both the prior art and the technology of the present invention can detect genes consistent with previous studies as DEGs. Therefore, it can be said that the technology of the present invention is a method for detecting DEGs with validity not inferior to the prior art.

[0150] Figure 25 shows the gene expression level distributions 30 of the genes C17orf64, CYP1A1, and SMG8, each of which was detected as a DEG in the prior art but not in the technology of the present invention. These genes C17orf64, CYP1A1, and SMG8 are related to Figure 24Compared with the genes ESR1, GATA3, and TBC1D9 shown, no significant differences were found in the gene expression level distributions 30 of the high-expression group HG and the low-expression group LG. Also, no literature reporting genes C17orf64, CYP1A1, and SMG8 as DEGs has been found so far. Therefore, it can be said that genes C17orf64, CYP1A1, and SMG8 are genes mismeasured as DEGs in the prior art. Thus, the technology of the present invention can exclude genes mismeasured as DEGs in the prior art.

[0151] Figure 26 shown compared with Figure 25 Conversely, in the technology of the present invention, the gene expression level distributions 30 of genes CLDN6, EN1, and TTLL4, which are detected as DEGs but not detected as DEGs in the prior art, are shown. These genes CLDN6, EN1, and TTLL4 are the same as Figure 24 the genes ESR1, GATA3, and TBC1D9 shown, and there are significant differences in the gene expression level distributions 30 of the high-expression group HG and the low-expression group LG. Also, in particular, multiple clusters appear in the gene expression level distributions 30 of genes CLDN6 and EN1.

[0152] In addition, genes CLDN6, EN1, and TTLL4 are reported as important genes in the following literatures B, C, and D, etc.

[0153] Literature B <Y. Liu, et al., “DNA methylation of claudin-6 promotes breast cancer cell migration and invasion by recruiting mecp2 and deacetylating h3ac and h4ac.” Journal of experimental & clinical cancer research, 35(1): 1-13, 2016.>

[0154] Literature C <G. Peluffo, et al., “En1 is a transcriptional dependency in triple-negative breast cancer associated with brain metastasis.” Cancer research, 79(16): 4173-4183, 2019.>

[0155] Document D <J. Arnold, et al., “Tubulin tyrosine ligase like 4 (ttl l4) overexpression in breast cancer cells is associated with brain metastasis and alters exosome biogenesis.” Journal of Experimental & Clinical Cancer Research, 39(1): 1 - 15, 2020.>

[0156] Therefore, the genes CLDN6, EN1, and TTLL4 can be considered as genes that were missed as DEGs in the prior art. Thus, the technology of the present invention can detect genes that were missed as DEGs in the prior art as DEGs.

[0157] From the above, it can be confirmed that the technology of the present invention is as follows: It can equally detect genes that were detected as DEGs in the prior art as DEGs, can exclude genes that were misdetected as DEGs in the prior art, and can detect genes that were missed as DEGs in the prior art as DEGs. Therefore, it is proved that compared with the prior art, the technology of the present invention can detect effective DEGs even when analyzing a cell population in which multiple subtypes co - exist.

[0158] The cell population is not limited to the exemplified cancer - tissue - derived cell population 24. It can also be a population of HEK293 cells (Human Embryonic Kidney 293 Cells) of human - derived cells used for synthesizing therapeutic proteins or viruses used in gene therapy. When it comes to HEK293 cells, the cell characteristic of concern is the contribution degree to the synthesis of therapeutic proteins or viruses used in gene therapy. And the cells constituting the cell population are not limited to human cells. For example, it can be a population of CHO cells. When it comes to CHO cells, the cell characteristic of concern is the contribution degree to antibody production.

[0159] As shown in the example, the information - processing device 10 can be a personal computer set up in a gene - analysis facility, or a server computer set up in a data center independent of the gene - analysis facility.

[0160] In the case where the information processing apparatus 10 is constituted by a server computer, a sample data group 11 is transmitted from a personal computer provided in each gene analysis facility to the server computer via a network such as the Internet. The server computer transmits various screens such as a gene analysis screen 80 in a format of screen data for web transmission generated by a markup language such as XML (Extensible Markup Language) to the personal computer. The personal computer reproduces the screen displayed on the web browser based on the screen data and displays it on the display. In addition, other data description languages such as JSON (Javascript (registered trademark) Object Notation) can be used instead of XML.

[0161] The hardware structure of the computer constituting the information processing apparatus 10 related to the technology of the present invention can be variously modified. For example, for the purpose of improving processing power and reliability, the information processing apparatus 10 can also be constituted by multiple computers separated as hardware. For example, the functions of the acquisition unit 50, the distribution generation unit 51, and the display control unit 55 and the functions of the distribution unit 52, the search unit 53, and the detection unit 54 are shared among two computers. In this case, the information processing apparatus 10 is constituted by two computers.

[0162] Thus, the hardware structure of the computer of the information processing apparatus 10 can be appropriately changed according to the required performance such as processing power, security, and reliability. In addition, not limited to hardware, for application programs such as the work program 45, for the purpose of ensuring security and reliability, of course, it can also be duplicated or distributed and stored in multiple storage devices.

[0163] In the above-described embodiments, for example, as the hardware structure of the processing units (Processing Unit) that execute various processes such as the acquisition unit 50, the distribution generation unit 51, the allocation units 52, 95, 100, and 105, the search units 53, 101, and 108, the detection unit 54, the display control unit 55, the function calculation unit 70, the score calculation unit 71, the determination unit 72, the evaluation unit 115, and the generation unit 120, various processors (Pro cessor) shown below can be used. As described above, in addition to the general-purpose processor, i.e., the CPU 37, which executes software (working program 45) to function as various processing units, various processors also include programmable logic devices (Programmable Logic Device: PLD) such as FPGA (Field Programmable Gate Array), which can change the circuit structure after manufacturing, and dedicated circuits such as ASIC (Application Specific Integrated Circuit), which have a circuit structure specifically designed to execute specific processes.

[0164] One processing unit can be composed of one of these various processors, or can be composed of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs and / or a combination of a CPU and an FPGA). Also, multiple processing units can be composed of one processor.

[0165] As an example of composing multiple processing units with one processor, first, there is the following method: as represented by computers such as clients and servers, a combination of one or more CPUs and software constitutes one processor, and this processor functions as multiple processing units. Second, there is the following method: as represented by a system on chip (System On Chip: SoC), etc., a processor that realizes the overall function of a system including multiple processing units through one IC (Integrated Circuit) chip is used. In this way, various processing units are composed using one or more of the above various processors as the hardware structure.

[0166] Furthermore, as the hardware structure of these various processors, more specifically, circuitry composed of circuit elements such as semiconductor elements can be used.

[0167] Based on the above description, the technology described in the following appended items can be grasped.

[0168] [Appended Item 1]

[0169] An information processing apparatus performs the following processing: based on gene expression data of a cell population in which multiple subtypes coexist, it detects expression-variant genes that are specifically expressed with respect to a cell characteristic of interest.

[0170] The information processing apparatus includes a processor.

[0171] The processor performs the following processing:

[0172] For each of a plurality of candidate genes that are candidates for the expression-variant genes, it assigns, to each sample that divides the cell population into two groups according to the cell characteristic of interest, a cluster that is presumed to belong in the gene expression distribution.

[0173] Based on the cluster assignment result, for each of the plurality of candidate genes, it separately searches for a first probability distribution that fits the gene expression distributions of the two groups.

[0174] [Supplementary Note 2]

[0175] The information processing apparatus according to Supplementary Note 1, wherein

[0176] The processor performs the following processing:

[0177] Based on the search result of the first probability distribution, for each of the plurality of candidate genes, it determines whether there is a statistically significant difference in the gene expression distributions of the two groups.

[0178] It detects, as the expression-variant genes, the candidate genes for which it is determined that there is a statistically significant difference in the gene expression distributions of the two groups.

[0179] [Supplementary Note 3]

[0180] The information processing apparatus according to Supplementary Note 2, wherein

[0181] The processor performs the following processing:

[0182] Based on the gene expression distribution and the searched first probability distribution, it calculates a representative index value for each of the two groups.

[0183] Based on the representative index values, it calculates a score indicating the degree of difference between the two groups.

[0184] It makes the determination based on the score.

[0185] [Supplementary Note 4]

[0186] The information processing apparatus according to any one of Supplementary Notes 1 to 3, wherein

[0187] The processor performs the following processing at least once:

[0188] Receives the search results of the first probability distribution and re - performs the allocation of the clusters, and based on the allocation results of the re - performed clusters, re - performs the search for the first probability distribution.

[0189] [Supplementary Note Item 5]

[0190] The information processing apparatus according to any one of Supplementary Note Items 1 to 4, wherein

[0191] The processor performs the following processing:

[0192] Changes the number of the allocated clusters, and performs the search for the first probability distribution for each changed number;

[0193] Based on the search results of the first probability distribution, determines the number of the first probability distributions that can search for the distribution most suitable for the gene expression level from the changed numbers.

[0194] [Supplementary Note Item 6]

[0195] The information processing apparatus according to any one of Supplementary Note Items 1 to 5, wherein

[0196] The processor performs the following processing:

[0197] Uses the second probability distribution to perform the allocation of the clusters.

[0198] [Supplementary Note Item 7]

[0199] The information processing apparatus according to any one of Supplementary Note Items 1 to 6, wherein

[0200] The first probability distribution is a negative binomial distribution.

[0201] [Supplementary Note Item 8]

[0202] The information processing apparatus according to any one of Supplementary Note Items 1 to 7, wherein

[0203] The processor performs the following processing:

[0204] Based on the searched first probability distribution, generates simulated gene expression level data through simulation.

[0205] [Supplementary Note Item 9]

[0206] The information processing apparatus according to Supplementary Note Item 8, wherein

[0207] The processor performs the following processing:

[0208] The gene expression level data of the expressed variant gene is generated only through simulation.

[0209] The technology of the present invention can also appropriately combine the above various embodiments and / or various modification examples. And, of course, without departing from the gist, it is not limited to the above embodiments, and various structures can be adopted. In addition, the technology of the present invention relates not only to programs but also to storage media for storing non-volatile programs.

[0210] The content described and the content shown above are detailed descriptions of parts related to the technology of the present invention, and are only examples of the technology of the present invention. For example, the description of the above structure, function, action and effect is an example of the structure, function, action and effect of the parts related to the technology of the present invention. Therefore, it should be understood that within the scope of not departing from the gist of the technology of the present invention, for the content described and the content shown above, unnecessary parts can be deleted, and new elements can be added or replaced. And, in order to avoid complication and facilitate understanding of the parts related to the technology of the present invention, in the content described and the content shown above, descriptions related to common technical knowledge that does not require special explanation are omitted on the basis that the technology of the present invention can be implemented.

[0211] In this specification, the meaning of "A and / or B" is the same as "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. And, in this specification, the same idea as "A and / or B" is also applicable when three or more cases are combined with "and / or".

[0212] All documents, patent applications, and technical standards described in this specification are incorporated into this specification by reference, in the same manner as if each individual document, patent application, and technical standard were specifically and individually described as being incorporated by reference.

Claims

1. An information processing apparatus performs the following processing: based on gene expression data of a cell population in which multiple subtypes coexist, it detects expression-altered genes that are specifically expressed with respect to a cell characteristic of interest. The information processing apparatus includes a processor. The processor performs the following processing: For each of a plurality of candidate genes that are candidates for the expression-altered genes, to each sample that divides the cell population into two groups according to the cell characteristic of interest, it assigns a cluster that is presumed to belong in the gene expression distribution. Based on the cluster assignment result, for each of the plurality of candidate genes, it separately searches for a first probability distribution that fits the gene expression distributions of the two groups.

2. The information processing apparatus according to claim 1, wherein the processor performs the following processing: Based on the search result of the first probability distribution, for each of the plurality of candidate genes, it determines whether there is a statistically significant difference in the gene expression distributions of the two groups; It detects as the expression-altered genes the candidate genes for which it is determined that there is a statistically significant difference in the gene expression distributions of the two groups.

3. The information processing apparatus according to claim 2, wherein the processor performs the following processing: According to the gene expression distribution and the searched first probability distribution, it calculates a representative index value for each of the two groups; According to the representative index values, it calculates a score indicating the degree of difference between the two groups; It makes the determination based on the score.

4. The information processing apparatus according to claim 1, wherein the processor performs the following processing at least once: Receives the search result of the first probability distribution and performs the cluster assignment again, and based on the cluster assignment result of the re - performed assignment, performs the search for the first probability distribution again.

5. The information processing apparatus according to claim 1, wherein the processor performs the following processing: Changes the number of the assigned clusters, and performs the search for the first probability distribution for each changed number; Based on the search result of the first probability distribution, it determines from the changed numbers the number that can search for the first probability distribution that best fits the gene expression distribution.

6. The information processing apparatus according to claim 1, wherein the processor performs the following processing: Uses a second probability distribution to perform the cluster assignment.

7. The information processing apparatus according to claim 1, wherein the first probability distribution is a negative binomial distribution.

8. The information processing apparatus according to claim 1, wherein the processor performs the following processing: Based on the searched first probability distribution, it generates simulated gene expression data through simulation.

9. The information processing apparatus according to claim 8, wherein the processor performs the following processing: Generates only the simulated gene expression data of the expression-altered genes through simulation.

10. A working method of an information processing device, the information processing device performing the following processing: based on gene expression data of a cell population in which multiple subtypes coexist, detecting expression-variable genes that exhibit specific expression with respect to a cell characteristic of interest, The working method of the information processing device includes the following steps: For each of a plurality of candidate genes that are candidates for the expression-variable genes, assigning, to each sample that divides the cell population into two groups according to the cell characteristic of interest, a cluster presumed to belong in the gene expression distribution; and Based on the cluster assignment result, for each of the plurality of candidate genes, respectively searching for a first probability distribution that fits the gene expression distributions of the two groups.

11. A working program of an information processing device, the information processing device performing the following processing: based on gene expression data of a cell population in which multiple subtypes coexist, detecting expression-variable genes that exhibit specific expression with respect to a cell characteristic of interest, The working program of the information processing device causes a computer to execute a process including the following steps: For each of a plurality of candidate genes that are candidates for the expression-variable genes, assigning, to each sample that divides the cell population into two groups according to the cell characteristic of interest, a cluster presumed to belong in the gene expression distribution; and Based on the cluster assignment result, for each of the plurality of candidate genes, respectively searching for a first probability distribution that fits the gene expression distributions of the two groups.