Information processing device, operation method for information processing device, and operation program for information processing device

The information processing device identifies super-enhancers as intervention sites for genome editing by analyzing histone modification and gene expression differences, improving the efficiency and reliability of cell phenotype transitions.

WO2025225426A1PCT designated stage Publication Date: 2025-10-30FUJIFILM CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/014548
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-26
Filing Date
2025-04-11
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Current techniques for genome editing or intervention on super-enhancers are impractical due to their large number, making it difficult to identify effective sites for controlling gene expression in cell phenotype transitions.

Method used

An information processing device and method that analyze differences in histone modification levels and gene expression between cell types to identify super-enhancers and their locations as intervention sites, using a processor to extract enhancers and super-enhancers as candidates for genome editing.

Benefits of technology

Facilitates the identification of more practical intervention sites for genome editing by narrowing down super-enhancers, enhancing the efficiency and reliability of cell phenotype transitions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025014548_30102025_PF_FP_ABST
    Figure JP2025014548_30102025_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device according to the present invention includes a processor. The processor acquires a first feature amount of a first type of cell having a first phenotype and a second feature amount of a second type of cell having a second phenotype different from the first type of cell, identifies a portion of a genomic region related to the difference between the first type of cell and the second type of cell on the basis of the difference between the first feature amount and the second feature amount, and presents information related to the identified portion.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, operating method for information processing device, and operating program for information processing device

[0001] The technology of the present disclosure relates to an information processing device, an operating method for an information processing device, and an operating program for an information processing device.

[0002] Recently, there has been active research into techniques for controlling cell phenotype, such as by causing immature differentiated cardiomyocytes to mature and differentiate into mature differentiated cardiomyocytes. Among enhancers within genome regions that function during gene transcription by RNA (ribonucleic acid) polymerase, superenhancers, which are enhancers that exist in close proximity to other enhancers, have attracted attention. [Violaine Saint-Andre et al. "Models of human core transcriptional regulatory circuits," Genome Research, February 3, 2016 (hereinafter referred to as Non-Patent Document 1) describes a technique for extracting super enhancers and identifying the network of super enhancers and core transcription factors that bind to them (referred to as CRC: Core Regulatory Circuitry in Non-Patent Document 1).

[0003] The number of super-enhancers is on the order of several thousand. For this reason, it is not practical to perform genome editing or other interventions on all super-enhancers. A more practical approach would be to further narrow down the super-enhancers, identify the so-called masterminds that control the expression of important genes in the genomic region, and present related information such as their locations as intervention sites for genome editing or other interventions.

[0004] One embodiment of the technology of the present disclosure provides an information processing device, an operating method for an information processing device, and an operating program for an information processing device that are capable of presenting more practical intervention sites for genome editing, etc.

[0005] The information processing device of the present disclosure includes a processor, which acquires a first feature of a first type of cell having a first phenotype and a second feature of a second type of cell having a second phenotype different from the first type of cell, and identifies a portion of a genomic region related to the difference between the first type of cell and the second type of cell based on the difference between the first feature and the second feature, and presents information related to the identified portion.

[0006] Preferably, the processor extracts enhancers within the genomic region as candidate portions.

[0007] Preferably, the processor identifies the portion based on a difference between a first activity level of the enhancer in the first type of cell as the first feature and a second activity level of the enhancer in the second type of cell as the second feature.

[0008] Preferably, the processor identifies the portion based on whether the enhancer contributes to the expression of the transcription factor, in addition to the difference between the first feature amount and the second feature amount.

[0009] Preferably, the processor extracts super-enhancers as candidate moieties based on the activity of the enhancer and the location of the enhancer in the genomic region.

[0010] Preferably, the processor identifies the portion based on the difference between the first feature amount and the second feature amount as well as the difference between the gene expression levels of the first type cells and the second type cells.

[0011] Preferably, the processor extracts enhancers within the genomic region as candidate segments, and the gene expression levels are those of target genes of the enhancers.

[0012] Preferably, the processor identifies the portion based on information about the topological domain in addition to the difference between the first feature amount and the second feature amount.

[0013] The first activity and the second activity are preferably quantified by any one of the amount of histone modification, the amount of a transcription factor or a transcription coactivator, and the amount of a chromatin structure regulator.

[0014] Preferably, the amount of histone modification involves acetylation of the lysine 27 residue of histone H3 or monomethylation of the lysine 4 residue of histone H3.

[0015] The method of operating the information processing device disclosed herein includes acquiring a first feature of a first type of cell having a first phenotype and a second feature of a second type of cell having a second phenotype different from the first type of cell, identifying a portion of a genomic region related to the difference between the first type of cell and the second type of cell based on the difference between the first feature and the second feature, and presenting information related to the identified portion.

[0016] The operating program of the information processing device disclosed herein causes a computer to execute processes including acquiring a first feature amount of a first type of cell having a first phenotype and a second feature amount of a second type of cell having a second phenotype different from the first type of cell, identifying a portion of a genomic region related to the difference between the first type of cell and the second type of cell based on the difference between the first feature amount and the second feature amount, and presenting information related to the identified portion.

[0017] According to the technology of the present disclosure, it is possible to provide an information processing device, an operating method for an information processing device, and an operating program for an information processing device that are capable of presenting more practical intervention sites for genome editing, etc.

[0018] 1 is a diagram showing an information processing system. FIG. 2 is a diagram showing transcription control. FIG. 3 is a diagram showing first activity information. FIG. 4 is a diagram showing second activity information. FIG. 5 is a diagram showing related information. FIG. 6 is a block diagram showing a computer constituting an information processing device and an operator terminal. FIG. 7 is a block diagram showing a processing unit of a CPU of the information processing device. FIG. 8 is a diagram showing a detailed configuration of an extraction unit. FIG. 9 is a diagram showing a detailed configuration of a related information generation unit. FIG. 10 is a block diagram showing a processing unit of a CPU of the operator terminal. FIG. 11 is a diagram showing an activity information input screen. FIG. 12 is a diagram showing a related information display screen. FIG. 13 is a flowchart showing a processing procedure of an information processing device. FIG. 14 is a diagram showing a second embodiment in which super enhancers that contribute to the expression of transcription factors in which differentially expressed genes are involved are identified as parts of genome regions related to the difference between first and second type cells. FIG. 15 is a diagram showing a third embodiment in which super enhancers whose target genes are differentially expressed genes are identified as parts of genome regions related to the difference between first and second type cells. FIG. 16 is a diagram showing a fourth embodiment in which super enhancers whose three-dimensional proximity based on topological domain information is within a sixth threshold range are identified as parts of genome regions related to the difference between first and second type cells. FIG. 17 is a diagram showing another example of first activity information. FIG. 18 is a diagram showing yet another example of first activity information. FIG. 19 is a diagram showing another example of related information.

[0023] Figure 1 shows how epithelial cells change into mesenchymal cells.

[0024] Figure 2 shows a table illustrating the structure of ChIP-Seq data in an example.

[0025] Figure 3 shows a table illustrating the structure of RNA-Seq data in an example.

[0026] Figure 4 shows a group of graphs illustrating the time course of expression levels of transcription factors identified in an example.

[0019] [First Embodiment] As shown in FIG. 1 as an example, an information processing system 10 is a system for processing information related to cells, and includes an information processing device 12 and an operator terminal 13. The information processing device 12 and the operator terminal 13 are connected via a network 14. The operator terminal 13 is installed in a cell research institution. The operator terminal 13 is operated by an operator OP who is involved in practical work at the research institution. The network 14 is, for example, a WAN (Wide Area Network) such as the Internet or a public communication network. Note that while only one operator terminal 13 is connected to the information processing device 12 in FIG. 1, in reality, multiple operator terminals 13 from multiple research institutions are connected to the information processing device 12.

[0020] The operator terminal 13 transmits a specific request 15 to the information processing device 12. The specific request 15 includes first activity information 161 and second activity information 162. The first activity information 161 is information on a first activity of an enhancer 21 (see FIG. 2 ) in a first type of cell 171 having a first phenotype. The second activity information 162 is information on a second activity of an enhancer 21 in a second type of cell 172 having a second phenotype different from that of the first type of cell 171. The first type of cell 171 and the second type of cell 172 may be any type of cell as long as they have different phenotypes, such as pluripotent stem cells, mesenchymal stem cells, skin cells, blood cells, endodermal cells, mesodermal cells, ectodermal cells, nerve cells, glial cells, cardiomyocytes, hepatocytes, retinal cells, etc. Although not shown in the figure, the specific request 15 also includes a terminal ID (Identification Data) for uniquely identifying the operator terminal 13 that is the sender of the specific request 15 .

[0021] As an example, as shown in Figure 2, enhancer 21 is present in DNA (deoxyribonucleic acid) 20, which constitutes the genome. Enhancer 21 is the binding site for activator 22. When RNA polymerase 23 transcribes gene 24 into mRNA (messenger RNA), transcription factor 25 binds to enhancer 21 via activator 22. Enhancer 21 has elements such as its location in the genome region, activity, the transcription factor 25 it binds to, and the target gene. At promoter 26, transcription factor 25 (more precisely, a general transcription factor) binds to RNA polymerase 23 to form a transcription initiation complex, initiating transcription of gene 24. Based on gene 24 thus transcribed into mRNA, various proteins essential for cellular activity are synthesized in ribosomes. The transcribed gene 24 differs depending on the cellular phenotype, and the enhancer 21 that functions during transcription also differs depending on the cellular phenotype. Therefore, if it is possible to identify the enhancer 21, which is the so-called mastermind that controls the expression of important genes in a genomic region, and / or the transcription factor 25 that binds to the enhancer 21, it will be possible to present an effective genome editing or epigenome editing intervention site for inducing or suppressing a change from the first type of cell 171 to the second type of cell 172. More specifically, genome editing involves, for example, partial or complete replacement or deletion of the enhancer 21. Alternatively, it is also possible to intervene by introducing a gene 24 under the control of the enhancer 21, thereby enhancing the expression of the gene 24.

[0022] As an example, as shown in FIGS. 3 and 4 , the first activity information 161 and the second activity information 162 include, as the first activity and the second activity, the amount of histone modification obtained by performing ChIP-Seq (Chromatin Immunoprecipitation-Sequencing) analysis on the first type of cell 171 and the second type of cell 172. As shown in graph 301, the amount of histone modification is the amount (vertical axis) related to acetylation of the 27th lysine residue of histone H3 (H3K27ac) at each position (horizontal axis) of DNA 20. Furthermore, as shown in graph 302, the amount of histone modification is the amount (vertical axis) related to monomethylation of the 4th lysine residue of histone H3 (H3K4me1) at each position (horizontal axis) of DNA 20. The amount of histone modification in the first activity information 161 is an example of a "first feature" according to the technology of the present disclosure. The amount of histone modification in the second activity information 162 is an example of a "second feature amount" according to the technology of the present disclosure.

[0023] Histones are proteins that make up nucleosomes in the chromatin structure. Histone modification is expressed by the type of modification at which amino acid, counting from the N-terminus of the histone, is modified. If the amount of acetylation at the 27th lysine residue of histone H3 or the amount of monomethylation at the 4th lysine residue of histone H3 is high, the chromatin structure is considered to be loose (active for transcription). Therefore, the amount of histone modification can indicate whether the part corresponding to each position in DNA 20 is functioning as an enhancer 21.

[0024] Returning to FIG. 1 , when the information processing device 12 receives the identification request 15, it identifies a portion of the genome region related to the difference between the first type cell 171 and the second type cell 172 (the change from the first type cell 171 to the second type cell 172) (hereinafter referred to as a master regulator MR, meaning a control factor that is the main cause of the difference between the first type cell 171 and the second type cell 172). The information processing device 12 distributes related information 18 of the identified master regulator MR to the operator terminal 13 that is the sender of the identification request 15. When the information processing device 12 receives the related information 18, the operator terminal 13 makes the related information 18 available for viewing by the operator OP.

[0025] As an example, as shown in FIG. 5, the association information 18 is a pair of an enhancer 21 identified as a master regulator MR and its location in DNA 20.

[0026] 6 , the computers that make up the information processing device 12 and the operator terminal 13 basically have the same configuration, and include a storage 35, a memory 36, a CPU (Central Processing Unit) 37, a communication unit 38, a display 39, and an input device 40. These are interconnected via a bus line 41.

[0027] The storage 35 is a hard disk drive built into the computer that constitutes the information processing device 12 and the operator terminal 13, or connected via a cable or network. Alternatively, the storage 35 is a disk array consisting of multiple hard disk drives. The storage 35 stores control programs such as an operating system, various application programs (hereinafter referred to as APs (Application Programs)), and various data associated with these programs. Note that a solid state drive may be used instead of a hard disk drive.

[0028] The memory 36 is a work memory for the CPU 37 to execute processing. The CPU 37 loads programs stored in the storage 35 into the memory 36 and executes processing in accordance with the programs. In this way, the CPU 37 comprehensively controls each part of the computer. The CPU 37 is an example of a "processor" according to the technology of the present disclosure. The memory 36 may be built into the CPU 37.

[0029] The communication unit 38 is a network interface that controls the transmission of various information via the network 14, etc. The display 39 displays various screens. The various screens are provided with an operation function using a GUI (Graphical User Interface). The computers that make up the information processing device 12 and the operator terminal 13 accept input of operation instructions from an input device 40 via the various screens. The input device 40 is a keyboard, a mouse, a touch panel, a microphone for voice input, etc.

[0030] In the following explanation, the parts of the computer that make up the information processing device 12 (storage 35 and CPU 37) are distinguished by adding the suffix "A" to their symbols, and the parts of the computer that make up the operator terminal 13 (storage 35, CPU 37, display 39, and input device 40) are distinguished by adding the suffix "B" to their symbols.

[0031] 7, an operating program 45 is stored in the storage 35A of the information processing device 12. The operating program 45 is an AP for causing a computer to function as the information processing device 12. In other words, the operating program 45 is an example of an "operating program of an information processing device" according to the technology of the present disclosure. The storage 35A also stores extraction conditions 46, specific conditions 47, and the like.

[0032] When the operating program 45 is started, the CPU 37A of the computer constituting the information processing device 12 works in cooperation with the memory 36, etc. to function as a request receiving unit 50, a read / write (hereinafter abbreviated as RW (Read Write)) control unit 51, an extraction unit 52, a related information generation unit 53, and a screen distribution control unit 54.

[0033] The request receiving unit 50 receives various requests from the operator terminal 13. In particular, the request receiving unit 50 receives a specific request 15 from the operator terminal 13. As described above, the specific request 15 includes the first activity level information 161 and the second activity level information 162. Therefore, by receiving the specific request 15, the request receiving unit 50 acquires the first activity level information 161 and the second activity level information 162. When the specific request 15 is received, the request receiving unit 50 outputs the first activity level information 161 and the second activity level information 162 included in the specific request 15 to the RW control unit 51. Furthermore, the request receiving unit 50 outputs the terminal ID of the operator terminal 13 included in the specific request 15 to the screen distribution control unit 54.

[0034] The RW control unit 51 controls the storage of various data in the storage 35A and the reading of various data from the storage 35A. For example, the RW control unit 51 stores the first activity information 161 and the second activity information 162 from the request receiving unit 50 in the storage 35A. The RW control unit 51 also reads the first activity information 161 and the second activity information 162 from the storage 35A and outputs the read first activity information 161 and the second activity information 162 to the extraction unit 52 and the related information generation unit 53.

[0035] Furthermore, the RW control unit 51 reads out the extraction conditions 46 from the storage 35A and outputs the read out extraction conditions 46 to the extraction unit 52. Furthermore, the RW control unit 51 reads out the specific conditions 47 from the storage 35A and outputs the read out specific conditions 47 to the related information generation unit 53.

[0036] The extraction unit 52 extracts enhancers 21 based on the first activity information 161, the second activity information 162, and the extraction conditions 46, and further extracts super enhancers (hereinafter abbreviated as SEs) 21S (see FIG. 8 ) from the extracted enhancers 21. The extraction unit 52 outputs SE extraction results 58 to the related information generation unit 53. The SE extraction results 58 are pairs of extracted SEs 21S and their positions in the DNA 20.

[0037] The related information generation unit 53 generates related information 18 based on the first activity level information 161, the second activity level information 162, the specific condition 47, and the SE extraction result 58. The related information generation unit 53 outputs the related information 18 to the screen distribution control unit 54.

[0038] The screen distribution control unit 54 controls the distribution of various screens to the operator terminal 13. Specifically, the screen distribution control unit 54 distributes and outputs various screens to the operator terminal 13 that has sent the various requests in the form of screen data for web distribution created using a markup language such as XML (Extensible Markup Language). At this time, the screen distribution control unit 54 identifies the operator terminal 13 that has sent the various requests based on the terminal ID from the request receiving unit 50. The various screens include an activity information input screen 75 (see FIG. 11 ) for inputting the first activity information 161 and the second activity information 162, and a related information display screen 80 (see FIG. 12 ) for displaying the related information 18. Note that other data description languages, such as JSON (Javascript (registered trademark) Object Notation), may be used instead of XML.

[0039] As an example, as shown in FIG. 8 , the extraction unit 52 includes an enhancer extraction unit 60 and an SE extraction unit 61. First activity information 161 and second activity information 162 are input to the enhancer extraction unit 60 and the SE extraction unit 61. The enhancer extraction unit 60 extracts enhancers 21 based on a first extraction condition 461 among the extraction conditions 46. The first extraction condition 461 specifies that the amount of histone modification in the first activity information 161, which represents the first activity, and / or the amount of histone modification in the second activity information 162, which represents the second activity, is equal to or greater than a first threshold. The enhancer extraction unit 60 extracts, as enhancers 21, portions of the DNA 20 that satisfy the first extraction condition 461. The enhancer extraction unit 60 outputs an enhancer extraction result 62 to the SE extraction unit 61. The enhancer extraction result 62 is a pair of the extracted enhancer 21 and its position in the DNA 20 .

[0040] The SE extraction unit 61 extracts SEs 21S based on a second extraction condition 462 among the extraction conditions 46. The second extraction condition 462 specifies that the amount of histone modification in the first activity information 161, which is the first activity level, and / or the amount of histone modification in the second activity information 162, which is the second activity level, is equal to or greater than a second threshold, and that the proximity to other enhancers 21 is within a third threshold range. The second threshold is set to a value higher than the first threshold of the first extraction condition 461. The proximity is, for example, the distance between each enhancer 21. The third threshold range is, for example, a distance between adjacent enhancers 21 within 15 kbp (base pair). The SE extraction unit 61 extracts enhancers 21 that satisfy the second extraction condition 462 as SEs from among the enhancers 21 extracted by the enhancer extraction unit 60. The SE extraction unit 61 outputs the SE extraction result 58 .

[0041] As an example, as shown in FIG. 9 , the related information generation unit 53 includes an activity difference calculation unit 65 and an identification unit 66. The activity difference calculation unit 65 and the identification unit 66 receive first activity information 161 and second activity information 162 as input. The activity difference calculation unit 65 also receives an SE extraction result 58. The activity difference calculation unit 65 calculates the difference between the amount of histone modification in the first activity information 161, which represents the first activity, and the amount of histone modification in the second activity information 162, which represents the second activity, for an SE 21S. The difference is the difference between the amounts of histone modification or the ratio of the amounts of histone modification. For example, if the amount of histone modification in the first activity information 161 for a certain SE 21S is 50 and the amount of histone modification in the second activity information 162 is 100, the difference is 50-100=-50 (in the case of a difference) or 50 / 100=0.5 (in the case of a ratio). The activity difference calculation unit 65 outputs a difference calculation result 67 to the identification unit 66. The calculation result 67 is a stored difference for each SE 21S.

[0042] The identification unit 66 identifies a master regulator MR from among the SEs 21S based on the identification condition 47. The identification condition 47 is that the difference between the first activity level and the second activity level is within a fourth threshold range. The fourth threshold range is, for example, a case where the absolute value of the difference in the amount of histone modification between the first activity level information 161 and the second activity level information 162 is 100 or more. Alternatively, the fourth threshold range is, for example, a case where the ratio of the amount of histone modification between the first activity level information 161 and the second activity level information 162 is 0.1 or less or 10 or more. The identification unit 66 extracts, as the master regulator MR, the SEs 21S that satisfy the identification condition 47 from among the SEs 21S extracted by the SE extraction unit 61. In other words, the identification unit 66 extracts, as the master regulator MR, the SEs 21S whose activity level is significantly different between the first type cell 171 and the second type cell 172. The identification unit 66 outputs the related information 18.

[0043] 10 as an example, a specific AP 70 is stored in the storage 35B of the operator terminal 13. The specific AP 70 is installed in the operator terminal 13 by the operator OP. The specific AP 70 is an AP for receiving a service that specifies the master regulator MR from the information processing device 12. When the specific AP 70 is activated, the CPU 37B of the operator terminal 13 functions as a browser control unit 72 in cooperation with the memory 36 and the like. The browser control unit 72 controls the operation of a web browser dedicated to the specific AP 70.

[0044] The browser control unit 72 reproduces various screens based on various screen data from the information processing device 12 and displays the reproduced various screens on the display 39B. The browser control unit 72 also accepts various operation instructions input by the operator OP from the input device 40B via the various screens. The browser control unit 72 transmits various requests, including the specific request 15, to the information processing device 12 in response to the operation instructions.

[0045] When the specific AP 70 is started and a login is performed, an activity information input screen 75 shown in Fig. 11 as an example is displayed on the display 39B under the control of the browser control unit 72. The activity information input screen 75 is provided with a first input box 761 for the first activity information 161 and a second input box 762 for the second activity information 162. Files of the first activity information 161 and the second activity information 162 can be dropped into the first input box 761 and the second input box 762.

[0046] The operator OP inputs the desired first activity level information 161 and second activity level information 162 into the first input box 761 and the second input box 762, and then selects the Identify button 77. When the Identify button 77 is selected, the browser control unit 72 generates a specific request 15 including the first activity level information 161 and the second activity level information 162 input into the first input box 761 and the second input box 762, and transmits the generated specific request 15 to the information processing device 12.

[0047] Furthermore, when the master regulator MR is identified in the information processing device 12 and the related information 18 is generated, an associated information display screen 80 shown in Fig. 12 as an example is displayed on the display 39B under the control of the browser control unit 72. The associated information 18 is displayed on the associated information display screen 80.

[0048] A save button 81 and an OK button 82 are provided at the bottom of the related information display screen 80. When the save button 81 is selected, the first activity information 161 and the second activity information 162 are associated with the related information 18 and stored in the storage 35B. When the OK button 82 is selected, the display of the related information display screen 80 is cleared.

[0049] Next, the operation of the above configuration will be described with reference to the flowchart shown in Fig. 13 as an example. When the operating program 45 is started in the information processing device 12, the CPU 37A functions as a request receiving unit 50, a RW control unit 51, an extraction unit 52, a related information generation unit 53, and a screen distribution control unit 54, as shown in Fig. 7. The extraction unit 52 functions as an enhancer extraction unit 60 and an SE extraction unit 61. The related information generation unit 53 functions as an activity difference calculation unit 65 and an identification unit 66. Furthermore, when a specific AP 70 is started in the operator terminal 13, the CPU 37B functions as a browser control unit 72, as shown in Fig. 10.

[0050] 11 is displayed on the display 39B of the operator terminal 13 under the control of the browser control unit 72. On the activity information input screen 75, the operator OP inputs the desired first activity information 161 and second activity information 162 into the first input box 761 and the second input box 762, and selects the Identify button 77. This causes a specific request 15 to be transmitted from the browser control unit 72 to the information processing device 12. As shown in FIG. 1, the specific request 15 includes the first activity information 161 and the second activity information 162.

[0051] In the information processing device 12, the request receiving unit 50 receives the specific request 15, thereby acquiring the first activity information 161 and the second activity information 162 included in the specific request 15 (YES in step ST100). The first activity information 161 and the second activity information 162 included in the specific request 15 are output from the request receiving unit 50 to the RW control unit 51 and stored in the storage 35A under the control of the RW control unit 51 (step ST110). In addition, the terminal ID of the operator terminal 13 included in the specific request 15 is output from the request receiving unit 50 to the screen distribution control unit 54.

[0052] The first activity information 161 and the second activity information 162 are read from the storage 35A by the RW control unit 51 (step ST120). The first activity information 161 and the second activity information 162 are output from the RW control unit 51 to the extraction unit 52 and the related information generation unit 53.

[0053] The RW control unit 51 reads the extraction conditions 46 from the storage 35A and outputs the read extraction conditions 46 to the extraction unit 52. In the extraction unit 52, as shown in FIG. 8 , the enhancer extraction unit 60 extracts enhancers 21 based on the first activity information 161, the second activity information 162, and the first extraction conditions 461 (step ST130). The enhancer extraction result 62 is output from the enhancer extraction unit 60 to the SE extraction unit 61. Next, the SE extraction unit 61 extracts SEs 21S based on the first activity information 161, the second activity information 162, and the second extraction conditions 462 (step ST140). The SE extraction result 58 is output from the SE extraction unit 61 to the related information generation unit 53.

[0054] The RW control unit 51 reads the specific condition 47 from the storage 35A and outputs the read specific condition 47 to the related information generation unit 53. In the related information generation unit 53, as shown in FIG. 9 , the activity difference calculation unit 65 calculates the difference between the amount of histone modification in the first activity information 161, which is the first activity, and the amount of histone modification in the second activity information 162, which is the second activity, for the SE 21S, based on the first activity information 161, the second activity information 162, and the SE extraction result 58 (step ST150). The difference calculation result 67 is output from the activity difference calculation unit 65 to the identification unit 66.

[0055] The identification unit 66 identifies a master regulator MR from among the SEs 21S based on the first activity level information 161, the second activity level information 162, and the identification condition 47 (step ST160). Then, the identification unit 66 generates the association information 18 shown in FIG. 5 , which stores a pair of the SE 21S identified as the master regulator MR and its position in the DNA 20 (step ST160). The association information 18 is output from the identification unit 66 to the screen distribution control unit 54.

[0056] The screen distribution control unit 54 generates screen data for a related information display screen 80 including the related information 18. Under the control of the screen distribution control unit 54, the screen data for the related information display screen 80 is distributed to the operator terminal 13 that is the sender of the specific request 15 (step ST170).

[0057] 12, in the operator terminal 13, under the control of the browser control unit 72, the screen data of the related information display screen 80 is reproduced, and the reproduced related information display screen 80 is displayed on the display 39B. As a result, the related information 18 is presented to the operator OP.

[0058] As described above, the CPU 37A of the information processing device 12 includes a request receiving unit 50, a related information generating unit 53, and a screen distribution control unit 54. By receiving an identification request 15, the request receiving unit 50 acquires first activity information 161 of a first type of cell 171 having a first phenotype and second activity information 162 of a second type of cell 172 having a second phenotype different from the first type of cell 171. The identification unit 66 of the related information generating unit 53 identifies the master regulator MR, which is a portion of the genome region related to the difference between the first type of cell 171 and the second type of cell 172, based on the difference between the amount of histone modification in the first activity information 161 and the amount of histone modification in the second activity information 162. The screen distribution control unit 54 presents the master regulator MR-related information 18 by distributing and outputting screen data of the related information display screen 80 to the operator terminal 13. This makes it possible to present more practical intervention sites, such as genome editing.

[0059] 8, the enhancer extraction unit 60 extracts enhancers 21 in a genome region as candidates for master regulator MRs. This allows the master regulator MR to be identified more efficiently than when the master regulator MR is identified without any guidelines.

[0060] 9 , the identification unit 66 identifies the master regulator MR based on the difference between the first activity level of the enhancer 21 in the first type cell 171 as the first feature amount and the second activity level of the enhancer 21 in the second type cell 172 as the second feature amount. The first activity level and the second activity level can be easily obtained, and the difference between the first activity level and the second activity level can also be easily calculated. Therefore, the master regulator MR can be easily identified.

[0061] 8, the SE extraction unit 61 extracts SE21S as candidates for the master regulator MR based on the activity of the enhancer 21 and the location of the enhancer 21 in the genomic region. This allows the master regulator MR to be identified more efficiently than when the master regulator MR is identified without any guidelines. SE21S on the order of several thousand can be further narrowed down.

[0062] As shown in Figures 3 and 4, the first activity and the second activity are quantified by the amount of histone modification. The amount of histone modification is related to acetylation of the 27th lysine residue of histone H3 (H3K27ac) or monomethylation of the 4th lysine residue of histone H3 (H3K4me1). The amount of histone modification, particularly the amount of histone modification related to acetylation of the 27th lysine residue of histone H3 and monomethylation of the 4th lysine residue of histone H3, is known to be an important parameter indicating whether transcription of gene 24 is activated. Therefore, the reliability of the master regulator MR identified based on the difference between the first activity and the second activity can be improved. The amount of histone modification may be both acetylation of the 27th lysine residue of histone H3 and monomethylation of the 4th lysine residue of histone H3, as exemplified, or only one of them.

[0063] [Second Embodiment] As an example, as shown in FIG. 14 , in the second embodiment, RNA-Seq (RNA-Sequencing) analysis is performed on a first type of cell 171 and a second type of cell 172 to obtain first gene expression level information 851 and second gene expression level information 852. The first gene expression level information 851 and the second gene expression level information 852 are included in a specific request 15 along with first activity information 161 and second activity information 162, etc., and are received by the request receiving unit 50. The first gene expression level information 851 stores the gene expression level of each gene 24 in the first type of cell 171. The second gene expression level information 852 stores the gene expression level of each gene 24 in the second type of cell 172. Note that the first gene expression level information 851 and the second gene expression level information 852 may be time-series data relating to the culture period of the first type of cell 171 and the second type of cell 172.

[0064] The related information generation unit of the second embodiment functions as a gene expression level difference calculation unit 86 in addition to the activity level difference calculation unit 65 and the identification unit 87 (only the identification unit 87 is shown in FIG. 14 ). First gene expression level information 851 and second gene expression level information 852 are input to the gene expression level difference calculation unit 86. The gene expression level difference calculation unit 86 calculates the difference between the gene expression level of the first type cell 171 in the first gene expression level information 851 and the gene expression level of the second type cell 172 in the second gene expression level information 852 for each gene 24. The difference is the difference in gene expression levels or the ratio of the gene expression levels. If the gene expression level of a certain gene 24 in the first type cell 171 is 100 and the gene expression level in the second type cell 172 is 20, the difference is 100−20=80 (in the case of a difference) or 100 / 20=5 (in the case of a ratio). The gene expression level difference calculation unit 86 outputs a difference calculation result 88 to the identification unit 87. The calculation result 88 is a stored difference for each gene 24. Although not shown, the identification unit 87 receives the first activity information 161, the second activity information 162, and the calculation result 67, similar to the identification unit 66 of the first embodiment.

[0065] The specific condition 89 in the second embodiment is that the difference between the first activity level and the second activity level is within a fourth threshold range, and that the differentially expressed gene contributes to the expression of the transcription factor 25 involved. The identifying unit 87 first extracts differentially expressed genes from among the genes 24 based on the calculation result 88. The differentially expressed genes are genes 24 whose difference in gene expression level is within a fifth threshold range. The fifth threshold range is, for example, a range in which the absolute value of the difference in gene expression level between the first gene expression level information 851 and the second gene expression level information 852 is 100 or greater. Alternatively, the fifth threshold range is, for example, a range in which the ratio of the gene expression levels between the first gene expression level information 851 and the second gene expression level information 852 is 0.1 or less or 10 or greater. The location of the differentially expressed gene in the genome region can be obtained by comparing it with known annotation information of the human genome.

[0066] The identification unit 87 extracts transcription factors 25 involved in expression-varying genes based on known relationships between genes 24 and transcription factors 25. Next, the identification unit 87 identifies, as master regulators MR, SE21S that contribute to the expression of the extracted transcription factors 25 and whose difference between the first activity and the second activity is within a fourth threshold range.

[0067] Transcription is highly activated if SE21S contributes to the expression of transcription factor 25. Therefore, according to the second embodiment in which the identification unit 87 identifies the master regulator MR based on whether or not SE21S contributes to the expression of transcription factor 25, in addition to the difference between the first activity level and the second activity level, it is possible to identify an SE21S that is more suitable as the master regulator MR.

[0068] 15 as an example, in the third embodiment, similar to the second embodiment, RNA-Seq analysis is performed on a first type of cell 171 and a second type of cell 172 to obtain first gene expression level information 851 and second gene expression level information 852. A gene expression level difference calculation unit 86 calculates the difference between the gene expression level of the first type of cell 171 in the first gene expression level information 851 and the gene expression level of the second type of cell 172 in the second gene expression level information 852 for each gene 24, and outputs a calculation result 88 to an identification unit 90. Although not shown in the figure, the identification unit 90 receives first activity information 161, second activity information 162, and a calculation result 67, similar to the identification unit 66 in the first embodiment.

[0069] The identifying condition 91 in the third embodiment is that the difference between the first activity level and the second activity level is within a fourth threshold range and the differentially expressed gene is a target gene. The identifying unit 90 first extracts differentially expressed genes from among the genes 24 based on the calculation result 88. The identifying unit 90 extracts SEs 21S whose differentially expressed genes are target genes based on the known relationship between the enhancer 21 and the gene 24. Next, the identifying unit 90 identifies the extracted SEs 21S whose difference between the first activity level and the second activity level is within the fourth threshold range as master regulators MR.

[0070] SE21S, whose target gene is an expression-variable gene, is thought to contribute greatly to transcriptional activity. Therefore, according to the third embodiment, in which the identification unit 90 identifies the master regulator MR based on the gene expression levels of the first type cells 171 and the second type cells 172, in particular the difference in the gene expression levels of the target genes of SE21S, in addition to the difference between the first activity level and the second activity level, it is possible to identify an SE21S that is more suitable as a master regulator MR.

[0071] Although the target gene of SE21S is not clear, it may be inferred that the target gene is one of the plurality of genes 24. In this case, the master regulator MR may be identified based on the difference in gene expression levels of the plurality of genes 24 (for example, the average value of the difference in gene expression levels of the plurality of genes 24).

[0072] 16 as an example, in the fourth embodiment, topological domain information 95 is acquired. The topological domain information 95 is included in a specification request 15 together with first activity information 161, second activity information 162, etc., and is received by a request receiving unit 50. The topological domain information 95 is input to a specification unit 96. The topological domain information 95 is an example of "topological domain information" according to the technology of the present disclosure.

[0073] The identifying condition 97 of the fourth embodiment is that the difference between the first activity level and the second activity level is within a fourth threshold range, and the three-dimensional proximity according to the topological domain information 95 is within a sixth threshold range. The identifying unit 96 extracts SEs 21S whose three-dimensional proximity according to the topological domain information 95 is within the sixth threshold range. Next, the identifying unit 96 identifies the extracted SEs 21S whose difference between the first activity level and the second activity level is within the fourth threshold range as master regulators MR.

[0074] In this way, if the identification unit 96 identifies the master regulator MR based on the topological domain information 95 in addition to the difference between the first activity level and the second activity level, it is possible to identify an SE 21S that is more suitable as the master regulator MR. Note that the SE extraction unit 61 may extract an SE 21S based on the topological domain information 95 in addition to the activity level and the position.

[0075] In the above embodiments, the amount of histone modification was used as an example of the activity of the enhancer 21, but this is not limited thereto. As an example, the amount of a transcription factor 25 or a transcription co-factor may be used, as shown in the first activity information 161X in FIG. 17 . A transcription co-factor is a factor that functions to facilitate the binding of multiple transcription factors 25, such as MED (Mediator) 1. While the amounts of both the transcription factor 25 and the transcription co-factor are used as the activity in FIG. 17 , the amount of either the transcription factor 25 or the transcription co-factor may be used as the activity.

[0076] As an example, the amount of a chromatin structure regulatory factor may be used, as in the first activity information 161Y shown in Fig. 18. The chromatin structure regulatory factor is a factor that determines the chromatin structure, such as CTCF (CCCTC-Binding Factor) and cohesin, which co-localizes with CTCF.

[0077] In addition, in each of the above embodiments, the association information 18 is a pair of the enhancer 21 identified as the master regulator MR and its position in the DNA 20, but this is not limiting. As an example, a transcription factor 25 may be identified as the master regulator MR and presented, as in the association information 100 shown in Figure 19. Note that both the enhancer 21 and the transcription factor 25 may be presented as association information.

[0078] The extraction condition 46 and the specific conditions 47, 89, 91, and 97 shown in the above embodiments are merely examples. Therefore, the extraction algorithm for enhancer 21, the extraction algorithm for SE21S, and the specific algorithm for master regulator MR are not limited to those described in the above embodiments, and various known algorithms may be used.

[0079] [Example] As an example, as shown in Figure 20, epithelial EC cells are known to transform into mesenchymal MC cells, particularly upon stimulation with TGF-β (Transforming Growth Factor-β). This transformation from epithelial EC cells to mesenchymal MC cells is called epithelial-mesenchymal transformation (EMT) and has been considered important in many studies, such as the malignant transformation of cancer cells. In the example, epithelial EC cells were treated as first-type cells 171, and mesenchymal MC cells were treated as second-type cells 172. Furthermore, in the example, A549 cells, which are human alveolar basal epithelial adenocarcinoma cells, were used as the epithelial EC cells.

[0080] First, as shown in Table 105 of Figure 21, a total of six public ChIP-Seq data sets (first activity information 161 and second activity information 162) were obtained, each showing the presence / absence of TGF-β stimulation and the amount of histone modification (control / H3K4me1 / H3K27ac). Furthermore, as shown in Table 106 of Figure 22, time-series RNA-Seq data (first gene expression level information 851 and second gene expression level information 852) were obtained for A549 cells cultured for 72 hours under conditions with / without TGF-β stimulation.

[0081] The ChIP-Seq data was subjected to three steps of preprocessing: quality control using fastp [v0.23.2], mapping using bwa [v0.7.17] with GRCh38.p13 (Genome Reference Consortium Human Build 38 patch release 13) as a reference, and conversion to a BAM file using samtools [v1.16.1]. Peak detection was then performed using MACS2 [v2.2.7.1]. ​​In MACS2, peak detection was performed by comparing H3K27ac and control, and H3K4me1 and control, with and without TGF-β stimulation, respectively. As an example, a total of four enhancers 21 were extracted, as shown in Table 107 of Figure 23. Hereinafter, the four types of enhancers will be referred to as me+ enhancers (number of extracts: 125,289), me- enhancers (number of extracts: 129,629), ac+ enhancers (number of extracts: 49,397), and ac- enhancers (number of extracts: 47,720).

[0082] Next, enhancers 21 that are active or repressed (differentially active) in epithelial cells EC and mesenchymal cells MC are identified by comparing ac+ enhancers and ac- enhancers. In this example, the read count for each enhancer 21 was normalized with respect to the total read count and the length of the enhancer 21, and enhancers 21 for which this ratio was 2-fold or more or 1 / 2 or less were defined as enhancers 21 with differential activity. As a result, differential activity was observed in 2,457 enhancers, or 5% of ac+ enhancers, and 1,892 enhancers, or 4% of ac- enhancers.

[0083] SE21S were defined according to the standard algorithm proposed by ROSE (RANK ORDERING OF SUPER-ENHANCERS). This algorithm aggregates enhancers 21 within 12.5 kbp of a region defined based on the amount of histone modification associated with H3K4me1, quantifies activity by the sum of the amount of histone modification associated with H3K27ac within the region of that enhancer 21, and defines the group of enhancers 21 that are particularly ranked high in terms of the amount of histone modification associated with H3K27ac as SE21S. Using this method, 1,361 and 1,428 SE21S activated with and without TGF-β stimulation, respectively, were extracted. In this example, we further evaluated the overlap of the extracted SE21S with genomic regions of ac enhancers that showed differential activity in the presence and absence of TGF-β stimulation, and identified SE21S containing enhancer 21 that showed differential activity based on the amount of histone modification related to H3K27ac. Through this process, the number of SE21S activated in the presence and absence of TGF-β stimulation (1,361 and 1,428, respectively) was narrowed down to 605 and 565.

[0084] On the other hand, in this example, differentially expressed genes with characteristic time series patterns were extracted by analyzing RNA-Seq data as follows: If a list of differentially expressed genes in epithelial cells (EC) and mesenchymal cells (MC) is publicly known, it is possible to use the publicly known list without performing the following analysis of the RNA-Seq data.

[0085] First, the RNA-Seq data was mapped to the GRCh38 (Genome Reference Consortium Human Build 38) reference using the DRAGEN (Dynamic Read Analysis for Genomics) Bio-IT Platform [v3.9.3], and gene expression levels were quantified using TPM (Transcripts Per Million) normalization. Furthermore, the difference in log2 TPM values ​​was calculated using RNA-Seq data with and without TGF-β stimulation, and then preprocessing was performed to extract the change components relative to the time average value for each gene. By analyzing this time-series data using the edge library, genes with temporal fluctuations, i.e., genes with altered expression, were extracted. Here, cubic spline function approximation was performed on the edge library, and genes with a q value of less than 0.001 were determined to be differentially expressed genes. As a result, 2086 differentially expressed genes were extracted.

[0086] To cluster the extracted differentially expressed genes based on their temporal patterns, we used a spline function with a coefficient k = 3 obtained from the edge library and performed hierarchical clustering using Euclidean distance and Ward's algorithm. As a result, 801 genes (24) were clustered into a linear increase class, 401 genes (24) into a linear decrease class, and 874 genes (24) into a nonlinear change class. Given the phenomenon of enhanced EMT under TGF-β stimulation, the linear increase class is expected to be a marker for mesenchymal MC cells, and the linear decrease class is expected to be a marker for epithelial EC cells. Indeed, CDH (Cadherin) 1 (E-cadherin), known as a marker for epithelial EC cells, belonged to the linear decrease class, while CDH2 (N-cadherin), known as a marker for mesenchymal MC cells, belonged to the linear increase class, suggesting the validity of the clustering. The nonlinear fluctuation class also included transcription factors 25, such as ETS (E26 Transformation-specific Sequence) 2, which are known as candidates for EMT induction.

[0087] To extract 25 transcription factors associated with differentially expressed genes extracted by analyzing the RNA-Seq data as described above, motif analysis was performed using the TRANSFAC (Transcription Factor Database). Since A549 cells are derived from the lung, "lung" was selected as the background tissue for motif analysis. As a result, 58 transcription factors were extracted from the linear increase class, 52 from the linear decrease class, and 66 from the nonlinear change class. After eliminating duplicates, the FDR (False Discovery Rate) was <10. -22 As a result, 112 transcription factors 25 were extracted.

[0088] The 1,170 SE21S with differential activity extracted with and without TGF-β stimulation and the 112 transcription factors 25 extracted by the analysis of differentially expressed genes described above were subjected to a process of coupling the nearest bases located on the same chromosome within 100 kbp of each other. As a result, 12 SE21S with TGF-β stimulation and 9 without TGF-β stimulation were identified. This demonstrates that the technology disclosed herein can further narrow down the SE21S.

[0089] There are 19 transcription factors 25 whose expression is contributed by 12 and 9 SE21S. In this example, the temporal changes in the expression of differentially expressed genes are revealed by analysis of RNA-Seq data. Therefore, among the 19 transcription factors 25, attention was focused on those that actually showed significant time changes. As an example, nine transcription factors 25 shown in the graph group 108 of FIG. 24 (BHLHE (Basic Helix-Loop-Helix) 40, EPAS (Endothelial PAS domain-containing protein) 1, ETS2, HIF (Hypoxia Inducible Factor) 1A, KLF (Kruppel-Like Factor) 6, MYB (Myeloblastosis) L1, SMAD (Small worm phenotype) and MAD (Mothers against The master regulators identified were SMAD7 (decapentaplegic)3, SMAD7, and ZBTB (zinc finger and BTB)7A. The activation pattern of SE21S correlated with the expression ratio of these identified transcription factors, confirming the cooperative behavior of SE21S and these transcription factors.

[0090] From the above, it was confirmed that the technology disclosed herein can provide more practical intervention sites for genome editing, etc. In particular, it was confirmed that it exhibits superior effectiveness when directional changes in cell phenotypes, such as EMT discussed in this example, are considered to be controlled by a small number of master regulators (MRs).

[0091] The information processing device 12 may be installed in a research institution, or may be installed in a data center independent of the research institution.

[0092] Instead of distributing the screen data of the related information display screen 80 to the operator terminal 13, the related information 18 itself may be distributed to the operator terminal 13. In this case, under the control of the browser control unit 72, the operator terminal 13 generates the related information display screen 80 based on the related information 18 and displays it on the display 39B.

[0093] The method of presenting the related information 18 to the operator OP is not limited to the example of delivering screen data. The related information 18 may be presented to the operator OP by printing it on a paper medium, or by attaching it to an e-mail and sending it to the operator terminal 13.

[0094] The hardware configuration of the computer constituting the information processing device 12 according to the technology of the present disclosure can be modified in various ways. For example, the information processing device 12 can be configured with multiple computers separated as hardware in order to improve processing power and reliability. For example, the functions of the request receiving unit 50 and the RW control unit 51 and the functions of the extraction unit 52, the related information generation unit 53, and the screen distribution control unit 54 can be distributed and performed by two computers. In this case, the information processing device 12 is configured with two computers. Furthermore, some or all of the functions of the information processing device 12 may be performed by the operator terminal 13.

[0095] In this way, the hardware configuration of the computer of the information processing device 12 can be changed as appropriate depending on the required performance such as processing power, safety, reliability, etc. Furthermore, not only the hardware but also APs such as the operating program 45 can be duplicated or stored in a distributed manner in multiple storage devices in order to ensure safety and reliability.

[0096] In each of the above embodiments, the hardware structure of the processing unit that executes various processes, such as the request receiving unit 50, RW control unit 51, extraction unit 52, related information generation unit 53, screen distribution control unit 54, enhancer extraction unit 60, SE extraction unit 61, activity difference calculation unit 65, identification units 66, 87, 90, and 96, browser control unit 72, and gene expression level difference calculation unit 86, can be any of the various processors shown below. As described above, the various processors include the CPUs 37A and 37B, which are general-purpose processors that execute software (the operating program 45 and the specific AP 70) and function as various processing units, as well as programmable logic devices (PLDs) that are processors whose circuit configuration can be changed after manufacture, such as a field programmable gate array (FPGA), and dedicated electrical circuits that are processors having a circuit configuration designed specifically for executing specific processing, such as an application specific integrated circuit (ASIC).

[0097] A single processing unit may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs and / or a combination of a CPU and an FPGA).Furthermore, multiple processing units may be configured with a single processor.

[0098] Examples of configuring multiple processing units with a single processor include, first, a form in which one processor is configured with a combination of one or more CPUs and software, as typified by computers such as client and server, and this processor functions as multiple processing units. Second, a form in which a processor is used to realize the functions of the entire system including multiple processing units with a single IC (Integrated Circuit) chip, as typified by systems on chips (SoCs). In this way, various processing units are configured using one or more of the above-mentioned various processors as a hardware structure.

[0099] Furthermore, more specifically, the hardware structure of these various processors can be an electric circuit (circuitry) that combines circuit elements such as semiconductor elements.

[0100] From the above description, the technology described in the following supplementary paragraphs can be understood.

[0101] [Supplementary Item 1] An information processing device comprising a processor, wherein the processor acquires a first feature amount of a first type of cell having a first phenotype and a second feature amount of a second type of cell having a second phenotype different from the first type of cell, identifies a portion of a genomic region related to the difference between the first type of cell and the second type of cell based on a difference between the first feature amount and the second feature amount, and presents information related to the identified portion. [Supplementary Item 2] The information processing device according to Supplementary Item 1, wherein the processor extracts enhancers within the genomic region as candidates for the portion. [Supplementary Item 3] The information processing device according to Supplementary Item 2, wherein the processor identifies the portion based on a difference between a first activity level of the enhancer in the first type of cell as the first feature amount and a second activity level of the enhancer in the second type of cell as the second feature amount. [Supplementary Item 4] The information processing device according to Supplementary Item 2 or Supplementary Item 3, wherein the processor identifies the portion based on whether the enhancer contributes to expression of a transcription factor, in addition to the difference between the first feature amount and the second feature amount. [Supplementary Item 5] The information processing device according to any one of Supplementary Item 2 to Supplementary Item 4, wherein the processor extracts a super enhancer as a candidate for the portion based on the activity of the enhancer and the position of the enhancer in the genomic region. [Supplementary Item 6] The information processing device according to any one of Supplementary Item 1 to Supplementary Item 5, wherein the processor identifies the portion based on the difference in gene expression level between the first type of cell and the second type of cell, in addition to the difference between the first feature amount and the second feature amount. [Supplementary Item 7] The information processing device according to Supplementary Item 6, wherein the processor extracts enhancers in the genomic region as the candidate for the portion, and the gene expression level is the gene expression level of a target gene of the enhancer. [Supplementary Item 8] The information processing device according to any one of Supplementary Items 1 to 7, wherein the processor identifies the portion based on information of a topological domain in addition to the difference between the first feature amount and the second feature amount.[Supplementary Item 9] The information processing device according to any one of Supplementary Items 3 to 8, wherein the first activity level and the second activity level are quantified by any one of an amount of histone modification, an amount of a transcription factor or a transcription coactivator, and an amount of a chromatin structure regulator. [Supplementary Item 10] The information processing device according to Supplementary Item 9, wherein the amount of the histone modification is related to acetylation of the 27th lysine residue of histone H3 or monomethylation of the 4th lysine residue of histone H3.

[0102] The technology of the present disclosure can be appropriately combined with the various embodiments and / or various modified examples described above. Furthermore, it is not limited to the above embodiments, and various configurations can be adopted without departing from the spirit of the present disclosure. Furthermore, the technology of the present disclosure extends not only to programs, but also to storage media that non-temporarily store programs, and computer program products that include programs.

[0103] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[0104] In this specification, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed by connecting them with "and / or."

[0105] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

Claims

1. An information processing device comprising a processor, which acquires a first feature of a first type of cell having a first phenotype and a second feature of a second type of cell having a second phenotype different from the first type of cell, identifies a portion of a genomic region related to the difference between the first type of cell and the second type of cell based on the difference between the first feature and the second feature, and presents information related to the identified portion.

2. The information processing device according to claim 1, wherein the processor extracts enhancers within the genome region as candidates for the portion.

3. The information processing device described in claim 2, wherein the processor identifies the portion based on the difference between a first activity level of the enhancer in the first type of cell as the first feature and a second activity level of the enhancer in the second type of cell as the second feature.

4. The information processing device described in claim 2, wherein the processor identifies the portion based on, in addition to the difference between the first feature and the second feature, whether the enhancer contributes to the expression of a transcription factor.

5. The information processing device according to claim 2, wherein the processor extracts super-enhancers as candidates for the portion based on the activity of the enhancer and the location of the enhancer in the genomic region.

6. The information processing device according to claim 1, wherein the processor identifies the portion based on the difference between the first feature amount and the second feature amount as well as the difference between the gene expression levels of the first type of cell and the second type of cell.

7. The information processing device according to claim 6, wherein the processor extracts enhancers within the genome region as candidates for the portion, and the gene expression level is the gene expression level of a target gene of the enhancer.

8. The information processing device according to claim 1, wherein the processor identifies the portion based on information of a topological domain in addition to the difference between the first feature amount and the second feature amount.

9. An information processing device according to claim 3, wherein the first activity level and the second activity level are quantified by any one of the amount of histone modification, the amount of transcription factor or transcription coactivator, and the amount of chromatin structure regulator.

10. The information processing device according to claim 9, wherein the amount of histone modification relates to acetylation of the 27th lysine residue of histone H3 or monomethylation of the 4th lysine residue of histone H3.

11. A method for operating an information processing device, comprising: acquiring a first feature of a first type of cell having a first phenotype and a second feature of a second type of cell having a second phenotype different from the first type of cell; identifying a portion of a genomic region related to the difference between the first type of cell and the second type of cell based on the difference between the first feature and the second feature; and presenting information related to the identified portion.

12. An operating program for an information processing device that causes a computer to execute processes including: acquiring a first feature amount of a first type of cell having a first phenotype and a second feature amount of a second type of cell having a second phenotype different from the first type of cell; identifying a portion of a genomic region related to the difference between the first type of cell and the second type of cell based on the difference between the first feature amount and the second feature amount; and presenting information related to the identified portion.

Citation Information

Patent Citations

  • Epigenetic profiling of cancer

    JP2019514344A