Gene set analysis device and gene set analysis method
The gene set analysis device integrates constituent gene feature vectors to identify biological phenomena through similarity with phenomenon vectors, addressing limitations of enrichment analysis by considering all gene information, enhancing analysis of both large and small gene sets.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- FRONTEO INC
- Filing Date
- 2024-11-29
- Publication Date
- 2026-06-04
AI Technical Summary
Enrichment analysis methods are limited in identifying unknown gene associations and are unsuitable for small gene sets, lacking information on all constituent genes and requiring pre-identified gene-GO term associations.
A gene set analysis device and method that calculates a gene set feature vector by integrating constituent gene feature vectors, identifying biological phenomena through similarity with phenomenon feature vectors, regardless of gene set size, using vector analysis that considers all gene information.
Enables identification of significant biological phenomena related to gene sets, even with unknown gene associations, by accounting for all genes, irrespective of set size, and enhances analysis of both large and small gene sets.
Smart Images

Figure JP2024042419_04062026_PF_FP_ABST
Abstract
Description
Gene set analysis apparatus and gene set analysis method
[0001] The present invention relates to a gene set analysis device and a gene set analysis method, and is particularly suitable for use in devices and methods for estimating biological phenomena related to gene sets.
[0002] Traditionally, enrichment analysis has been known as one method for analyzing the function of a gene set consisting of multiple genes. Enrichment analysis is an analysis that examines whether the proportion of genes with a specific function (GO term) in a gene set of interest is statistically significantly higher than the proportion of genes with a specific function in the entire population from which data was collected. GO (Gene Ontology) is information that systematically organizes and defines concepts related to gene function, and each concept is associated with a term (GO term) and a number (GO ID).
[0003] Enrichment analysis makes it possible to determine whether genes with functions related to a specific GO term are enriched within a set of genes of interest. However, as mentioned above, since enrichment analysis statistically calculates the proportion of genes with GO terms included in a set of genes, it requires a pre-identified list of which genes are associated with which GO terms, and therefore has the problem of not being able to analyze unknown associations not included in the list.
[0004] Furthermore, since enrichment analysis statistically calculates the proportion of genes related to a specific GO term among the constituent genes included in a gene set, it can be said that the analysis lacks information on genes that were not focused on. Therefore, there was a problem in that it was not possible to perform GO term analysis that took into account the information of all constituent genes included in the gene set. In addition, enrichment analysis, which statistically calculates proportions, is suitable for analyzing large gene sets with a large number of constituent genes, but is unsuitable for analyzing small gene sets.
[0005] In addition to enrichment analysis, other techniques for analyzing GO terms are known (see, for example, Patent Document 1). In the system described in Patent Document 1, an item vector is created for each considered GO term based on the score calculated for each subset of data considered with respect to the GO term, and an analysis is performed to find GO terms that receive similar scores in different datasets.
[0006] According to the disclosure in Patent Document 1, item vectors can be used to identify GO terms that exhibit similar behavior (receiving similar significant scores across various datasets). For example, if two or more pathways have similar scores across multiple datasets, they are analyzed to be potentially related to each other with respect to the biological process under consideration.
[0007] Japanese Patent Publication No. 2006-120143
[0008] This invention was made to solve the problems of the aforementioned enrichment analysis, and aims to enable the identification of more significant biological phenomena for a gene set by vector analysis that takes into account information of all genes constituting the gene set, regardless of the size of the gene set of interest.
[0009] To solve the above-mentioned problems, the present invention identifies biological phenomena that are presumed to be related to a gene set, based on the similarity between a gene set feature vector, which is a feature vector of the gene set to be analyzed and is formed by integrating the constituent gene feature vectors for each of the multiple genes that make up the gene set, and phenomenon feature vectors, which are feature vectors of various biological phenomena.
[0010] According to the present invention configured as described above, the relationship between a gene set and a biological phenomenon is estimated based on the similarity between the gene set feature vector related to the gene set under analysis and the phenomenon feature vector related to various biological phenomena. Therefore, even if the gene set includes genes whose relationship to biological phenomena is unknown, and regardless of whether the gene set is large or small, it is possible to identify biological phenomena that are estimated to be related to the gene set under analysis. Furthermore, since the present invention performs vector analysis that takes into account the information of all genes constituting the gene set of interest, it is possible to identify biological phenomena that are more significant to the gene set.
[0011] This is a diagram illustrating an example of the processing performed by the gene set analysis device according to the first embodiment. This is a diagram illustrating an example of the processing performed by the gene set analysis device according to the first embodiment. This is a block diagram showing an example of the functional configuration of the gene set analysis device according to the first embodiment. This is a diagram showing an example of the hardware configuration of the gene set analysis device according to the first embodiment. This is a block diagram showing a specific example of the functional configuration of the gene feature vector calculation unit. This is a diagram showing an example of a feature vector. This is a block diagram showing a specific example of the functional configuration of the phenomenon feature vector calculation unit. This is a diagram showing an example of the output of analysis results by the association estimation unit. This is a flowchart showing an example of the operation of the gene set analysis device according to the first embodiment. This is a block diagram showing an example of the functional configuration of the gene set analysis device according to the first modified example. This is a block diagram showing an example of the functional configuration of the gene set analysis device according to the second modified example. This is a block diagram showing an example of the functional configuration of the gene set analysis device according to the second embodiment. This is a diagram illustrating the processing content of the cluster analysis unit according to the second embodiment. This is a block diagram showing an example of the functional configuration of the gene set analysis device according to other modified examples.
[0012] (First Embodiment) Hereinafter, a first embodiment of the present invention will be described with reference to the drawings. Figure 1 is a diagram illustrating an example of the processing performed by the gene set analysis device according to the first embodiment.
[0013] Figure 1 shows an example of a process to identify biological phenomena presumed to be associated with a gene set, using multiple molecules (genes) constituting one pathway 111 within an intermolecular interaction network (also called a pathway) 110, which represents the intermolecular interactions of related molecules as a pathway diagram. Note that pathway 110 in Figure 1 shows only a portion of the pathways included in the overall network.
[0014] In the pathway 110 shown in Figure 1, diamond-shaped nodes represent causative genes, square-shaped nodes represent responsive genes, and elliptical nodes represent connecting genes. Causality refers to the property that a gene may cause a disease due to its presence or mutation. Responsiveness refers to the property that a gene may mutate (change) as a result of the onset of a disease. Pathway 110 is generated so that causative genes are placed upstream of the pathway, responsive genes are placed downstream of the pathway, and other connecting genes are placed between the causative and responsive genes.
[0015] The pathway 111 that constitutes the gene set is a network consisting of one causative gene as the starting point, one responsive gene as the ending point, and one or more connecting genes linking them together. Other causative or responsive genes may be included between the starting and ending points. In Figure 1, a small-scale pathway 111 consisting of five genes from the starting point to the ending point is used as the gene set for analysis.
[0016] In this embodiment, the biological phenomenon that is estimated to be associated with a gene set is the biological function of the gene set. As an example, the gene set analysis device of this embodiment identifies GO terms that are estimated to be associated with a gene set from among various GO terms that define concepts related to the biological function of genes.
[0017] The gene set analysis device of this embodiment calculates a feature vector (hereinafter referred to as the gene set feature vector) GV1 for the gene set of one pathway 111 to be analyzed, and also calculates feature vectors (hereinafter referred to as the phenomenon feature vectors) PV1, PV2, ... for various GO terms, and identifies GO terms that are estimated to be related to the gene set based on the similarity between the gene set feature vector GV1 and the phenomenon feature vectors PV1, PV2, .... The phenomenon feature vector PV is data that represents the features of a GO term (features that can identify a GO term) as a combination of values of multiple elements.
[0018] For example, the gene set analysis device identifies the phenomenon feature vector PVmax with the greatest similarity to the gene set feature vector GV1, and estimates that the GO term corresponding to this phenomenon feature vector PVmax is predominantly associated with the gene set. In the example in Figure 1, the similarity between the gene set feature vector GV1 and the phenomenon feature vector PV3 is highest at "0.7", and it is estimated that the gene set of pathway 111 is predominantly associated with the GO term of "DNA Damage Response".
[0019] Here, the gene set analysis device calculates feature vectors (hereinafter referred to as "constituent gene feature vectors") GV11 to GV15 for each constituent gene included in the gene set to be analyzed, and calculates the gene set feature vector GV1 by integrating multiple constituent gene feature vectors GV11 to GV15. Constituent gene feature vectors GV11 to GV15 are data that represent the features of the constituent genes (features that can identify the constituent genes) as a combination of values of multiple elements. The integration of constituent gene feature vectors GV11 to GV15 is performed, for example, by calculating the sum of constituent gene feature vectors GV1 to GV5. However, it is not limited to this. For example, integration by averaging or integration by weighted averaging may also be used.
[0020] The similarity between the gene set feature vector GV1 and the phenomenon feature vectors PV1, PV2, ... is, for example, expressed as cosine similarity. However, it is not limited to this, and the similarity of the vectors may also be calculated using methods such as Euclidean distance, Manhattan distance, Mahalanobis distance, edit distance, Hamming distance, or Pearson correlation coefficient.
[0021] In this embodiment, the constituent gene feature vectors GV11 to GV15 are calculated for all constituent genes included in the gene set to be analyzed, and the similarity to the phenomenon feature vectors PV1, PV2, ... is calculated using the gene set feature vector GV1, which is formed by integrating these vectors. In other words, by performing vector analysis that takes into account the information of all genes constituting the gene set of interest, the GO term that is estimated to be related to the gene set is identified.
[0022] Here, we have shown an example of a process to identify the associated GO term by analyzing a single gene set corresponding to one pathway 111 contained in pathway 110. However, it is also possible to identify the associated GO term for each of the multiple gene sets that are included as multiple pathways in the overall pathway. All pathways included in the overall pathway may be analyzed, or multiple pathways that form part of the overall pathway may be analyzed.
[0023] Figure 2 shows an example of a process to identify GO terms that are presumed to be associated with a set of genes, using a set of genes constituting another pathway 112 included in the pathway 110 shown in Figure 1 as the set of genes to be analyzed. The pathway 112 shown in Figure 2 also represents a small set of genes consisting of six genes from the starting point to the ending point.
[0024] In the example shown in Figure 2, the gene set analysis device calculates constituent gene feature vectors GV21 to GV26 for each constituent gene included in the gene set of pathway 112, and calculates the gene set feature vector GV2 by integrating multiple constituent gene feature vectors GV21 to GV26. Then, based on the similarity between the gene set feature vector GV2 and the phenomenon feature vectors PV1, PV2, ..., it identifies the GO term that is presumed to be associated with the gene set. In the example shown in Figure 2, the similarity between the gene set feature vector GV2 and the phenomenon feature vector PV3 is highest at "0.6", and it is presumed that the gene set of pathway 112 is predominantly associated with the GO term of "DNA Damage Response".
[0025] When analyzing multiple pathways included in the overall pathway, the overall similarity may be calculated for each GO term of the analyzed pathway by adding up the similarities calculated for multiple gene sets that are estimated to be related to the same GO term.
[0026] In the example shown in Figures 1 and 2, since both the gene set of pathway 111 and the gene set of pathway 112 are estimated to be related to the GO term of "DNA Damage Response," the overall similarity of "1.3" is calculated by adding their respective similarities of "0.7" and "0.6."
[0027] Similarly, for other pathways not shown in Figures 1 and 2, the similarity between the gene set feature vector GV and the phenomenon feature vector PV is calculated to identify GO terms that are presumed to be associated with the gene set. Then, if there are multiple gene sets that are presumed to be associated with the same GO term, the overall similarity of that GO term is calculated by adding the similarities calculated for each of them.
[0028] In this way, it is possible to identify GO terms that are presumed to be predominantly associated with multiple gene sets corresponding to multiple pathways that have been analyzed. For example, by analyzing all pathways included in a pathway and identifying the phenomenon feature vector with the highest overall similarity, it is possible to estimate the GO term that is most predominantly associated with the pathway as a whole. Alternatively, by identifying phenomenon feature vectors with an overall similarity greater than a threshold, it is also possible to estimate several GO terms that are predominantly associated with the pathway as a whole.
[0029] Figure 3 is a block diagram showing an example of the functional configuration of the gene set analysis device 1A according to the first embodiment. Figure 4 is a diagram showing an example of the hardware configuration of the gene set analysis device 1A according to the first embodiment.
[0030] The gene set analysis device 1A according to the first embodiment is composed of a general-purpose computer such as a workstation or personal computer. As shown in Figure 4, the gene set analysis device 1A of this embodiment has a hardware configuration that includes a control unit 101, a storage unit 102, an input unit 103, an output unit 104, and a communication interface 105.
[0031] The control unit 101 is composed of, for example, a microcomputer processor equipped with a CPU, RAM, and ROM, and operates according to the operating system and application programs stored in the storage unit 102, and performs information processing using various data stored in the storage unit 102. In addition to the microcomputer, it may also be equipped with a DSP (Digital Signal Processor) or the like.
[0032] The storage unit 102 is composed of a computer-readable storage medium. For example, the storage unit 102 may include ROM, RAM, a hard disk, or semiconductor memory. Various programs executed by the control unit 101 are stored in the storage unit 102.
[0033] Furthermore, the storage unit 102 also stores various types of data, including data used in the processing of the control unit 101 and data generated during the processing. These various types of data include multiple document data sets. Alternatively, document data may be stored in external storage such as a data server, and the control unit 101 may retrieve the necessary data from the external storage via the communication interface 105 when performing analysis.
[0034] The input unit 103 is comprised of, for example, a keyboard, mouse, or touch panel, and supplies data to the control unit 101 in accordance with the user's input. The output unit 104 is comprised of, for example, a liquid crystal display, and displays information according to instructions input from the control unit 101. The output unit 104 may also include a printer or a speaker.
[0035] The communication interface 105 is an interface for connecting to a communication network, and may include, for example, a modem for connecting to a public communication network or a public telephone network. In addition, it may include at least one of the following: an adapter for connecting to a LAN (Local Area Network), a wireless communication device for wireless communication, and a USB (Universal Serial Bus) connector or an RS232C connector for serial communication.
[0036] The gene set analysis device 1A may be a server device logically implemented by, for example, cloud computing. The server device may consist of a single device or a combination of multiple devices. In this case, the gene set analysis device 1A does not need to have an input unit 103 and an output unit 104 in the hardware configuration shown in Figure 4. That is, the gene set analysis device 1A may input data necessary for analysis from an external source via a communication interface 105 and output analysis result data to an external source via the communication interface 105.
[0037] As shown in Figure 3, the gene set analysis device 1A according to the first embodiment includes, as a functional configuration, a gene set acquisition unit 11, a gene feature vector calculation unit 12, a phenomenon feature vector calculation unit 13, a relationship estimation unit 14, and a total similarity calculation unit 15. These functional blocks 11 to 15 perform the processes described below through the cooperation of hardware and software. For example, the processing of the above functional blocks 11 to 15 is executed by the operation of a program stored in the storage unit 102 under the control of the control unit 101 shown in Figure 4.
[0038] Furthermore, as shown in Figure 3, the gene set analysis device 1A according to the first embodiment is connected to a first document data storage unit 21 and a second document data storage unit 22 as storage media. Here, we show an example configuration in which the two document data storage units 21 and 22 are external storage for the gene set analysis device 1A, but the two document data storage units 21 and 22 may also be configured as the storage unit 102 shown in Figure 4.
[0039] The gene set acquisition unit 11 acquires data for the gene set to be analyzed. For example, it acquires information (hereinafter referred to as gene information) that can identify multiple genes that constitute one pathway 111 included in the pathway 110 shown in Figure 1. Gene information that can identify individual genes is, for example, the gene name. For the pathway 110 to be analyzed, data that has been pre-generated in a device other than the gene set analysis device 1A can be used.
[0040] The method for acquiring genetic information related to a gene set by the gene set acquisition unit 11 is arbitrary. For example, the gene information of the gene set may be input by operating the input unit 103. Alternatively, the pathway 110 may be displayed on the output unit 104's display, and the input unit 103 may be used to specify the route 111 to be analyzed from among the pathways 110, thereby acquiring the genetic information recorded in association with each node included in the specified route 111.
[0041] When analyzing multiple pathways (all or some of the pathways included in a pathway), including pathway 111 shown in Figure 1 and pathway 112 shown in Figure 2, the gene set acquisition unit 11 acquires data for multiple gene sets corresponding to the multiple pathways. When analyzing all pathways included in a pathway, it is not necessary to display and specify the pathway on the display; data for the pathway, with gene information linked to each node, may be acquired.
[0042] The above example illustrates how to obtain gene set data from pathways pre-generated by a device separate from the gene set analysis device 1A. However, the gene set analysis device 1A may also be equipped with a pathway generation function. The pathway generation method is arbitrary. Details regarding this will be described later.
[0043] The gene feature vector calculation unit 12 calculates a gene set feature vector GV, which is a feature vector of a gene set acquired as the target of analysis by the gene set acquisition unit 11, and is formed by integrating the constituent gene feature vectors for each of the multiple genes that make up the gene set. In this embodiment, the gene feature vector calculation unit 12 calculates the gene set feature vector GV using text data stored in the first text data storage unit 21. The first text data storage unit 21 stores data of multiple texts that contain various genes as words (gene names).
[0044] In other words, the gene feature vector calculation unit 12 calculates the gene set feature vector GV by utilizing natural language processing on the text data stored in the first text data storage unit 21. Here, the gene feature vector calculation unit 12 calculates a constituent gene feature vector for each of the multiple genes that make up the gene set, and then calculates the gene set feature vector by integrating the multiple constituent gene feature vectors that have been calculated.
[0045] As described above, in the case of the gene set of pathway 111 illustrated in Figure 1, the gene feature vector calculation unit 12 calculates the constituent gene feature vectors GV11 to GV15 for each of the five constituent genes included in the gene set, and calculates the gene set feature vector GV1 by summing the multiple constituent gene feature vectors GV11 to GV15. Furthermore, when multiple pathways included in a pathway are to be analyzed, the gene feature vector calculation unit 12 calculates the gene set feature vector GV for each of the multiple gene sets corresponding to those multiple pathways.
[0046] While the method for calculating gene set feature vectors (GV) from text data is arbitrary, one example is the configuration shown in Figure 5. Figure 5 is a block diagram showing a specific functional configuration example of the gene feature vector calculation unit 12.
[0047] As shown in Figure 5, the gene feature vector calculation unit 12 is configured as follows: a word extraction unit 121, a vector calculation unit 122, an index value calculation unit 123, a constituent gene feature vector identification unit 124, and a gene set feature vector calculation unit 125. The vector calculation unit 122 is configured as follows: a sentence vector calculation unit 122A and a word vector calculation unit 122B.
[0048] The word extraction unit 121 analyzes m sentences (where m is any integer greater than or equal to 2) contained in the sentence data input from the first sentence data storage unit 21, and extracts n words (where n is any integer greater than or equal to 2) from these m sentences. For sentence analysis, for example, known morphological analysis can be used. Here, the word extraction unit 121 may extract all morphemes of parts of speech that are divided by morphological analysis as words, or it may extract only morphemes of specific parts of speech as words.
[0049] Note that among the m sentences, there may be multiple occurrences of the same word. In this case, the word extraction unit 121 does not extract multiple occurrences of the same word, but only extracts one. That is, the n words extracted by the word extraction unit 121 mean n types of words. Among the words extracted by the word extraction unit 121, there are disease names, molecule names such as genes and proteins, and various other words.
[0050] The vector calculation unit 122 calculates m sentence vectors and n word vectors from the m sentences and the n words. Here, the sentence vector calculation unit 122A calculates m sentence vectors composed of q axis components (q is an arbitrary integer of 2 or more) by vectorizing the m sentences to be analyzed by the word extraction unit 121 into q dimensions according to a predetermined rule. Further, the word vector calculation unit 122B calculates n word vectors composed of q axis components by vectorizing the n words extracted by the word extraction unit 121 into q dimensions according to a predetermined rule.
[0051] In the present embodiment, as an example, sentence vectors and word vectors are calculated as follows. Now, consider a set S = <d ∈ D, w ∈ W> consisting of m sentences and n words. Here, each sentence d i (i = 1, 2, ···, m) and each word w j (j = 1, 2, ···, n) are respectively associated with a sentence vector d i → and a word vector w j → (hereinafter, the symbol "→" indicates that it is a vector). Then, for any word w j and any sentence d i , the probability P(w j | d i ) shown in the following formula (1) is calculated.
[0052]
[0053] Note that this probability P(w j | d iThis value can be calculated by following the probability p disclosed in the publicly available document "Distributed Representations of Sentences and Documents" by Quoc Le and Tomas Mikolov, Google Inc, Proceedings of the 31st International Conference on Machine Learning Held in Beijing, China on 22-24 June 2014. This publicly available document, for example, states that when there are three words, "the," "cat," and "sat," the fourth word to be predicted to be "on," and the formula for calculating that prediction probability p is provided.
[0054] The probability p(wt | wt-k, ..., wt+k) described in publicly available literature is the probability of correctly predicting one word wt from multiple words wt-k, ..., wt+k. In contrast, the probability P(wt|wt-k|wt-k|wt+k) shown in equation (1) used in this embodiment is j |d i ) is one sentence d out of m sentences i So, one of the n words lol j This represents the expected probability of getting the correct answer. One sentence d i One word from lol j Predicting means, specifically, a certain sentence d i When it appears, there is a word lol j This means predicting the possibility that it will be included.
[0055] Note that this equation (1) is d i lol j Since it's symmetrical, one out of n words is lol j From m sentences, one sentence d i The expected probability P(d) i |w j You may calculate () one word lol j From one sentence d i To predict is a certain word lol j When it appears, it is sentence d iThis means predicting the possibilities that are included within that.
[0056] Equation (1) uses an exponential function with base e and exponent being the dot product of the word vector w→ and the sentence vector d→. The sentence d to be predicted is... i and word lol j The exponential function value calculated from the combination of and text d i and n words lol k The ratio of the sum of n exponential function values calculated from each combination of (k = 1, 2, ..., n) to one sentence d i One word from lol j This is calculated as the expected probability of getting the correct answer.
[0057] Here, word vector lol j → and text vector d i →The dot product of this and the word vector w j → text vector d i The scalar value when projected in the direction of →, i.e., the word vector lol j →The text vector d that possesses i It could also be called the component value in the direction of →. This is a word lol j is text d i It can be thought of as representing the degree to which it contributes. Therefore, using the exponential function value calculated using such an inner product, n words w k A single word for the sum of exponential function values calculated for (k = 1, 2, ..., n) j Finding the ratio of the exponential function values calculated for is one sentence d i One of n words from lol j This is equivalent to calculating the expected probability of getting the correct answer.
[0058] Here, we have shown an example of calculation using an exponential function whose exponent is the dot product of the word vector w→ and the sentence vector d→, but it is not mandatory to use an exponential function. Any calculation formula that utilizes the dot product of the word vector w→ and the sentence vector d→ will suffice; for example, the probability could be calculated using the ratio of the dot product values themselves.
[0059] Next, the vector calculation unit 122 calculates the probability P(w) calculated by equation (1) as shown in the following equation (2). j |d i The text vector d that maximizes the value L obtained by summing the texts over the set S. i → and word vector w j → calculates the probability P(w) calculated by the above formula (1). That is, the sentence vector calculation unit 122A and the word vector calculation unit 122B calculate the probability P(w) calculated by the above formula (1). j |d i The value obtained by calculating the sum of the values for all combinations of m sentences and n words is used as the target variable L, and the sentence vector d that maximizes this target variable L is determined. i → and word vector w j → Calculate.
[0060]
[0061] The probability P(w) calculated for all combinations of m sentences and n words. j |d i Maximizing the sum L of ) is equivalent to a sentence d i (i = 1, 2, ..., m) A word that is a part of this lol j This means that (j = 1, 2, ..., n) maximizes the expected probability of being the correct answer. In other words, the vector calculation unit 122 calculates the text vector d that maximizes this probability of being the correct answer. i → and word vector w j → can be said to be something that calculates →.
[0062] In this embodiment, as described above, the vector calculation unit 122 calculates m sentences d i By vectorizing each of these into q dimensions, we obtain m sentence vectors d consisting of q axis components. i By calculating the → and vectorizing each of the n words into q dimensions, we obtain n word vectors w consisting of q axis components. j →Calculate the text vector d that maximizes the above-mentioned target variable L, with q axis directions being variable. i → and word vector w j This is equivalent to calculating →.
[0063] The index value calculation unit 123 calculates m sentence vectors d calculated by the vector calculation unit 122. i → and n word vectors w j →By calculating the inner product of each, m sentences d i and n words lol j An index value that reflects the relationship between them is calculated. In this embodiment, the index value calculation unit 123 calculates m text vectors d as shown in the following equation (3). i → Each q axis component (d 11 ~d mq A sentence matrix D whose elements are ) and n word vectors w j → Each q axis component (w 11 ~w nq By calculating the product with a word matrix W whose elements are ), an index value matrix DW whose elements are m × n index values is calculated. Here, W t This is the transpose of the word matrix.
[0064]
[0065] Each element of the index matrix DW calculated in this way can be said to represent how much each word contributes to each sentence, and how much each sentence contributes to each word. For example, the element dw in a 1x2 grid. 12 This value represents the extent to which word w2 contributes to sentence d1, and similarly, it represents the extent to which sentence d1 contributes to word w2. Therefore, each row of the index matrix DW can be used to evaluate sentence similarity, and each column can be used to evaluate word similarity.
[0066] The constituent gene feature vector identification unit 124 identifies a group of word index values consisting of m index values for each of the multiple gene names included in the gene set acquired by the gene set acquisition unit 11 as the target of analysis, from among the n words, as a constituent gene feature vector. That is, as shown in Figure 6, the constituent gene feature vector identification unit 124 identifies a group of word index values corresponding to the words of the multiple genes that make up the gene set, from among the n sets of word index value groups (m index values per column) that make up each column of the index value matrix DW, as a constituent gene feature vector for each gene name.
[0067] The gene set feature vector calculation unit 125 calculates a gene set feature vector by integrating multiple constituent gene feature vectors identified by the constituent gene feature vector identification unit 124.
[0068] Returning to Figure 3, the explanation continues. The phenomenon feature vector calculation unit 13 calculates the phenomenon feature vector PV, which is the feature vector of various biological phenomena (GO term in this embodiment). In this embodiment, the phenomenon feature vector calculation unit 13 calculates the phenomenon feature vector PV using the text data stored in the second text data storage unit 22.
[0069] The second text data storage unit 22 stores data for multiple sentences that contain various GO terms as words. A GO term is written as, for example, "DNA Damage Response," and strictly speaking, it can be understood as a compound word consisting of multiple words in sequence, but since it is a technical term that refers to a unified concept, it is treated as a single word.
[0070] In other words, the phenomenon feature vector calculation unit 13 calculates the phenomenon feature vector PV using natural language processing on the text data stored in the second text data storage unit 22. The method for calculating the phenomenon feature vector PV from the text data is arbitrary, but as an example, it can be calculated using the configuration shown in Figure 7. Figure 7 is a block diagram showing a specific example of the functional configuration of the phenomenon feature vector calculation unit 13.
[0071] As shown in Figure 7, the phenomenon feature vector calculation unit 13 is configured as follows: a word extraction unit 131, a vector calculation unit 132, an index value calculation unit 133, and a phenomenon feature vector identification unit 134. The vector calculation unit 132 is configured as follows: a sentence vector calculation unit 132A and a word vector calculation unit 132B.
[0072] In Figure 7, the processes performed by the word extraction unit 131, the vector calculation unit 132, and the index value calculation unit 133 are the same as the processes performed by the word extraction unit 121, the vector calculation unit 122, and the index value calculation unit 123 shown in Figure 5. In other words, only the text data to be processed is different, and the functions of the word extraction unit 131, the vector calculation unit 132, and the index value calculation unit 133 shown in Figure 7 are the same as the functions of the word extraction unit 121, the vector calculation unit 122, and the index value calculation unit 123 shown in Figure 5.
[0073] The phenomenon feature vector identification unit 134 identifies a group of word index values consisting of m index values for each of the multiple GO terms among the n words extracted by the word extraction unit 131 as the phenomenon feature vector PV. That is, as shown in Figure 6, the phenomenon feature vector identification unit 134 identifies the group of word index values corresponding to the word that corresponds to the GO term from the n sets of word index value groups (m index values per column) that constitute each column of the index value matrix DW as the phenomenon feature vector PV for each GO term.
[0074] Returning to Figure 3, the explanation continues. The association estimation unit 14 identifies GO terms that are estimated to be related to the gene set obtained by the gene set acquisition unit 11, based on the similarity between the gene set feature vector GV calculated by the gene feature vector calculation unit 12 and the phenomenon feature vector PV of various GO terms calculated by the phenomenon feature vector calculation unit 13.
[0075] When analyzing multiple pathways included in a pathway, the association estimation unit 14 identifies GO terms that are estimated to be associated with multiple gene sets based on the similarity between the multiple gene set feature vectors GV calculated by the gene feature vector calculation unit 12 for each of the multiple gene sets corresponding to the multiple pathways, and the phenomenon feature vectors PV of various GO terms calculated by the phenomenon feature vector calculation unit 13.
[0076] The association estimation unit 14 outputs the estimated GO term information to the output unit 104, or outputs it externally via the communication interface 105. Here, the association estimation unit 14 may output the GO term information in a manner that allows recognition of the association between the analyzed pathway (gene set) and the estimated GO term.
[0077] For example, as shown in Figure 8(a), the association estimation unit 14 may output table information that associates gene sets with GO terms. Alternatively, as shown in Figure 8(b), the association estimation unit 14 may output the analyzed pathway as a pathway diagram, representing the analyzed pathway in an identifiable manner, and indicating the GO terms associated with that pathway. The similarity score for each pathway calculated by the association estimation unit 14 may also be indicated.
[0078] When analyzing multiple pathways included in a pathway, the overall similarity calculation unit 15 calculates the overall similarity for each GO term by adding the similarity between the gene set feature vector GV and the phenomenon feature vector PV calculated for multiple gene sets that the association estimation unit 14 has estimated to be related to the same GO term.
[0079] The overall similarity calculation unit 15 outputs the calculated overall similarity information along with the corresponding GO term to the output unit 104, or outputs it externally via the communication interface 105. Here, the overall similarity calculation unit 15 may output the overall similarity information in a manner that allows recognition of the link between the multiple routes analyzed and the calculated overall similarity.
[0080] Figure 9 is a flowchart showing an example of the operation of the gene set analysis device 1A according to the first embodiment configured as described above. First, the gene set acquisition unit 11 acquires data for one or more gene sets to be analyzed (step S1).
[0081] Next, the gene feature vector calculation unit 12 calculates the gene set feature vector GV, which is the feature vector of the gene set acquired as the target of analysis by the gene set acquisition unit 11 (step S2). In addition, the phenomenon feature vector calculation unit 13 calculates the phenomenon feature vector PV, which is the feature vector of various GO terms (step S3).
[0082] Next, the association estimation unit 14 identifies GO terms that are estimated to be related to the gene set obtained by the gene set acquisition unit 11, based on the similarity between the gene set feature vector GV of the gene set calculated by the gene feature vector calculation unit 12 and the phenomenon feature vector PV of the GO term calculated by the phenomenon feature vector calculation unit 13 (step S4).
[0083] Subsequently, the overall similarity calculation unit 15 determines whether or not there are multiple gene sets obtained by the gene set acquisition unit 11 in step S1 (step S5). If there are multiple gene sets obtained as targets for analysis, the overall similarity calculation unit 15 calculates the overall similarity for each GO term by adding the similarities calculated for the multiple gene sets that the association estimation unit 14 has estimated to be related to the same GO term (step S6). This completes the processing shown in the flowchart in Figure 9. On the other hand, if there is only one gene set obtained as a target for analysis, the processing in step S6 is skipped, and the processing shown in the flowchart in Figure 9 is completed.
[0084] As explained in detail above, in the first embodiment, GO terms that are estimated to be related to the gene set are identified based on the similarity between a gene set feature vector, which is formed by integrating the constituent gene feature vectors for each of the multiple genes constituting the gene set to be analyzed, and the phenomenon feature vectors, which are the feature vectors of various GO terms.
[0085] ff According to the first embodiment configured in this way, the relationship between a gene set and GO terms is estimated based on the similarity between the gene set feature vector related to the gene set to be analyzed and the phenomenon feature vector related to various GO terms. Therefore, even if the gene set includes genes whose relationship with GO terms is unknown, and regardless of whether the gene set is large or small, it is possible to identify GO terms that are estimated to be related to the gene set. Furthermore, in the first embodiment, since vector analysis is performed taking into account information of all genes constituting the gene set of interest, it is possible to identify GO terms that are more significant to the gene set.
[0086] The following describes an example configuration of a gene set analysis device equipped with a pathway generation function. Figure 10 is a block diagram showing an example of the functional configuration of a gene set analysis device 1B according to a first modified example equipped with a pathway generation function. In Figure 10, components with the same reference numerals as those shown in Figure 3 have the same function, so redundant explanations are omitted here.
[0087] As shown in Figure 10, the gene set analysis device 1B according to the first modified example further includes a pathway generation unit 16 in addition to the configuration shown in Figure 3. The pathway generation unit 16 is connected to a connection information storage unit 23, a property information storage unit 24, and a similarity information storage unit 25, which serve as storage media.
[0088] The pathway generation unit 16 also performs the following processes through the cooperation of hardware and software. Specifically, the processes of the functional blocks 11 to 16 shown in Figure 10 are executed by the operation of a program stored in the storage unit 102, under the control of the control unit 101 shown in Figure 4.
[0089] The connection information storage unit 23 stores connection information that represents the connection relationships between molecules. The connection information includes known information such as interactome information. Interactome information is known information about intermolecular interactions, such as when the expression level of one molecule (gene) increases (or decreases), the expression level of another molecule increases (or decreases) in conjunction with it.
[0090] It should be noted that this interactome information only indicates which molecules have relationships with each other, and does not include information indicating the magnitude of these connections. Furthermore, interactome information is a collection of information indicating the connections between two molecules, and does not indicate sequential connections between three or more molecules.
[0091] The property information storage unit 24 stores property information representing the properties of molecules that act on a disease or symptom. The properties of a molecule indicate whether it acts causally or responsively on the disease or symptom.
[0092] Here, for example, when generating a pathway related to a specific disease (hereinafter referred to as a disease pathway), it is possible to use known information recorded in various literatures and databases as information about multiple molecules related to that disease and information about the properties (causal or responsive) of those molecules that affect the disease. It is also possible to estimate molecules related to a specific disease using a predetermined algorithm and use information about the properties of those molecules that affect the disease estimated using a predetermined algorithm.
[0093] Any algorithm can be used to estimate molecules associated with a disease or to estimate the properties of those molecules. For example, it is possible to estimate new molecules associated with a disease using a machine learning model that utilizes known information about the relationship between diseases and molecules. Similarly, it is possible to estimate the properties of new molecules using a machine learning model that utilizes known information about the properties of molecules that act on diseases.
[0094] The same applies when generating pathways related to specific symptoms (hereinafter referred to as symptom pathways). That is, it is possible to use known information recorded in various literatures and databases as information about multiple molecules related to a symptom and information about the properties of those molecules that affect the symptom. It is also possible to use information obtained by estimating molecules related to a specific symptom using a predetermined algorithm, and by estimating the properties of those molecules that affect the symptom using a predetermined algorithm.
[0095] The similarity information storage unit 25 stores similarity information representing the similarity between molecules. Here, the similarity information storage unit 25 stores information representing the similarity between each molecule for multiple molecules stored as interactome information in the connection information storage unit 23. The similarity between molecules can be evaluated in various ways. For example, it is possible to apply a method in which predetermined feature quantities are extracted for each of the information about multiple molecules (e.g., a description of the molecule, the structure or function of the molecule, etc.) and the similarity of the feature quantities is evaluated. As an example, it is possible to use the similarity of pairs of molecules represented by the fingerprint method.
[0096] As another example, it is possible to calculate molecular feature vectors by analyzing multiple texts (such as literature information) that describe information about a molecule, and then use the similarity of these molecular feature vectors. A molecular feature vector is data that represents the features that a molecule possesses (features that can identify a molecule) as a combination of values of multiple elements. As an example, a vector that shows how much each sentence to which a molecule name, which is included as a word in multiple sentences, contributes to the other sentences can be used as a molecular feature vector.
[0097] Molecular names, as individual words, tend to be used in texts describing molecules, but not in texts unrelated to molecules. Furthermore, even within texts describing molecules, sentences containing a particular molecular name as a word are likely to be texts describing that specific molecule, and are unlikely to contain that same molecular name in texts describing other types of molecules. In other words, the texts containing molecular names as words tend to differ depending on the type of molecule the text is about. Therefore, a vector representing the extent to which a molecular name contributes to a given text can be used as a feature vector capable of identifying molecules.
[0098] Such molecular feature vectors can be calculated using an algorithm similar to the algorithm used to calculate constituent gene feature vectors by the functional blocks 121-124 shown in Figure 5. In other words, molecular feature vectors can be calculated by replacing the constituent gene feature vector identification unit 124 and the gene set feature vector calculation unit 125 shown in Figure 5 with a molecular feature vector identification unit.
[0099] The molecular feature vector identification unit identifies a group of word index values consisting of m index values for each of the multiple molecular names from among n words as the molecular feature vector. That is, as shown in Figure 6, the molecular feature vector identification unit identifies the group of word index values corresponding to the word corresponding to the molecular name from among the n sets of word index value groups (m index values per column) that make up each column of the index value matrix DW as the molecular feature vector for each molecular name. In the constituent gene feature vector identification unit 124 in Figure 5, constituent gene feature vectors were calculated only for the genes included in the gene set to be analyzed, but the molecular feature vector identification unit calculates molecular feature vectors for all molecules (genes).
[0100] The pathway generation unit 16 generates pathways representing intermolecular interactions as pathway diagrams by optimization processing using a minimum flow algorithm, based on property information representing the properties of molecules that act on a specific disease or symptom specified by the user, connection information (interactome information) representing the connection relationships between molecules with respect to the specified disease or symptom, and similarity information representing the similarity between molecules. The interactome information, property information, and similarity information are read and used from the connection information storage unit 23, property information storage unit 24, and similarity information storage unit 25, respectively.
[0101] At this time, the pathway generation unit 16 generates a pathway by arranging the causative molecules (hereinafter referred to as causative molecules) indicated by the property information to be on the upstream side and the responsive molecules (hereinafter referred to as responsive molecules) to be on the downstream side, and by arranging the other molecules (hereinafter referred to as connecting molecules) between the causative molecules and responsive molecules, reflecting the connection relationships indicated by the interactome information, and by arranging the pathway so that molecules with high similarity indicated by the similarity information are more likely to connect with each other. Here, the pathway generation unit 16 sets one starting point further upstream of all causative molecules and one ending point further downstream of all responsive molecules, and applies a minimum flow algorithm to the section from the starting point to the ending point.
[0102] Here, when a specific disease is specified by the user, the pathway generation unit 16 generates an intermolecular network representing the intermolecular interactions of molecules related to the specified disease as a pathway diagram, which is called a disease pathway. Also, when a specific symptom is specified by the user, the pathway generation unit 16 generates an intermolecular network representing the intermolecular interactions of molecules related to the specified symptom as a pathway diagram, which is called a symptom pathway.
[0103] Alternatively, pathways may be generated using the configuration shown in Figure 11 instead of Figure 10. Figure 11 is a block diagram showing an example of the functional configuration of a gene set analysis device 1C according to a second modified example equipped with a pathway generation function. In Figure 11, components with the same reference numerals as those shown in Figure 3 have the same function, so redundant explanations are omitted here. Here, the case of generating disease pathways will be used as an example for explanation.
[0104] As shown in Figure 11, the gene set analysis device 1C according to the second modified example further includes a disease feature vector calculation unit 17, a related molecule estimation unit 18, a molecular property estimation unit 19, and a pathway generation unit 20, in addition to the configuration shown in Figure 3. Furthermore, the gene set analysis device 1C is connected to a third text data storage unit 26, a first model storage unit 27, a second model storage unit 28, and a knowledge DB storage unit 29, which serve as storage media.
[0105] The disease feature vector calculation unit 17 calculates a disease feature vector corresponding to the disease name. The disease feature vector is data that represents the characteristics of a disease (characteristics that can identify a disease) as a combination of values of multiple elements. The disease feature vector calculation unit 17 calculates the disease feature vector using the text data stored in the third text data storage unit 26. That is, the disease feature vector calculation unit 17 calculates the disease feature vector using natural language processing on multiple texts that contain various diseases as words (disease names).
[0106] Such disease feature vectors can be calculated using an algorithm similar to the algorithm used to calculate the constituent gene feature vectors using the functional blocks 121-124 shown in Figure 5. That is, in the index value matrix DW calculated as shown in Figure 6, the group of word index values corresponding to the word of the disease name from among the n groups of word index values constituting each column is identified as the disease feature vector for each disease name.
[0107] The related molecule estimation unit 18 estimates multiple molecules associated with a disease by inputting the disease feature vector calculated by the disease feature vector calculation unit 17 into a first pre-trained model pre-stored in the first model storage unit 27. Here, the first pre-trained model is machine-trained to output information about molecules corresponding to molecular feature vectors similar to disease feature vectors, based on the similarity between disease feature vectors and molecular feature vectors, when a disease feature vector is input.
[0108] The first pre-trained model is generated as follows: Disease feature vectors are calculated for multiple disease names, and molecular feature vectors are calculated for multiple molecule names. Then, machine learning is performed on the first pre-trained model using these datasets, and the first pre-trained model, which has been trained based on the similarity between the disease feature vectors and molecular feature vectors, is stored in the first model storage unit 27.
[0109] The molecular property estimation unit 19 inputs the disease feature vector identified by the disease feature vector calculation unit 17 and the molecular feature vectors identified for multiple molecules estimated by the related molecule estimation unit 18 into the second trained model stored in the second model storage unit 28, thereby estimating the probability that each of the multiple molecules estimated to be related to the disease is causal or responsive as a property of the molecule acting on the disease.
[0110] Here, the second pre-trained model is machine-trained to output the probability that a molecular property is causal or responsive, given disease feature vectors and molecular feature vectors as input, using a dataset of disease feature vectors, molecular feature vectors, and property information representing the properties of molecules that act on the disease as training data.
[0111] The pathway generation unit 20 uses the molecular properties estimated by the molecular property estimation unit 19 and known interactome information indicating the connection relationships between molecules stored in the knowledge DB storage unit 29 to generate pathways that represent intermolecular interactions as pathway diagrams for multiple molecules whose association with a disease has been estimated by the related molecule estimation unit 18. These pathways are arranged so that causative molecules are upstream and responsive molecules are downstream, and other connecting molecules are positioned between the causative molecules and responsive molecules, reflecting the connection relationships indicated by the interactome information.
[0112] At this time, the pathway generation unit 20 generates a pathway using a minimum flow algorithm, such that, for example, causative molecules whose probability value estimated by the molecular property estimation unit 19 is greater than the first threshold Th1 are placed on the upstream side of the pathway, responsive molecules whose probability value is less than the second threshold Th2 (Th1 > Th2) are placed on the downstream side of the pathway, and connecting molecules whose probability value is greater than or equal to the second threshold Th2 and less than or equal to the first threshold Th1 are placed between the causative molecules and the responsive molecules.
[0113] (Second Embodiment) Next, a second embodiment will be described based on the drawings. Figure 12 is a block diagram showing an example of the functional configuration of the gene set analysis device 2 according to the second embodiment. The hardware configuration of the gene set analysis device 2 according to the second embodiment is the same as in Figure 4.
[0114] In Figure 12, components with the same reference numerals as those shown in Figure 3 have the same function, so redundant explanations are omitted here. As shown in Figure 12, the gene set analysis device 2 according to the second embodiment further includes a cluster analysis unit 31 in addition to the configuration shown in Figure 3.
[0115] The cluster analysis unit 31 also performs the following processes through the cooperation of hardware and software. Specifically, the processes of the functional blocks 11 to 15 and 31 shown in Figure 12 are executed by the operation of a program stored in the storage unit 102, under the control of the control unit 101 shown in Figure 4.
[0116] In the second embodiment, by performing processing by the gene set acquisition unit 11, gene feature vector calculation unit 12, phenomenon feature vector calculation unit 13, association estimation unit 14, and overall similarity calculation unit 15 for multiple intermolecular interaction networks (pathways), GO terms that are estimated to be associated with multiple heritability sets included in the pathway are identified for each of the multiple pathways, and the overall similarity of each GO term is calculated.
[0117] The cluster analysis unit 31 classifies the multiple GO terms identified in multiple pathways by the association estimation unit 14 based on their commonalities. For example, the cluster analysis unit 31 classifies the multiple GO terms identified from multiple paths included in multiple pathways into GO terms common to all of the multiple pathways, GO terms common to some of the multiple pathways, and GO terms that have no commonalities.
[0118] Figure 13 is a diagram illustrating the processing content of the cluster analysis unit 31. In Figure 13, the vertical axis is shown as GT 1 ~GT x GT x+1 ~GT y GT y+1 ~GT z (x, y, z are arbitrary numbers) are multiple GO terms identified by the association estimation unit 14 from multiple gene sets (multiple pathways) contained in multiple pathways. Also, the PW shown on the horizontal axis 1 ~PW 5 These are the five pathways that were analyzed.
[0119] Each value shown in the matrix between the vertical and horizontal axes represents the overall similarity calculated by the overall similarity calculation unit 15 for each GO term in each pathway. A total similarity value of "0" means that the GO term was not identified by the association estimation unit 14 as being related to a gene set.
[0120] For example, the first pathway PW 1 Regarding the first pathway PW1 GT is obtained from a plurality of gene sets included in 1 ~GT x , GT x+1 ~GT y while the GO term represented by GT y+1 ~GT z is not identified, indicating that the GO term represented by GT
[0121] On the other hand, regarding the third pathway PW3, GT is obtained from a plurality of gene sets included in the third pathway PW3 1 ~GT x , GT y+1 ~GT z while the GO term represented by GT x+1 ~GT y is not identified, indicating that the GO term represented by GT 5 The same applies to the fourth and fifth pathways PW4 and PW
[0122] Based on the analysis results of the correlation estimation unit 14 as described above, the cluster analysis unit 31 selects a plurality of GO terms (GT 1 ~GT x , GT x+1 ~GT y , GT y+1 ~GT z ), and classifies them into common GO terms (GT 1 ~PW 5 ), which are identified by the correlation estimation unit 14 in all of the five pathways PW 1 ~GT x ), GO terms (GT 1 ), which are common only in the two pathways PW y+1 ~GT z ), and GO terms (GT y+1 ~GT z ), which are common only in the three pathways PW3 to PW5
[0123] According to the classification results in FIG. 13, the GO term (GT 1 ~GT x ) exists in the five pathways PW 1 ~PW5 It can be presumed that this relates to the general properties common to them, and GO term (GT y+1 ~GT z ) has two pathway PW 1 It can be presumed that this is related to a unique property common to PW2, and GO term (GT y+1 ~GT z It can be inferred that this relates to a unique property common to the three pathways PW3 to PW5.
[0124] Note that Figure 13 only shows GO terms common to all or some of the five pathways PW1 to PW5, but it may also include GO terms identified from only one pathway (GO terms that do not share commonality). In that case, the cluster analysis unit 31 classifies the common GO terms identified by the association estimation unit 14 in all or some of the multiple pathways, as well as the non-common GO terms.
[0125] Figure 13 shows the overall similarity score for multiple GO terms, but it is not necessarily required to calculate the overall similarity score. For example, a flag value of "1" or "0" may be set depending on whether or not a GO term has been identified by the association estimation unit 14, and the GO term may be classified based on the flag value.
[0126] The second embodiment configured as described above can be used for various purposes. For example, by performing analysis according to the second embodiment using multiple pathways related to multiple different diseases, the multiple GO terms whose association with a gene set is estimated by the association estimation unit 14 can be classified into GO terms estimated to be related to properties common to all diseases and GO terms estimated to be related to properties common to some diseases. In some cases, it is also possible to identify GO terms estimated to be related to properties of only one disease.
[0127] Furthermore, by performing the analysis according to the second embodiment using multiple pathways related to multiple different symptoms, the multiple GO terms whose association with the gene set is estimated by the association estimation unit 14 can be classified into GO terms estimated to be related to properties common to all symptoms and GO terms estimated to be related to properties common to some symptoms. In some cases, it is also possible to identify GO terms estimated to be related to properties of only one symptom.
[0128] Furthermore, by performing the analysis according to the second embodiment using multiple pathways related to disorders induced by different types of drugs, it is possible to classify the multiple GO terms whose association with a gene set is estimated by the association estimation unit 14 into GO terms estimated to be related to all drugs, GO terms estimated to be related to some drugs, GO terms estimated to be related to only one drug, etc. This makes it possible to classify GO terms estimated to be related to each type of toxicity expression caused by drug ingestion, and to utilize the analysis according to the second embodiment for predicting the toxicity expression mechanism.
[0129] In the second embodiment described above, an example of adding the cluster analysis unit 31 to the configuration shown in Figure 3 was shown, but it is also possible to add the cluster analysis unit 31 to the configuration shown in Figure 10 or Figure 11.
[0130] Furthermore, the embodiments described above (including modifications) are merely examples, and the present invention is not limited thereto. For example, the above embodiments described an example in which GO term is used as an example of a biological phenomenon presumed to be associated with a gene set, but the invention is not limited thereto. For example, one may identify a disease, symptom, or phenotype that is presumed to be associated with a gene set. In this case, the phenomenon feature vector is calculated as data representing the characteristics of the disease, symptom, or phenotype as a combination of values of multiple elements.
[0131] Furthermore, while the above embodiment described an example in which multiple genes included in a small pathway within a pathway are used as the gene set for analysis, it is also possible to use multiple genes included in a large pathway as the gene set for analysis.
[0132] Furthermore, while the above embodiment described an example in which multiple constituent genes included in a pathway are used as the gene set to be analyzed, the method is not limited to this. That is, it is possible to analyze any combination of multiple genes of interest. For example, a set of genes that shows quantitative changes in expression in relation to a specific disease or symptom, or a set of genes grouped by arbitrary cluster analysis, may be used as the target of analysis.
[0133] Furthermore, although the above embodiment describes a configuration in which the gene set analysis devices 1A to 1C and 2 are equipped with a phenomenon feature vector calculation unit 13 to calculate phenomenon feature vectors, the invention is not limited to this configuration. For example, as shown in Figure 14(a), the phenomenon feature vectors may be calculated in advance in a device (not shown) separate from the gene set analysis device 1D and stored in a phenomenon feature vector storage unit 42, and the phenomenon feature vector acquisition unit 41 may acquire and use the phenomenon feature vectors stored in the phenomenon feature vector storage unit 42.
[0134] Furthermore, although the above embodiment describes a configuration in which the gene set analysis devices 1A to 1C and 2 are equipped with a gene feature vector calculation unit 12 to calculate gene set feature vectors, the invention is not limited to this configuration. For example, as shown in Figure 14(b), the gene set feature vectors may be calculated in a device (not shown) separate from the gene set analysis device 1E and stored in a gene feature vector storage unit 44, and the gene feature vector acquisition unit 43 may acquire and use the gene feature vectors stored in the gene feature vector storage unit 44.
[0135] Although Figure 14 shows a modified configuration of the gene set analysis device 1A, it is not limited to this. That is, the gene set analysis devices 1B, 1C, and 2 may also be configured to include a phenomenon feature vector acquisition unit 41 and a phenomenon feature vector storage unit 42 as modified configurations, or to include a gene feature vector acquisition unit 43 and a gene feature vector storage unit 44 as modified configurations of the gene set analysis devices 1B, 1C, and 2.
[0136] Finally, examples of various configurations that can be applied as embodiments of the present invention are summarized below.
[0137] [Configuration 1] A gene set analysis device characterized by comprising a gene set feature vector, which is a feature vector of a gene set to be analyzed, formed by integrating the constituent gene feature vectors for each of the multiple genes constituting the gene set, and a relation estimation unit that identifies biological phenomena that are estimated to be related to the gene set based on the similarity between the gene set feature vector and phenomenon feature vectors, which are feature vectors of various biological phenomena.
[0138] [Configuration 2] The gene set analysis device according to Configuration 1, further comprising a gene feature vector calculation unit that calculates the gene set feature vector using natural language processing on multiple sentences containing various genes as words, wherein the gene feature vector calculation unit calculates the constituent gene feature vector for each of the multiple genes constituting the gene set, and calculates the gene set feature vector by integrating the calculated multiple constituent gene feature vectors.
[0139] [Configuration 3] The gene set analysis device according to Configuration 1 or 2, further comprising a phenomenon feature vector calculation unit that calculates the phenomenon feature vector using natural language processing on multiple sentences containing the above biological phenomena as words.
[0140] [Configuration 4] The gene set analysis device according to any one of Configurations 1 to 3, characterized in that the gene set is a set of genes that constitute one pathway included in a pathway, which is an intermolecular interaction network representing the intermolecular interactions of related molecules as a pathway diagram.
[0141] [Configuration 5] The gene set analysis device according to Configuration 4, characterized in that the association estimation unit identifies the biological phenomena that are estimated to be associated with the multiple gene sets, based on the similarity between the multiple gene set feature vectors and the phenomenon feature vectors for each of the multiple gene sets included as multiple pathways in the pathway.
[0142] [Configuration 6] The gene set analysis device according to Configuration 2 and Configuration 5, characterized in that the gene feature vector calculation unit calculates the gene set feature vector for each of the multiple gene sets included in the pathway as one of the multiple pathways.
[0143] [Configuration 7] The gene set analysis device according to Configuration 5, further comprising a comprehensive similarity calculation unit that calculates a comprehensive similarity for each of the biological phenomena for the multiple pathways targeted for analysis by adding the similarity between the gene set feature vectors calculated for multiple gene sets estimated to be related to the same biological phenomenon by the association estimation unit, and the phenomenon feature vectors.
[0144] [Configuration 8] A gene set analysis device according to any one of Configurations 5 to 7, further comprising a cluster analysis unit that performs the association estimation unit's processing on multiple pathways to identify, for each of the multiple pathways, the biological phenomena that are estimated to be associated with the multiple heritability sets included in the pathway, and classifies the multiple biological phenomena identified by the association estimation unit based on the commonalities of the biological phenomena identified in the multiple pathways by the association estimation unit.
[0145] [Configuration 9] A gene set analysis device according to any one of Configurations 1 to 8, characterized in that the biological phenomenon presumed to be associated with the above gene set is a GO term, disease, symptom, or phenotype.
[0146] Furthermore, the above embodiments are merely examples of how the present invention may be implemented, and the technical scope of the invention should not be interpreted as being limited by them. In other words, the present invention can be implemented in various ways without departing from its gist or its main features.
[0147] 1A-1E,2 Gene Set Analysis Device 11 Gene Set Acquisition Unit 12 Gene Feature Vector Calculation Unit 13 Phenomenon Feature Vector Calculation Unit 14 Relationship Estimation Unit 15 Overall Similarity Calculation Unit 16 Pathway Generation Unit 17 Disease Feature Vector Calculation Unit 18 Related Molecules Estimation Unit 19 Molecular Property Estimation Unit 20 Pathway Generation Unit 21 First Text Data Storage Unit 22 Second Text Data Storage Unit 23 Connection Information Storage Unit 24 Property Information Storage Unit 25 Similarity Information Storage Unit 26 Third Text Data Storage Unit 27 First Model Storage Unit 28 Second Model Storage Unit 29 Knowledge DB Storage Unit 31 Cluster Analysis Unit 41 Phenomenon Feature Vector Acquisition Unit 42 Phenomenon Feature Vector Storage Unit 43 Gene Feature Vector Acquisition Unit 44 Gene Feature Vector Storage Unit
Claims
1. A gene set analysis device characterized by comprising a gene set feature vector, which is a feature vector of a gene set to be analyzed, formed by integrating the constituent gene feature vectors for each of the multiple genes constituting the gene set, and a phenomenon feature vector, which is a feature vector of various biological phenomena, and an association estimation unit that identifies biological phenomena that are presumed to be related to the gene set.
2. The gene set analysis device according to claim 1, further comprising a gene feature vector calculation unit that calculates the gene set feature vector using natural language processing on multiple sentences containing various genes as words, wherein the gene feature vector calculation unit calculates the constituent gene feature vector for each of the multiple genes constituting the gene set, and calculates the gene set feature vector by integrating the calculated multiple constituent gene feature vectors.
3. The gene set analysis device according to claim 1 or 2, further comprising a phenomenon feature vector calculation unit that calculates the phenomenon feature vector using natural language processing on multiple sentences containing the above biological phenomena as words.
4. The gene set analysis device according to claim 1, characterized in that the gene set described above consists of multiple genes that constitute one pathway included in a pathway, which is an intermolecular interaction network representing the intermolecular interactions of related molecules as a pathway diagram.
5. The gene set analysis device according to claim 4, characterized in that the association estimation unit identifies the biological phenomena that are estimated to be associated with the plurality of gene sets, based on the similarity between the plurality of gene set feature vectors and the phenomenon feature vectors for each of the plurality of gene sets included as multiple pathways in the pathway.
6. The gene set analysis device according to claim 2 and claim 5, characterized in that the gene feature vector calculation unit calculates the gene set feature vector for each of the plurality of gene sets included in the pathway as a plurality of pathways.
7. The gene set analysis apparatus according to claim 5, further comprising a comprehensive similarity calculation unit that calculates a comprehensive similarity for each of the biological phenomena for the multiple pathways targeted for analysis by adding the similarity between the gene set feature vectors calculated for multiple gene sets estimated to be related to the same biological phenomenon by the related estimation unit described above and the phenomenon feature vectors described above.
8. The gene set analysis device according to claim 5 or 7, further comprising a cluster analysis unit that performs the association estimation unit processing on multiple pathways to identify, for each of the multiple pathways, the biological phenomena that are estimated to be associated with the multiple heritability sets included in the pathway, and classifies the multiple biological phenomena identified by the association estimation unit based on the commonalities of the biological phenomena identified in the multiple pathways by the association estimation unit.
9. The gene set analysis device according to claim 1, characterized in that the biological phenomenon presumed to be associated with the above gene set is a GO term, disease, symptom, or phenotype.
10. A gene set analysis method characterized in that a computer association estimation unit performs a process to identify biological phenomena that are estimated to be associated with the gene set, based on the similarity between a gene set feature vector, which is a feature vector of the gene set to be analyzed and is formed by integrating the constituent gene feature vectors for each of the multiple genes constituting the gene set, and phenomenon feature vectors, which are feature vectors of various biological phenomena.