Pathway analysis device, pathway analysis method, and pathway analysis program

By applying a minimum flow algorithm to generate pathways based on molecular properties and similarities, the method overcomes the limitations of existing technologies, enabling effective drug discovery by calculating dominance scores for molecular interactions.

JP7755274B2Active Publication Date: 2025-10-16FRONTEO INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024543976
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-06-29
Filing Date
2024-05-15
Publication Date
2025-10-16
Estimated Expiration
2044-05-15

AI Technical Summary

Technical Problem

Existing pathway analysis technologies, such as those described in Patent Document 1, are biased and limited to known intermolecular interactions, making it difficult to analyze beyond the scope of known molecular connections and predict drug discovery targets effectively.

Method used

The use of a minimum flow algorithm to generate pathways based on molecular properties, connection relationships, and similarity information, calculating dominance scores for molecules to represent the strength of their relationship with diseases or symptoms.

Benefits of technology

This approach allows for the generation of unbiased pathways that provide valuable insights for drug discovery by evaluating molecular interactions beyond known intermolecular relationships, identifying potential drug targets and predicting their effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007755274000002
    Figure 0007755274000002
  • Figure 0007755274000003
    Figure 0007755274000003
  • Figure 0007755274000004
    Figure 0007755274000004
Patent Text Reader

Abstract

This pathway analysis device is provided with: a pathway generation unit 11 for generating a pathway by applying a minimum flow algorithm on the basis of property information indicating the property of a molecule acting on a disease or symptom, connection information indicating the connection relationship between molecules, and similarity information indicating the similarity between molecules; and a score calculation unit 12 which, with respect to molecules included in the generated pathway, adds together flow rate values attached to one or more molecules connected directly under one molecule on the basis of the flow rate values attached to the individual molecules, to thereby calculate a superiority score representing the strength of the relationship between a molecule and a disease or symptom for drug discovery. On the basis of the pathway generated by the analysis of an intermolecular relationship exceeding the range of a known relationship with respect to the interaction between two molecules, the superiority score is calculated by effectively utilizing the flow rate values attached to the individual molecules at the time of the analysis, making it possible to obtain useful knowledge about drug discovery beyond the range of the known intermolecular interaction.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a pathway analysis device, a pathway analysis method, and a pathway analysis program, and is particularly suitable for use in a technique for generating pathways that represent intermolecular interactions as a route diagram. [Background technology]

[0002] Pathways (also called molecular networks) are known in the past, which represent molecular interactions as a route diagram. Pathways are expressed by representing molecules such as genes and proteins with symbols such as circles and squares, and connecting the symbols with arrows that represent the molecular interactions. Visualizing molecular interactions in this way makes it easier to understand biological phenomena, for example by identifying which pathways contain genes whose expression levels have changed. Pathways are widely used, for example, in the fields of disease treatment and drug discovery.

[0003] There is a known technology that allows specifying any item such as pathological events, pharmaceutical molecules, genes, or disease-related information, and generating a molecular network that represents the required range of functional and biosynthetic molecular connections through a search, thereby making it possible to estimate bioevents that are directly or indirectly involved in the expression of any biomolecule (see, for example, Patent Document 1).

[0004] Patent Document 1 discloses that an intermolecular network generated by a connect search using a biomolecule-linkage database is scored using a combination of relationship codes, relationship function codes, reliability codes, and several other codes, thereby narrowing down the intermolecular network (for example, a specific group of biomolecules or biomolecule pairs are highlighted and displayed on the intermolecular network).It also discloses that a desired intermolecular network can be found by scoring the molecular connections found using a single or multiple combinations of relationship codes, relationship function codes, reliability codes, directionality of the acting organ or biomolecule pairs, and other information.

[0005] The technology described in Patent Document 1 can be used to select molecular-function networks that are likely to be related to a disease, for example, based on knowledge of bioevents and pathological events that are characteristic of the disease, changes in the amount of biomolecules, etc., and to infer the molecular mechanisms of the disease. It also makes it possible to develop drug discovery strategies, such as determining which process in the network is effective in inhibiting to treat a specific disease or symptom, which molecules in the network are promising as drug discovery targets (proteins and other biomolecules targeted in drug development), what side effects are expected from the drug discovery targets, and what assay systems should be used to select drug development candidates to avoid these side effects.

[0006] However, the technology described in Patent Document 1 calculates scores based on a code indicating the known relationship between two molecules that make up a biomolecule pair, and a code indicating the level of reliability of direct binding for each biomolecule pair and the experimental method that formed the basis for the direct binding, etc., and therefore has problems such as the presence of various biases and the difficulty of performing analysis beyond the scope of known intermolecular interactions.

[0007] [Patent Document 1] WO2003 / 77159 publication Summary of the Invention [Problem to be solved by the invention]

[0008] The present invention has been made to solve such problems, and aims to enable the acquisition of unbiased and useful knowledge for drug discovery that goes beyond the scope of known intermolecular interactions. [Means for solving the problem]

[0009] To solve the above-mentioned problems, the present invention applies a minimum flow algorithm to generate a pathway that represents intermolecular interactions as a path diagram based on property information that represents the properties of molecules that act on a disease or symptom, connection information that represents the connection relationships between molecules, and similarity information that represents the similarity between molecules. Then, for each molecule included in the generated pathway, a dominance score that represents the strength of the relationship between the molecule and the disease or symptom being the target of drug discovery is calculated based on the flow values ​​assigned to each molecule when the minimum flow algorithm is applied. The dominance score of a molecule is calculated based on the flow values ​​assigned to one or more molecules directly connected to the molecule. [Effects of the Invention]

[0010] According to the present invention configured as described above, pathways are generated using a minimum flow algorithm based on property information describing molecular properties, connectivity information describing intermolecular connections, and similarity information describing the similarity between molecules. A molecular advantage score is calculated based on the flow values ​​assigned to each molecule when the minimum flow algorithm is applied. Therefore, based on pathways generated by analyzing intermolecular relationships that go beyond the scope of known interactions between two molecules, the flow values ​​assigned to each molecule during that analysis can be effectively utilized to calculate advantage scores that represent the strength of the relationship between the molecule and a disease or symptom targeted for drug discovery. This advantage score can provide useful insights into drug discovery that go beyond the scope of known intermolecular interactions. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a block diagram showing an example of the functional configuration of a pathway analysis device according to this embodiment. [Figure 2] FIG. 2 is a block diagram illustrating an example of a functional configuration of a feature vector calculation device. [Figure 3] FIG. 10 is a diagram illustrating an example of a molecular feature vector. [Figure 4]FIG. 1 is a schematic diagram for explaining conditions (1) and (2) when executing a minimum flow algorithm. [Figure 5] FIG. 1 is a diagram showing an example of a pathway generation method. [Figure 6] FIG. 10 is a diagram illustrating an example of calculation of a superiority score by the score calculation unit of the present embodiment. [Figure 7] FIG. 10 is a diagram for explaining a first analysis example. [Figure 8] FIG. 10 is a diagram for explaining a second analysis example. [Figure 9] FIG. 10 is a diagram for explaining a third analysis example. [Figure 10] FIG. 10 is a diagram for explaining a fourth analysis example. [Figure 11] FIG. 10 is a block diagram showing another example of the functional configuration of the pathway analysis device according to the present embodiment. [Figure 12] FIG. 10 is a block diagram showing an example of the functional configuration of a pathway analysis device according to another embodiment. [Figure 13] FIG. 2 is a diagram schematically illustrating an example of disease genome analysis executed by a disease genome analysis unit. DETAILED DESCRIPTION OF THE INVENTION

[0012] An embodiment of the present invention will be described below with reference to the drawings. Fig. 1 is a block diagram showing an example of the functional configuration of a pathway analysis device 10 according to this embodiment. As shown in Fig. 1, the pathway analysis device 10 of this embodiment includes, as functional components, a pathway generation unit 11 and a score calculation unit 12. In addition, a connection information storage unit 13, a property information storage unit 14, and a similarity information storage unit 15 are connected to the pathway analysis device 10 of this embodiment as storage media.

[0013] The functional blocks 11 and 12 can be configured using any of hardware, a DSP (Digital Signal Processor), and software. For example, the functional blocks 11 and 12 are realized by the operation of a pathway analysis program stored in a storage medium such as RAM, ROM, a hard disk, or a semiconductor memory under the control of a microcomputer configured with a CPU, RAM, ROM, etc.

[0014] The connection information storage unit 13 stores connection information that represents the connection relationships between molecules. The connection information includes known information such as interactome information. The interactome information is known information regarding molecular interactions, such as, for example, when the expression level of a certain molecule (gene or protein) increases (or decreases), the expression level of another molecule increases (or decreases) in conjunction with this.

[0015] Note that this interactome information only indicates which molecules are related to each other, and does not include information indicating the magnitude of the connection. Also, interactome information is a collection of information indicating the connection between two molecules, not information indicating the sequential connection between three or more molecules.

[0016] The property information storage unit 14 stores property information that indicates the properties of molecules that act on diseases or symptoms. The property of a molecule indicates whether the molecule acts as a causative or responsive molecule on a disease or symptom. Causative property is the property that the presence or mutation of the molecule may cause a disease or symptom. Responsive property is the property that the molecule may mutate due to the onset of a disease or symptom.

[0017] When generating a pathway related to a specific disease (hereinafter referred to as a disease pathway), it is possible to use known information recorded in various literature, databases, etc. as information on multiple molecules related to the disease and information on the properties (causality or responsiveness) of the actions of these molecules on the disease. It is also possible to estimate molecules related to a specific disease using a predetermined algorithm and use information on the properties of the actions of specific molecules on the disease estimated using the predetermined algorithm.

[0018] Any algorithm can be used as an algorithm for predicting disease-related molecules or molecular properties. For example, a new molecule related to a disease can be predicted by a prediction model trained by machine learning using known information about the association between a disease and a molecule. Furthermore, the properties of a new molecule can be predicted by a prediction model trained by machine learning using known information about the properties of molecules that act on a disease.

[0019] The same applies to generating a pathway related to a specific symptom (hereinafter referred to as a symptom pathway). That is, known information recorded in various literature, databases, etc. can be used as information on multiple molecules related to the symptom and information on the properties of the actions of those molecules on the symptom. It is also possible to estimate molecules related to a specific symptom using a predetermined algorithm, and use information on the properties of the actions of the specific molecules on the symptom estimated using the predetermined algorithm.

[0020] The similarity information storage unit 15 stores similarity information representing the similarity between molecules. Here, the similarity information storage unit 15 stores information representing the similarity between each of a plurality of molecules stored as interactome information in the connection information storage unit 13. The similarity between molecules can be evaluated using various methods. For example, a method can be applied in which predetermined features are extracted from information about a plurality of molecules (e.g., descriptions of the molecules, structures or functions of the molecules, etc.) and the similarity of the features is evaluated. As an example, the similarity between pairs of molecules expressed by fingerprinting can be used.

[0021] As another example, it is possible to calculate molecular feature vectors by analyzing multiple sentences (such as literature information) that describe information about molecules, and use the similarity of the molecular feature vectors. A molecular feature vector is data that represents the features (features that can identify molecules) possessed by molecules such as proteins and genes as a combination of multiple element values. As an example, a vector that represents the degree to which a molecule name contained as a word in multiple sentences contributes to each sentence is used as the molecular feature vector.

[0022] Molecular names as words tend to be used in sentences describing molecules, but not in sentences unrelated to molecules. Furthermore, among sentences describing molecules, sentences containing a certain molecular name as a word are likely to describe that molecule, while sentences describing other types of molecules are likely to not contain that molecular name. In other words, sentences containing molecular names as a word tend to differ depending on the type of molecule they are about. Therefore, a vector that represents the degree to which a molecular name contributes to each sentence can be used as a feature vector capable of identifying molecules.

[0023] Such a molecular feature vector can be calculated, for example, using the algorithm described in Japanese Patent No. 6915818. A method for calculating a molecular feature vector using the algorithm described in Japanese Patent No. 6915818 will be briefly described below. Figure 2 is a block diagram showing an example of the functional configuration of a feature vector calculation device 200.

[0024] In FIG. 2, the word extraction unit 201 analyzes m sentences (m is any integer equal to or greater than 2) and extracts n words (n is any integer equal to or greater than 2) from the m sentences. For example, known morphological analysis can be used to analyze the sentences. Note that the m sentences may contain multiple instances of the same word. In this case, the word extraction unit 201 does not extract multiple instances of the same word, but extracts only one. In other words, the n words extracted by the word extraction unit 201 mean n types of words.

[0025] The vector calculation unit 202 calculates m sentence vectors and n word vectors from m sentences and n words. Here, the sentence vector calculation unit 202A calculates m sentence vectors d consisting of q axis components (q is an arbitrary integer equal to or greater than 2) by vectorizing each of the m sentences that are the subject of analysis by the word extraction unit 201 into q dimensions according to a predetermined rule. i →(i=1,2,...,m) (the symbol "→" indicates a vector). In addition, the word vector calculation unit 202B vectorizes each of the n words extracted by the word extraction unit 201 into q dimensions according to a predetermined rule, thereby generating n word vectors w consisting of q axis components. j → (j=1,2,...,n) is calculated. i → and word vector w j →We will not explain the specific calculation method here.

[0026] The index value calculation unit 203 calculates m sentence vectors d i → and n word vectors w j By taking the dot product of each of them, m sentences di and n words w j Here, the index value calculation unit 203 calculates an index value that reflects the relationship between m sentence vectors d i →each q axis component (d 11 ~d mq ) and n word vectors w j →each q axis component (w 11 ~w nq ) as elements of the word matrix W, and calculates an index value matrix DW having m×n index values ​​as elements. t is the transpose of the word matrix.

[0027]

number

[0028] Each element of the index value matrix DW calculated in this way can be said to represent the degree to which a word contributes to a sentence, and the degree to which a sentence contributes to a word. For example, the element dw in the first row and second column 12 can be said to be a value that represents the degree to which word w2 contributes to sentence d1, and also a value that represents the degree to which sentence d1 contributes to word w2. As a result, each row of the index value matrix DW can be used to evaluate the similarity of sentences, and each column can be used to evaluate the similarity of words.

[0029] The feature vector identification unit 204 identifies, for each of a plurality of molecular names among the n words, a word index group consisting of m indexes for one molecular name as a molecular feature vector. That is, as shown in Fig. 3, the feature vector identification unit 204 identifies, as a molecular feature vector for each molecular name, a word index group for a word corresponding to a molecular name among n sets of word index groups (m indexes per column) constituting each column of the index matrix DW.

[0030] The similarity of molecular feature vectors can be evaluated in various ways. For example, a method can be applied in which a feature amount is extracted for each molecular feature vector using a predetermined function and the similarity of the feature amount is evaluated. Alternatively, the Euclidean distance or cosine similarity between molecular feature vectors may be used, or the edit distance may be used.

[0031] Returning to Figure 1, the pathway generation unit 11 generates a pathway that represents intermolecular interactions as a path diagram by an optimization process using a minimum flow algorithm based on property information that represents the properties of molecules that act on a specific disease or symptom specified by, for example, a user, connection information (interactome information) that represents the connection relationships between molecules related to the specified disease or symptom, and similarity information that represents the similarity between molecules. The interactome information, property information, and similarity information are read out and used from the connection information storage unit 13, property information storage unit 14, and similarity information storage unit 15, respectively.

[0032] In this case, the pathway generation unit 11 generates a pathway by positioning causative molecules (hereinafter referred to as causative molecules) indicated by the property information upstream and responsive molecules (hereinafter referred to as responsive molecules) downstream, and by positioning other molecules (hereinafter referred to as connecting molecules) between the causative molecules and responsive molecules in a manner that reflects the connection relationships indicated by the interactome information, and by making it easier for molecules with high similarity indicated by the similarity information to connect to each other. Here, the pathway generation unit 11 sets one start point further upstream of all causative molecules and one end point further downstream of all responsive molecules, and applies the minimum flow algorithm to the section from the start point to the end point.

[0033] Here, when a specific disease is designated by the user, the pathway generation unit 11 generates an intermolecular network representing, as a disease pathway, the intermolecular interactions of molecules related to the designated disease as a pathway diagram.Furthermore, when a specific symptom is designated by the user, the pathway generation unit 11 generates, as a symptom pathway, an intermolecular network representing, as a pathway diagram, the intermolecular interactions of molecules related to the designated symptom.

[0034] Furthermore, in this embodiment, the pathway generation unit 11 applies the minimum flow algorithm so as to satisfy at least the following two conditions. (1) Generate a path by connecting only molecules whose flow rate is greater than the minimum threshold and less than the maximum threshold. (2) Generate a route by connecting only molecules whose flow rate is below a threshold value related to similarity.

[0035] Here, the flow rate value is an evaluation value calculated for each molecule when connecting the molecules using the minimum flow algorithm. This evaluation value is calculated so as to satisfy the following conditions: (A) The flow rate value of the starting point set upstream of the causative molecule is distributed sequentially to each causative molecule, connecting molecule, and responsive molecule, and the flow rate value aggregated at the end point set downstream of the responsive molecule is the same as the flow rate value of the starting point. The flow rate value of each molecule is smaller than the flow rate values ​​of the starting and end points. (B) The magnitude of the flow value of each molecule is determined taking into account the similarity between the molecules. For example, the flow value is determined based on a predetermined function or calculation algorithm that uses the similarity as a variable. This function or calculation algorithm is designed so that the flow value increases as the similarity increases. (C) The flow value is normalized to a value between 0 and 1. The magnitude of the similarity between molecules is also normalized to a value between 0 and 1. In this case, the flow value may be set to the similarity between molecules.

[0036] Regarding condition (1), for example, the minimum threshold for flow values ​​can be set to "0." This means that pathways containing molecules with a flow value of "0" are not adopted, and the flow values ​​of each molecule included in the pathway to be generated will not be "0." The maximum threshold for flow values ​​in condition (1) may be a value different from or the same as the threshold for similarity in condition (2). If the values ​​are the same, this means that the maximum threshold for flow values ​​in condition (1) is determined by the threshold for similarity (condition (2) is incorporated into condition (1)).

[0037] Figure 4 is a schematic diagram for explaining the above conditions (1) and (2). In the pathway shown in Figure 4, diamond-shaped nodes represent causal molecules, square-shaped nodes represent responsive molecules, and oval-shaped nodes represent connecting molecules (the same applies to other figures described later).

[0038] As shown in Figure 4(a), two possible routes are possible based on interactome information for the route from the causative molecule 41 to the responsive molecule 46. That is, the first route is causative molecule 41 → connecting molecules 42, 43, 44 → responsive molecule 46, and the second route is causative molecule 41 → connecting molecules 42, 45 → responsive molecule 46. Here, for example, if the flow value of the connecting molecule 44 is calculated to be zero, the first route is eliminated and the second route is adopted.

[0039] 4(b), in a situation where the first and second paths are possible, as in FIG. 4(a) (however, the flow value of the connecting molecule 44 is ≠ 0), the flow value of one connecting molecule 43 branching from the connecting molecule 42 is calculated to be 0.2, and the flow value of the other connecting molecule 45 is calculated to be 0.6. If the threshold for similarity is 0.4, the flow value of the connecting molecule 45 exceeds the upper limit of similarity, so the second path is eliminated and the first path is adopted. Note that if the flow values ​​of the connecting molecules 43 and 45 are both less than 0.4, both the first and second paths are adopted.

[0040] Figure 5 is a diagram showing an example of a pathway generation method by the pathway generation unit 11. First, as shown in Figure 5(a), the pathway generation unit 11 places causative molecules upstream and responsive molecules downstream. Then, it sets a start point further upstream of all causative molecules and a finish point further downstream of all responsive molecules.

[0041] Next, as shown in Figure 5(b), the pathway generation unit 11 applies a minimum flow algorithm to the section from the start point to the end point, and generates a pathway by placing connecting molecules between the causative molecules and the responsive molecules based on the connection relationships indicated by the interactome information. At this time, the pathway generation unit 11 calculates the flow value of each molecule so as to satisfy conditions (A) to (C), and performs optimization processing based on the calculated flow value so as to satisfy conditions (1) and (2).

[0042] Returning to Figure 1, the score calculation unit 12 calculates a dominance score representing the strength of the relationship between the molecule and the disease or symptom for which drug discovery is targeted, for each molecule included in the pathway generated by the pathway generation unit 11, based on the flow value assigned to each molecule when applying the minimum flow algorithm. Here, the score calculation unit 12 calculates the dominance score of a molecule based on the flow values ​​assigned to one or more molecules connected directly below the molecule.

[0043] 6 is a diagram showing an example of calculation of dominance scores by the score calculation unit 12. FIG. 6 shows an example of calculation of dominance scores Sc61 to Sc65 for five causative molecules 61 to 65 included in a pathway. For example, the dominance scores Sc61 and Sc65 of the leftmost causative molecule 61 and the rightmost causative molecule 65 are calculated based on the flow values ​​assigned to the two connecting molecules connected directly below the causative molecules 61 and 65, respectively. For example, the sum of the flow values ​​assigned to the two connecting molecules is calculated as the dominance scores Sc61 and Sc65.

[0044] Furthermore, the dominance score Sc62 of the second causative molecule 62 from the left is calculated by adding up the flow values ​​assigned to the three connecting molecules connected directly below the causative molecule 62. The dominance scores Sc63 and Sc64 of the third and fourth causative molecules 63 and 64 are calculated by adding up the flow value assigned to one connecting molecule connected directly below the causative molecules 63 and 64, respectively. In this case, the flow value = dominance score.

[0045] Here, an example of calculating the dominance scores Sc61 to Sc65 of the causal molecules 61 to 65 is shown, but the dominance scores of the connecting molecules are calculated in the same way. Note that here, an example has been described in which the sum of the flow values ​​assigned to one or more molecules connected directly below is calculated as the dominance score, but this is not limiting. For example, the multiplication or average value of the flow values ​​assigned to one or more molecules connected directly below may also be calculated as the dominance score.

[0046] Furthermore, the dominance scores of each molecule may be calculated as a relative value using the dominance scores calculated for multiple pathways. For example, the dominance score of a molecule calculated as described above within a pathway may be divided by the maximum dominance score of all molecules included in multiple pathways, and the resulting value may be used as the dominance score of the molecule.

[0047] As described above, in this embodiment, pathways are generated using a minimum flow algorithm based on property information that represents the properties of molecules and connection information (interactome information) that represents the connection relationships between molecules, as well as similarity information that represents the similarity between molecules, and molecular dominance scores are calculated based on the flow values ​​assigned to each molecule when applying the minimum flow algorithm.

[0048] Therefore, based on pathways generated by analyzing intermolecular relationships that go beyond the scope of known interactions between two molecules, it is possible to effectively utilize the flow values ​​assigned to individual molecules during that analysis to calculate a dominance score that represents the strength of the relationship between the molecule and the disease or symptom that is the target of drug discovery. This dominance score makes it possible to obtain useful insights for drug discovery that go beyond the scope of known intermolecular interactions.

[0049] An example of analysis using the superiority score calculated as above will be described below.

[0050] <First analysis example> The first analysis example involves analyzing pathways from the perspective of determining which disease or symptom should be targeted for a certain molecule (e.g., a drug discovery candidate gene). Figure 7 is a diagram for explaining the first analysis example.

[0051] In the first analysis example, the pathway generation unit 11 generates multiple pathways that represent, as a pathway diagram, intermolecular interactions between molecules related to multiple diseases or multiple symptoms. Figure 7 shows an example in which multiple disease pathways related to diseases A to Z are generated.

[0052] The score calculation unit 12 calculates, for each of the multiple pathways, the dominance scores Sc70A to Sc70Z of the target molecule 70 that is commonly included in the multiple pathways. The score calculation unit 12 may further rank the multiple dominance scores Sc70A to Sc70Z calculated for each of the multiple pathways in descending order of their values.

[0053] By calculating the dominance score of a molecule as in the first analysis example, it is possible to rank a disease or symptom for a molecule based on the dominance score in terms of the likelihood that the disease or symptom is a candidate for the target disease or symptom. It is also possible to identify the molecule with the highest dominance score as the target disease or symptom.

[0054] <Second analysis example> In the second analysis example, pathway analysis is performed from the perspective of determining which molecules should be targeted for drug discovery for a certain disease or symptom. Figure 8 is a diagram for explaining the second analysis example.

[0055] In the second analysis example, the pathway generation unit 11 generates a pathway that represents, as a path diagram, the intermolecular interactions of molecules related to a disease or a symptom. Figure 8 shows an example of generating a disease pathway related to disease A.

[0056] The score calculation unit 12 calculates a dominance score for each of a plurality of molecules included in one pathway. FIG. 8 shows an example in which dominance scores are calculated for all causative molecules and all connecting molecules included in a pathway (for ease of viewing the figure, only some molecules and their corresponding dominance scores are labeled 81-89, Sc81-Sc89). Note that dominance scores may be calculated only for causative molecules, or only for connecting molecules.

[0057] The score calculation unit 12 may further rank the multiple dominance scores calculated for each of the multiple molecules included in one pathway in descending order of their values.

[0058] By calculating the dominance score of a molecule as in the second analysis example, it is possible to rank the molecules by their dominance score in terms of the likelihood that they will be candidate target molecules for a particular disease or symptom. It is also possible to identify the molecule with the highest dominance score as the target molecule.

[0059] <Third analysis example> In the third analysis example, pathway analysis is performed from the viewpoint of predicting suppressor genes (suppressor genes) for a certain disease or symptom. Figure 9 is a diagram for explaining the third analysis example.

[0060] In the third analysis example, the pathway generation unit 11 generates a pathway that represents molecular interactions as a path diagram, starting from only one causative molecule among molecules related to one disease or one symptom. Figure 9 shows a pathway generated by specifying a loss-of-function causative molecule 90 among the five causative molecules (◇) included in the pathway of disease A shown in Figure 7, and setting only the causative molecule 90 as the starting point and the seven responsive molecules (□) included in the pathway of disease A shown in Figure 7 as the end points.

[0061] The score calculation unit 12 calculates a dominance score for each of the multiple molecules included in the pathway of disease A, excluding one causative molecule 90. FIG. 9 shows an example in which dominance scores are calculated for all connecting molecules. Note that dominance scores may be calculated for only some of the connecting molecules (for example, connecting molecules connected directly below the causative molecule 90).

[0062] The score calculation unit 12 may further rank the multiple dominance scores calculated for each of the multiple molecules included in one pathway in descending order of their values.

[0063] By calculating the dominance score of a molecule as in the third analytical example, it is possible to rank the likelihood that a molecule will be a candidate suppressor molecule for a particular disease or symptom based on the dominance score. Because there is a high probability that a suppressor gene is found among molecules with a high dominance score, it is possible to select candidate molecules for analysis as suppressors based on the dominance score.

[0064] <4th analysis example> In the fourth analysis example, a pathway is analyzed from the viewpoint of what effect occurs on the entire pathway when a specific molecule is knocked out in the pathway once it has been generated. Figure 10 is a diagram for explaining the fourth analysis example.

[0065] FIG. 10 shows an example of regenerating a pathway by knocking out one connecting molecule 50 contained in a pathway that has already been generated. FIG. 10(a) shows the initially generated pathway, in which one connecting molecule 50 is knocked out. The molecule to be knocked out can be specified arbitrarily. For example, it is possible to knock out a molecule that is a target gene for drug discovery.

[0066] Knocking out a molecule 50 is equivalent to forcibly setting the flow value of that molecule 50 to zero. Therefore, as shown in FIG. 10(b), the pathway that passes through the knocked-out molecule 50 is eliminated. When a pathway including the knocked-out molecule 50 disappears, the flow value that was allocated to that pathway is allocated to another pathway. At that time, the pathway generation unit 11 recalculates the flow value of each molecule so as to satisfy conditions (A) to (C), and re-executes the optimization process based on the recalculated flow values ​​so as to satisfy conditions (1) and (2).

[0067] Therefore, in addition to the pathway containing the knocked-out molecule 50, a pathway that was adopted before knockout because it satisfied conditions (1) and (2) may disappear after knockout because it no longer satisfies these conditions. Conversely, a pathway that was not adopted before knockout because it did not satisfy conditions (1) and (2) may be adopted after knockout because it satisfies these conditions. As a result, the pathway before knockout shown in Figure 10(a) changes to the pathway shown in Figure 10(c). In Figure 10(c), dotted lines indicate the eliminated pathways, and bold lines indicate the newly emerged pathways.

[0068] The score calculation unit 12 calculates a dominance score for each of the multiple molecules included in the pathway regenerated by knocking out a specific molecule 50. Figure 10(c) shows an example in which a dominance score is calculated for each of all causal molecules (◇), but dominance scores may also be calculated for all connecting molecules (oval symbols). The score calculation unit 12 may further rank the multiple dominance scores calculated for each of the multiple molecules included in the regenerated pathway in descending order of their values.

[0069] Furthermore, the score calculation unit 12 may calculate a dominance score for each of the molecules included in the pathway (pathway in FIG. 10(a)) before knocking out the specific molecule 50, and may also calculate a dominance score for each of the molecules included in the pathway (pathway in FIG. 10(c)) regenerated by knocking out the specific molecule 50, and calculate the difference between the dominance scores calculated for the same molecule before and after knocking out. The score calculation unit 12 may further rank the molecules based on the magnitude of the calculated difference.

[0070] Pathways regenerated by knocking out specific molecules are thought to be pathways that show drug resistance (treatment resistance) when drug treatment targeting that specific molecule is administered. By calculating the dominance score for those pathways, it is possible to rank them by their likelihood of becoming candidate drug target molecules for treatment-resistant diseases.

[0071] In the above embodiment, in addition to the conditions (1) and (2), one or both of the following conditions (3) and (4) may be added. (3) The network size of the pathway must be within a predetermined range. The network size may be, for example, the number of molecules on the path from the causative molecule to the responsive molecule, or the total number of molecules in the entire pathway. (4) When multiple paths are possible from a causative molecule to a responsive molecule, the shortest path (the path containing the fewest molecules) is selected. Alternatively, the path with the largest total flow rate from the causative molecule to the responsive molecule is selected.

[0072] In the above embodiment, the flow value of each molecule is a value determined taking into account the similarity between molecules, but the present invention is not limited to this. For example, the flow value of each molecule may be a value calculated independently of the similarity between molecules, and the flow value of each molecule may not be the similarity between molecules.

[0073] Furthermore, the disease pathway or symptom pathway may be generated using the algorithm described in Japanese Patent No. 6915818, for example.

[0074] Below, we will briefly explain a pathway generation method using the algorithm described in Japanese Patent No. 6915818. Figure 11 is a block diagram showing an example of the functional configuration of a pathway generation device 100 when this algorithm is used. Note that the explanation will be given here using an example of generating a disease pathway.

[0075] The disease feature vector identification unit 101 identifies a feature vector corresponding to the disease name (hereinafter referred to as a disease feature vector). The disease feature vector is data that represents the features of a disease (features that can identify the disease) as a combination of multiple element values. As an example, a vector that represents the degree to which a disease name contained as a word in multiple sentences contributes to each sentence is used as the disease feature vector.

[0076] Such a disease feature vector can be calculated, for example, by the feature vector calculation device 200 shown in Fig. 2. The feature vector calculation device 200 inputs sentence data related to a sentence and calculates a disease feature vector that reflects the relationship between the sentence and the words contained therein. That is, in the index value matrix DW calculated as shown in Fig. 3, of the n sets of word index values ​​constituting each column, a word index value group related to a word corresponding to a disease name is identified as the disease feature vector for each disease name.

[0077] The associated molecule inference unit 102 infers a plurality of molecules associated with the disease by inputting the disease feature vector identified by the disease feature vector identification unit 101 into a first trained model pre-stored in the first model storage unit 111. Here, the first trained model has been machine-trained so that when a disease feature vector is input, it outputs information about molecules corresponding to a molecular feature vector similar to the disease feature vector, based on the similarity between the disease feature vector and the molecular feature vector.

[0078] The feature vector calculation device 200 calculates disease feature vectors for a plurality of disease names and molecule feature vectors for a plurality of molecule names. Then, machine learning of a first trained model is performed in advance using these data sets, and the first trained model trained based on the similarity between the disease feature vectors and the molecule feature vectors is stored in the first model storage unit 111.

[0079] Here, the similarity between the disease feature vector and the molecule feature vector can be evaluated by various methods. For example, a method can be applied in which features are extracted from each of the disease feature vector and the molecule feature vector using a predetermined function, and the similarity of the features is evaluated. Alternatively, the Euclidean distance or cosine similarity between the word index value group of the disease feature vector and the word index value group of the molecule feature vector may be used, or the edit distance may be used.

[0080] The molecular property estimation unit 103 inputs the disease feature vector identified by the disease feature vector identification unit 101 and the molecular feature vector identified for multiple molecules estimated by the related molecule estimation unit 102 into a second trained model stored in the second model memory unit 112, and estimates the probability that each of the multiple molecules estimated to be associated with the disease is causative or reactive as a molecular property acting on the disease.

[0081] Here, the second trained model has been machine-trained to output the probability that the molecular property is causative or responsive when a disease feature vector and a molecular feature vector are input, using a dataset of disease feature vectors, molecular feature vectors, and property information representing the properties of molecules that act on the disease as training data.

[0082] The pathway generation unit 104 uses the molecular properties estimated by the molecular property estimation unit 103 and known interactome information indicating the connection relationships between molecules stored in the knowledge DB storage unit 113 to generate a pathway (intermolecular network) that represents intermolecular interactions as a route diagram, by arranging the causative molecules upstream and the responsive molecules downstream for multiple molecules whose association with a disease has been estimated by the related molecule estimation unit 102, and by placing other connecting molecules between the causative molecules and the responsive molecules in a manner that reflects the connection relationships indicated by the interactome information.

[0083] In the pathway analysis device 10 shown in Fig. 1, when a pathway is generated using the algorithm described in Japanese Patent No. 6915818 described above, the pathway generation unit 104 corresponds to the pathway generation unit 11 in Fig. 1, and the knowledge DB storage unit 113 corresponds to the connection information storage unit 13 in Fig. 1. Furthermore, the pathway generation unit 11 generates a pathway using the estimation results by the molecular property estimation unit 103 instead of the property information storage unit 14 in Fig. 1.

[0084] At this time, the pathway generation unit 11 calculates the flow value of each molecule to satisfy conditions (A) to (C), and generates a pathway using a minimum flow algorithm based on the calculated flow values ​​to satisfy conditions (1) and (2), for example, so that causative molecules whose probability value estimated by the molecular property estimation unit 103 to be causative is greater than the first threshold Th1 are placed upstream of the pathway, responsive molecules whose probability value is less than the second threshold Th2 (Th1>Th2) are placed downstream of the pathway, and connecting molecules whose probability value is greater than the second threshold Th2 and less than the first threshold Th1 are placed between the causative molecule and the responsive molecule.

[0085] Alternatively, the pathway generation unit 11 may place causative molecules whose probability value estimated by the molecular property estimation unit 103 to be causative is greater than a first threshold Th1 upstream of the pathway, and responsive molecules whose probability value estimated to be responsive is greater than a third threshold Th3 (either Th1=Th3 or Th1≠Th3) downstream of the pathway, and place other connecting molecules between the causative molecules and the responsive molecules.

[0086] Here, as the similarity used in condition (B) when calculating the flow rate value, instead of the similarity information stored in the similarity information storage unit 15, for example, a value of the similarity between the disease feature vector and the molecular feature vector identified by the associated molecule inference unit 102 may be used. Also, as a threshold value for the similarity used in condition (2), a threshold value for the similarity between the disease feature vector and the molecular feature vector may be used.

[0087] As another example, instead of placing a causative molecule estimated by the molecular property estimation unit 103 upstream of a pathway, a gene whose expression level has been found to fluctuate by a predetermined disease genome analysis may be placed upstream as a causative molecule. Figure 12 is a block diagram showing an example of the functional configuration of the pathway analysis device 10' in this case. In Figure 12, components with the same reference numerals as those in Figures 1 and 11 have the same functions, and therefore a redundant description will be omitted here.

[0088] In the example shown in FIG. 12, a pathway analysis device 10′ includes a pathway generation unit 11′ instead of the pathway generation unit 104 (corresponding to the pathway generation unit 11) shown in FIG. 11. A disease genome analysis unit 20 is also connected to the pathway analysis device 10′. For example, a first information processing device including the pathway analysis device 10′ and a second information processing device including the disease genome analysis unit 20 are configured as separate entities, and the first information processing device and the second information processing device are connected via a communication network.

[0089] The pathway generation unit 11' of the pathway analysis device 10' obtains and uses information on the analysis results by the disease genome analysis unit 20 via a communication network. Note that it is not essential that the pathway analysis device 10' and the disease genome analysis unit 20 are connected via a communication network. For example, the pathway analysis unit 10' may be configured to obtain information on the analysis results by the disease genome analysis unit 20 via removable media. Furthermore, the pathway analysis device 10' may be configured to have the processing functions of the disease genome analysis unit 20.

[0090] The disease genome analysis unit 20 identifies genes whose expression levels are found to be fluctuating by analyzing gene loci where gene mutations associated with a disease cause fluctuations. Figure 13 is a diagram schematically illustrating an example of disease genome analysis performed by the disease genome analysis unit 20. Figure 13(a) shows the results of analyzing a portion of the gene sequence of a chromosome extracted from a healthy individual and the amount of each gene contained in the chromosome. Figure 13(b) shows the results of analyzing a portion of the gene sequence of a chromosome extracted from a patient suffering from a specific disease and the amount of each gene contained in the chromosome.

[0091] FIG. 13 shows that a mutation in a certain gene A associated with a disease affects the expression level of another gene F. A gene whose expression level fluctuates in this way is called an eGene. The disease genome analysis unit 20 performs disease genome analysis to identify such eGene. This disease genome analysis can be, for example, a known eQTL (expression Quantitative Trait Locus) analysis, but is not limited to this.

[0092] It is known that analyzing eGene, whose expression varies depending on the disease, may lead to the discovery of drug discovery targets. However, despite the vast amount of disease genomic information that has been obtained through conventional disease genome analysis, there are many diseases for which no treatment has been developed due to the difficulty of identifying complex disease mechanisms. One reason for this is thought to be the difficulty in linking genetic information with the onset of disease using conventional genome analysis methods.

[0093] To address these conventional problems, the pathway generation unit 11' of the pathway analysis device 10' shown in Figure 12 places eGenes whose expression levels have been found to fluctuate through disease genome analysis by the disease genome analysis unit 20 as causative molecules upstream of the pathway. Furthermore, among multiple molecules whose association with the disease has been inferred by the related molecule inferring unit 102, molecules inferred to be responsive by the molecular property inferring unit 103 are placed downstream, and connecting molecules other than the molecules inferred to be causative and responsive are placed between the causative molecule and the responsive molecule, reflecting the connection relationships indicated by the interactome information stored in the connection information storage unit 13.

[0094] Then, with the above-described arrangement, the pathway generation unit 11' calculates the flow value of each molecule so as to satisfy conditions (A) to (C), and generates a pathway by applying a minimum flow algorithm based on the calculated flow values ​​so as to satisfy conditions (1) and (2). In the configuration example shown in FIG. 12, the similarity value between the disease feature vector and the molecular feature vector identified by the related molecule inference unit 102 is used as the similarity used in condition (B) when calculating the flow value. Note that a similarity information storage unit 15 may be provided, and the similarity information stored in the similarity information storage unit 15 may be used.

[0095] As described above, by configuring the pathway analysis device 10' as shown in Figure 12 and performing the processing of, for example, the second analysis example described above, it is possible to utilize the disease genome information analyzed by the disease genome analysis unit 20 to search for genes that are potential drug discovery targets.

[0096] 12 illustrates an example in which a molecule estimated to be responsive by the molecular property estimation unit 103 is placed downstream, but the present invention is not limited to this. For example, a property information storage unit 14 may be provided as in FIG. 1, and the responsive molecule indicated by the property information may be placed downstream of the pathway. If a similarity information storage unit 15 is also provided in addition to this, the disease feature vector identification unit 101, the associated molecule estimation unit 102, the molecular property estimation unit 103, the first model storage unit 111, and the second model storage unit 112 can be omitted.

[0097] Furthermore, the above-described embodiments are merely examples of specific embodiments for carrying out the present invention, and the technical scope of the present invention should not be construed as being limited thereby. In other words, the present invention can be carried out in various forms without departing from the gist or main characteristics thereof. [Explanation of symbols]

[0098] 10,10' Pathway Analysis Device 11,11' Pathway generation section 12 Score calculation section 13 Connection information storage unit 14 Property information storage unit 15 Similarity information storage unit 20 Department of Genomics

Claims

1. a pathway generation unit that generates a pathway that represents intermolecular interactions as a path diagram by applying a minimum flow algorithm based on property information that represents the properties of molecules that act on a disease or symptom, connection information that represents the connection relationships between molecules, and similarity information that represents the similarity between molecules; a score calculation unit that calculates a dominance score representing the strength of a relationship between a molecule and a disease or symptom for which drug discovery is targeted, based on a flow rate value assigned to each molecule when the minimum flow algorithm is applied, for each molecule included in the pathway generated by the pathway generation unit; The score calculation unit calculates the dominance score of one molecule based on flow values ​​assigned to one or more molecules connected directly below the one molecule. A pathway analysis device characterized by:

2. the pathway generation unit generates a plurality of pathways that represent, as a path diagram, intermolecular interactions of molecules related to a plurality of diseases or a plurality of symptoms; The score calculation unit calculates the dominance score of a molecule commonly included in the plurality of pathways for each of the plurality of pathways. The pathway analysis device according to claim 1 .

3. The pathway analysis device according to claim 2, wherein the score calculation unit ranks the plurality of dominance scores calculated for molecules commonly included in the plurality of pathways.

4. the pathway generation unit generates a pathway that represents, as a path diagram, intermolecular interactions of molecules related to a disease or a symptom; The score calculation unit calculates the dominance score for each of the plurality of molecules included in the one pathway. The pathway analysis device according to claim 1 .

5. The pathway analysis device according to claim 4, wherein the score calculation unit ranks the plurality of dominance scores calculated for each of the plurality of molecules included in the one pathway.

6. the pathway generation unit generates a pathway that represents molecular interactions as a path diagram, starting from only one causative molecule among molecules related to one disease or one symptom; The score calculation unit calculates the dominance score for each of the plurality of molecules included in the one pathway, excluding the one causative molecule. The pathway analysis device according to claim 1 .

7. The pathway analysis device according to claim 6, wherein the score calculation unit ranks the plurality of dominance scores calculated for each of the plurality of molecules included in the one pathway.

8. When a specific molecule is designated to be knocked out from a pathway that has already been generated, the pathway generation unit regenerates the pathway by applying the minimum flow algorithm with the molecule designated to be knocked out deleted; The score calculation unit calculates the dominance score for each of the plurality of molecules included in the pathway regenerated by the pathway generation unit. The pathway analysis device according to claim 1 .

9. The pathway analysis device according to claim 8, characterized in that the score calculation unit ranks the multiple dominance scores calculated for each of the multiple molecules included in the regenerated pathway.

10. The pathway analysis device of claim 8, characterized in that the score calculation unit calculates the dominance score for each of the multiple molecules included in the pathway once generated, and also calculates the dominance score for each of the multiple molecules included in the regenerated pathway, and calculates the difference between the dominance scores calculated for the same molecule before and after knockout.

11. the pathway generation unit is configured to arrange the causative molecule indicated by the property information so that the causative molecule is located upstream and the responsive molecule is located downstream, and to arrange other molecules between the causative molecule and the responsive molecule so as to reflect the connection relationship indicated by the connection information, and to generate the pathway by applying the minimum flow algorithm based on this arrangement; By analyzing the gene locus where the gene expression level fluctuates due to the influence of the gene mutation associated with the disease, the gene where the expression level fluctuates is located upstream as the causative molecule. The pathway analysis device according to claim 1 .

12. a related molecule estimation unit that estimates a plurality of molecules associated with a disease to be analyzed by inputting a disease feature vector identified for the disease to be analyzed into a first trained model; a molecular property estimation unit that estimates, for each of the plurality of molecules, whether the property acting on the disease is causal or responsive by inputting a disease feature vector identified for the disease to be analyzed and a molecular feature vector identified for the plurality of molecules estimated by the associated molecule estimation unit into a second trained model; The pathway generation unit arranges the molecule that is estimated to be responsive by the molecular property estimation unit on the downstream side. The pathway analysis device according to claim 11.

13. a step in which a pathway generation unit of the computer applies a minimum flow algorithm to generate a pathway that represents intermolecular interactions as a path diagram based on property information that represents the properties of molecules that act on a disease or symptom, connection information that represents the connection relationships between molecules, and similarity information that represents the similarity between molecules; and a step in which the score calculation unit of the computer calculates, for each molecule included in the pathway generated by the pathway generation unit, a dominance score that indicates the strength of the relationship between the molecule and the disease or symptom for which drug discovery is targeted, based on the flow value assigned to each molecule when the minimum flow algorithm is applied, and the dominance score of one molecule is calculated based on the flow value assigned to one or more molecules connected directly below the molecule. A pathway analysis method characterized by:

14. a pathway generation means for applying a minimum flow algorithm to generate a pathway that represents intermolecular interactions as a route diagram based on property information that represents the properties of molecules that act on a disease or symptom, connection information that represents the connection relationships between molecules, and similarity information that represents the similarity between molecules; and A score calculation means for calculating a dominance score representing the strength of the relationship between the molecule and a disease or symptom for which drug discovery is targeted, based on the flow rate value assigned to each molecule when the minimum flow algorithm is applied, for the molecules included in the pathway generated by the pathway generation means, and for calculating the dominance score of one molecule based on the flow rate values ​​assigned to one or more molecules connected directly below the molecule. Make your computer function as A pathway analysis program characterized by:

Citation Information

Patent Citations

  • Method for identifying driver gene from differential network

    CN111816246A

  • Human Metabolic Models and Methods

    JP2005521929A

  • Method and system for predicting protein-protein interaction as drug target

    JP2010165230A

  • JPP6915818B

  • JPP7034453B