Target search device and target search method
The target search device uses natural language processing to calculate influence relationship evaluation values, expanding the scope of drug target and biomarker identification beyond experimental findings, enabling a broader search for relevant molecules.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-19
- Publication Date
- 2026-03-26
AI Technical Summary
Existing methods for identifying drug targets and biomarkers are limited by the scope of experimental findings, making it difficult to link genomic information with disease mechanisms and identify relevant molecules for drug development.
A target search device and method that utilizes influence relationship evaluation values, including causality score, responsiveness score, relevance score, and cosine similarity, calculated through natural language processing of textual data to identify molecules beyond experimental scope.
Enables the discovery of drug targets and biomarkers beyond experimental limitations by leveraging textual data analysis, facilitating a broader search for molecules associated with diseases or symptoms.
Smart Images

Figure JP2024033483_26032026_PF_FP_ABST
Abstract
Description
Target search device and target search method
[0001] The present invention relates to a target discovery device and a target discovery method, and is particularly suitable for use in the discovery of molecules that may be candidates for drug targets or biomarkers.
[0002] Traditionally, drug development has involved the search for drug targets and biomarkers. Target search is the process of identifying disease-related factors (molecules such as genes and proteins) and exploring whether they can serve as suitable targets for treatment or biomarkers for diagnosis. The primary objective of drug target search is to identify factors involved in the onset and progression of disease, and factors that are effective in treating that disease. The primary objective of biomarker search is to identify appropriate evaluation indicators related to the presence or absence of disease, changes in symptoms, and the effectiveness of treatment during the drug development phase.
[0003] Traditionally, much research has focused on disease genome analysis with the aim of identifying drug targets based on the causes of diseases. Examples of disease genome analysis include GWAS (Genome-Wide Association Study) and eQTL (expression Quantitative Trait Locus) analysis. GWAS is an analytical method for identifying genetic characteristics associated with specific diseases. For example, by comparing genomic information between a group of people with a specific disease and a general population, it is possible to identify gene mutations associated with a specific disease. eQTL analysis is an analytical method that identifies chromosomal loci that are affected by changes in gene expression levels due to a specific disease, rather than identifying gene mutations. Genes whose expression has changed quantitatively are called eGenes. GWAS and eQTL analysis are analytical methods that can identify eGenes.
[0004] On the one hand, despite the fact that a vast amount of disease genomic information has been obtained through disease genomic analysis, there are still many diseases for which drug and treatment development has not been achieved because it is difficult to identify complex disease mechanisms. One of the reasons is considered to be that it is difficult to link the obtained genomic information with the onset of diseases by conventional genomic analysis methods. That is, the problem is that it is impossible to narrow down genes that can serve as drug discovery targets or biomarkers only from the genomic information obtained by disease genomic analysis.
[0005] In addition, analysis methods have been proposed for the purpose of extracting high-quality lead compounds (compounds that exhibit activity and pharmacological effects against drug discovery targets and can be a starting point for further optimization) and selecting drug discovery targets (see, for example, Patent Document 1). In the scatter diagram generation device described in this Patent Document 1, symbols representing compounds are arranged on a two-dimensional plane according to a plurality of characteristics of the compounds for a plurality of compounds to create a scatter diagram, and lead compounds are extracted from the compounds represented by the symbols arranged within a predetermined region on the scatter diagram.
[0006] However, in the technology described in Patent Document 1, the two characteristic values that determine the arrangement positions of the symbols on the two-dimensional plane are the activity value of the compound and the selectivity (the ratio of the activity of the lead compound at the target drug discovery target to the activity of the lead compound at molecular targets other than the target drug discovery target), and both are characteristic values obtained from experiments. Therefore, there was a problem that the search for drug discovery targets could only be carried out within the scope of the findings obtained from the actually conducted experiments.
[0007] Japanese Unexamined Patent Application Publication No. 2016-204376
[0008] The present invention has been made to solve the above problems, and an object thereof is to enable the search for drug discovery targets and biomarkers beyond the scope of findings grasped only by experiments.
[0009] In order to solve the above problems, in the present invention, for a plurality of molecules to be analyzed, a plurality of influence relationship evaluation values representing the characteristics of the influence relationship between a disease or symptom and the molecules are respectively obtained, and molecules whose combinations of the plurality of influence relationship evaluation values satisfy a predetermined condition are extracted from the plurality of molecules to be analyzed. Here, as at least one of the plurality of influence relationship evaluation values, an evaluation value specified using a disease-related feature vector calculated for a disease or symptom and a molecule feature vector calculated for a molecule by natural language processing for a plurality of sentences is used.
[0010] According to the present invention configured as described above, when searching for drug discovery targets and biomarkers using a plurality of influence relationship evaluation values representing the characteristics of the influence relationship between a disease or symptom and a molecule, an evaluation value specified by natural language processing for a plurality of sentences is also used. Therefore, it is possible to search for drug discovery targets and biomarkers beyond the range of knowledge grasped only by experiments.
[0011] It is a diagram showing a functional configuration example of the target search device according to the present embodiment. It is a diagram showing a hardware configuration example of the target search device according to the present embodiment. It is a block diagram showing a functional configuration example of the feature vector calculation unit according to the present embodiment. It is a diagram showing an example of a feature vector. It is a block diagram showing a functional configuration example of the score calculation unit according to the present embodiment. It is a diagram for explaining an example of the processing by the molecule extraction unit of the present embodiment. It is a diagram for explaining an example of the processing by the molecule extraction unit of the present embodiment. It is a diagram for explaining an example of the processing by the molecule extraction unit of the present embodiment. It is a diagram for explaining an example of the processing by the molecule extraction unit of the present embodiment. It is a diagram for explaining an example of the processing by the molecule extraction unit of the present embodiment. It is a diagram for explaining an example of the processing by the molecule extraction unit of the present embodiment. It is a diagram for explaining an example of the processing by the molecule extraction unit of the present embodiment. It is a diagram for explaining an example of the processing by the molecule extraction unit of the present embodiment. It is a diagram for explaining an example of the processing by the molecule extraction unit of the present embodiment. It is a diagram for explaining an example of the processing by the molecule extraction unit of the present embodiment.
[0012] Hereinafter, one embodiment of the present invention will be described with reference to the drawings. Figure 1 is a diagram showing an example of the functional configuration of the target search device 1 according to this embodiment. Figure 2 is a diagram showing an example of the hardware configuration of the target search device 1 according to this embodiment.
[0013] The target search device 1 of this embodiment is configured using a general-purpose computer such as a workstation or personal computer. As shown in Figure 2, the target search device 1 has a hardware configuration comprising a control unit 101, a storage unit 102, an input unit 103, and a display unit 104.
[0014] The control unit 101 is composed of, for example, a microcomputer processor equipped with a CPU, RAM, and ROM, and operates according to the operating system and application programs stored in the storage unit 102, and performs information processing using various data stored in the storage unit 102. In addition to the microcomputer, it may also be equipped with a DSP (Digital Signal Processor) or the like.
[0015] The storage unit 102 is composed of a computer-readable storage medium. For example, the storage unit 102 may include ROM, RAM, a hard disk, or semiconductor memory. Various programs executed by the control unit 101 are stored in the storage unit 102.
[0016] Furthermore, the memory unit 102 also stores various types of data, including data used in the processing of the control unit 101 and data generated during the processing. This data includes multiple textual data, such as descriptions of diseases and their symptoms, and descriptions of molecules like genes and proteins. Alternatively, the textual data may be stored in external storage such as a data server, and the control unit 101 may retrieve and use the necessary data from the external storage when performing analysis.
[0017] Furthermore, the various data stored in the memory unit 102 also include the trained model, which will be described later. Alternatively, the trained model may be stored on an external server device, and the control unit 101 may access the server device and utilize the trained model when performing analysis.
[0018] The input unit 103 is composed of, for example, a keyboard, mouse, or touch panel, and supplies data to the control unit 101 in accordance with the user's operation input. The display unit 104 is composed of, for example, a liquid crystal display, and displays information according to instructions input from the control unit 101.
[0019] The target search device 1 may be a server device logically implemented by, for example, cloud computing. The server device may consist of a single device or a combination of multiple devices. In this case, the target search device 1 further includes a communication unit in the hardware configuration shown in Figure 2. On the other hand, the target search device 1 does not need to include the input unit 103 and display unit 104 shown in Figure 2. The communication unit consists of a communication module that exchanges data with a client device via a communication network. The communication network may be, for example, the internet and may include public communication lines or mobile phone lines. The communication network may also include wireless communication channels or LANs (Local Area Networks).
[0020] As shown in Figure 1, the target search device 1 of this embodiment includes an evaluation value acquisition unit 11 and a molecular extraction unit 12 as its functional configuration. The evaluation value acquisition unit 11 specifically includes a feature vector calculation unit 111 and a score calculation unit 112 as its functional configuration. Furthermore, the target search device 1 of this embodiment includes a text data storage unit 13, a first model storage unit 14, and a second model storage unit 15 as its storage medium.
[0021] Functional blocks 11 to 12 perform the processes described below through the cooperation of hardware and software. For example, the processes of functional blocks 11 to 12 are executed by the operation of a program stored in the storage unit 102 under the control of the control unit 101 shown in Figure 2. The evaluation value acquisition unit 11 and molecular extraction unit 12 shown in Figure 1 also represent the processing procedure of the target search method. In addition, the text data storage unit 13, the first model storage unit 14, and the second model storage unit 15 are composed of the storage unit 102 shown in Figure 2.
[0022] The evaluation value acquisition unit 11 acquires multiple influence relationship evaluation values for each of the multiple molecules to be analyzed, each representing the characteristics of the influence relationship between the disease or symptom and the molecule. There are, for example, two of the multiple influence relationship evaluation values. The two influence relationship evaluation values are, for example, any two of the following: causality score, responsiveness score, relevance score, cosine similarity, and eGene expression level.
[0023] In this embodiment, at least one of the multiple influence relationship evaluation values is an influence relationship evaluation value identified using feature vectors calculated for diseases or symptoms by natural language processing of multiple sentences stored in the text data storage unit 13 (hereinafter referred to as disease-related feature vectors) and feature vectors calculated for molecules (hereinafter referred to as molecular feature vectors). With respect to disease-related feature vectors, the feature vector calculated for diseases is called a disease feature vector, and the feature vector calculated for symptoms is called a symptom feature vector.
[0024] The causality score, responsiveness score, relevance score, and cosine similarity are influence assessment values identified using disease-related feature vectors and molecular feature vectors, and are not evaluation values obtained from experiments. On the other hand, eGen expression levels are influence assessment values identified using experimental values through known GWAS or eQTL analyses.
[0025] A disease-related feature vector is data that represents the characteristics of a disease or symptom (features that can identify a disease or symptom) as a combination of values of multiple elements. In this embodiment, as an example, a vector representing the extent to which disease names or symptom names, which are included as words in multiple sentences, contribute to each sentence is used as the disease-related feature vector.
[0026] Furthermore, molecular feature vectors are data that represent the characteristics (features that allow a molecule to be identified) of molecules such as proteins and genes as a combination of values of multiple elements. In this embodiment, as an example, a vector representing how much each molecular name, which is included as a word in multiple sentences, contributes to each sentence is used as the molecular feature vector.
[0027] In this embodiment, the text in question may consist of a single sentence (a unit separated by a period) or multiple sentences. A text consisting of multiple sentences may be part or all of the text contained in a single document. The text is not limited to descriptions of diseases, but may also include descriptions of various other themes.
[0028] Disease names, as individual words, tend to be used in texts describing diseases, but not in texts unrelated to diseases. Furthermore, even within texts describing diseases, sentences containing a particular disease name as a word are likely to be texts describing that specific disease, and are unlikely to contain that disease name in texts describing other types of diseases. In other words, the inclusion of disease names as words in texts tends to vary depending on the type of disease the text is about. Therefore, a vector representing the extent to which a disease name contributes to a given text can be used as a feature vector that can identify diseases. The same can be said for symptom names and molecular names.
[0029] The following describes an example of a method for calculating disease-related feature vectors and molecular feature vectors. Figure 3 is a block diagram showing an example of the configuration of the feature vector calculation unit 111, which is specifically included in the evaluation value acquisition unit 11. The feature vector calculation unit 111 shown in Figure 3 receives text data from the text data storage unit 13, calculates and outputs feature vectors that reflect the relationship between the text and the words contained therein.
[0030] In this example, the evaluation value acquisition unit 11 of the target search device 1 is shown to include a feature vector calculation unit 111, but the device is not limited to this configuration. For example, a feature vector calculation unit may be provided separately from the target search device 1, and the feature vector calculation unit may calculate feature vectors from text data. In this case, the evaluation value acquisition unit 11 receives disease-related feature vectors and molecular feature vectors from the feature vector calculation unit and calculates influence relationship evaluation values using these feature vectors.
[0031] As shown in Figure 3, the feature vector calculation unit 111 is configured as follows: a word extraction unit 31, a vector calculation unit 32, an index value calculation unit 33, and a feature vector identification unit 34. The vector calculation unit 32 is configured as follows: a sentence vector calculation unit 32A and a word vector calculation unit 32B.
[0032] The word extraction unit 31 analyzes m sentences (where m is any integer greater than or equal to 2) contained in the sentence data input from the sentence data storage unit 13, and extracts n words (where n is any integer greater than or equal to 2) from these m sentences. For sentence analysis, for example, known morphological analysis can be used. Here, the word extraction unit 31 may extract all morphemes of parts of speech that are divided by morphological analysis as words, or it may extract only morphemes of specific parts of speech as words.
[0033] Note that there may be multiple occurrences of the same word in the m sentences. In this case, the word extraction unit 31 does not extract multiple occurrences of the same word, but only extracts one occurrence. That is, the n words extracted by the word extraction unit 31 mean n types of words. Among the words extracted by the word extraction unit 31, there are disease names, symptom names, molecular names such as genes and proteins, and various other words.
[0034] The vector calculation unit 32 calculates m sentence vectors and n word vectors from the m sentences and the n words. Here, the sentence vector calculation unit 32A calculates m sentence vectors composed of q axis components (q is an arbitrary integer of 2 or more) by vectorizing each of the m sentences to be analyzed by the word extraction unit 31 into q dimensions according to a predetermined rule. Further, the word vector calculation unit 32B calculates n word vectors composed of q axis components by vectorizing each of the n words extracted by the word extraction unit 31 into q dimensions according to a predetermined rule.
[0035] In this embodiment, as an example, the sentence vector and the word vector are calculated as follows. Now, consider a set S = <d ∈ D, w ∈ W> consisting of m sentences and n words. Here, each sentence d i (i = 1, 2,..., m) and each word w j (j = 1, 2,..., n) are respectively associated with a sentence vector d i → and a word vector w j → (hereinafter, the symbol "→" shall indicate that it is a vector). Then, for any word w j and any sentence d i , the probability P(w j | d i ) shown in the following formula (1) is calculated.
[0036]
[0037] Note that this probability P(w j | d iThis value can be calculated by following the probability p disclosed in the publicly available document "Distributed Representations of Sentences and Documents" by Quoc Le and Tomas Mikolov, Google Inc, Proceedings of the 31st International Conference on Machine Learning Held in Beijing, China on 22-24 June 2014. This publicly available document, for example, states that when there are three words, "the," "cat," and "sat," the fourth word to be predicted to be "on," and the formula for calculating that prediction probability p is provided.
[0038] The probability p(wt | wt-k, ..., wt+k) described in publicly available literature is the probability of correctly predicting one word wt from multiple words wt-k, ..., wt+k. In contrast, the probability P(wt|wt-k|wt-k|wt+k) shown in equation (1) used in this embodiment is j |d i ) is one sentence d out of m sentences i So, one of the n words lol j This represents the expected probability of getting the correct answer. One sentence d i One word from lol j Predicting means, specifically, a certain sentence d i When it appears, there is a word lol j This means predicting the possibility that it will be included.
[0039] Note that this equation (1) is d i lol j Since it's symmetrical, one out of n words is lol j From m sentences, one sentence d i The expected probability P(d) i |w j You may calculate () one word lol j From one sentence d i To predict is a certain word lol j When it appears, it is sentence d iThis means predicting the possibilities that are included within that.
[0040] Equation (1) uses an exponential function with base e and exponent being the dot product of the word vector w→ and the sentence vector d→. The sentence d to be predicted is... i and word lol j The exponential function value calculated from the combination of and text d i and n words lol k The ratio of the sum of n exponential function values calculated from each combination of (k = 1, 2, ..., n) to one sentence d i One word from lol j This is calculated as the expected probability of getting the correct answer.
[0041] Here, word vector lol j → and text vector d i →The dot product of this and the word vector w j → text vector d i The scalar value when projected in the direction of →, i.e., the word vector lol j →The text vector d that possesses i It could also be called the component value in the direction of →. This is a word lol j is text d i It can be thought of as representing the degree to which it contributes. Therefore, using the exponential function value calculated using such an inner product, n words w k A single word for the sum of exponential function values calculated for (k = 1, 2, ..., n) j Finding the ratio of the exponential function values calculated for is one sentence d i One of n words from lol j This is equivalent to calculating the expected probability of getting the correct answer.
[0042] Here, we have shown an example of calculation using an exponential function whose exponent is the dot product of the word vector w→ and the sentence vector d→, but it is not mandatory to use an exponential function. Any calculation formula that utilizes the dot product of the word vector w→ and the sentence vector d→ will suffice; for example, the probability could be calculated using the ratio of the dot product values themselves.
[0043] Next, the vector calculation unit 32 calculates the probability P(w) calculated by equation (1) as shown in equation (2) below. j |d i The text vector d that maximizes the value L obtained by summing the texts over the set S. i → and word vector w j → calculates the probability P(w) calculated by the above formula (1). That is, the sentence vector calculation unit 32A and the word vector calculation unit 32B calculate the probability P(w) calculated by the above formula (1). j |d i The value obtained by calculating the sum of the sum of the sums for all combinations of m sentences and n words is used as the target variable L, and the sentence vector d that maximizes the target variable L is determined. i → and word vector w j → Calculate.
[0044]
[0045] The probability P(w) calculated for all combinations of m sentences and n words. j |d i Maximizing the sum L of ) is equivalent to a sentence d i (i = 1, 2, ..., m) A word that is a part of this lol j This means maximizing the expected probability of getting the correct answer (j = 1, 2, ..., n). In other words, the vector calculation unit 32 calculates the text vector d that maximizes this probability of getting the correct answer. i → and word vector w j → can be said to be something that calculates →.
[0046] In this embodiment, as described above, the vector calculation unit 32 calculates m sentences d i By vectorizing each of these into q dimensions, we obtain m sentence vectors d consisting of q axis components. i By calculating the → and vectorizing each of the n words into q dimensions, we obtain n word vectors w consisting of q axis components. j →Calculate the text vector d that maximizes the above-mentioned target variable L, with q axis directions being variable. i → and word vector w j This is equivalent to calculating →.
[0047] The index value calculation unit 33 calculates m sentence vectors d calculated by the vector calculation unit 32. i → and n word vectors w j →By calculating the inner product of each, m sentences d i and n words lol j An index value that reflects the relationship between them is calculated. In this embodiment, the index value calculation unit 33 calculates m text vectors d as shown in the following equation (3). i → Each q axis component (d 11 ~d mq A sentence matrix D whose elements are ) and n word vectors w j → Each q axis component (w 11 ~w nq By calculating the product with a word matrix W whose elements are ), an index value matrix DW whose elements are m × n index values is calculated. Here, W t This is the transpose of the word matrix.
[0048]
[0049] Each element of the index matrix DW calculated in this way can be said to represent how much each word contributes to each sentence, and how much each sentence contributes to each word. For example, the element dw in a 1x2 grid. 12 This value represents the extent to which word w2 contributes to sentence d1, and similarly, it represents the extent to which sentence d1 contributes to word w2. Therefore, each row of the index matrix DW can be used to evaluate sentence similarity, and each column can be used to evaluate word similarity.
[0050] The feature vector identification unit 34 identifies a group of word index values consisting of m index values for each of the multiple disease names from among the n words as a disease feature vector. That is, as shown in Figure 4, the feature vector identification unit 34 identifies the group of word index values corresponding to the word corresponding to the disease name from among the n sets of word index value groups (m index values per column) that make up each column of the index value matrix DW as the disease feature vector for each disease name. When searching for drug targets or biomarkers for a specific disease, it is sufficient to identify the group of word index values corresponding to the word of that disease name as the disease feature vector.
[0051] Furthermore, the feature vector identification unit 34 identifies a group of word index values consisting of m index values for each of the multiple symptom names among the n words as a symptom feature vector. Also, the feature vector identification unit 34 identifies a group of word index values consisting of m index values for each of the multiple molecule names among the n words as a molecular feature vector.
[0052] Returning to Figure 1, the explanation continues. The evaluation value acquisition unit 11 uses the disease-related feature vector and molecular feature vector calculated as described above to calculate at least one of the causality score, responsiveness score, relevance score, and cosine similarity as an influence relationship evaluation value. The evaluation value acquisition unit 11 may also acquire the eGene expression level obtained by disease genome analysis such as GWAS or eQTL analysis as one of the influence relationship evaluation values. Disease genome analysis may be performed by the evaluation value acquisition unit 11, or it may be performed by an analysis device other than the target search device 1. In the latter case, the evaluation value acquisition unit 11 acquires information on eGene identified by the other analysis device from that other analysis device.
[0053] The causality score is a value that indicates the probability that a molecule's properties in relation to a disease or symptom are causal. Causality is the property that suggests the presence or mutation of the molecule may cause a disease or symptom. The responsiveness score is a value that indicates the probability that a molecule's properties in relation to a disease or symptom are responsive. Responsiveness is the property that suggests the molecule may mutate as a result of the onset of a disease or symptom.
[0054] The relevance score is a value that indicates the degree to which a molecule is associated with a disease or symptom. Relevance refers to the likelihood that a molecule is associated with a disease or symptom, regardless of whether its nature is causal or responsive. Cosine similarity is a value that indicates the similarity between the disease-associated feature vector and the molecular feature vector. eGene expression level is a value that indicates the increase or decrease in the expression of eGene, a gene whose expression level has changed quantitatively due to the effects of the disease.
[0055] The following describes an example of how to calculate the causality score, responsiveness score, relevance score, and cosine similarity. Figure 5 is a block diagram showing an example of the configuration of the score calculation unit 112, which is specifically included in the evaluation value acquisition unit 11. The score calculation unit 112 shown in Figure 5 calculates the causality score, responsiveness score, relevance score, and cosine similarity using the disease-related feature vector and molecular feature vector calculated by the feature vector calculation unit 111.
[0056] As shown in Figure 5, the score calculation unit 112 includes, as a functional configuration, a related molecule estimation unit 51, a molecular property estimation unit 52, and a similarity calculation unit 53. Below, we will describe an example of calculating four influence relationship evaluation values using disease feature vectors and molecular feature vectors, but the same procedure applies when calculating four influence relationship evaluation values using symptom feature vectors and molecular feature vectors.
[0057] The related molecule estimation unit 51 estimates multiple molecules associated with a disease by inputting disease feature vectors calculated by the feature vector calculation unit 111 for a specific disease to be analyzed into a first pre-trained model pre-stored in the first model storage unit 14. Here, the first pre-trained model is machine-trained to estimate molecules corresponding to molecular feature vectors similar to disease feature vectors when a disease feature vector is input, based on the similarity between the disease feature vector and the molecular feature vector, and outputs information about the estimated molecules (e.g., molecular name) and a relevance score. The relevance score is a value that indicates the reliability of the estimation by the first pre-trained model (a value corresponding to the likelihood that the molecules output from the first pre-trained model are estimated to be associated with the disease corresponding to the disease feature vector input to the first pre-trained model) for molecules that are estimated to be associated with the disease corresponding to the disease feature vector input to the first pre-trained model and output from the first pre-trained model.
[0058] The related molecule estimation unit 51 supplies molecular information output from the first trained model to the molecular property estimation unit 52, and also supplies the relatedness score to the molecular extraction unit 12 in Figure 1.
[0059] The form of the first trained model stored in the first model storage unit 14 can be any of the following: a regression model, a tree model, a neural network model, a Bayesian model, a clustering model, etc. Note that the models listed here are merely examples and are not limited to them. For example, it could be a function model that calculates the similarity between a disease feature vector and a molecular feature vector, and outputs information about molecules corresponding to molecular feature vectors whose similarity to the disease feature vector is greater than or equal to a predetermined value.
[0060] Machine learning of the first pre-trained model can be performed, for example, as follows: The feature vector calculation unit 111 or a feature vector calculation device with similar functionality is used to calculate disease feature vectors for multiple disease names and molecular feature vectors for multiple molecule names. Multiple sets of training data are generated, with each set of one disease feature vector and one molecular feature vector forming a single dataset. Machine learning of the first pre-trained model is then performed using this training data, and the first pre-trained model, trained based on the similarity between the disease feature vectors and molecular feature vectors, is stored in the first model storage unit 14.
[0061] Here, the similarity between the disease feature vector and the molecular feature vector can be evaluated in various ways. For example, it is possible to extract features using a predetermined function for both the disease feature vector and the molecular feature vector, and then evaluate the similarity of the features. Alternatively, the Euclidean distance or cosine similarity between the word index values of the disease feature vector and the word index values of the molecular feature vector may be used, or the edit distance may be used.
[0062] The similarity between disease feature vectors and molecular feature vectors means that the way a word as a disease name contributes to a given text is similar to the way a word as a molecular name contributes to a given text. Since texts are written along specific themes, disease names and molecular names with similar disease feature vectors and molecular feature vectors contribute similarly to multiple texts written in relation to each theme, making it possible to infer that there is some kind of relationship between the disease and the molecule.
[0063] When a disease name and a molecule name are described within a single sentence, it is clear that the disease and the molecule are related. On the other hand, when a disease name and a molecule name are described across multiple sentences, it is unclear whether there is a relationship between the disease described in one sentence and the molecule described in another sentence. Even if a medical professional reads those sentences, it is difficult for them to immediately understand that there is a relationship.
[0064] In contrast, according to this embodiment, even when disease names and molecule names are described across multiple sentences, it is possible to estimate that there may be some kind of relationship between the disease and the molecule. As a result, if a disease feature vector corresponding to a certain disease name is input into the first trained model, even molecules whose relationship to the disease was unknown may be output as related molecules through learning-based estimation.
[0065] The molecular property estimation unit 52 inputs the disease feature vector calculated by the feature vector calculation unit 111 for the specific disease to be analyzed, and the multiple molecular feature vectors identified for multiple molecules estimated by the related molecule estimation unit 51, into the second trained model stored in the second model storage unit 15. This allows the unit to estimate the probability that each of the multiple molecules estimated to be related to the disease is causal or responsive as a property of the molecule acting on the disease.
[0066] Here, the molecular property estimation unit 52 may, for example, supply the information list of molecules output from the related molecule estimation unit 51 to the feature vector calculation unit 111, causing the feature vector calculation unit 111 to calculate molecular feature vectors corresponding to the molecule names in the information list. Alternatively, a database may be prepared in advance in which molecule names and their corresponding molecular feature vectors are associated and stored, and when the information list of molecules is output from the related molecule estimation unit 51, the molecular feature vectors corresponding to the molecule names in the information list may be read from the database.
[0067] As another example, the first pre-trained model used in the related molecule estimation unit 51 may be one that has been machine-trained to output a molecular feature vector similar to a disease feature vector when a disease feature vector is input, based on the similarity between the disease feature vector and the molecular feature vector. In this case, the molecular property estimation unit 52 can directly input the disease feature vector output from the feature vector calculation unit 111 and the molecular feature vector output from the related molecule estimation unit 51 into the second pre-trained model.
[0068] The second trained model stored in the second model memory unit 15 can take the form of any of the following: a regression model, a tree model, a neural network model, a Bayesian model, a clustering model, etc. Note that the models listed here are merely examples and are not limited to these.
[0069] Here, the second pre-trained model is machine-trained to output the probability that a molecule's properties are causal or responsive, given disease feature vectors and molecular feature vectors as input. This machine learning is performed using multiple sets of training data, with one set of training data consisting of one disease feature vector, one molecular feature vector, and a dataset of property information representing the properties of molecules acting on the disease. As an example, the second pre-trained model includes a causality estimation model that outputs the probability that a molecule's properties with respect to the disease are causal, and a responsiveness estimation model that outputs the probability that a molecule's properties with respect to the disease are responsive.
[0070] For known diseases, there is known information about which molecules are causative and which are responsive. The second pre-trained model is created using machine learning with a dataset of disease feature vectors, molecular feature vectors, and molecular property information generated from this known information as training data (the molecular property information is the ground truth data). Therefore, for molecules known to be causative for known diseases, the causality estimation model outputs a high probability value, while the responsiveness estimation model outputs a low probability value. On the other hand, for molecules known to be responsive for known diseases, the responsiveness estimation model outputs a high probability value, while the causality estimation model outputs a low probability value.
[0071] Furthermore, among the multiple molecules whose association with the disease is estimated by the related molecule estimation unit 51, there is a possibility that some molecules whose association with the disease was previously unknown to human knowledge. For such molecules as well, the causality estimation model outputs a probability value indicating that the molecule may exhibit causal properties with respect to the disease, and the responsiveness estimation model outputs a probability value indicating that the molecule may exhibit responsive properties.
[0072] In other words, when there is a high degree of similarity between the combination of a disease feature vector corresponding to a certain disease and a molecular feature vector corresponding to a molecule whose association with that disease was unknown (the resulting feature), and the combination of a disease feature vector corresponding to that disease and a molecular feature vector corresponding to a molecule known to be causative (the resulting feature), the causality estimation model tends to output a relatively high probability value. This probability value output by the causality estimation model is the causality score.
[0073] On the other hand, if there is a high degree of similarity between the combination of a disease feature vector corresponding to a certain disease and a molecular feature vector corresponding to a molecule whose association with that disease was unknown (the resulting feature), and the combination of a disease feature vector corresponding to that disease and a molecular feature vector corresponding to a molecule known to be responsive (the resulting feature), the responsiveness model tends to output a relatively high probability value. The probability value output from this responsiveness estimation model is the responsiveness score.
[0074] The molecular property estimation unit 52 supplies the causality score and responsiveness score output from the second trained model to the molecular extraction unit 12 shown in Figure 1.
[0075] Here, we have described an example in which the second pre-trained model includes both a causality estimation model and a responsiveness estimation model, but it is not limited to this. For example, the second pre-trained model may consist only of a causality estimation model. For example, if the probability values output from the causality estimation model take values from "0" to "1", the probability values output from the causality estimation model may be used as the causality score, while the value obtained by subtracting the probability values output from the causality estimation model from "1" may be used as the responsiveness score. Alternatively, the second pre-trained model may consist only of a responsiveness estimation model, and the probability values output from the responsiveness estimation model may be used as the responsiveness score, while the value obtained by subtracting the probability values output from the responsiveness estimation model from "1" may be used as the causality score.
[0076] As described above, the molecular property estimation unit 52 estimates the properties of the molecule's effect on disease by inputting the disease feature vector and molecular feature vector into the second trained model. Here, the molecular feature vector used corresponds to the molecular feature vectors of multiple molecules estimated by the related molecule estimation unit 51, that is, the molecular information list output from the first trained model. In this case, the multiple molecules estimated by the related molecule estimation unit 51 become the target of analysis when searching for molecules that can serve as drug targets or biomarkers for specific diseases or symptoms.
[0077] In contrast, the molecular property estimation unit 52 can also use molecular feature vectors corresponding to multiple eGenes identified by disease genome analysis. In this case, the multiple eGenes identified by disease genome analysis become the targets of analysis when exploring whether they can serve as drug targets or biomarkers for specific diseases or symptoms.
[0078] When using molecular feature vectors corresponding to eGenes identified by disease genome analysis, for example, molecular feature vectors corresponding to multiple eGenes are calculated by the feature vector calculation unit 111. The molecular property estimation unit 52 inputs the disease feature vector calculated by the feature vector calculation unit 111 for a specific disease and the molecular feature vectors calculated by the feature vector calculation unit 111 for multiple eGenes identified by disease genome analysis into a second trained model to estimate the probability that the properties of the eGenes are causal or responsive.
[0079] The similarity calculation unit 53 calculates the similarity between the disease feature vector calculated by the feature vector calculation unit 111 for a specific disease and the molecular feature vector calculated for multiple molecules to be analyzed. The similarity is, for example, the cosine similarity between the word index value group of the disease feature vector and the word index value group of the molecular feature vector. The similarity calculation unit 53 supplies the calculated cosine similarity to the molecular extraction unit 12 in Figure 1. Note that Euclidean distance or edit distance may be used instead of cosine similarity.
[0080] When multiple molecules whose association with a disease is estimated by the association molecule estimation unit 51 are used as the multiple molecules to be analyzed, the similarity calculation unit 53 calculates the similarity between the disease feature vector calculated by the feature vector calculation unit 111 for a specific disease and the molecular feature vectors corresponding to the multiple molecules whose association with the disease is estimated by the association molecule estimation unit 51. On the other hand, when multiple eGenes identified by disease genome analysis are used as the multiple molecules to be analyzed, the similarity calculation unit 53 calculates the similarity between the disease feature vector calculated by the feature vector calculation unit 111 for a specific disease and the molecular feature vectors calculated by the feature vector calculation unit 111 for the multiple eGenes.
[0081] Returning to Figure 1, the explanation continues. The molecular extraction unit 12 extracts molecules from among the multiple molecules to be analyzed in which the combination of multiple influence relationship evaluation values obtained by the evaluation value acquisition unit 11 satisfies predetermined conditions. As described above, the influence relationship evaluation values used in this embodiment are any two of the following: causality score, responsiveness score, relevance score, cosine similarity, and eGene expression level.
[0082] Here, the causality score is a suitable influence assessment value to use when searching for drug targets. The responsiveness score is a suitable influence assessment value to use when searching for biomarkers. The relevance score and cosine similarity are suitable influence assessment values to use when searching for drug targets or biomarkers. eGen expression level is a suitable influence assessment value to use when narrowing down drug targets or biomarkers from among multiple eGenes identified by disease genome analysis such as eQTL analysis or GWAS.
[0083] Figures 6 to 15 illustrate an example of processing by the molecular extraction unit 12. Of these, Figures 6 to 8 show an example of processing in which molecules that satisfy predetermined conditions are extracted using the causality score and eGene expression level. Figure 6 shows an example of processing by the molecular extraction unit 12 when "bladder cancer" is specified as the specific disease and multiple eGenes identified by eQTL analysis are used as molecules to be analyzed.
[0084] Figure 6 shows a two-dimensional plane with the causality score on the vertical axis and the eGene expression level on the horizontal axis. For each of the eGenes to be analyzed, circular symbols are plotted at coordinate positions identified by the combination of causality scores and eGene expression levels obtained by the evaluation value acquisition unit 11. The molecular extraction unit 12 extracts eGenes corresponding to symbols located in a specific region AR1 that satisfies predetermined conditions on this two-dimensional plane as candidate genes for drug targets.
[0085] A specific region AR1 that meets predetermined conditions is a region where the eGene expression level (increase or decrease) is above a predetermined value and the causal score is above a predetermined value. The plotted position of the symbol corresponding to the molecule "FGFR3," which is known as a drug target for approved drugs for bladder cancer, is located within the specific region AR1. This confirms that the analytical method according to this embodiment is effective in the search for drug targets for bladder cancer.
[0086] Figure 7 shows an example of processing in the molecular extraction unit 12 when "schizophrenia" is specified as a specific disease and multiple eGenes identified by eQTL analysis are used as the molecules to be analyzed. Figure 8 shows an example of processing in the molecular extraction unit 12 when "nicotine addiction" is specified as a specific disease and multiple eGenes identified by eQTL analysis are used as the molecules to be analyzed.
[0087] In both Figure 7 and Figure 8, similar to Figure 6, symbols are plotted on a two-dimensional plane with the causality score on the vertical axis and the eGene expression level on the horizontal axis. These symbols are located at coordinate positions identified by the combination of causality scores and eGene expression levels obtained by the evaluation value acquisition unit 11 for multiple eGenes to be analyzed. The molecular extraction unit 12 extracts eGenes corresponding to symbols located in specific regions AR2 and AR3 that satisfy predetermined conditions on this two-dimensional plane as candidate genes for drug targets.
[0088] The specific regions AR2 and AR3 that meet the predetermined conditions are regions where the eGene expression level (increase or decrease) is above a predetermined value and the causality score is above a predetermined value. The plotted position of the symbol corresponding to the molecule "DRD2," which is known as a drug target for approved drugs for schizophrenia, is located within specific region AR2. The plotted position of the symbol corresponding to the molecule "CHRNA4," which is known as a drug target for approved drugs for nicotine dependence, is located within specific region AR3. This confirms that the analytical method according to this embodiment is also effective in the search for drug targets for schizophrenia and nicotine dependence.
[0089] Figures 6 to 8 all show examples where two indicators, the causality score and the eGene expression level, are used as evaluation values for the relationship between the effects. Specific regions AR1 to AR3 are all regions where the causality score is 0.7 or higher, and the eGene expression level is 0.4 or higher, or -0.4 or lower.
[0090] Figures 6 to 8 show an example where two indicators, the causality score and the eGene expression level, are used as evaluation values for the relationship between the causality and the relationship between the two factors. However, the responsiveness score may be used instead of the causality score.
[0091] Figure 9 shows an example of processing by the molecular extraction unit 12 when "bladder cancer" is specified as the specific disease and multiple eGenes identified by eQTL analysis are used as the molecules to be analyzed. Figure 10 shows an example of processing by the molecular extraction unit 12 when "schizophrenia" is specified as the specific disease and multiple eGenes identified by eQTL analysis are used as the molecules to be analyzed.
[0092] Figures 9 and 10 show examples of a process for extracting molecules that satisfy predetermined conditions using cosine similarity and eGene expression levels. Specifically, Figures 9 and 10 show symbols plotted on a two-dimensional plane with cosine similarity on the vertical axis and eGene expression level on the horizontal axis, with each symbol representing a coordinate position determined by the combination of cosine similarity and eGene expression levels obtained by the evaluation value acquisition unit 11 for multiple eGene molecules to be analyzed.
[0093] The molecular extraction unit 12 extracts eGenes corresponding to symbols located in specific regions AR4 and AR5 that satisfy predetermined conditions on this two-dimensional plane, as candidate genes for drug targets. The specific regions AR4 and AR5 that satisfy the predetermined conditions are regions in which the eGen expression level (increase or decrease) is above a predetermined value and the cosine similarity is above a predetermined value.
[0094] The plotted position of the symbol corresponding to the molecule "FGFR3," known as a drug target for approved drugs for bladder cancer, falls within the specific region AR4. Similarly, the plotted position of the symbol corresponding to the molecule "DRD2," known as a drug target for approved drugs for schizophrenia, falls within the specific region AR5. This confirms that the analysis method according to this embodiment is effective even when cosine similarity is used instead of causality scores in the search for drug targets for bladder cancer and schizophrenia.
[0095] Figures 9 and 10 both show examples where cosine similarity and eGene expression level are used as evaluation values for the relationship between influences. In both specific regions AR4 to AR5, the cosine similarity is 0.15 or higher, and the eGene expression level is 0.4 or higher or -0.4 or lower.
[0096] Figure 11 shows an example of processing by the molecular extraction unit 12 when "schizophrenia" is specified as a specific disease and multiple molecules whose association with the disease is estimated by the related molecule estimation unit 51 are used as analysis targets. Figure 12 also shows an example of processing by the molecular extraction unit 12 when "hypertension" is specified as a specific disease and multiple molecules whose association with the disease is estimated by the related molecule estimation unit 51 are used as analysis targets.
[0097] Figures 11 and 12 show the state in which symbols are plotted on a two-dimensional plane with cosine similarity on the vertical axis and causality score on the horizontal axis, at coordinate positions identified by the combination of cosine similarity and causality score values obtained by the evaluation value acquisition unit 11 for multiple molecules to be analyzed.
[0098] The molecular extraction unit 12 extracts molecules corresponding to symbols located in specific regions AR6 and AR7 that satisfy predetermined conditions on this two-dimensional plane, as candidate genes for drug targets. The specific regions AR6 and AR7 that satisfy the predetermined conditions are regions in which the cosine similarity is above a predetermined value (0.15 or higher) and the causality score is above a predetermined value (0.7 or higher).
[0099] The plotted symbols corresponding to molecules known as drug targets for approved drugs for schizophrenia, such as "DRD1," "DRD2," "DRC3," "HTR1A," "HTR2A," and "HTR2C," are located within the specific region AR6. Similarly, the plotted symbols corresponding to molecules known as drug targets for approved drugs for hypertension, such as "ACE," "REN," "AGTR1," "ADRB1," "CACNA1D," and "NA3C2," are located within the specific region AR7. This confirms that the analysis method according to this embodiment, using causality scores and cosine similarity, is effective in the search for drug targets for schizophrenia and hypertension.
[0100] Figures 11 and 12 show examples where two metrics, the causality score and cosine similarity, are used as evaluation values for the relationship between influences. However, the responsiveness score may be used instead of the causality score, and the relevance score may be used instead of the cosine similarity score.
[0101] Figure 13 shows an example of processing by the molecular extraction unit 12 when "pancreatic cancer" is specified as a specific disease, "proliferation" is specified as a specific symptom, and multiple molecules that are estimated by the related molecule estimation unit 51 to have both a relationship with the disease and a relationship with the symptom are used as the multiple molecules to be analyzed.
[0102] Figure 13 shows a two-dimensional plane with the causality score calculated for the disease "pancreatic cancer" on the vertical axis and the causality score calculated for the symptom "proliferation" on the horizontal axis. Symbols are plotted at coordinate positions identified by the combination of causality scores for the disease and the causality scores for the symptoms obtained by the evaluation value acquisition unit 11 for multiple molecules to be analyzed.
[0103] The molecular extraction unit 12 extracts molecules corresponding to symbols placed in specific regions AR8 that satisfy predetermined conditions on this two-dimensional plane, as candidate genes for drug targets. Specific regions AR8 that satisfy predetermined conditions are regions where the causality score for disease is above a predetermined value (0.7 or higher) and the causality score for symptoms is above a predetermined value (0.7 or higher).
[0104] The plotted positions of symbols corresponding to molecules known as drug targets for pancreatic cancer, such as "TUBB1," "KRAS," and "RRM1," are located within the specific region AR8. This confirms that the analytical method according to this embodiment is effective in the search for drug targets for pancreatic cancer.
[0105] In Figure 13, an example is shown where a causality score (related to symptoms) is used as one of the influence relationship evaluation values, and the same type of causality score (related to disease) is used as the other influence relationship evaluation value. However, the analysis may also be performed using a combination of a symptom-related causality score and other types of influence relationship evaluation values (responsiveness score, relevance score, cosine similarity, or eGene expression level).
[0106] Figure 14 shows an example of processing by the molecular extraction unit 12 when "acute kidney injury" is specified as a specific disease and multiple molecules whose association with the disease is estimated by the related molecule estimation unit 51 are used as the target of analysis. In Figure 14, symbols are plotted on a two-dimensional plane with the association score on the vertical axis and cosine similarity on the horizontal axis, at coordinate positions identified by the combination of association scores and cosine similarity values obtained by the evaluation value acquisition unit 11 for each of the multiple molecules to be analyzed.
[0107] The molecular extraction unit 12 extracts molecules corresponding to symbols placed in a specific region AR9 that satisfies predetermined conditions on this two-dimensional plane, as candidate genes for drug targets or biomarkers. A specific region AR9 that satisfies predetermined conditions is a region in which the relevance score is above a predetermined value (0.7 or higher) and the cosine similarity is above a predetermined value (0.15 or higher).
[0108] The plotted positions of symbols corresponding to several molecules known as biomarkers for acute kidney injury are located within the specific region AR9. This confirms that the analytical method according to this embodiment is effective in the search for biomarkers for acute kidney injury.
[0109] Figure 15 shows another example of processing by the molecular extraction unit 12 when "acute kidney injury" is specified as a specific disease and multiple molecules whose association with the disease is estimated by the related molecule estimation unit 51 are used as the target of analysis. In Figure 15, symbols are plotted on a two-dimensional plane with the association score on the vertical axis and the responsiveness score on the horizontal axis, at coordinate positions identified by the combination of association scores and responsiveness scores obtained by the evaluation value acquisition unit 11 for each of the multiple molecules to be analyzed.
[0110] The molecular extraction unit 12 extracts molecules corresponding to symbols placed in a specific region AR10 that satisfies predetermined conditions on this two-dimensional plane, as candidate biomarker genes. A specific region AR10 that satisfies predetermined conditions is a region in which the relevance score is above a predetermined value (0.7 or higher) and the responsiveness score is above a predetermined value (0.7 or higher).
[0111] The plotted positions of symbols corresponding to several molecules known as biomarkers for acute kidney injury are located within the specific region AR10. This confirms that the analytical method according to this embodiment is effective in the search for biomarkers for acute kidney injury.
[0112] Note that the two-dimensional scatter plots shown in Figures 6 to 15 are illustrated to illustrate the processing performed by the molecular extraction unit 12, and it is not necessarily required to generate scatter plots. Alternatively, the molecular extraction unit 12 may generate scatter plots as illustrated in Figures 6 to 15 and display the generated scatter plots on the display unit 104. In this case, specific regions may be displayed in a way that allows for identification.
[0113] Alternatively, the user may specify an arbitrary rectangular area on the two-dimensional plane of the scatter plot displayed on the display unit 104, and the molecular extraction unit 12 may analyze the rectangular area specified by the user as a specific region AR that satisfies predetermined conditions. In this case, predetermined conditions are identified by the combination of two influence relationship evaluation values based on the boundary of the rectangular area specified by the user, and the molecular extraction unit 12 extracts molecules contained within the rectangular area specified by the user.
[0114] As explained in detail above, in this embodiment, for each of the multiple molecules to be analyzed, multiple influence relationship evaluation values representing the characteristics of the relationship between the disease or symptom and the molecule are obtained, and molecules whose combination of multiple influence relationship evaluation values satisfies predetermined conditions are extracted from among the multiple molecules to be analyzed. Here, at least one of the multiple influence relationship evaluation values is an evaluation value identified using disease-related feature vectors and molecular feature vectors.
[0115] According to this embodiment, when searching for drug targets or biomarkers using multiple influence relationship evaluation values that represent the characteristics of the influence relationship between a disease or symptom and a molecule, not only evaluation values obtained from experiments but also evaluation values identified by natural language processing of multiple sentences are used. Therefore, it is possible to search for drug targets and biomarkers beyond the scope of knowledge that can be grasped by experiments alone.
[0116] In the above embodiment, an example was described in which two of the multiple influence evaluation values used for analysis were used, but the invention is not limited to this. For example, any three or four of the causality score, responsiveness score, relevance score, cosine similarity, and eGene expression level may be used, or all five may be used. In this case, the molecular extraction unit 12 may, for example, extract molecules from among the multiple molecules to be analyzed that satisfy the condition that all of the multiple influence evaluation values used are above a predetermined value.
[0117] Furthermore, in the above embodiment, the example given for a predetermined condition identified by a combination of multiple influence evaluation values is that all influence evaluation values are greater than or equal to a predetermined value (the predetermined value may differ for each influence evaluation value), but the embodiment is not limited to this. For example, the condition may be to extract a predetermined number of molecules starting with the one with the larger influence evaluation value.
[0118] Furthermore, for multiple influence evaluation values, a combination of conditions may be applied: one where the value is above a predetermined value, and another where a predetermined number of molecules are extracted from the values with the highest values. For example, in the case where two influence evaluation values, the causality score and the eGene expression level, are used as shown in Figures 6 to 8, the condition that the causality score is 0.7 or higher may be applied, while the condition that a predetermined number of molecules are extracted from the values with the highest eGene expression levels may be applied.
[0119] Furthermore, the above embodiments are merely examples of how the present invention may be implemented, and the technical scope of the invention should not be interpreted as being limited by them. In other words, the present invention can be implemented in various ways without departing from its gist or its main features.
[0120] 1 Target Search Device 11 Evaluation Value Acquisition Unit 12 Molecular Extraction Unit 13 Text Data Storage Unit 14 First Model Storage Unit 15 Second Model Storage Unit 31 Word Extraction Unit 32 Vector Calculation Unit 32A Text Vector Calculation Unit 32B Word Vector Calculation Unit 33 Index Value Calculation Unit 34 Feature Vector Identification Unit 51 Related Molecules Estimation Unit 52 Molecular Properties Estimation Unit 53 Similarity Calculation Unit 111 Feature Vector Calculation Unit 112 Score Calculation Unit
Claims
1. A target search device comprising: an evaluation value acquisition unit that acquires multiple influence relationship evaluation values for each of the multiple molecules to be analyzed, each representing the characteristics of the influence relationship between a disease or symptom and the molecule; and a molecule extraction unit that extracts molecules from the multiple molecules to be analyzed in which the combination of the multiple influence relationship evaluation values satisfies predetermined conditions, wherein at least one of the multiple influence relationship evaluation values is an evaluation value identified using disease-related feature vectors calculated for the disease or symptom and molecular feature vectors calculated for the molecule by natural language processing of multiple sentences.
2. The target search device according to claim 1, wherein at least one of the above-mentioned multiple influence relationship evaluation values is a causality score indicating the probability that the molecule is causal in that it has properties that act on the above-mentioned disease or symptom, and the evaluation value acquisition unit inputs the disease-related feature vector and the molecular feature vector into a trained model and outputs the causality score from the trained model.
3. The target search device according to claim 1, wherein at least one of the above-mentioned multiple influence relationship evaluation values is a responsiveness score indicating the probability that the molecule is responsive as a property that acts on the above-mentioned disease or symptom, and the evaluation value acquisition unit inputs the disease-related feature vector and the molecular feature vector into a trained model and outputs the responsiveness score from the trained model.
4. The target search device according to claim 1, characterized in that at least one of the above-mentioned multiple influence relationship evaluation values is the similarity between the disease-related feature vector and the molecular feature vector.
5. The target search device according to claim 1, wherein at least one of the above-mentioned multiple influence relationship evaluation values is a relevance score indicating the degree to which the molecule is related to the disease or symptom, and the evaluation value acquisition unit inputs the disease-related feature vector to a trained model, estimates molecules corresponding to the molecular feature vector similar to the disease-related feature vector as molecules related to the disease or symptom, and outputs the relevance score indicating the likelihood of the estimation from the trained model.
6. The target discovery device according to claim 5, characterized in that the plurality of molecules to be analyzed are plurality of molecules whose association with the disease or symptoms is estimated by the evaluation value acquisition unit.
7. The target discovery device according to any one of claims 1 to 4, characterized in that the multiple molecules to be analyzed are eGenes identified by disease genome analysis.
8. The target discovery device according to claim 7, characterized in that the multiple influence relationship evaluation values include, in addition to evaluation values identified using the disease-related feature vector and the molecular feature vector, the amount of increase or decrease in eGene expression.
9. A target search method comprising the steps of: a computer evaluation value acquisition unit acquiring multiple influence relationship evaluation values for multiple molecules to be analyzed, each representing the characteristics of the influence relationship between a disease or symptom and the molecule; and a computer molecule extraction unit extracting molecules from the multiple molecules to be analyzed in which the combination of the multiple influence relationship evaluation values satisfies predetermined conditions, wherein at least one of the multiple influence relationship evaluation values is an evaluation value identified using disease-related feature vectors calculated for the disease or symptom and molecular feature vectors calculated for the molecule by natural language processing of multiple sentences.
Citation Information
Patent Citations
Drug relocation method and device, electronic equipment and storage medium
CN115579053A
Pathway generation device, pathway generation method, and pathway generation program
JP6915818B1
Information analysis device, information analysis method, and information analysis program
JP7034453B1