Computer-assisted drug target selection

A computer-based method for drug target selection using machine learning to analyze publication data optimizes the drug discovery process by efficiently identifying relevant targets, reducing time and costs.

JP7911002B2Active Publication Date: 2026-08-25EXSCIENTIA AI LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023550727
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-29
Filing Date
2021-10-29
Publication Date
2026-08-25
Estimated Expiration
2041-10-29

AI Technical Summary

Technical Problem

The vast amount of scientific literature and data makes it impossible for humans to efficiently identify and select promising drug targets, leading to inefficiencies and increased costs in the drug discovery process.

Method used

A computer-based method for drug target selection that retrieves and analyzes publication data using machine learning algorithms to classify and rank potential drug targets based on their relevance and ambiguity, enabling informed selection for drug discovery projects.

Benefits of technology

This method reduces the time and cost associated with drug discovery by optimizing the identification and selection of drug targets, improving the efficiency of identifying candidate compounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007911002000001
    Figure 0007911002000001
  • Figure 0007911002000002
    Figure 0007911002000002
  • Figure 0007911002000003
    Figure 0007911002000003
Patent Text Reader

Abstract

The present invention provides a method for computational drug target selection. The method includes retrieving published data related to a plurality of published documents, including historical published documents and current published documents, from at least one published data source. The method includes searching the published data to provide, for each of the published documents, an indication of whether the respective published document is related to one or more drug targets. The method includes determining expected published parameters for each of the one or more drug targets based on the retrieved published data from the historical published documents and determining actual published parameters for each of the one or more drug targets based on the retrieved published data from the current published documents; and evaluating each of the one or more drug targets for selection based on its actual published parameters relative to its expected published parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to methods and systems for selecting a target molecule or gene, such as a drug target, by computer, and for designing a molecule, such as a drug, to interact in an optimal manner.

Background Art

[0002] Drug discovery is the process of identifying candidate compounds for the next stage of pharmaceutical development, such as preclinical trials. Such candidate compounds need to meet certain criteria for further development. Modern drug discovery involves the identification and optimization of "hit" compounds in an initial screening. In particular, such compounds need to be optimized against the required criteria, which includes the optimization of a number of different properties. Properties to be optimized can include, for example: activity against a desired biological target; selectivity against an undesired biological target; low toxicity potential; and good pharmacokinetic and pharmacodynamic properties (ADME). Only compounds that meet the specified requirements can become candidate compounds and proceed to the drug discovery process.

[0003] Therefore, the identification and selection of the biological or drug target against which a hit compound is to be next optimized is an important step in the drug discovery process, and in fact, target identification and prioritization is the first important step in the drug discovery process and the development of new pharmaceuticals. A drug target is something that exists within an organism with which a drug interacts, e.g., binds, and is usually, for example, a protein or nucleic acid. Such interaction with a drug causes a change in the behavior of the drug target. A promising drug target can be related to a particular disease under consideration, e.g., the drug target can modify the disease or play a role in the pathophysiology of the disease.

[0004] The process of selecting a drug target is complex due to the vast number of possible drug targets available. For example, in the case of human diseases, there are tens of thousands of genes that express proteins that could potentially be targeted by new drugs. Furthermore, just as there are thousands of human diseases classified by medicine, there are millions, and sometimes even hundreds of millions, of possible target-disease combinations. Therefore, the search space for solutions is so vast that it is impossible to experimentally test every combination and hypothesis.

[0005] Traditionally, drug targets have been identified on a case-by-case basis by medicinal chemists interpreting published scientific literature such as academic journals and public databases. In other words, a considerable amount of target identification has traditionally been performed by individual scientists using their expertise to interpret scientific literature. However, a growing problem with this approach is the sheer volume of public data, such as searchable academic papers. In the life sciences, tens of millions of scientific papers are published, hundreds of thousands of genomes, and hundreds of databases exist. In fact, thousands of peer-reviewed papers are published daily, without considering other data sources such as preprints and clinical trial reports. Therefore, it is clear that it is impossible for humans to keep track of all available data sources when selecting drug targets. In other words, the increasing rate of publication makes it difficult to maintain an overview for identifying promising new or existing drug targets. [Overview of the project] [Problems that the invention aims to solve]

[0006] Optimizing the identification and selection of drug targets is crucial for optimizing the entire drug discovery process. In particular, optimally selecting drug targets for a specific drug discovery project increases the likelihood of identifying candidate compounds in a shorter time, i.e., with fewer design cycles for the project. This reduces the time and / or costs associated with that particular project.

[0007] This invention is established based on the above background. [Means for solving the problem]

[0008] Summary of the Invention The present invention provides an improved method for identifying biological targets of drugs, possibly involved in a particular disease, to reduce the overall time and / or cost associated with the drug discovery process, for example, to increase the efficiency of identifying candidate compounds as part of a particular drug discovery project. Furthermore, the present invention provides a method for drug discovery. In particular, a method comprising selecting at least one drug target, the method may include initiating a drug discovery project based on the at least one drug target, and optionally selecting and / or synthesizing and / or testing potential therapeutic compounds against the at least one selected drug target.

[0009] According to one aspect of the present invention, a method for computer-based drug target selection is provided. The method includes retrieving publication data related to a plurality of publication documents, including past and present publication documents, from at least one publication data source. The method includes retrieving the publication data and providing an indicator for each of the publication documents of whether each publication document is related to one or more drug targets. The method includes determining expected publication parameters for each of the one or more drug targets based on the retrieved publication data from the past publication documents, and determining actual publication parameters for each of the one or more drug targets based on the retrieved publication data from the present publication documents. The method includes evaluating each of the one or more drug targets for selection based on its actual publication parameters relative to its expected publication parameters.

[0010] This method may include defining one or more character representations for each drug target as referring to the drug target, and searching the published data includes searching the published data for the one or more character representations for each drug target.

[0011] This method may include classifying each of the one or more character representations for each drug target as either a safe or unsafe character representation. This classification may be based on the likelihood that instances of the character representation in the published data refer to the drug target.

[0012] In some embodiments, if the retrieved published data from one of the published documents contains safe character representations, that published document is determined to be related to the drug target.

[0013] One or more user-defined character expressions can be classified as safe character expressions.

[0014] Users can define one or more unsafe features of a character representation to indicate that the corresponding character representation is unsafe. Character representations in retrieved public data that exhibit one or more unsafe features may be classified as unsafe character representations.

[0015] The unsafe features of the one or more user-defined character representations mentioned above may include one or more of the following: character representations corresponding to specific natural language words; character representations having fewer than a predetermined number of characters, where the predetermined number is optionally three; and character representations defined as referring to at least two different drug targets.

[0016] Ambiguity features can be defined for one or more character representations, and an ambiguity score can be assigned to one or more character representations. Each of these character representations can be classified as either a safe or unsafe character representation based on the corresponding assigned ambiguity score.

[0017] Users can define ambiguity features for one or more character representations.

[0018] If the aforementioned character representation has an ambiguity score greater than a predetermined threshold ambiguity score, it may be classified as an unsafe character representation.

[0019] The ambiguity feature of the one or more character representations may include, for each drug target: the total number of published documents in the published data that include the one or more defined character representations that refer to the drug target; the number of published documents in the published data that include one of the defined character representations that refer to the drug target, relative to the total number of published documents in the published data that include the one or more defined character representations that refer to the drug target; the number of characters in one of the defined character representations that refer to the drug target; the frequency with which each character in one of the defined character representations that refer to the drug target appears in the published data, optionally the sum of the frequencies of each of the characters in that character representation, optionally the logarithm of that sum; the number of defined character representations for the one or more drug targets that include the one defined character representation; the probability that the selected character representation is also included in a published document in the published data that includes one of the defined character representations other than the selected character representation that is a safe character representation from the defined character representation that refers to the drug target; and the probability that one of the defined character representations other than the selected character representation is also included in a published document in the published data that includes the selected character representation.

[0020] This method can include applying a machine learning algorithm to assign an ambiguity score to each of the one or more character representations based on the ambiguity characteristics of the one or more character representations.

[0021] The machine learning algorithm can use non-safe characteristics of the one or more character representations to assign an ambiguity score to each of the one or more character representations.

[0022] The machine learning algorithm may include positive unlabeled learning techniques.

[0023] The machine learning algorithm can include applying a random forest classifier.

[0024] In some embodiments, after each iteration of the machine learning algorithm, a subset of the assigned ambiguity scores is inspected by a user and a determination is made as to whether to manually change any of the subset of the assigned ambiguity scores.

[0025] The subset may correspond to a predetermined number of character representations having the highest assigned ambiguity scores.

[0026] The published data for at least a portion of the published document may include citation data indicating a citation made by one published document to one or more other published documents from the plurality of published documents. Searching the published data may include using the citation data to identify pairs of published documents cited by the same published document.

[0027] This method may include determining, for each identified pair of published documents, a co-citation value representing the number of published documents that cite both of the pair of published documents.

[0028] This method can include assigning a pair of published documents to one of a plurality of communities of published documents based on the determined co-citation value and the published document that cites the pair of the published documents.

[0029] In some embodiments, assigning a pair of published documents to one of a plurality of communities includes applying a powerful optimization algorithm.

[0030] This method may include, for each of the plurality of communities of published documents, determining whether to associate the community with one of the drug targets.

[0031] The determination may include determining which of the defined character representations that reference the one drug target are present in the published data of each of the published documents within the community.

[0032] The determination may include determining the proportion of published documents within the community whose published data includes at least one safe character representation. In some embodiments, if the proportion is greater than a predetermined threshold proportion, the community is determined to be associated with the one of the drug targets.

[0033] In some embodiments, searching for a pair of published documents includes searching for a pair of published documents that includes at least one of the character representations defined as each referencing one of the drug targets.

[0034] In some embodiments, the published data for at least some of the published documents does not include citation data. For each of the published documents, the method may include determining whether to assign the published document to one of the communities associated with one of the drug targets, based on the published data, in particular on one or more of the defined character representations that refer to the drug targets within the published data.

[0035] In some embodiments, if the published data of the published document includes at least one instance of a safe character representation, it is decided to assign the published document to one of the communities associated with one of the drug targets.

[0036] In some embodiments, if the published data of the published document does not include at least one instance of a safe character representation, the decision of whether to assign the published document to one of the communities associated with one of the drug targets is performed using a machine learning algorithm.

[0037] The aforementioned machine learning algorithm may include positive unlabeled learning techniques.

[0038] The machine learning algorithm may include the application of a machine learning classifier, which may optionally include at least one of the following: logistic regression classifier; extra-tree classifier; Gaussian process classifier; k-nearest neighbor classifier; ridge classifier; random forest classifier; and support vector machine classifier.

[0039] For each drug target, the expected publication parameter is the expected number of published documents related to the drug target, and the actual publication parameter may be the actual number of published documents related to the drug target.

[0040] For each drug target, the expected publication parameter may be one of the following: the expected number of clinical trials related to the drug target; the expected number of review publications related to the drug target; and the expected number of publications related to the defined company size; and the actual publication parameter may be one of the following: the actual number of clinical trials related to the drug target; the actual number of review publications related to the drug target; and the actual number of publications related to the defined company size.

[0041] In some embodiments, determining the expected publication parameters involves using a machine learning algorithm trained with the retrieved publication data from the past publication documents.

[0042] The aforementioned machine learning algorithm may also be a recurrent neural network algorithm.

[0043] In some embodiments, evaluating the drug targets for selection includes ranking the drug targets based on a comparison of their actual published parameters with expected published parameters.

[0044] The aforementioned drug targets may be ranked according to a parameter that indicates the difference between their actual published parameters and their expected published parameters.

[0045] The method may include determining target-target co-occurrence parameters between the drug target pairs, which are determined based on an index from the retrieved published data regarding which published documents are associated with both drug targets of the pair. Each target-target co-occurrence parameter may indicate the number of published documents in which both drug targets of the pair appear. The method may include evaluating one or more drug targets for selection based on the determined target-target co-occurrence parameters.

[0046] This method may include searching the published data and providing an indicator for each of the published documents of whether each published document is associated with one or more diseases.

[0047] This method may include defining one or more character representations that refer to each disease. Searching the published data may include searching the published data for the one or more character representations for each disease.

[0048] This method may include determining a target-disease co-occurrence parameter between each of the drug targets and each of the diseases. This target-disease co-occurrence parameter may be determined based on the indicator from the retrieved published data regarding which published documents each drug target and each disease is associated with. Each target-disease co-occurrence parameter may indicate the number of published documents in which one of the drug targets and one of the diseases appear. This method may include evaluating one or more of the drug targets for selection based on the determined target-disease co-occurrence parameter.

[0049] This method may include applying a topic modeling algorithm to the published data of the published documents relating to each of the drug targets to obtain one or more topics relating to each drug target. This method may also include evaluating the one or more drug targets for selection based on the one or more topics obtained.

[0050] This method may include determining, for each drug target, the error in the relationship between one or more published documents and the drug target based on one or more topics obtained.

[0051] Topic modeling algorithms may include at least one of the following: latent Dirichlet assignment algorithms; and non-negative matrix factorization algorithms.

[0052] The published data relating to one or more of the aforementioned published documents may include one or more of the following: the title of the published document; a summary of the published document; and one or more keywords relating to the published document.

[0053] The aforementioned published data may include the publication date for each of the aforementioned multiple published documents.

[0054] The publication date can define whether each of the aforementioned publication documents is a past publication or a current publication.

[0055] In some embodiments, a published document having a publication date prior to a predetermined threshold date is defined as a past published document.

[0056] In some embodiments, a published document having a publication date after the predetermined threshold date is defined as the current published document.

[0057] In some embodiments, a published document having a publication date within a predetermined threshold date range is defined as a current published document.

[0058] The aforementioned at least one published data source may include at least one online published data source.

[0059] The aforementioned one or more drug targets may include one or more genes, optionally one or more human genes, and optionally one or more proteins encoded by such genes.

[0060] This method may include using the evaluation of one or more drug targets to inform the selection of at least one of the drug targets for use in a drug discovery project.

[0061] This method may include designing the drug discovery project by selecting at least one of the drug targets for use in the drug discovery project based on the evaluation.

[0062] This method may include initiating the drug discovery project using the at least one selected drug target.

[0063] In some embodiments, commencing the drug discovery project includes selecting compounds for at least one selected drug target, optionally synthesizing and testing them (in computer, in a test tube, and / or in vivo).

[0064] According to another aspect of the present invention, a method is provided for identifying a drug / compound having binding affinity to a drug target / target molecule, the method comprising commencing a drug discovery project (e.g., based on the method for identifying drug targets according to the aspects and embodiments disclosed herein) and optionally selecting and / or synthesizing and / or testing compounds against at least one selected drug target in order to identify compounds having therapeutic activity against the drug target, where “therapeutic activity” may include, but is not limited to, desirable binding properties (e.g., affinity, selectivity); inhibitory properties; and agonist or antagonist properties.

[0065] Any feature of any aspect or embodiment disclosed herein may be combined with any feature of any other aspect or embodiment disclosed herein, and it will be understood that all such combinations of features are conceivable and disclosed herein, provided that such combinations do not obviously contradict each other.

[0066] According to another aspect of the present invention, a non-temporary computer-readable storage medium is provided that stores instructions causing a computer processor to perform the above-described method when executed by the computer processor.

[0067] According to another aspect of the present invention, a computer device for selecting drug targets is provided. The computer device is configured to take in, receive, or download publication data relating to a plurality of publication documents, including past and current publication documents, from at least one publication data source. The computer device is configured to search the publication data and provide an index for each of the publication documents regarding whether each publication document relates to one or more drug targets. The computer device is configured to determine expected publication parameters for each of the one or more drug targets based on the retrieved publication data from the past publication documents, and to determine actual publication parameters for each of the one or more drug targets based on the retrieved publication data from the current publication documents. The computer device is configured to evaluate each of the one or more drug targets for selection based on its actual publication parameters relative to its expected publication parameters.

[0068] Next, embodiments of the present invention will be described with reference to the following drawings. [Brief explanation of the drawing]

[0069] [Figure 1] This diagram summarizes the steps of the computer-based drug selection method according to the present invention. [Figure 2] This figure shows a graph database illustrating the relationships between published documents determined to be related to a specific target gene using the method described in Figure 1. [Figure 3] This figure shows a comparison of predicted and actual publication dynamics associated with different genes, determined using the method in Figure 1. [Figure 4] This figure shows the predicted dynamics relative to actual published dynamics, related to different genes associated with different disease groups, as determined using the method in Figure 1. [Figure 5] This figure shows the co-occurrence network indicating gene-gene relationships and gene disease associations determined using the method in Figure 1. [Figure 6] This figure shows a timeline of the number of publications mentioning different genes of interest for different groups of publications related to different extracted topics, determined using the method in Figure 1. [Modes for carrying out the invention]

[0070] Detailed explanation The design of molecules or drugs can be considered a multidimensional optimization problem that uses a cycle of hypothesis generation and experimentation to advance knowledge. Each compound design can be seen as a hypothesis disproven by experiment. Experimental results are expressed as structure-activity relationships, building a hypothesis landscape regarding which chemical structures are likely to contain the desired features. The drug design process is also an optimization problem, as each project requires a defined product profile of desired specified attributes—i.e., drug target function—for analyzing hit compounds.

[0071] The drug discovery process is typically carried out in iteratives known as design cycles. In each iteration, a set of molecules or compounds are synthesized, and their biological properties are measured. Their activity is analyzed, and based on the results from previous iterations, a new set of compounds is proposed. This process is repeated until a clinical candidate is found. In addition to activity, the biological properties measured may include one or more of selectivity, toxicity, absorption, distribution, metabolism, and excretion.

[0072] The drug discovery process is generally time-consuming and expensive. Therefore, improving efficiency at any stage of the process helps reduce the time and costs associated with drug discovery projects. Pharmaceutical companies are actively seeking ways to reduce employee turnover, the time required for drug development, and the associated development costs.

[0073] The selection of drug targets for new drug development is the first, and perhaps most important, decision in the drug discovery process. Historically, target identification has been widely carried out on a case-by-case basis, relying on the scientific interpretation of available literature. However, thousands of peer-reviewed papers are published daily, in addition to preprints, patent document data, and clinical trial reports. The online resource PubMed® allows searching and access to published documents in the fields of life sciences and biomedical information. PubMed® alone provides access to tens of millions of published documents, particularly over 30 million publications, with scientific output doubling approximately every nine years. This clearly makes it impossible for humans, such as medicinal chemists, to keep up with all the developments in published literature, thus creating a corpus of "undiscovered public knowledge." As a result, it becomes more difficult for humans to make informed decisions regarding the identification and selection of drug targets based on available literature.

[0074] The vast search space for potential drug targets also makes it difficult for humans to make optimal choices. In the case of human diseases, theoretically there are tens of thousands of genes that could potentially play a role in the behavior of a particular disease, and millions of gene-disease combinations that could be studied, but in reality, this is clearly not feasible.

[0075] The problems described above mean that the use of computational techniques to analyze the vast amount of available information and the enormous number of possible gene-disease combinations becomes an attractive proposal. In particular, there is a high demand for machine learning (ML), artificial intelligence (AL), and other computer techniques to leverage current knowledge and facilitate the maintenance of an overview of this overwhelming amount of literature, with the aim of optimizing the identification and selection of drug targets.

[0076] The present invention recognizes that computer-aided methods can be used to identify trends in published literature concerning potential drug targets, such as genes, which can be used, for example, to inform the selection of drug targets for a particular drug discovery project. The present invention is advantageous in that it provides a computer-aided method for drug target selection that can detect, for example, changes in fundamental assumptions about a particular gene, such as trends in change about a particular gene that are interpreted as indicating scientific progress regarding that gene.

[0077] According to the present invention, a first step of a computer-based drug target selection method includes incorporating published data from at least one published data source. For example, the published data source may be an online published data source, or it may include a database such as PubMed® that provides access to millions of published documents in the form of academic papers or journal articles in a particular area of ​​interest. The published data may also be incorporated from other sources, such as published clinical trial report data, published patent document data, and / or published preprints of papers, as an addition or alternative. It will be understood that the published data can be obtained and incorporated from any suitable source of such published data, and from any number of these suitable sources.

[0078] The information contained in the published data incorporated for a given published document may depend on the specific type of document or the specific source from which the published data is obtained. For example, the published data may be limited to the data available as open-source data for a given published document, with further information only accessible behind a paywall.

[0079] The incorporated published data is related to multiple different published documents; that is, each published document has related published data. To detect literature trends over time, the published documents from which the published data is incorporated can be considered to be divided into historical published documents and current published documents. Such a division is useful for comparing the trends that would be expected to be observed based on historical literature with the trends actually observed based on current publications, as will be explained in detail below. It will be understood that the incorporation of historical and current published data can be performed separately or simultaneously.

[0080] Publication data relating to at least some of multiple published documents may include the publication dates of those published documents, or publication dates related to those published documents. This can be used to determine or define which documents are identified as historical published documents and which are identified as current published documents. The publication date may be, for example, the specific day, month, or year in which the related documents were published. As merely one example, published documents with publication dates prior to a given threshold date may be defined as historical published documents, while published documents with publication dates after a given threshold date may be defined as current published documents. However, the publication date of a document may be used in any appropriate way to define whether a document is historical or current; for example, published documents with publication dates within a given threshold date range may be defined as current published documents.

[0081] The next step of the present invention includes searching the incorporated published data to provide an indicator for each published document of whether each published document is related to one or more drug targets, such as genes. That is, searching the information contained in the published data of each published document to identify the potential relevance or link that each published document has to one or more drug targets. Such a search can be performed in any suitable way. One option is to search for mentions of drug targets of interest within the published data. In this regard, the published data may include one or more data such as the title, abstract, and one or more keywords related to the published document. Such information is generally readily available as open-source information from online published databases that store journal articles, for example, and therefore this information can be readily incorporated as part of the published data related to various published documents.

[0082] In one example, the names of one or more target drug targets, such as target genes, are defined (e.g., by the user), and the published data is automatically searched for the defined drug target names. For example, if the defined name of one drug target is found in published data related to a particular published document, that drug target can be considered associated with or linked to that particular published document. The defined name of a drug target may also be an approved symbol, such as an approved gene symbol, according to an approved nomenclature. Hereafter, "gene symbol" is used to refer to the approved symbol of a particular gene from any of the 19,084 human protein-coding genes approved by the HUGO Gene Nomenclature Committee, but it should be understood that this is for illustrative purposes only and is not limited to this.

[0083] A major obstacle in automated analysis of biomedical literature using computational methods is the use of non-overlapping synonyms, symbols, and acronyms for alternative genes (drug targets) from different competing information sources, which may have different meanings in other research fields. That is, a single drug target, such as a gene, may be referenced in various ways accepted in the field within the literature. Furthermore, one or more synonyms for a particular drug target may correspond to, be part of, terms or expressions of entirely different concepts, or have entirely different meanings in different contexts. These factors make it difficult for automated (computer) analysis to definitively determine which published documents containing references matching a drug target name actually refer to that drug target.

[0084] To identify references to a drug target mentioned in various ways, by various names, or in various languages ​​within the published data, this method may include defining one or more character representations or synonyms for each drug target that refer to that drug target. These character representations can be defined by the user and may include any suitable character that can be computationally retrieved. For example, suitable characters may include one or more letters used in natural languages ​​or other types of symbols. Then, for each drug target, searching the published data may include searching the published data for one or more defined character representations or synonyms for each drug target.

[0085] In this context, when the drug target is a gene, a “gene synonym” may refer to any variation of a gene name that the scientific community refers to, or that may have referred to, a specific gene. The approved gene symbols defined above are also included in gene synonyms. As an example, “EGFR” is an approved gene symbol, but “EGFR,” “epidermal growth factor receptor,” “ERBB1,” “ErbB-1,” “c-erbB1,” “HER1,” and “ERBB” may be gene synonyms.

[0086] Various characterizations, defined as potentially referring to specific drug targets, may be obtained from various information sources. For example, if one is interested in human genes, various characterizations, or synonyms, of different human genes may be collected from different sources in order to sample published documents that may refer to human gene names.

[0087] As mentioned earlier, another problem with automated analysis of biomedical literature is that instances of accepted synonyms for drug targets in published data—that is, instances of defined characterizations—may or may not actually refer to the drug target in a particular document. In fact, it is not uncommon for a particular characterization to be used to refer to two different genes in different contexts, for example. Therefore, clarifying the biomedical entities within the scientific literature is essential for accurately analyzing the publication trends of various drug targets.

[0088] Each character expression defined for a drug target, such as different synonyms for a gene, may be considered to have different levels of ambiguity. For example, several synonyms found in the literature that have meanings other than the gene of a particular interest may be considered to have a greater level of ambiguity because the examples of those synonyms in the literature are likely to refer to something other than the gene of interest. Conversely, if a particular character expression or synonym has no (common) meaning other than referring to the gene of a particular interest, such a synonym may be considered to have a lower level of ambiguity because the examples of that synonym in the literature are likely to refer to the gene of interest.

[0089] The level of ambiguity associated with gene synonyms (characterizations of drug targets) can arise for a variety of reasons. For example, so-called "indistinguishable gene names" can be considered any gene name that is synonymous with multiple genes. This may include previous official gene symbols (according to accepted nomenclature) because they have not been removed from the literature. As a helpful example, "CDH3" and "cadherin3" are indistinguishable from the gene symbols "CHD15" and "CHD3". Also, "ARP1" is a gene synonym for the gene symbols "NR2F2", "ACTR1A", "ACTR1B", "ANGPTL1", "APOBEC2", "ARFRP1", and "PITX2". Another example is so-called "nested gene synonyms", which can be considered gene synonyms that are part of another gene synonym. For example, "insulin" is a synonym for the nested gene "insulin receptor". Furthermore, "TNF" is a synonym for the nested genes "TNF receptor superfamily member 1A" (gene symbol "TNFRSF1A") and "TNF receptor-related factor 2" (gene symbol "TRAF2").

[0090] Therefore, to enable a more accurate analysis of the retrieved published data, the method of the present invention may include classifying each of the one or more defined character expressions (e.g., synonyms for the gene) found in the published data for each drug target (e.g., a gene) as either a safe character expression or an unsafe character expression. This classification is based on the likelihood that instances of the character expression in the published data refer to a drug target, i.e., the level of ambiguity associated with the character expression. In particular, safe character expressions may be associated with a relatively low level of ambiguity, while unsafe character expressions may be associated with a relatively high level of ambiguity. For example, if a character expression is a synonym for a gene, an "unsafe gene synonym" may include synonyms for the gene that have different meanings in other fields of study or different contexts, such as words found in an English dictionary. As an example, the gene symbol "STAR" may be considered an unsafe character expression, in contrast to its gene synonym, "steroidogenic acute regulatory protein." As another example, "CCP4" may be considered unsafe because it is both a gene synonym and the name of crystal analysis software.

[0091] If published data retrieved from one of the published documents is determined to contain safe character expressions (for a specific drug target), then that published document may be determined to be associated with that drug target. In other words, that published document is considered to be linked to or associated with the specific drug target under consideration, e.g., the gene of interest. For a particular drug target, at least some of the defined character expressions that may refer to that drug target, i.e., those appearing in the literature, can certainly be considered safe character expressions. In particular, one or more character expressions can be user-defined to be classified as safe character expressions. That is, because there are certain character expressions that are unambiguous or have a very low level of ambiguity, it is a priori known that instances of those character expressions in published data actually refer to the relevant drug target, regardless of the context in which those instances appear in the published data. Therefore, such character expressions can be automatically classified as safe character expressions when found in published data. This means that if such character expressions are found in the published data of a published document, that published document may be automatically determined to be associated with or linked to the drug target for which the character expression is a defined synonym.

[0092] One or more characteristics of a character expression may be defined, for example by a user, to indicate that the character expression exhibiting such characteristics has a level of ambiguity that is unsafe. In particular, character expressions exhibiting “unsafe features of character expressions” included in retrieved published data may be automatically classified as unsafe character expressions. An example of a user-defined unsafe feature of a character expression is any character expression corresponding to a specific natural language word, for example, a word in an English dictionary (see the “STAR” example above). Another example of a user-defined unsafe feature of a character expression may be a character expression having fewer characters than a given number. For example, the given number may be 3, or it may be any other appropriately defined number. A further example of a user-defined unsafe feature of a character expression may be a character expression defined to refer to at least two different drug targets (see “indiscriminate gene name” above). Thus, a character expression containing at least one defined unsafe feature is considered obviously unsafe. It will be understood that any appropriate feature of a character expression may be defined as indicating that the character expression is highly ambiguous and unsafe.

[0093] According to the above definition, there may be a considerable number of defined character representations that are neither definitively safe nor definitively unsafe. In this regard, the level of ambiguity associated with the remaining character representations can be determined or calculated to classify these character representations as safe or unsafe. One option for determining which character representations or synonyms have a high level of ambiguity that would make them potentially unsafe is to perform feature engineering to obtain variables that characterize unsafe synonyms and then assign a level of ambiguity to each of the synonyms based on the obtained variables. For example, the longer a gene name is, the less likely it is to be ambiguous. More generally, the ambiguity feature of a character representation can be defined in any suitable way to attribute an ambiguity score to one or more character representations. This may be user-defined, for example, based on feature engineering, or otherwise. Each of these character representations can be classified as either a safe or unsafe character representation based on the correspondingly attributed ambiguity score. For example, a character representation may be classified as an unsafe character representation if its ambiguity score is greater than a given threshold ambiguity score.

[0094] By applying machine learning algorithms, for example through feature engineering, ambiguity scores can be assigned to character representations or synonyms based on the ambiguity features of the acquired character representations. In particular, machine learning algorithms can use the unsafe features of character representations to assign ambiguity scores to each character representation that has not yet been classified as safe or unsafe. That is, ambiguity scores are assigned to synonyms in the set of unlabeled synonyms, i.e., synonyms that are not included in the set of synonyms previously labeled as safe or previously labeled as unsafe. The ambiguity scores are then used to label the unlabeled synonyms as safe or unsafe. To achieve this, machine learning algorithms may include the application of positive unlabeled learning techniques, such as positive unlabeled bagging strategies, and classification schemes, such as random forest classifiers.

[0095] Machine learning algorithms can be run as iterative processes. With each iteration of the algorithm, a subset of the returned ambiguity scores can be examined by the user to decide whether to manually modify any of them—that is, to correct the classifications made by the algorithm to train it and improve the accuracy of subsequent iterations. For example, the subset might correspond to a predetermined number of synonyms or character representations that have the highest returned ambiguity scores (and are therefore considered the least safe by the algorithm).

[0096] For example, the ambiguity feature of a character representation obtained through feature engineering may include the total number of published documents in the published data that contain a defined representation referring to a particular drug target. The ambiguity feature may include the number of published documents in the published data that contain one of the defined character representations referring to the drug target, compared to the total number of published documents in the published data that contain a defined character representation referring to a particular drug target. The ambiguity feature may include the frequency with which each character in one of the defined character representations referring to a drug target appears in the published data. More specifically, the sum of the frequencies of each character in a particular character representation, i.e., the frequency score of the representation as a whole, may be considered. Any appropriate metric using this overall frequency score, for example, the logarithm of this overall score, can be used. For example, a synonym or character representation that includes a less common character may be less ambiguous than a synonym consisting only of characters that are commonly found (or considered more commonly found) in the published data. Further ambiguity features for a particular character expression or synonym may be based on the number of defined character expressions for a drug target that contains a particular defined character expression. In other words, the ambiguity feature may be based on the number of nested synonyms (as defined above) associated with a particular synonym, i.e., the number of other gene synonyms that contain the particular gene synonym under consideration. Another ambiguity feature may be the probability that a published document in published data containing one of the defined character expressions other than the selected safe character expression also contains the selected safe character expression. In other words, the ambiguity feature may be a conditional probability of finding the gene synonym of interest in the published data of a particular published document, assuming that one of the (other) gene synonyms (as defined above) of the same gene symbol appears in the text.A further ambiguity feature could be the inverse probability described above, i.e., the probability that a published document in the published data containing the selected character expression (i.e., the synonym under consideration) also contains another defined character expression for that drug target. In other words, the ambiguity feature could be the conditional probability of finding another (other) gene synonym for the same gene symbol in the published data of a particular published document, assuming that the gene synonym of interest appears in the text. As a final example, the ambiguity feature might be based on whether the character expression under consideration is an acceptable character expression for a particular drug target, for example, whether the gene synonym under consideration is a gene symbol.

[0097] The method of the present invention allows each drug target, such as a human gene, to be clearly associated with or linked to a subset of published documents into which the published data has been incorporated, using labeled character expressions, i.e., safe or unsafe labeling depending on the ambiguity of the associated expression. To do this, a co-citation network-based approach can be used. That is, citations of published documents are used to more accurately determine whether a mention of a character expression of a particular drug target in the published data actually means that the particular drug target is linked to the published document (or whether the character expression is mentioned in a different context and therefore does not actually refer to the drug target). Specifically, using the co-citation approach (described in more detail below), "false positives" can be reduced or eliminated from the retrieved published data, i.e., published documents that mention the character expressions (gene synonyms) defined in the published data, which indicate that the published document may be associated with a drug target (gene) related to the defined character expression, but is not actually associated with or linked to the drug target. This approach may be seen as based on the assumption that published documents containing "false positives" tend to belong to different communities of publications related to different research areas than published documents containing "true positives," i.e., published documents containing defined characterizations in text that actually refer to the drug target in question. In this way, the identified communities of published documents can be determined to be linked to or not linked to the gene in question (as a whole).

[0098] To analyze specific citations made by different published documents, the published data, which is incorporated for at least a portion of the published documents, may include citation data showing that one published document cites one or more other published documents from multiple published documents. This method may involve identifying so-called “co-citations” within the published data. Co-citation can be considered when two published documents are both cited in a third document. That is, if “Published Document A” and “Published Document B” are in the bibliography of “Published Document C”, then there is a co-citation between “Published Document A” and “Published Document B”. Therefore, the process of searching the published data may involve using the incorporated citation data to identify pairs of (first and second) published documents that are cited by the same (third) published document. In particular, since it is desirable to obtain communities of publications, each referring to a specific drug target, this process of searching for pairs of published documents may involve searching for pairs of published documents, each containing at least one of the character expressions (gene synonyms) defined as referring to one of the drug targets (genes).

[0099] A co-citation network can be obtained using identified co-citations, i.e., pairs of published documents. For each identified pair of published documents, a co-citation value can be determined that represents the number of (different) published documents that cite both documents in the pair. That is, a weighted co-citation graph can be obtained in which the edge weights represent the frequency of two publications being cited (co-cited) simultaneously by a third publication. If two publications are repeatedly co-cited, this is considered to strongly suggest that both belong to the same field of research. This means that both publications in the co-cited pair are assumed to be true positives or false positives.

[0100] Once a co-citation network is established, pairs of published documents are assigned to different communities of published documents. Each community contains published documents that include instances of character representations defined for a particular drug target. However, not all communities may actually contain published documents related to a specific drug target; that is, some communities may consist of documents in which instances of character representations are in a different context than that of a particular drug target.

[0101] Therefore, this method may involve assigning pairs of published documents to one of several communities of published documents, based on the determined co-citation values ​​and the published documents that cite those pairs of published documents. This can be done automatically using appropriate community detection techniques. For example, assigning pairs of published documents to one of the communities may involve applying a (fast) powerful optimization algorithm.

[0102] Once a large number of communities of published documents have been obtained, it is necessary to distinguish the identified communities from one another. In particular, this method may involve determining, for each of the multiple communities of published documents, whether that community should be associated with one of the drug targets. To do this, the relative "safety" of the character representations (determined as described above) present in the published documents of a particular community can be used. This may involve determining or identifying which of the defined character representations that refer to a particular drug target are present in the published data of each published document within that particular community. The determination of how many safe character representations there are in a community can be used to determine whether that community is associated with the relevant drug target. For example, the percentage of published documents in the community under consideration that contain at least one safe character representation in their published data can be determined. Then, if the determined percentage is greater than a predetermined threshold percentage, it may be decided to associate that community with the drug target of interest. Alternatively, one or more communities with the highest percentage of safe character representations can be considered associated with the relevant drug target.

[0103] A potential problem with the co-citation approach described above is that the published data for some published documents may not include citation data, i.e., details of the citations made by specific published documents. This can be particularly problematic when published data from open access publications is incorporated, as citation data is often unavailable from such sources.

[0104] Therefore, for each published document whose published data does not include citation data, the method may include determining whether to assign the published document to one of the communities associated with one of the drug targets, based on the published data, in particular on one or more defined character representations that refer to the drug target within the published data. For example, if the published data of a published document includes at least one instance of a safe character representation, it may be decided to assign the published document to one of the communities associated with the relevant drug target. On the other hand, if the published data of a published document does not include a safe character representation, the decision on whether to assign the published document to one of the communities associated with the relevant drug target can be performed using a machine learning algorithm, such as a positive unlabeled learning technique. The machine learning algorithm can be one of the following machine learning classifiers: a logistic regression classifier, an extra-tree classifier, a Gaussian process classifier, a k-nearest neighbor classifier, a ridge classifier, a random forest classifier, and a support vector machine classifier. That is, using a positive unlabeled bagging approach, multiple classifiers can be trained to associate fragmented publications (without citation data) with previously computed co-citation network components using words / representations contained in the published data, such as titles and summaries.

[0105] The above process of searching the incorporated published data makes it possible to accurately indicate which published documents within the literature are associated with a specific drug target (such as a gene). This enables more accurate and reliable analysis of long-term publication trends for one or more drug targets, such as the long-term publication rate of a particular gene.

[0106] Accordingly, according to the present invention, the next step of the method includes determining expected and actual publication parameters for each drug target of interest based on the retrieved publication data. In particular, expected publication parameters are determined based on publication data from past publication documents. Specifically, for example, the past publication trends of a particular gene are calculated using past publication documents, and these past publication trends are then used, for example, by extrapolation to determine or predict expected publication parameters. As an exemplary example, the publication trends of a given gene in each of several consecutive years can be calculated using past publication data, for example, using publication dates related to past publication documents in the publication data, and these calculated (past) publication trends can be used to predict the current publication trends of that given gene. The determination of expected publication parameters may be performed using a machine learning algorithm, for example, a recurrent neural network algorithm, trained using publication data retrieved from past publication documents. Actual publication parameters are determined based on publication data from current publication documents.

[0107] Expected and actual publication parameters may be measures or indicators of any one or more aspects of publication trends related to a particular drug target. For example, expected and actual publication parameters may be the expected number of published documents and the actual number of published documents, e.g., the number of published documents in a given year. Alternatively, or in addition to these, expected and actual publication parameters may include the expected and actual number of clinical trials related to a particular drug target under consideration, the expected and actual number of review published documents related to a particular drug target, and the expected and actual number of published documents associated with a defined company size. In any case, relevant information must be available in the incorporated publication data in order to determine the relevant parameters. For example, publication data for some published documents may indicate whether the published document is related to a large or medium-sized pharmaceutical company. As an illustrative example, if a manuscript whose authors belong to a large pharmaceutical company cites other publications, these citations may be classified as “large pharmaceutical company” citations. Conversely, publications citing this manuscript whose authors belong to a large pharmaceutical company may not be classified as “large pharmaceutical company” citations.

[0108] According to the present invention, in order to detect new trends in the literature, this method includes evaluating each of the target drug targets for selection based on its actual published parameters compared to its expected published parameters. For example, the evaluation of drug targets for selection may include ranking the drug targets in a (prioritized) list based on a comparison of their actual published parameters with their expected published parameters. In particular, if there is a (large) difference between each actual published parameter and its expected published parameter, the drug target may be considered potentially interesting for selection. This is because it may mean that interest in the drug target has shifted gradually compared to what would be expected according to past published data.

[0109] Typically, the described methods are found to generate accurate predictions of publication dynamics; that is, actual publication parameters generally coincide with predicted publication parameters. However, for small subsets of drug targets, such as genes, the actual number of publications or citations may be significantly higher than predicted. When the actual number of publications or citations exceeds predictions, this can be interpreted as a significant shift in publication trends that cannot be explained solely by the publication history of the gene of interest, suggesting, for example, that a meaningful discovery in that field has been made recently. The term "trendiness" can be defined as the probability of a multiplier change between the predicted and actual numbers of publications and citations for a particular gene. This metric can be used to identify the "most trending" genes in the academic community (using all publications) or the pharmaceutical industry (using publications from pharmaceutical companies).

[0110] This method may include using drug target evaluation to inform the selection of at least one drug target to be used in a drug discovery project, for example, based on a ranked list of the most trending genes. In particular, this method may include designing a drug discovery project by selecting at least one drug target to be used in the drug discovery project based on evaluation. This method may include carrying out a drug discovery project using at least one drug target selected at least in part based on the above evaluation. Such a drug discovery project may include, for example, selecting compounds and testing them against at least one selected drug target to identify compounds that have potential therapeutic activity against a disease target. The method of this disclosure may include synthesizing at least one compound that has potential binding activity against the selected drug target.

[0111] When analyzing and selecting specific drug targets, it is also beneficial to consider the relationships between two different drug targets. In particular, it may be found that the target gene can cluster within an associated network. Therefore, it can be useful to gain insight into the relationships between one gene and others in the literature. This is because identifying one gene whose publication trends have changed significantly may mean that one or more other genes whose publication trends have changed in a manner of interest may be found.

[0112] In this regard, analyzing the publication trends of drug targets may involve determining target-target co-occurrence parameters between pairs of drug targets. Such parameters can be determined based on indicators from the retrieved publication data, specifically, which publication documents are associated with both drug targets in a pair, i.e., from the publication documents in which two different drug targets are associated. Each target-target co-occurrence parameter may indicate the number of publication documents in which both drug targets of the pair appear. The evaluation of drug targets for selection can then be carried out based on the determined target-target co-occurrence parameters.

[0113] Potential drug targets of interest to the pharmaceutical industry may be those potentially associated with specific diseases. The described method of searching published data to associate drug targets with publications is also applicable to associating specific diseases with publications. In this regard, the method may include searching the incorporated published data to provide an indicator for each published document of whether that document is associated with one or more diseases. Similar to the drug target method described above, this may include defining one or more character representations that refer to the disease for each disease and searching published data for each character representation of the disease.

[0114] As a non-limiting example, disease names and their synonyms can be obtained from the BioPortal's Medical Subject Headings (MeSH) ontology. The MeSH ontology contains 4818 different disease nodes at various levels of the ontology. For example, a dictionary for each disease can be created using preferred and alternative names. Then, using the corresponding techniques described above for genes, the disease can be deambiguated within published data (e.g., title, abstract, etc.).

[0115] Next, this method may involve determining target-disease co-occurrence parameters between each drug target and each disease. Such parameters can be determined based on indicators of which published documents from the retrieved published data are relevant to each drug target and each disease, with each target-disease co-occurrence parameter indicating the number of published documents in which one of the drug targets and one of the diseases appears. The evaluation of drug targets for selection can then be carried out based on the determined target-disease co-occurrence parameters.

[0116] To evaluate drug targets to select, it may be desirable to gain deeper insights into why publications of a particular drug target have changed, perhaps in relation to a specific disease. In this way, groups of publications referencing the gene of interest can be analyzed. For example, this method may involve applying a topic modeling algorithm to publication data of publication documents related to a drug target of interest to obtain one or more topics related to the drug target, and then evaluating drug targets for selection based on the obtained topics. A topic can be considered a set of similar words specific to a group of documents. A set of latent topics for each query can be generated using non-negative matrix factorization. In particular, topic modeling algorithms may include latent Dirichlet assignment algorithms and / or non-negative matrix factorization algorithms. Topic detection can also be used to determine errors in the association between one or more publication documents and drug targets, according to the search publication data based on one or more obtained topics, thus further improving the accuracy of drug target relevance in the literature.

[0117] Figure 1 summarizes the steps of the computer-based drug target selection method 10 according to the present invention. In step 101, published data is received, ingested, or downloaded from at least one published data source, such as an online database storing published documents, e.g., articles, journal papers, etc. Published documents include past and current published documents. The published data may include publication date, author name, title, abstract, keywords, citations, etc., related to the published document.

[0118] In step 102, the received published data is searched to provide an indicator for each published document regarding whether each published document is associated with one or more potential drug targets, such as genes. In particular, this may involve searching the published data for mentions or instances of one or more defined character expressions for each drug target. If one of these character expressions is mentioned in the published data of one of the published documents, it indicates that the published document may be associated with a particular potential drug target. Further steps can be performed to determine whether the published document is indeed associated with a potential drug target. For example, the relative "safety" of a character expression in the published data may be established (as described above) to indicate confidence that the character expression in the published data does indeed refer to the target potential drug target in question. Further steps can be performed to cluster the published documents into communities based on the searched published data to determine whether the cluster of published documents is indeed associated with or linked to the target drug target in question. Generally, the search of published data is used to establish groups of published documents linked to each of one or more potential drug targets.

[0119] In step 103, expected publication parameters for each potential drug target are determined based on retrieved publication data from (i.e., related to) past publication documents. Actual or true publication parameters for each potential drug target are also determined based on retrieved publication data from current publication documents. Publication parameters can be any appropriate parameters describing the publication dynamics (over time) for each potential drug target. For example, publication parameters could represent the number of publication documents per calendar year related to a particular drug target. Expected or predicted publication parameters can be determined by determining past publication trends for drug targets based on past publication data and extrapolating these trends to predict current or future publication trends.

[0120] In step 104, each potential drug target may be evaluated for selection based on its actual published parameters versus its expected published parameters. In particular, the difference between the expected and actual parameters of a particular potential drug target may indicate a change in assumptions about the drug target and may indicate interest in further investigation for selection as a drug target. The evaluation includes creating a target list of potential drug targets based on the above analysis (i.e., based on the difference between predicted and actual values, and possibly the confidence of the predictions). This is to prioritize potential drug targets for selection, for example, by considering their relevance to any disease or biological selection mechanism. This evaluation provides insights into the selection of drug targets for various applications, such as the design and execution of a specific drug discovery project.

[0121] The method of the present invention can be implemented on any suitable computing device by one or more functional units or modules implemented, for example, on one or more computer processors. Such functional units may be provided by suitable software running on any suitable computing board using conventional or customer processors and memory. One or more functional units may use a common computing board (for example, they may run on the same server) or separate boards, or one or both themselves may be distributed across multiple computing devices. Computer memory can store instructions for performing this method, and a processor can execute the stored instructions to perform the method.

[0122] Many modifications can be made to the above example without departing from the scope of the attached claims.

[0123] Below, we will describe some non-limiting specific examples of the computer-based drug target selection methods outlined above. [Examples]

[0124] [Example 1] The PubMed® baseline, released in December 2019, contains over 30 million publications, approximately 170 million citations from open-source data, approximately 9 million authors, and approximately 300 million MeSH annotations. PubMed® was converted into a graph database using the Neo4J graph database platform, allowing for efficient querying of relationships such as author names, references, and annotations. The resulting database included five distinct node types: Publications; Authors; Genes encoding human proteins; Human diseases; and Medical SubHeadings (MeSH) terms. Publication nodes had multiple attributes extracted from the PubMed® baseline: PubMed ID; Title; Abstract; Keywords; Authors; Affiliation; Publication Date; Journal; and Article Type (e.g., Article, Review, or Clinical Trial). An attribute aggregating affiliation data was also included to determine whether pharmaceutical companies were involved with the authors of the publications. There are five types of relationships (edges): cited (publisher to publication); published (author to publication); MeSH annotation (MESH term to publication); gene annotation (gene to publication); and disease annotation (disease to publication). In preparing this database, a disambiguation pipeline was implemented to clearly link symbols of human protein-coding genes and human diseases to individual publications.

[0125] Synonyms for human genes were collected from various sources (Ensembl, UniProt, HGCN, Entrez, OpenTargets) to sample publications that may refer to human gene names. Note that, on average, each human gene has approximately ten synonyms, and many of these synonyms are ambiguous (when considered in isolation from context). Over 30% of gene symbols have at least one indifference synonym, about 10% have different meanings in different contexts, at least one gene synonym exists in an English dictionary, and nearly 50% of gene symbols have nested synonyms. Combining these issues, nearly 60% of the 19082 gene symbols have at least one of these types of ambiguities. To determine which synonyms are potentially ambiguous, feature engineering was performed to obtain variables that characterize unsafe synonyms (e.g., longer gene names are less likely to be ambiguous). Next, we calculated the probability that a gene synonym is "unsafe" using a positive unlabeled bagging (PU) strategy with a random forest classifier equipped with manipulated features.

[0126] More specifically, 19,082 human genes encoding proteins, annotated by the HUGO Gene Nomenclature Committee (HGNC), were used. Gene synonyms identical to disease names in the Medical Subject Headings (MeSH) database were removed. This mainly occurs when genes are named after associated diseases, such as "Li-Fraumeni syndrome" as a gene synonym for gene TP53, or "Marfan syndrome" for "FBN1."

[0127] Genetic synonyms were categorized as either “safe” or “unsafe” using a modified version of positive unlabeled (PU) learning with bootstrap aggregation. PU learning is a form of semi-supervised learning that iteratively finds positive examples within unlabeled data. To construct a binary classifier that can distinguish unlabeled classes (U) into unsafe (P, positive) and safe (N, negative) classes, a set of features were manipulated, such as the frequency of character binding in genetic synonyms (e.g., “ZNF” is safer than “EDA” because the characters “Z” and “F” are less frequent than “E”, “D”, and “A” in the PubMed® corpus), or the probability of a genetic synonym given other genetic synonyms in the text (given “steroid-producing acute regulatory protein”, the probability of “STAR” is higher, but given “STAR”, the probability of “steroid-producing acute regulatory protein” is lower because “STAR” is more ambiguous).

[0128] PU learning was performed over five iterations using a random forest classifier. The pure positive class (unsafe) was constructed by combining gene synonyms from the English dictionary, gene synonyms with fewer than three characters, and indiscriminate gene synonyms. In an active learning manner, after each iteration, the top 1000 most likely unsafe examples that were misclassified were manually relabeled. For example, true positive unsafe synonyms such as gene families (e.g., "G protein-coupled receptor"), phenotypes (e.g., "Williams Buren syndrome"), and other biological entities (e.g., "cell surface antigen") were included in the true positive set for the next iteration. False positives such as "thymopoietin" and "tubulin alpha-1C chain" were incorporated into a new true negative class for the remaining iterations.

[0129] After 5 iterations, a genetic synonym was considered unsafe if: (i) it is included in an English dictionary; (ii) it is a word with fewer than 3 characters; (iii) it has a prediction score higher than 0.5 for a random forest classifier; and (iv) it is an indiscriminate genetic synonym.

[0130] To link all human genes to a subset of publications, a co-citation network and a machine learning-based disambiguation pipeline were implemented. Titles, summaries, and keywords of publications matching any synonyms were collected in Elasticsearch using regular expressions. Specifically, the Elasticsearch API search engine was used to retrieve the PubMed® ID of publications whose titles, summaries, or keywords contained synonyms for genes or diseases. These PubMed® IDs were later used to retrieve publication attributes from Neo4J using the Cypher language via a Python driver. Regular expressions were used to account for punctuation and character variations and to avoid ambiguity in nested names due to fuzzy matching (e.g., "ErbB-1", "erbB1", "ERBB1", "ErbB1").

[0131] To detect publication communities, a co-citation network, a weighted graph where edge weights represent the frequency with which two publications are simultaneously cited (co-cited) by a third publication, was used. Using iGraph's fast and powerful modulation algorithm, communities within the co-citation network were determined, and publication communities focused on the target gene were distinguished by detecting the presence of "safe gene synonyms" in the title and abstract. If the ratio of publications mentioning at least one safe synonym was higher than 0.1% compared to publications mentioning only unsafe synonyms, each publication within the community was labeled with the target gene symbol.

[0132] Finally, because only citations from open-access publications included in PubMed Central (PMC) were used, 46% of the publications were decoupled in the PubMed® cocitation graph. Decoupled publications that mentioned safe synonyms were automatically linked to the target gene symbol. The remaining decoupled publications were linked to the target gene using a PU approach bagging strategy with a binary logistic regression classifier based on words in the text corpus (keywords, titles, and summaries) of communities already linked to the target gene and discarded communities. All machine classifiers available in Scikit-Learn were used, but logistic regression was chosen due to its speed-to-accuracy ratio.

[0133] Each corpus was preprocessed in the following ways: (i) removal of non-alphanumeric characters; (ii) tokenization or splitting by whitespace; (iii) removal of stop words from NLTK (Natural Language Toolkit); (iv) transformation of small characters; (v) removal of tokens shorter than 3 characters; (vi) removal of tokens representing integers; and (vii) stemming (e.g., "disambiguated", "disambiguations", and "disambiguating" are transformed into "disambiguat"). Lists of tokens (uni, bi, tri, tetragram) with at least 2 counts and a frequency of less than 0.6 in the complete corpus were vectorized using TF IDF (Term Frequency - Inverse Document Frequency). If there were fewer than 1000 unlabeled publications in the training set for the gene of interest, an auxiliary negative class was generated to increase the number of negative examples in the training data. This auxiliary negative class consisted of a random sample of 1000 publications that mentioned genes different from the gene of interest.

[0134] To test the performance of the disambiguation method, the disambiguation results were compared to gene publication annotations from GeneRif (manually curated annotations), DISEASES (computer-annotated annotations), and UniProt (computer- and manually curated annotations). On average, the disambiguation method restores over 85% of all publications contained in these databases. Since annotations in both GeneRif and Uniprot do not necessarily include gene synonyms in their titles or abstracts, these publications are outside the scope of the described method. The disambiguation results show an average accuracy of 70% in UniProt, the only collection of disambiguated publications of a similar scale. Finally, the clarified gene publication annotations were incorporated into a graph database.

[0135] These vectors were fed into all machine learning classifiers available from the Python library sklearm (extra-tree classifier, Gaussian process classifier, K-nearest neighbor classifier, logistic regression classifier, ridge classifier, random forest classifier, and support vector machine). All classifiers were trained using hyperparameter tuning and three-fold cross-validation to avoid overfitting every 50 PU bagging iterations. The loss function was modified to account for class imbalance. The logistic regression (LOG) classifier was chosen as the deambiguation method, considering the balance between accuracy and speed.

[0136] The same procedure used for gene entity recognition was used for disease entity, co-citation network, and machine learning detection. The Medical Subject Headings (MeSH) ontology was downloaded by querying the Rest-API available in BioOntology. Each disease was a node in the ontology. Disease synonyms were obtained from the ontology's "Concept List Terms" field to collect preferred and alternative ways of representing the disease. Further synonyms for the disease were generated by reversing the order of synonyms with commas, such as from "Insipidus, Diabetes" to "Diabetes Insipidus".

[0137] Co-occurrence of genes and diseases was calculated using the co-occurrence of gene / disease tags in publications after ambiguity resolution, and normalized by the total number of publications presenting those tags. Mutual information metrics for gene-gene and gene-disease associations were also calculated.

[0138] All disease MeSH terms were associated with the lowest-level ancestor in the MeSH ontology under the Disease node. After calculating gene-disease co-occurrence, each gene was associated with the most frequent ancestral disease term.

[0139] To detect new trends in the literature, publication trends for specific human genes were collected from an unambiguous graph database. These time series include publications, clinical trials, reviews, and the number of publications by large and medium-sized pharmaceutical companies, as well as the number of citations of publications from the mentioned categories by calendar year.

[0140] For most genes, this model accurately predicts publication trends, but for a small subset of genes, the actual number of publications or citations is significantly higher than expected. Gene tendency can be considered the probability of observing a multiple change between the predicted and actual number of publications for that gene. For genes associated with a small number of publications, the prediction error is inevitably large. To correct this, five bins were generated based on initial publication numbers (percentiles 20, 40, 60, 80, and 100). The distribution of the multiple change between prediction and observed reality in each of the five bins was calculated using the Gaussian kernel density estimator available in Scikit-Learn (bandwidth = 0.1, remaining parameters at default values). The area under the obtained probability density function is equal to 1. The tendency is the right-tail region of the probability density function enclosed on the left by the observed multiple change. This provides an estimate of how extreme the multiple change for that gene was within a particular bin.

[0141] Using time series data from 1980 to 2013, we predicted gene-specific publication trends for each category from 2014 to 2019 using a recurrent neural network model with an encoder-decoder architecture where both the encoder and decoder consist of five hidden layers of Gated Recurrent Units (GRUs) and an attention layer precedes them. The model was implemented in Keras using the Tensorflow-GPU backend. The time series was rescaled before training using Min-Max normalization. The optimizer was RMSprop, and the loss was calculated as log error. 30% of the time series was saved for validation during training.

[0142] The input data was in both cumulative and differential formats. Multiple normalizations ("none", "minmax", "log", "standard", and combinations thereof) were used. Similar results were obtained with different normalizations, and minmax was ultimately selected. Multiple recurrent neural network (RNN) architectures were used in the form of encoder-decoders with different numbers of neurons (1, 5, 10, 20, 50) (GRU, LSTM). The models were compared to Mean Precision Scale Error (MASE), an unbiased method for comparing time series forecasting models by comparing how well each model performs compared to a simple model that repeats the last value. The 5-neuron GRU was selected because it is the most economical model with the smallest MASE.

[0143] To identify pharmaceutically interesting genes, normalized cross-information values ​​of genes and diseases included in published titles and abstracts were calculated. When obtaining gene-gene and gene-disease related networks, many trending genes cluster together to form trending pathways. Using enhanced gene ontology (GO) terminology for biological processes, common pathways among the top 100 most trending genes are revealed. Among the most abundant GO terms in both academia and the pharmaceutical industry are the execution phases of T-cell costimulation, necroptosis, and pyroptosis. These biological processes are rich in trending genes, which may reflect that these areas of study are generating the most innovation and promise in current biomedical research.

[0144] The next step after detecting gene trends was to understand why those genes were trending and to organize potential errors in disambiguation. For this purpose, a topic detection pipeline was implemented as an automated, high-speed detection tool for studying groups of publications that mention the target genes. In this context, topic modeling algorithms were used. A topic is a set of similar words specific to a group of documents. Two different topic detection algorithms were used: Latent Dirichlet Allocation (LDA) and Non-Negative Matrix Factorisation (NMF). Both algorithms factorize a non-negative matrix 'A' of size NxM, where N is the number of publications and M is the dimension of the TF IDF vector obtained for named entity recognition, into a non-negative factor matrix W of size NxK and a matrix H of size KxM, where WxH is an approximation of matrix A. Matrix W contains the strength of association of specific publications belonging to a latent topic, and H contains the strength of association between the latent topic and a given n grams. Using Scikit-Learn implementations of both algorithms, we generated a user-defined "K" topics (with a tolerance of 1e-12) using default parameters until convergence was achieved. The topic timeline was obtained by calculating the mean and standard deviation of the probabilities of all published topics mentioning the target gene for each calendar year.

[0145] A review recommendation system can also be designed to accelerate the screening of publications that cover most of the information within the network. On average, there are 2.9 reviews that cite publications that refer to at least one gene name. The goal was to minimize read time and maximize information within the gene subnetwork. This algorithm aggregates both topic and network information from the citation subgraph of publications that mention the gene of interest and retrieves the most query-centric reviews. Topic information is obtained from latent topics obtained from a topic detection algorithm. Topic probabilities of publications and aggregated PageRank scores of the citation network were used. Network information was obtained from the PageRank scores of the subgraphs. Users can select the interval between reviews (R) they want to read, between 2 and 3 or 3 and 50. Next, three matrices are defined for each group of publications: (i) a binary sparse matrix of size NxR containing the N publications and R reviews that constitute the citation adjacency network; (ii) an Nx1 weight matrix that constitutes the PageRank score; and (iii) an NxK matrix containing the topic probabilities of N publications and K user-defined topics. The score for each review is defined as the sum of the PageRank scores of its references, and the score for a combination of reviews is defined as the sum of the rows obtained by multiplying an indexed NxR matrix by an Nx1PageRank vector, plus the sum of the resulting vectors. The results were then normalized by the total maximum score, which was defined as a hypothetical review that cited all gene publications. In this method, the best review is one that cites the publication with the highest PageRank score. Finally, to minimize the number of reviews, a combination was found that maximized the cumulative PageRank score while simultaneously minimizing overlap in combined citations.

[0146] In this way, you can obtain a small number of reviews that cover the major topics and publications in this field. Using this recommendation system, you can select the best subset of reviews to assess why genes are trending.

[0147] The total number of publications per gene is generally quite predictable. However, some genes may have significantly more publications than expected; this means that recent breakthroughs have occurred that cannot be explained by publication dynamics alone. Trend metrics can identify new targets from the literature for rapid profiling at the genome scale. Combining trends with gene-disease associations prioritizes potential drug targets: emerging genes that are disease-related but still included in pharmaceutical publications are worth investigating as potential targets. Trending genes are typically observed to cluster along the same biological pathways.

[0148] In summary, the described exemplary method involves downloading published data from a PubMed® baseline and creating a graph database using the retrieved information. A comprehensive collection of human coding gene names and synonyms is obtained, and this method includes automated determination of potentially ambiguous (unsafe) gene names. The graph database is annotated with clear gene symbols by combining a co-citation network topology and a binary classifier. This method uses a recurrent neural network to predict publication trends for each gene. If a gene has significantly more publications or citations than predicted by the model, that gene is considered "trendy." This method optionally includes automated topic detection of the publication collection, and this algorithm was used to quantify the evolution of topics over time in trending gene publications. Optionally, a review recommendation system can be implemented that recommends the most efficient set of reviews for searching the literature using information from the citation network and topic detection.

[0149] Figure 2 shows examples of graph databases created for a specific gene using different techniques or processes described above. In particular, Figure 2(a) shows a citation network of a subset of published documents from PubMed® that reference any of the gene synonyms for the gene symbol LRWD1, including ORCA. Nodes represent published documents, and the size of the node represents the number of citations. Edges indicate citations between documents, including the direction of citation. Figure 2(b) shows a co-citation network of the same subset of published documents as in Figure 2(a). The thickness of the edges represents the number of times a pair of documents have been co-cited. Figure 2(c) shows various communities of published documents obtained using iGraph's fast and powerful algorithm as described above. Each community is associated with a different topic from which it was obtained. For example, there is the so-called "Orca" community 201, the "Orca Plant" cluster or community 202, the "LRWD1 in Drosophila" community 203, and the "LRWD1 in Heterochromatin" community 204. Figure 2(d) shows the number of safe synonyms in the title or abstract of each published document within the same co-citation network. Figure 2(e) shows a citation network to which review documents have been added to indicate citations of published documents by review documents. Figure 2(f) shows the review information defined by the recommendation system on a scale from 0 to 1.

[0150] Figure 3 shows trends across various genes and the detection of gene-gene-disease co-occurrences. In particular, Figure 3(a) shows a logarithmic scatter plot of predicted publication numbers for various genes against the actual number of publications in 2019. Similarly, Figures 3(b), 3(c), and 3(d) show, for various genes, the predicted number of review documents, citations, and citations from "major" pharmaceutical companies for 2019, against the actual number of review documents, citations, and citations from "major" pharmaceutical companies, respectively. For example, genes with a higher number of actual publications than predicted values ​​(i.e., their nodes lie on a line showing a log-linear relationship) are in a trending state and are considered interesting as potential drug targets.

[0151] Figure 4 shows the log2(predicted / actual) trends of various genes associated with different disease groups (by MeSH parent category). Specifically, Figure 4(a) shows the average trends of publication, review, citation, and citation from review for all (general) published documents, and Figure 4(b) shows the average trends of publication, review, citation, and citation from review from large and medium-sized pharmaceutical companies.

[0152] Figure 5 shows the gene-gene-disease co-occurrence network of the first neighboring gene of CD274. Disease and gene nodes are given defined names, and the size of the gene node represents its "tendency" according to a defined metric. Edges indicate gene-disease and gene-gene associations, and the width of the edges reflects the number of co-occurrences in each case.

[0153] Figure 6 shows a topic timeline relating to the number of publications mentioning various genes in question, i.e., the evolution of topics related to some of the most trending genes is investigated. Specifically, Figures 6(a), 6(b), and 6(c) show topic timelines of publications mentioning either immune checkpoint inhibitors, necroptosis, or pyroptosis pathway genes, respectively. In each case, timelines for four topics are shown. The four potential topics were obtained using non-negative factorization of all gene-annotated publications after disambiguation. All timelines show topics that have risen since 2013, illustrating why these genes have become “trending.”

[0154] Referring to Figure 6(a), for immune checkpoint inhibitors (CD274, PDCD1, TGIT, and CTLA4), the topic timeline suggests a rapid decline in the likelihood of publications discussing the biological roles of these immune checkpoint inhibitors since 2010 (shown in the topic timeline labeled 601), which coincides with a significant increase in topics (labeled 602) discussing cancer therapies and monoclonal antibodies targeting these four different transmembrane immunoglobulins. In this way, the topic detection pipeline can capture the evolution of research from biological explanation to clinical application.

[0155] Referring to Figure 6(b), the topic timeline for members of the necroptosis pathway (RIPK1, RIPK3, and MLKL) suggests that over the past decade, there has been a decrease in publications discussing these genes in relation to apoptosis (indicated by the topic timeline labeled 611), and a dominance of publications discussing newly discovered forms of cell death, the necroptosis pathway, and translational medicine aspects of this pathway, as suggested in terms such as mouse, therapeutic and active or cancer (indicated by the topic timeline labeled 612).

[0156] Referring to Figure 6(c), the topic timeline for members of the pyroptosis pathway (CGAS, TMEM173, GSDMA, and GSDMAD) shows a rapid increase since 2013 in publications discussing therapeutic opportunities in cancer immunotherapy with TMEM173 agonists (indicated by the topic timeline labeled 621), while the remaining topics appear to include information on the biochemistry and biological roles of the genes.

[0157] The following are some case studies illustrating the methods described above.

[0158] <Immune checkpoint inhibitors: CTLA4, CD274, PDCD1, TIGIT> CTLA4, PDCD1 (PD-1), CD274 (PD-L1), and TIGIT were among the most trending genes in academia and the pharmaceutical industry in 2019. The CTLA4, PDCD1, CD274, and TIGIT genes encode four different transmembrane immunoglobulins that act as co-inhibitory receptors: checkpoints or "breakpoints" in the adaptive immune response that prevent T cells from performing their function. CTLA4 competes with the similar CD28 for CD80 and CD86 to prevent premature activation of T cells. The PDCD1-CD274 interaction counteracts positive signals that may have already activated effector T cells. TIGIT interacts with CD155 to downregulate natural killer cells and T lymphocytes. Cancer cells attempt to disrupt these checkpoints, and currently there are seven FDA-approved monoclonal antibodies targeting three proteins (CTLA4: ipilimumab, PDCD1: nivolumab, pembrolizumab, cemiprimab, CD274: atezolizumab, avelumab) and several candidates targeting TIGIT (BGB-A1217, OMP-313M32, MTIG7192A, AB154).

[0159] <Neurodegeneration: TREM2 and C9orf72> Recent discoveries are revolutionizing our understanding of neurodegenerative diseases. C9orf72 encodes a guanine nucleotide exchange factor involved in endosomal transport and autophagy. Expansion of hexanucleotide repeats in the promoter or intronic region of C9orf72 is a major cause of both sporadic and familial forms of both amyotrophic lateral sclerosis (ALS) and frontotemporal dementia. Antisense oligonucleotides have been used to interfere with the transcription of C9orf72 or the CRISPR-Cas9 system, targeting GGGGCC repeats in DNA or RNA.

[0160] The TREM2 gene encodes a transmembrane immunoglobulin receptor expressed in macrophages, osteoclasts, dendritic cells, and brain microglia. TREM2 variants are thought to be associated with Nasu Hakora disease, late-onset Alzheimer's disease, frontotemporal dementia, amyotrophic lateral sclerosis, and Parkinson's disease. TREM2 activates a pathway via TYROBP / DAP12 that promotes inflammation and facilitates the phagocytosis of cellular waste, apoptotic cell debris, and pathogens. Currently, two independent groups are producing anti-TREM2 antibodies that stimulate microglia to remove amyloid plaques. Furthermore, an mAb produced by Alenco, one of these groups, in collaboration with AbbVie, is in Phase I clinical trials.

[0161] <cGAS-STINGによるDNAセンシング:cGAS,TMEM173,GSDMD,GSDMA> The cytosolic nucleic acid sensing pathway triggers pyroptosis, a type of lytic pro-inflammatory cell death involved in antiviral, antibacterial, and anticancer responses. cGAS is a nucleotidyltransferase that catalyzes the production of cyclic GMP-AMP (cGAMP) upon recognition of double-stranded DNA. TMEM173 (STING) binds to cGAMP, promoting the activation of both TBK1 and IRF3 and increasing the transcription of genes encoding type I interferons. GSDMA and GSDMD are intramembrane pore-forming effector proteins that release pro-inflammatory interleukins such as IL-1β and IL-18. The cGAS-STING pathway is associated with several autoimmune and chronic inflammatory diseases, including non-alcoholic fatty liver disease, systemic lupus erythematosus, vascular and pulmonary syndrome, macular degeneration, Bloom syndrome, Eicardi-Goutier syndrome, cancer, DNA damage, and neurodegeneration. Currently, clinical trials are underway for TMEM173 and GSDMD, but no trials have been reported for GSDMA or cGAS.

[0162] <Necroptosis: RIPK1, RIPK3, and MLKL> RIPK1, RIPK3, and MLKL form part of the tumor necrosis factor-induced necroptosis pathway. This pathway is associated with several conditions, including systemic inflammatory response syndromes, ulcerative colitis, psoriasis, rheumatoid arthritis, neurodegenerative diseases, and even cancer. TNFR1, FasL, TRAIL, and TLR all activate RIPK1, determining cell fate (inflammation, apoptosis, or necrosis). When caspase-8 is inhibited, RIPK1 and RIPK3 form necrosomes, which then phosphorylate MLKL. MLKL forms homotrimers, migrates to the cell membrane, binds to highly phosphorylated inositol phosphate, creates pores in the membrane, and disrupts cellular integrity. The discovery of RIPK1 dates back to 1995. Since then, four inhibitor programs have progressed through Phase II safety trials in humans. The first publication mentioning MLKL is recent, and despite the lack of kinase activity, pharmaceutical companies have cited that publication more than 60 times since 2013. Although clinical trials have not yet been conducted, at least three different chemoinhibitors are known.

[0163] <Mechanobiology: YAP1 / WWTR1, PIEZO1 and PIEZO2> Cells use mechanical signals from the environment to guide behaviors such as proliferation and migration. Forces act as signals transmitted to the nucleus, where they regulate gene expression. Mechanical forces are important regulators of organ and tissue homeostasis, morphogenesis, and regeneration, and are a crucial aspect of diseases such as cancer, metastasis, fibrosis, and cardiac hypertrophy. YAP1 / WWTR1 (TAZ) is a transcriptional coactivator and mechanotransducer. YAP / TAZ is overactivated in cancer, and its inhibition reduces atherosclerosis and fibrosis, causes pulmonary hypertension, and is necessary for intestinal epithelial regeneration. PIEZO1 and PIEZO2 are two mechanosensitive cation channels that play important roles in cell number regulation and migration, hearing, nerve and vascular development, somatosensory function, and proprioception. Piezo channels have recently been thought to be associated with several pathological conditions, including arthrodesis, apnea, congenital lymphangiopathy, hyperalgesia, malaria, pancreatitis, psoriasis, Gordon syndrome, Marden-Walker syndrome, and distal arthrodesis type 5. The discovery of mechanotransduction signaling pathways has attracted considerable attention in recent years and may open the door to new therapeutic strategies for treating these diseases.

Claims

1. A method for selecting drug targets using a computer, The process of retrieving published data related to multiple published documents, including past and current published documents, from at least one published data source; A step of searching the published data to provide an indicator of whether each of the published documents is related to one or more drug targets; A step of determining expected publicly available parameters for each of the one or more drug targets based on the retrieved publicly available data from the past publicly available documents, and determining actual publicly available parameters for each of the one or more drug targets based on the retrieved publicly available data from the current publicly available documents; and The step of evaluating each of the one or more drug targets for selection based on the actual published parameters against the expected published parameters; For each drug target, the expected publication parameter and the actual publication parameter are, respectively, one or more parameters selected from the group consisting of the expected number and actual number of published documents related to a particular drug target, the expected and actual number of clinical trials related to a particular drug target, the expected and actual number of review published documents related to a particular drug target, and the expected and actual number of published documents associated with a defined company size, in the method.

2. The method according to claim 1, comprising defining one or more character representations that refer to the drug target for each drug target, wherein the step of searching for the published data comprises searching for the published data for each of the one or more character representations for each drug target.

3. The method of claim 2, comprising the step of classifying each of the one or more character representations for each drug target as either a safe character representation or an unsafe character representation, wherein the classification is based on the possibility that instances of the character representation in the published data refer to the drug target, and if the retrieved published data from one of the published documents contains a safe character representation, the published document is determined to be related to the drug target.

4. One or more character representations have user-defined insecure features that indicate the corresponding character representation is unsafe. The method according to claim 2 or 3, wherein a character representation in the retrieved public data that exhibits one or more of the unsafe characteristics of the character representation is classified as an unsafe character representation.

5. The ambiguity feature of one or more character representations is defined such that the ambiguity score is attributed to one or more of the aforementioned character representations. The method according to any one of claims 2 to 4, wherein each of the character representations is classified as either a safe or unsafe character representation based on the correspondingly returned ambiguity score.

6. The method according to claim 5, comprising applying a machine learning algorithm to attribute the ambiguity score to each of the one or more character representations based on the ambiguity features of the one or more character representations, wherein the machine learning algorithm uses the unsafe features of the one or more character representations to attribute the ambiguity score to each of the one or more character representations.

7. The method according to claim 6, wherein after each iteration of the machine learning algorithm, a subset of the returned ambiguity scores is examined by the user to determine whether to manually modify any of the subsets of returned ambiguity scores, the subsets corresponding to a predetermined number of the character representations having the highest returned ambiguity score.

8. The published data for at least a portion of the published documents includes citation data indicating that one published document made a citation of one or more other published documents from the plurality of published documents, The method according to any one of claims 1 to 7, wherein the step of searching the published data includes using the cited data to identify pairs of published documents cited by the same published document.

9. The method according to claim 8, comprising determining a co-citation value for each identification pair of published documents, which represents the number of published documents that cite both of the pair of published documents.

10. The method according to claim 9, comprising assigning a pair of published documents to one of a plurality of communities of published documents based on their determined co-citation values ​​and the published documents that cite the pair of published documents.

11. The method according to claim 10, comprising defining one or more character representations that refer to the drug target for each drug target, the step of searching the published data comprising searching the published data for the one or more character representations for each drug target, and determining whether to associate each of the plurality of communities of the published document with one of the drug targets, the determination comprising determining which of the defined character representations that refer to the drug target is present in the published data of each of the published documents in the community.

12. The method of claim 11, as dependent on claim 3, comprising defining one or more character representations that refer to the drug target for each drug target, the step of searching the published data comprising searching the published data for the one or more character representations for each drug target, and the decision to associate the community with one of the drug targets comprising determining the proportion of the published documents in the community that contain at least one safe character representation in the published data.

13. The method according to any one of claims 8 to 12 as dependent on claim 2, comprising defining one or more character representations that refer to the drug target for each drug target, the step of searching the published data comprising searching the published data for the one or more character representations for each drug target, and the step of searching the pair of published documents comprising searching for a pair of published documents each containing at least one of the character representations defined as referring to one of the drug targets.

14. The published data for at least a portion of the aforementioned published documents does not include cited data, Furthermore, with respect to each of the published documents, the method includes determining whether to assign the published document to one of the communities associated with one of the drug targets, based on the published data, and in particular on one or more of the defined character representations that refer to the drug target within the published data. The method according to any one of claims 10 to 13, wherein if the published data of the published document includes at least one instance of a safe character representation, it is decided to assign the published document to one of the communities associated with one of the drug targets.

15. The method according to any one of claims 1 to 14, wherein determining the expected publication parameters includes using a machine learning algorithm trained with the retrieved publication data from the past publication documents.

16. This includes determining the target-target co-occurrence parameters between the pair of drug targets, The aforementioned target-target co-occurrence parameter is determined based on an indicator from the retrieved published data of which published documents are associated with both drug targets of the pair. Each target-target co-occurrence parameter indicates the number of published documents in which both of the paired drug targets appear. The method according to any one of claims 1 to 15, comprising evaluating one or more drug targets for selection based on the determined target-target co-occurrence parameters.

17. The method according to any one of claims 1 to 16, comprising searching the published data and providing an indicator for each of the published documents whether each of the published documents relates to one or more diseases, defining for each disease one or more character representations as referring to that disease, and the searching of the published data comprising searching the published data for each of the one or more character representations for each disease.

18. This includes determining target-disease co-occurrence parameters between each of the drug targets and each of the diseases, The aforementioned target-disease co-occurrence parameters are determined based on indicators from the retrieved published data of which published documents are associated with each drug target and each disease. Each target-disease co-occurrence parameter indicates the number of published documents in which one of the drug targets and one of the diseases appear. The method according to claim 17, comprising evaluating one or more drug targets for selection based on the determined target-disease co-occurrence parameters.

19. The method according to any one of claims 1 to 18, comprising: applying a topic modeling algorithm to the published data for the published documents relating to each of the drug targets to obtain one or more topics relating to each drug target; evaluating the one or more drug targets for selection based on the obtained one or more topics; determining, for each drug target, an error in the relationship between the one or more published documents and the drug target based on the obtained one or more topics; and determining, for each drug target, an error in the relationship between the one or more published documents and the drug target based on the obtained one or more topics.

20. The method according to any one of claims 1 to 19, wherein the published data includes a publication date for each of the plurality of published documents, and the publication date defines whether each of the published documents is a past published document or a current published document.

21. The method according to any one of claims 1 to 20, comprising using the evaluation of one or more drug targets to inform the selection of at least one of the drug targets for use in a drug discovery project, and designing the drug discovery project by selecting at least one of the drug targets for use in the drug discovery project based on the evaluation.

22. The method according to claim 21, comprising commencing the drug discovery project using the at least one selected drug target, wherein commencing the drug discovery project comprises selecting and testing compounds against the at least one selected drug target.

23. A computer device for drug target selection, Incorporate published data related to multiple published documents, including past and current published documents, from at least one published data source; The aforementioned published data is searched, and for each of the aforementioned published documents, an indicator is provided of whether each of the aforementioned published documents is related to one or more drug targets; Based on the retrieved public data from the aforementioned past published documents, predictable public parameters are determined for each of the one or more drug targets; and based on the retrieved public data from the aforementioned current published documents, actual public parameters are determined for each of the one or more drug targets; and The system is configured to evaluate each of the one or more drug targets for selection based on the actual published parameters relative to the expected published parameters; For each drug target, the expected published parameters and the actual published parameters are as follows: Each of these, the expected number of published documents related to a specific drug target and the actual number of published documents. The number of mentations, the expected and actual number of clinical trials related to a specific drug target, and the specific drug target The expected and actual number of review publication documents related to this, as well as defined company regulations One of the groups consisting of the expected and actual number of publicly available documents associated with the model is selected. A computer device with more than one parameter.

Citation Information

Patent Citations

  • Inter-document relation analyzing device, and program and method of the same

    JP2011076254A

  • System for predicting efficacy of targeted drugs to treat disease - Patent Application 20070122997

    JP2019522256A

  • Determining drug effectiveness ranking for a patient using machine learning

    US20200227176A1

  • Blood pressure measurement device

    WO2020137479A1