Data processing system and method for redeveloping existing drug

The data processing system uses machine learning to analyze patient records, grouping patients into clusters to efficiently identify new drug indications, addressing the limitations of existing methods in drug redevelopment by reducing computational time and improving accuracy.

JP2025108516AActive Publication Date: 2025-07-23SANOFI SA(FR)
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025064018
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-06-09
Filing Date
2025-04-09
Publication Date
2025-07-23
Estimated Expiration
2040-12-09

AI Technical Summary

Technical Problem

Existing methods for drug redevelopment, such as network-based and text-mining approaches, face challenges in accurately identifying new indications for existing drugs due to incomplete data, ambiguous language, and limited accuracy, making it difficult to distinguish positive and negative associations and verify biological relationships.

Method used

A data processing system using machine learning techniques, including unsupervised clustering and dimensionality reduction, to analyze large medical records of patients, identify patient characteristics, and group them into distinct clusters based on relevant criteria, thereby identifying new indications for existing drugs.

Benefits of technology

This approach reduces computational time and dependence on human input, enabling the discovery of new drug indications that conventional methods miss, by accurately grouping patients and identifying relevant patient characteristics for drug redevelopment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025108516000001_ABST
    Figure 2025108516000001_ABST
Patent Text Reader

Abstract

To provide a data processing system for redeveloping an existing drug.SOLUTION: One or more operations include: receiving data representing medical records of a plurality of patients; selecting a set of patients; determining a plurality of patient characteristics of the set of patients; grouping, in accordance with the plurality of patient characteristics, the set of patients to generate a plurality of distinct groups, each of the distinct groups including at least one patient of the set of patients; selecting, on the basis of one or more group selection criteria, a set of distinct groups; and identifying one or more relevant patient characteristics by analyzing each distinct group of the set of distinct groups.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Claims of Priority This application claims the benefit of European Patent Application No. 20315299.6, filed on Jun. 9, 2020, and U.S. Provisional Patent Application No. 62 / 945,814, filed on Dec. 9, 2019. The entire contents of the above are hereby incorporated by reference into this specification.

[0002] The present disclosure generally relates to data processing systems and methods for redeveloping existing drugs.

Background Art

[0003] The redevelopment (drug repurposing) (sometimes referred to as repositioning) of clinical existing drugs can be a relatively low-cost and can provide high efficiency, and may refer to a strategy for drug discovery. The redevelopment of existing drugs usually involves analyzing whether a drug approved to treat one type of medical condition (e.g., disease) can be used to treat other types of medical conditions (e.g., common diseases and / or rare diseases). There are several therapeutic areas that show high potential for the redevelopment of existing drugs, including oncology, immunology, infectious diseases, and orphan diseases.

Summary of the Invention

Means for Solving the Problems

[0004] In at least one aspect of the present disclosure, a data processing system is provided. The data processing system includes a computer-readable memory including computer-executable instructions, and at least one processor configured to execute executable logic including the computer-executable instructions and at least one machine learning model. When the at least one processor executes the computer-executable instructions, the at least one processor is configured to perform one or more operational methods. The one or more operational methods include receiving data representing medical records of a plurality of patients. The one or more operational methods include, based on the medical records: determining at least one target signaling pathway related to a drug; and selecting a set of patients by determining one or more indicators based on one or more factors corresponding to a diagnosis associated with the target signaling pathway. The one or more operational methods include determining a plurality of patient characteristics of the set of patients, each patient of the set of patients exhibiting at least one of the plurality of patient characteristics. The one or more operational methods include grouping the set of patients to generate a plurality of distinct groups according to the plurality of patient characteristics using the machine learning model, each of the distinct groups including at least one patient of the set of patients. The one or more operational methods include selecting a set of distinct groups of the plurality of distinct groups based on one or more group selection criteria. The one or more operational methods include identifying one or more associated patient characteristics by analyzing each distinct group of the set of distinct groups.

[0005] The machine learning model can be trained to group a set of patients using one or more unsupervised clustering techniques. The one or more unsupervised clustering techniques can include a bisecting k-means clustering technique. Grouping the set of patients can include performing multiple correspondence analysis to reduce the dimensionality of the plurality of patient characteristics.

[0006] Selecting a set of distinct groups can include determining, for each distinct group of a plurality of distinct groups, a feature score for each patient characteristic represented by that distinct group. Selecting a set of distinct groups can include comparing the feature scores of each distinct group of a plurality of distinct groups to a feature score threshold. Selecting a set of distinct groups can include at least one of: determining a stability measure for each distinct group of a plurality of distinct groups; determining a purity measure for each distinct group of a plurality of distinct groups; and determining the number of patients in a set of patients included in each distinct group of a plurality of distinct groups.

[0007] Identifying one or more relevant patient characteristics can include generating a plurality of potentially relevant patient characteristics by, for each distinct group of a set of distinct groups, selecting, from among a plurality of patient characteristics, the patient characteristics represented by that distinct group and corresponding to the theme of that group. Identifying one or more relevant patient characteristics can include ranking each of the potentially relevant patient characteristics of the plurality of potentially relevant patient characteristics. Ranking each of the potentially relevant patient characteristics can include assigning a rank value to each potentially relevant patient characteristic based on the frequency of co-occurrence of that potentially relevant patient characteristic and at least one reference indication related to the drug. Co-occurrence can be measured by determining, for each of the potentially relevant patient characteristics, the ratio of a set of distinct groups that includes both that potentially relevant patient characteristic and at least one reference indication. Identifying one or more relevant patient characteristics can include, for each of the potentially relevant patient characteristics: determining at least one of clinical feasibility and commercial feasibility.

[0008] One or more operating methods can include identifying at least one of one or more associated patient characteristics as a target indication for redeveloping an existing drug.

[0009] These and other aspects, configurations, and embodiments can be expressed as a method, apparatus, system, component, program product, means or steps for performing a function, and in other ways.

[0010] Embodiments of the present disclosure can provide one or more of the following advantages. Compared with the prior art, the systems and methods described herein can improve computational efficiency by reducing the computational time for processing large amounts of data with different levels of complexity for identifying potentially relevant indications (which may be referred to herein as patient characteristics) for use in redeveloping existing drugs. New indications that cannot be identified by the prior art can be discovered using specific machine learning techniques. Compared with the prior art, the systems and methods described herein can reduce dependence on human input and skills.

[0011] These and other aspects, configurations, and embodiments will become apparent from the following description, including the claims.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

DETAILED DESCRIPTION OF THE INVENTION

[0013] The redevelopment of existing drugs can be used to find new clinical indications (e.g., reasons to use a drug) for clinically approved drugs. When hundreds of new indications are explored, key opinion leaders (KOLs) may not be able to handle the high levels of complexity associated with these new indications. Processing large amounts of data and using advanced analytics can support a KOL-centered approach. Guided by relevant clinical questions, powerful advanced analytics techniques can draw out clinically relevant information hidden in large amounts of data that can later be useful for clinical decision-making.

[0014] Computational methods for the redevelopment of existing drugs can detect new drug-disease relationships using similarity metrics (chemical similarity, molecular activity similarity, gene expression similarity, or side effect similarity), molecular docking, or shared molecular pathology. Existing drug redevelopment techniques can be classified as network-based, text mining (literature search), and semantic techniques.

[0015] Network-based approaches can involve creating an integrated network by combining multiple data sources such as drugs, proteins, genes, and diseases. For example, the Connectivity Map (C-Map) approach can utilize transcriptomics by relating biology, chemistry, and clinical phenotypes using gene expression profiling to facilitate the discovery of new disease-gene-drug connections. Network-based clustering techniques can be used to find biological modules. This approach is inspired by the fact that biological entities (diseases, drugs, proteins, etc.) in the same module of a biological network usually share similar properties. Clustering can involve using the topological structure of the network to find drug-disease, drug-drug, disease-disease, or drug-target relationships. Implementing clustering techniques can have several difficulties. For example, the edges associating drugs and diseases can be determined by the incomplete collection of drug-disease associations, which may require integrating multiple databases to improve prediction accuracy. Additionally, by incorporating different data sources that provide information on drug side effects, potential safety signals can be collected. Furthermore, it can be difficult to distinguish negative and positive associations, and there is no "absolute" reference method to verify the relationships between biological modules.

[0016] Text-mining-based approaches can use keyword co-occurrence and semantic inference for new drug-disease associations. Some approaches can be based on Swanson's analogical inference method (ABC model), which assumes that if "B" is one of the characteristics of disease "C" and substance "A" affects "B", an implicit link "A:C" can be inferred through the B connection. However, relying solely on literature-mining-based methods may have limitations due to the ambiguous nature of language, the limited scope of biomedical relationships, and the limited accuracy of text-mining techniques.

[0017] Embodiments of the data processing systems and methods described herein use a real-world data-driven protocol for the redevelopment of existing drugs to identify indications This can mitigate the above-mentioned disadvantages. In some embodiments, the data processing systems and methods described herein combine analytics and real-world data with KOL clinical outputs. In some embodiments, the data processing systems and methods described herein use real-world data and analytics in an unsupervised manner. In some embodiments, patients are clustered using machine learning techniques, and groups of medical conditions that occur across multiple clusters are identified. In some embodiments, the identified groups of medical conditions may correspond to common biological pathways. Potential indications that were not previously identified using conventional techniques can be identified using the data processing systems and methods described herein.

[0018] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. It will be apparent, however, that the present disclosure may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form in order not to unnecessarily obscure the present disclosure.

[0019] In the drawings, for ease of explanation, specific arrangements or orderings of schematic elements such as those representing devices, modules, instruction blocks, and data elements are shown. However, one of ordinary skill in the art will understand that the specific orderings or arrangements of schematic elements in the drawings are not intended to imply that a particular order or sequence of processing, or separation of processes, is required. Further, including schematic elements in the drawings is not intended to imply that such elements are essential in all embodiments, or that the configurations represented by such elements cannot be included in or combined with other elements in some embodiments.

[0020] Furthermore, in the drawings, when using connecting elements such as solid or dashed lines or arrows to indicate the connection, relationship, or association between two or more other schematic elements, the absence of any of these connecting elements is not intended to mean that there is no possibility of a connection, relationship, or association. In other words, in the drawings, some connections, relationships, or associations between elements are not shown so as not to obscure the present disclosure. In addition, for ease of illustration, a single connecting element is used to represent multiple connections, relationships, or associations between elements. For example, when the connecting element represents the communication of signals, data, or instructions, those skilled in the art should understand that, if necessary, such elements may represent one or more signal paths (e.g., buses) in order to affect the communication.

[0021] Here, reference will be made in detail to the embodiments illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various described embodiments. However, it will be apparent to those skilled in the art that the various described embodiments can be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0022] Hereinafter, several configurations will be described that can each be used independently of one another or in any combination with other configurations. However, any individual configuration may not address any of the above-described problems, or may only address one of the above-described problems. Some of the above-described problems may not be fully addressed by any of the configurations described herein. Although headings are provided, data related to a particular heading but not found in the section having that heading may also be found elsewhere in this specification.

[0023] Examples of Data Processing Systems and Methods Figure 1 shows an example of a data processing system 100. In some embodiments, the data processing system 100 is configured to process data that can represent the medical records of multiple patients to identify new indications for drugs (existing drugs to be redevelopment). The system 100 includes a computer processor 110. The computer processor 110 includes a computer-readable memory 111 and computer-readable instructions 112. The system 100 also includes a machine learning system 150. The machine learning system 150 includes a machine learning model 120. The machine learning model 120 can be separated from the computer processor 110 or integrated with the computer processor 110.

[0024] The computer-readable medium 111 (or computer-readable memory) can include any type of data storage technology suitable for the local technical environment, including but not limited to semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, removable memory, disk memory, flash memory, dynamic random access memory (DRAM), static random access memory (SRAM), electronically erasable programmable read-only memory (EEPROM), and the like. In some embodiments, the computer-readable medium 111 includes a code segment having executable instructions.

[0025] In some embodiments, computer processor 110 includes a general-purpose processor. In some embodiments, computer processor 110 includes a central processing unit (CPU). In some embodiments, computer processor 110 includes at least one application-specific integrated circuit (ASIC). Computer processor 110 can also include a general-purpose programmable microprocessor, a graphics processing unit, a dedicated programmable microprocessor, a digital signal processor (DSP), a programmable logic array (PLA), a field programmable gate array (FPGA), a dedicated electronic circuit, etc. or a combination thereof. Computer processor 110 is configured to execute program code such as computer-executable instructions 112 and is configured to execute executable logic including machine learning model 120.

[0026] Computer processor 110 is configured to receive data representing the medical records of multiple patients. For example, computer processor 110 can receive data from a database including electronic medical records (EMRs) for approximately 94 million (or more) patients that can be identified by a key identifier (ID) that enables patient matching between different data tables. In some embodiments, the data indicates diagnoses, clinical tests, treatments, medications, patient events, insurance, biomarkers, measurements, clinical status, lifestyle parameters, microbiology, prescriptions, etc. In some embodiments, the data includes natural language processing-driven data. The data can be received through any of a variety of techniques such as wireless communication, fiber optic communication, USB, CD-ROM, etc.

[0027] The machine learning system 150 can apply machine learning techniques to train the machine learning model 120. As part of the training of the machine learning model 120, the machine learning system 150 can form a training set of input data by identifying a positive training set of input data items that are determined to have the characteristics of interest, and in some embodiments, can form a negative training set of input data items that do not have the characteristics of interest.

[0028] The machine learning system 150 extracts feature values from the input data of the training set, and this feature is a variable that may be considered relevant to whether the input data item has one or more characteristics to which it is related. In this specification, an ordered list of features for the input data is referred to as the feature vector of the input data. In some embodiments, the machine learning system 150 applies dimensionality reduction to reduce the amount of data in the feature vector of the input data to a smaller , more representative set of data. For example, the machine learning system 150 can apply multiple correspondence analysis (MCA), linear discriminant analysis (LDA), principal component analysis (PCA), etc.

[0029] In some embodiments, the machine learning system 150 trains the machine learning model 120 using unsupervised machine learning. Typically, unsupervised machine learning techniques make inferences from a dataset using input vectors without referring to known, i.e., labeled, results. In some embodiments, the machine learning system 150 performs clustering to divide data points into multiple groups such that data points within the same group are more similar to other data points within the same group and not similar to data points within other groups. In some embodiments, performing clustering includes performing K-means clustering, where the dataset is repeatedly divided to create a one-level non-nested partition of the data points. That is, if K is the desired number of clusters, in each iteration, the dataset is divided into K independent clusters. This process can continue until a specified clustering criterion function value is optimized. In some embodiments, the machine learning system 150 is configured to perform bisecting K-means clustering. Bisecting k-means clustering generally involves dividing one cluster into two sub-clusters in each bisection step (e.g., by using k-means) until k clusters are obtained. Bisecting K-means clustering may be more beneficial compared to K-means clustering, as bisecting K-means clustering can reduce the computation time when K is a relatively large value, can generate clusters of similar sizes, and can generate clusters with lower entropy.

[0030] The computer processor 110 is configured to execute computer-executable instructions 112 to implement one or more operational methods. In some embodiments, the one or more operational methods include receiving data representing the medical records of a plurality of patients. For example, the computer processor 110 can receive data from a database containing electronic medical records (EMRs) of approximately 94 million (or more) patients that can be identified by a key identifier (ID) that enables matching of patients across different data tables. In some embodiments, the data indicates diagnoses, clinical tests, treatments, medications, patient events, insurance, biomarkers, measurements, clinical status, lifestyle parameters, microbiology, prescriptions, etc. In some embodiments, the data includes natural language processing-driven data. The data can be received through any of a variety of techniques, such as wireless communication, fiber optic communication, USB, CD-ROM, etc.

[0031] In some embodiments, one or more of the methods of operation include selecting a set of patients based on medical records. Selecting the set of patients includes determining at least one target signaling pathway associated with an existing drug to be redeployed. For example, if the existing drug to be redeployed is Dupilumab, the computer processor 110 can determine, based on the known function of the drug, that the drug regulates the interleukin 4 (IL-4) and interleukin 13 (IL-13) signaling pathways. Selecting the set of patients also includes determining one or more indicators based on one or more factors corresponding to a diagnosis associated with the target signaling pathway. For example, factors such as pathway mechanism, related clinical symptoms, therapeutic analogs, data and epidemiology, and the integrity of pharmaceutical lifecycle management can be used to search sources including medical databases and medical evidence software to identify diseases associated with the determined signaling pathway. These diseases can be categorized based on the strength of the link to the determined signaling pathway. The categories include a focused group, a medium group, and a broad group It is possible. For example, returning to the IL-4 / IL-13 example, the concentrated group of diseases can include diseases directly related to the IL-4 / IL-13 mechanism of action on the Th2 pathway, the median group of diseases can include diseases indirectly related to the IL-4 / IL-13 mechanism of action on the Th2 pathway, and the broad group of diseases can include diseases related to a broader inflammatory response. By moving from the concentrated group to the broad group, the number of indicators to be considered when selecting a set of patients can be increased, and the potential for molecular influence can be reduced. Thus, in some embodiments, only the concentrated group, or the concentrated group and the median group, are used to select a set of patients. In some embodiments, only patients having at least one diagnosis, pharmaceutical, clinical test, and / or treatment related to the determined signaling pathway are selected to be included in a set of patients. Detailed examples of factors and indicators are provided later with reference to Table 1.

[0032] In some embodiments, one or more operating methods include determining a plurality of patient characteristics (which may be referred to herein as features) for a set of patients, where each patient in the set of patients exhibits at least one of the plurality of patient characteristics. Determining the plurality of patient characteristics can include analyzing the initially received data to identify a broad range of patient characteristics for incorporating all or a substantial portion of the received data. For example, the broad range of patient characteristics can correspond to diagnoses (e.g., immune diseases, diabetes), prescriptions (e.g., immune drugs, other drug classes), treatments (e.g., human leukocyte antigen typing), and test results (e.g., IgE abnormally high / low). In some embodiments, determining the plurality of patient characteristics includes receiving user input (e.g., via a user interface communicating with computer processor 110). For example, the user can input patient characteristics based on clinical input, demographics, pharmaceuticals, co - morbidities, treatments, and clinical test data specific to immunology. Customized feature classes may also be added to improve the completeness and representativeness of the data and to gather more information about diseases and drug reactions. In some embodiments, determining the plurality of patient characteristics includes validating the plurality of patient characteristics. Validation can include calculating the proportion of selected patients having at least one of each feature family (e.g., the proportion of patients having prescription records) and comparing this proportion to the proportion of patients in the initially received data having at least one of each feature family to determine whether the patient characteristics in the initially received data are accurately mapped to the selected set of patients. The closer the values of the two numbers, the more accurately the mapping has been done. Validation can include identifying a plurality of patients included in both the initially received data and the selected set of patients and validating the same mapping of patient characteristics between the patients in the initially received data and the selected set of patients to determine whether the patient characteristics are accurately mapped to the correct patients.

[0033] In some embodiments, one or more operating methods include grouping a set of patients to generate a plurality of distinct groups, each including at least one patient of a set of patients, according to a plurality of patient characteristics (e.g., as defined by characteristics associated with a determined signaling pathway). For example, one or more computer processors 110 can execute a machine learning model 120 to perform a clustering technique such as the binary k-means clustering technique described above. Clustering can result in a plurality of clusters (e.g., distinct groups) of patients, where patients in one cluster are more similar to each other with respect to their corresponding patient characteristics than patients in other clusters. In some embodiments, the generated clusters may exhibit correlations between patient characteristics even if they do not exist in the same patient. Clinical input can be received and used at various stages of the clustering process to ensure the clinical relevance of the resulting clusters. For example, clinical input from disease experts can facilitate the creation of clinically relevant cohorts in the inclusion and grouping of clinically relevant features and in the validation and evaluation of clusters. Patient characteristics can be identified as being characteristic in a cluster if they occur more frequently in the cluster than in the general population (e.g., overall in a selected set of patients).

[0034] In some embodiments, multiple correspondence analysis (MCA) is used to reduce the dimensionality of patient characteristics. Binary K-means facilitates the proper and effective separation of patients into sufficiently "tight" but stable clusters, enabling the use of a large number of clusters that exhibit immune-relatedness for scoring patient characteristics (described in more detail later). The resulting clusters can be presented to a user (e.g., a clinical expert) for validation and evaluation (e.g., via a user interface). This can reduce the risk of non-interpretability of the clusters and ensure that there are no overlapping features between different clusters.

[0035] In some embodiments, one or more operating methods include selecting a set of distinct groups out of a plurality of distinct groups based on one or more group selection criteria. In some embodiments, selecting a set of distinct groups includes ranking the groups and selecting a number of the most highly ranked groups (e.g., the top 60 ranked groups). The groups can be ranked based on immunological enrichment, stability, purity, and size. In some embodiments, for ranking the clusters, one or more metrics (which may be referred to herein as feature scores) are calculated for each patient characteristic. Examples of the one or more metrics can include, for example, distinctiveness (which may be referred to herein as a “lift score”), the number of patients within the cluster presenting the patient characteristic, and an immunological score. The distinctiveness score measures how characteristic the patient characteristic is within the cluster relative to the rest of the population (e.g., if males make up 50% of the population and 75% of the cluster, the “lift score” could be equal to 1.5). In some embodiments, only patient characteristics that have a lift score exceeding a threshold lift score (e.g., 1) and that occur at a rate exceeding a threshold percentage of patients (e.g., 10%) are considered to define the cluster and correspond to the theme of the cluster. Patient characteristics considered to define the cluster may be referred to herein as potentially relevant patient characteristics. Next, an immunological score can be assigned to the patient characteristic (e.g., either the patient characteristic considered to define the cluster or any of all patient characteristics), and this immunological score scores the patient characteristic according to its type (e.g., disease, drug, clinical test, treatment, etc.) and immunological relevance. Then, the patient characteristic scores within each cluster can be aggregated (e.g., summed) and normalized. Then, clusters that meet a threshold cluster score (e.g., 50%) can be considered immunologically distinct.

[0036] Selecting a set of distinct groups can include evaluating one or more of stability, purity, and the number of patients within each cluster. Stability can be evaluated using one or more of the following methods: (1) reproducing clusters with different sizes of data; (2) changing the initialization seeds of the clusters; (3) changing the number of generated clusters; and (4) applying a training-test method. For each cluster in the training set, stability can be defined as the maximum ratio of patients who were originally grouped together in the test set. Purity can be measured by the within-cluster variance of the MCA components of the patients within the cluster, which can result in homogeneous and high-density clusters. In some embodiments, a cluster is selected if it exceeds a threshold stability ratio (e.g., 50%) and a threshold purity ratio (e.g., the cluster is in the highest 20% purity among all clusters).

[0037] In some embodiments, one or more operating methods are for a set of distinct groups Including identifying one or more associated patient characteristics (e.g., indication) by analyzing individual groups. Identifying one or more associated patient characteristics can include ranking the patient characteristics presented by each selected cluster (e.g., all of the patient characteristics, or those considered to define the cluster). The ranking can be based on the frequency of co-occurrence with each of a number of (reference) established (reference) characteristics of the existing drug to be redeveloped (e.g., if the drug is dupilumab, the reference characteristics can include asthma, atopic dermatitis, IgE allergy, and composite immunological scores). Co-occurrence can be measured by calculating the ratio of the patient-weighted cluster that includes both the patient characteristic and the reference. In some embodiments, one or more patient characteristics determined by a content area expert to be relevant to the core cluster theme (such as indicated by user input received via a user interface) can also be considered for evaluation regardless of the number of patients in which these characteristics appear (which can be less than 10%).

[0038] Identifying one or more patient characteristics can include assessing the clinical and commercial viability of the patient characteristics. For example, patient characteristics indicating distinct clinical diagnoses can be identified. Commercial evaluation can be based on available sales forecasts and data indicating the assets of competing companies, determined links to target signaling pathways (regardless of whether described in publications), the worldwide prevalence of the patient characteristic, and the disability-adjusted life years (DALY) of the patient characteristic (e.g., for 100,000 life years). In some embodiments, at least one of the one or more patient characteristics can be identified as a target indication for redeveloping an existing drug. As a result, in some embodiments, one or more methods of operation generally output one or more new indications for the existing drug to be redeveloped.

[0039] This specification generally describes the patient as a human patient, but the embodiments are not limited thereto. For example, the patient can refer to a non-human animal, a plant, or a human replication system.

[0040] For illustrative purposes, this specification describes receiving data corresponding to approximately 94 million patients, but it is understood that the data can correspond to fewer or more patients.

[0041] For illustrative purposes, this specification describes dupilumab as an existing drug to be redeveloped, but it is understood that the existing drug to be redeveloped can be any drug.

[0042] Figure 2 is a flowchart showing an example method 200 for redeveloping an existing drug. Method 200 can be implemented by the data processing system 100 described above with reference to Figure 1. Method 200 includes receiving data representing medical records (block 210), selecting a set of patients (block 220), determining a plurality of patient characteristics (block 230), grouping a set of patients to generate a plurality of distinct groups (block 240), selecting a set of distinct groups (block 250), and identifying one or more relevant patient characteristics (block 260).

[0043] In block 210, data representing the medical records of a plurality of patients is received. The data can be received, for example, from a database including the EMRs of approximately 94 million (or more) patients that can be identified by a key ID that enables matching of patients across different data tables. In some embodiments, the data indicates diagnoses, clinical tests, treatments, pharmaceuticals, patient events, insurance, biomarkers, measurements, clinical states, lifestyle parameters, microbiology, prescriptions, medical images, etc. In some embodiments, the data includes natural language processing-driven data. The data can be received via any of a variety of techniques such as wireless communication, fiber optic communication, USB, CD-ROM, etc. including natural language processing-driven data. The data can be received via any of a variety of techniques such as wireless communication, fiber optic communication, USB, CD-ROM, etc.

[0044] In block 220, at least one target signaling pathway associated with the existing drug to be redeployed is determined. For example, if the drug is dupilumab, based on the known function of the drug, it can be determined that the drug regulates the interleukin-4 (IL-4) and interleukin-13 (IL-13) signaling pathways. In some embodiments, one or more indicators can be determined based on one or more factors corresponding to a diagnosis associated with the target signaling pathway. For example, using factors such as pathway mechanism, related clinical symptoms, therapeutic analogs, data and epidemiology, and the integrity of pharmaceutical life cycle management, sources including medical databases and medical evidence software can be searched to identify diseases associated with the determined signaling pathway. These diseases can be categorized based on the strength of the link to the determined signaling pathway. The categories can include an intensive group, a median group, and an extensive group. For example, returning to the IL-4 / IL-13 example, the intensive group of diseases can include diseases directly related to the IL-4 / IL-13 mechanism of action on the Th2 pathway, the median lens group of diseases can include diseases indirectly related to the IL-4 / IL-13 mechanism of action on the Th2 pathway, and the extensive group of diseases can include diseases related to a broader inflammatory response. Moving from the intensive group to the extensive group can increase the number of indicators to be considered when selecting a set of patients and can reduce the potential for molecular influence. Thus, in some embodiments, only the intensive group, or the intensive group and the median group, are used to select a set of patients. In some embodiments, only patients having at least one diagnosis, pharmaceutical, clinical test, and / or treatment associated with the determined signaling pathway are selected to be included in a set of patients.

[0045] In block 230, a plurality of patient characteristics of a set of patients are determined, where each patient in the set of patients exhibits at least one of the plurality of patient characteristics. Determining the plurality of patient characteristics can include analyzing the initially received data to identify a wide range of patient characteristics for incorporating all or a substantial portion of the received data. For example, the wide range of patient characteristics can correspond to diagnosis (e.g., immune disease, diabetes), prescription (e.g., immune drugs, other drug classes), treatment (e.g., human leukocyte antigen typing), and test results (e.g., IgE abnormally high / low). In some embodiments, determining the plurality of patient characteristics includes receiving user input (e.g., via a user interface). For example, the user can input patient characteristics based on clinical input and demographics, pharmaceuticals, co-morbidities, treatments, and clinical test data specific to immunology. Customized characteristic classes may also be added to improve data integrity, representativeness, and collect more information about diseases and drug reactions. In some embodiments, determining the plurality of patient characteristics includes validating the plurality of patient characteristics. Validation can include calculating the percentage of selected patients having at least one of each characteristic family (e.g., the percentage of patients having prescription records) and comparing this percentage to the percentage of patients in the initially received data having at least one of each characteristic family to determine whether the patient characteristics in the initially received data are accurately mapped to the selected set of patients. The closer the values of the two numbers are, the more accurately the mapping has been performed. Validation can include identifying a plurality of patients included in both the initially received data and the selected set of patients and validating the same mapping of patient characteristics between the patients in the initially received data and the selected set of patients to determine whether the patient characteristics are accurately mapped to the correct patients.

[0046] In block 240, (e.g., by features related to the determined signaling pathway) According to the definition, a set of patients is grouped according to multiple patient characteristics, and a plurality of distinct groups are generated, each containing at least one patient of the set of patients. For example, for a set of patients, multiple patient characteristics can be used to perform a clustering technique such as the binary k-means clustering technique described above. Clustering can result in multiple clusters (e.g., distinct groups) of patients, where patients in one cluster are more similar to each other with respect to their corresponding patient characteristics than patients in other clusters. In some embodiments, the generated clusters may show correlations between patient characteristics even if they do not exist in the same patient. Clinical input can be received and used at various stages of the clustering process to ensure the clinical relevance of the resulting clusters. For example, clinical input from disease experts can facilitate the creation of clinically relevant cohorts in the inclusion and grouping of clinically relevant features and in the validation and evaluation of clusters. Patient characteristics can be identified as characteristic in a cluster if they occur more frequently in the cluster than in the general population (e.g., overall in a selected set of patients).

[0047] In some embodiments, multiple correspondence analysis (MCA) is used to reduce the dimensionality of patient characteristics. Binary K-means facilitates the proper and effective separation of patients into sufficiently "tight" but stable clusters, and enables the use of a number of clusters that show immunological relevance for scoring patient characteristics (to be described in more detail later). The resulting clusters can be presented to a user (e.g., a clinical expert) for validation and evaluation (e.g., via a user interface). This can reduce the risk of non-interpretability of the clusters and ensure that there are no overlapping features between different clusters.

[0048] In block 250, a set of distinct groups out of a plurality of distinct groups is selected based on one or more group selection criteria. In some embodiments, selecting a set of distinct groups includes ranking the groups and selecting a number of the most highly ranked groups (e.g., the top 60 ranked groups). The groups can be ranked based on immunological enrichment, stability, purity, and size. In some embodiments, for each patient characteristic, one or more metrics are calculated to rank the clusters. Examples of the one or more metrics can include, for example, distinctiveness (which may be referred to herein as the "lift score"), the number of patients within the cluster presenting the patient characteristic, and the immunological score. The distinctiveness score measures how characteristic the patient characteristic is within the cluster relative to the rest of the population (e.g., if men make up 50% of the population and 75% of the cluster, the "lift score" could be equal to 1.5). In some embodiments, only patient characteristics that have a lift score exceeding a threshold lift score (e.g., 1) and that occur at a rate exceeding a threshold percentage of patients (e.g., 10%) are considered to define the cluster and correspond to the theme of the cluster. Patient characteristics considered to define the cluster may be referred to herein as potentially relevant patient characteristics. Next, an immunological score can be given to the patient characteristic (e.g., the patient characteristic considered for defining the cluster, or any of all patient characteristics), which scores the patient characteristic according to its type (e.g., disease, drug, clinical test, treatment, etc.) and immunological relevance. Next, the patient characteristic scores within each cluster can be aggregated (e.g., summed) and normalized. Next, clusters that meet a threshold cluster score (e.g., 50%) can be considered immunologically distinct.

[0049] Selecting a set of distinct groups can include evaluating one or more of stability, purity, and the number of patients within each cluster. Stability can be evaluated using one or more of the following methods: (1) reproducing the clusters with data of different sizes; (2) changing the initialization seeds of the clusters; (3) changing the number of generated clusters; and (4) applying a training-test method. For each cluster in the training set, stability can be defined as the maximum ratio of patients who were originally grouped together in the test set. Purity can be measured by the within-cluster variance of the MCA components of the patients within the cluster, which can result in homogeneous and high-density clusters. In some embodiments, a cluster is selected if it exceeds a threshold stability ratio (e.g., 50%) and a threshold purity ratio (e.g., the cluster is in the top 20% purity among all clusters). Selecting a set of distinct groups can include evaluating one or more of stability, purity, and the number of patients within each cluster. Stability can be evaluated using one or more of the following methods: (1) reproducing the clusters with data of different sizes; (2) changing the initialization seeds of the clusters; (3) changing the number of generated clusters; and (4) applying a training-test method. For each cluster in the training set, stability can be defined as the maximum ratio of patients who were originally grouped together in the test set. Purity can be measured by the within-cluster variance of the MCA components of the patients within the cluster, which can result in homogeneous and high-density clusters. In some embodiments, a cluster is selected if it exceeds a threshold stability ratio (e.g., 50%) and a threshold purity ratio (e.g., the cluster is in the top 20% purity among all clusters).

[0050] In block 260, one or more relevant patient characteristics are identified by analyzing each distinct group of a set of distinct groups. Identifying one or more relevant patient characteristics can include ranking the patient characteristics presented by each selected cluster (e.g., all of the patient characteristics, or the patient characteristics considered to define the cluster). The ranking can be based on the frequency of co-occurrence with each of a number of (reference) established (reference) characteristics of an existing drug to be redeveloped (e.g., if the drug is dupilumab, the reference characteristics may include asthma, atopic dermatitis, IgE allergy, and composite immunological scores). Co-occurrence can be measured by calculating the ratio of the patient-weighted cluster that includes both the patient characteristic and the reference. In some embodiments, one or more patient characteristics determined by a content area expert to be relevant to a core cluster theme (e.g., as indicated by user input received via a user interface) can also be considered for evaluation regardless of the number of patients in which these characteristics appear (which may be less than 10%).

[0051] Identifying one or more patient characteristics can include evaluating the clinical and commercial viability of the patient characteristics. For example, patient characteristics indicating distinct clinical diagnoses can be identified. Commercial evaluations can be based on available sales forecasts and data indicating the assets of competing companies, determined links to target signaling pathways (whether described in publications or not), the worldwide prevalence of the patient characteristics, and the disability-adjusted life years (DALYs) of the patient characteristics (e.g., for 100,000 life years). In some embodiments, at least one of the one or more patient characteristics can be identified as a target indication for redeveloping an existing drug.

[0052] Experimental results Figure 3 is a diagram showing an experiment using the systems and methods described herein. This experiment was conducted to validate an RWD-driven protocol for redeveloping an existing drug of the drug dupilumab to identify a new indication for the anti-IL4 / IL13 drug. One objective of this experiment was to reduce the drug development cost and shorten the time to market while minimizing dropouts and risks. A hybrid approach using scientific and clinical capabilities was applied through an analysis combined with KOL expertise, commercial evaluations, and real-world data.

[0053] Data source: The Optum Humedica dataset from 2014 to 2018 was used. This database contained electronic medical records for 94 million patients identifiable by key identifiers, enabling patient matching across different data tables. This database collected information on EMR data such as diagnoses, clinical tests, procedures, pharmaceuticals, patient events, insurance, biomarkers, measurements, clinical status and lifestyle parameters, microbiology, and prescriptions. Natural language processing (NLP)-driven tables were not included because the data scope and clinical relevance were limited. Additionally, data tables containing incomplete or irrelevant information were excluded. A total of five data tables were included, reducing the data source to 40 million patients.

[0054] Patient selection: The patient selection indicator was based on a clinical framework related to the underlying immunological pathway, as shown in Table 1.

[0055] [Table 1]

[0056] Only adult patients (18 years and older) who had received at least one diagnosis, pharmaceutical, clinical test, and procedure related to a diagnosis associated with the IL4 / 13 pathway (i.e., signaling pathway) were selected. Using the criteria for these immunological conditions and data completeness, a cohort consisting of 17 million unique patients was obtained as a result.

[0057] To identify the indicators, five factors (as shown in Table 1) were considered. These factors consist of several platforms (DOC Library, DOC Data, Doctor Evidence, DOC Label, DOC Search) and were searched through the sources of the Doctor Evidence engine database, a medical evidence software and service company that includes PubMed, ClinicalTrials.gov, WHO, etc. Then, the information within these factors was classified according to three lenses based on the Th2 response: intensive, median, and broad. For example, diseases were assigned to the intensive or median lens based on their direct and indirect relationships, respectively, to the mechanism of action of IL4 / IL13 in the Th2 pathway, and to the broad lens if they were related to a more extensive inflammatory response. By moving from the intensive lens to the broad lens, the range of indicators to be considered widened, and the potential for molecular influence decreased. The indicators in the broad lens were ultimately excluded from the analysis because their mechanistic links were not specific enough to meet the criteria for identifying patient populations with similar characteristics. Therefore, only the indicators in the intensive and median lenses were included in the experiment. For the analysis across 17 broad systems, the final list of 208 indicators was included. The indicators in the broad lens were ultimately excluded from the analysis because their mechanistic links were not specific enough to meet the criteria for identifying patient populations with similar characteristics. Therefore, only the indicators in the intensive and median lenses were included in the experiment. For the analysis across 17 broad systems, the final list of 208 indicators was included.

[0058] Feature Selection: To incorporate the information available in the Optum dataset, a wide range of features (patient characteristics) were selected; subsequently, the features selected by clinical experts were prioritized and verified to ensure that all essential variables were included and that the variable values were clinically meaningful. Where appropriate, some features were retained and others (certain demographics) were newly created (as shown by V1 as described later with reference to Figure 4). Based on clinical inputs and demographic, pharmaceutical, comorbid disease, treatment, and clinical laboratory data specific to immunology, new features were added. To improve data completeness and representativeness and to collect more information regarding the severity of disease and drug reactions, customized feature classes were generated (as shown by V2 and V3 as described later with reference to Figure 4) and iteratively added to the analysis.

[0059] To ensure the completeness of features in the final database, robust methods were used. Two confirmation steps were performed based on the mapping of patients and features across the Optum database and the generated tables to verify that the features were accurately generated. First, to confirm whether the Optum patient features were accurately mapped to the data tables of the present invention, the proportion of patients having at least one feature family was calculated and it was determined whether this number was the same in Optum Humedica and the data tables of the present invention. Second, to confirm whether the features were mapped to the correct patients, ten random patients were tracked from the raw data of Optum Humedica to the generated dataset to ensure the same mapping of features in the two datasets. After accurate mapping was demonstrated by feature confirmation, the algorithm was run on 17 million patients including 2700 features.

[0060] Clustering: Using clustering techniques, patients sharing similar characteristics as defined by features related to the IL4 / 13 pathway were grouped together. Clustering searched for similarities between patients based on their characteristics. The generated clusters found correlations between symptoms even when they did not exist in the same patient. To ensure the clinical relevance of the results, clinical input was incorporated at various stages of the process. Thus, clinical input from disease experts helped in creating clinically relevant cohorts, including and grouping clinically relevant features, and ultimately validating and evaluating the clusters. Features were identified as characteristic within a cluster if they occurred more frequently than in the general population.

[0061] Multicorrespondence analysis (MCA) was used to reduce the dimensionality of the features. Then, binary K-means was utilized to divide the data into 500 clusters, providing an appropriate and effective separation of patients in clusters that were sufficiently "tight" but stable, and enabling the use of a number of clusters showing immune-relatedness for indication scoring. Clinical experts verified and evaluated the clusters identified through this process. This step facilitated the reduction of the risk of non-interpretability of the clusters and ensured that there were no overlapping features between different clusters. This clustering method was run on data representing 2,700 features (e.g., MCA components) and 17 million patients. The number of clusters generated at the end of this algorithm was 500.

[0062] Identification of new indications (i.e., related patient characteristics): For further evaluation, clinical and commercial judgments were carried out to obtain a short list of priority signals, and the The most clinically relevant indications across clusters were identified based on the STA output. Four methodological steps were used. In the first step, the top 60 clusters were selected from 500 ranked clusters based on immunological enrichment, stability, purity, and size. For each feature included in each cluster, three metrics were calculated: specificity, the number of patients presenting that feature within each cluster, and the immunological score, to determine the selection. Specificity, also referred to as the "lift score," measured how characteristic a feature was within a cluster relative to the rest of the population (e.g., if men make up 50% of the population and 75% of the cluster, the "lift score" is equal to 1.5). Only features with a lift score greater than 1 (meaning occurring within the cluster at a higher rate than expected compared to the broader population) and present in more than 10% of the patients were considered to define (and name) the clusters and develop the cluster themes. Additionally, each selected feature was scored according to its type (disease, drug, clinical test, or procedure) and immunological relevance. The feature scores within each cluster were summed and normalized. Then, if a cluster met a pre-defined threshold score of 50%, it was considered immunologically distinct. As a second step, clusters were selected according to stability, purity, and the number of patients. Stability was evaluated using four methods: 1. reproducing the cluster with different-sized data, 2. changing the initialization seed of the cluster, 3. changing the number of generated clusters, and 4. applying a training-test method. For each cluster in the training set, stability was defined as the maximum proportion of patients who were grouped together in the test set. Purity could be measured by the within-cluster variance of the MCA components of the patients within that cluster, which could result in a homogeneous and high-density cluster. Clusters were included in the analysis if they had a stability greater than 50% and a purity in the top 20%. Additionally, all indications that content area experts determined to be relevant to the core cluster themes were considered for evaluation regardless of the number of patients in whom these features occurred (which could be less than 10%).Then, in the third step, these new indications were ranked based on the frequency of co-occurrence with each of the four established (reference) indications of dupilumab (asthma, atopic dermatitis, IgE allergy, and composite immunology score). Co-occurrence was measured by calculating the ratio of patient-weighted clusters that included both the indication and the reference. In the final step, the final list of indications was further characterized by clinical and commercial viability. Clinical evaluation retained indications that represented distinct clinical diagnoses. Conditions that did not appear in the top 30 based on ranking with the IL4 / IL13 reference were removed as they were considered to have low relevance to IL4 / IL13 modulation. Commercial evaluation was only possible for a subset of indications that were considered clinically valid when data on sales projections and competitor assets were available. In addition, for commercial evaluation, multiple factors were also considered: the link to the IL4 / 13 pathway, whether or not it was described in the literature, the worldwide prevalence of the indication, and the disability-adjusted life years (DALYs) of the indication (per 100,000 life years).

[0063] Results: Figure 4 is a diagram showing the experimental results resulting from the use of the systems and methods described in this specification. The final cohort of 17 million patients extracted from Optum Humedica was analyzed to evaluate the completeness and representativeness of the data across the selected characteristics. In the median lens population, the patient cohort consisted of 59% females, mostly white (77%), with a mean age at last activity of 53 years (SD = 7 years) and a mean follow-up period of 7.1 years. As immunological conditions, patients most frequently showed ICD10 codes for acute sinusitis (25.2%), allergic rhinitis of unknown detail (20.6%) and other asthma of unknown detail (19.4%). The most frequently used immunology-related pharmaceuticals were prednisone (28.0%), fluticasone furoate (22.2%) and methylprednisolone (13.3%), and 0.4% of the patients received allergen immunotherapy injections and β2 glycoprotein antibody measurements. Most of the patients had white blood cell counts (70.8%), absolute Received tests for the absolute neutrophil count (ANC) (64.3%) and absolute lymphocyte count (ALC) (63.8%). By the clustering procedure, 500 clusters were generated, 125 of which were classified as being enriched and stable for immune diseases. Of these 125, 110 were also classified as being of high purity. Of these, 60 clusters containing the largest number of patients per cluster were retained and analyzed for clinically relevant signals. Following the validation process, using the training-test method, 84% of the clusters were considered highly stable and reproduced 90% and 99% of the top 20 clusters, respectively, regardless of the seed position and number of patients in the data table. Based on the characteristics included in the clusters, six cluster themes were identified: multi-organ immunological effects, neoplasms, asthma and other allergies, musculoskeletal dysfunction, cardiometabolic spectrum, and gynecological and obstetric conditions, and were also partially reproduced in the V2 and V3 iterations. Of these, 250 indications were selected for cluster evaluation and ranked by co-occurrence with each reference. Based on clinical and commercial viability, a list of approximately 85 indications was further characterized: approximately 20 of these did not represent distinct clinical diagnoses or had insufficient clinical justification for IL4 / 13 modulation, and the others had no readily available commercial evaluation information. The final list of indications from the hybrid approach identified approximately 90% of the indications already in life cycle management and approximately 60% of the additional possible new indications.

[0064] FIG. 5 is a block diagram of an example computer system 600 used to provide computer functionality related to the algorithms, methods, functions, processes, flows, and procedures described in the present disclosure (such as method 200 etc. described above with reference to FIG. 2). The illustrated computer 602 is intended to encompass any computing device, including a server, desktop computer, laptop / notebook computer, wireless data port, smartphone, personal digital assistant (PDA), tablet computing device, or one or more processors within these devices, including physical instances, virtual instances, or both. The computer 602 can include input devices such as a keypad, keyboard, and touch screen that can accept user information. Also, the computer 602 can include output devices that can communicate information related to the operation of the computer 602. The information can include digital data, visual data, audio information, or a combination of information. The information can be presented in a graphical user interface (UI) (or GUI).

[0065] The computer 602 can serve as a client, network component, server, database, persistence, or a component of a computer system that implements the subject matter described in the present disclosure. The illustrated computer 602 is communicatively coupled to a network 630. In some embodiments, one or more components of the computer 602 can be configured to operate in different environments, including cloud computing-based environments, local environments, global environments, and combinations of environments.

[0066] Broadly speaking, computer 602 is an electronic computing device operable to receive, transmit, process, store, and manage data and information related to the described subject matter. According to some embodiments, computer 602 may also include, or be communicatively coupled with, an application server, an email server, a web server, a cache server, a streaming data server, or a combination of servers.

[0067] Computer 602 can receive requests from a client application (e.g., running on another computer 602) via network 630. Computer 602 can respond to received requests by processing the received requests using software applications. Requests can also be sent to computer 602 from internal users (e.g., from a command console), external (or third parties), automated applications, entities, individuals, systems, and computers.

[0068] Each of the components of computer 602 can communicate using system bus 603. In some embodiments, any or all of the components of computer 602, including hardware or software components, can interface with each other or with interface 604 (or a combination of both) via system bus 603. The interface can use an application programming interface (API) 612, a service layer 613, or a combination of API 612 and service layer 613. API 612 can include specifications for routines, data structures, and object classes. API 612 may or may not be dependent on a computer language. API 612 can refer to a complete interface, a single function, or a set of APIs.

[0069] The service layer 613 can provide software services to the computer 602 and other components (regardless of whether they are shown in the figure) communicatively coupled to the computer 602. The functions of the computer 602 can be accessible to all service consumers using this service layer. Software services, such as those provided by the service layer 613, can provide reusable defined functions via a defined interface. For example, the interface can be software written in a language that provides data in the form of JAVA, C++, or Extensible Markup Language (XML). Although shown as an integrated component of the computer 602, in an alternative embodiment, the API 612 or the service layer 613 can be a stand-alone component in relation to other components of the computer 602 and other components communicatively coupled to the computer 602. Further, any or all parts of the API 612 or the service layer 613 can be implemented as a child or sub-module of another software module, enterprise application, or hardware module without departing from the scope of the present disclosure.

[0070] Computer 602 includes an interface 604. Although shown as a single interface 604 in FIG. 5, two or more interfaces 604 can be used according to the specific needs, requirements, or specific embodiments of computer 602, as well as the described functions. Interface 604 can be used by computer 602 to communicate with other systems connected to network 603 in a distributed environment (regardless of whether shown). Generally, interface 604 can include or implement logic encoded in software or hardware (or a combination of software and hardware) operable to communicate with network 630. More specifically, interface 604 can include software that supports one or more communication protocols related to communication. Thus, network 630 or the hardware of the interface can be operable to communicate physical signals either inside or outside the illustrated computer 602.

[0071] Computer 602 includes a processor 605. Although shown as a single processor 605 in FIG. 5, two or more processors 605 can be used according to the specific needs, requirements, or specific embodiments of computer 602, as well as the described functions. Generally, processor 605 can execute instructions and manipulate data to implement the operating method of computer 602, including an operating method using algorithms, methods, functions, processes, flows, and procedures as described in this disclosure.

[0072] Computer 602 also includes a database 606 that can hold data for computer 602 and other components connected to network 630 (whether shown or not). For example, database 606 can be an in-memory, traditional, or database that stores data consistent with the present disclosure. In some embodiments, database 606 can be a combination of two or more different database types (e.g., a hybrid in-memory data database and a traditional database) according to the specific needs, requirements or specific embodiments of computer 602, as well as the described functionality. Although illustrated as a single database in FIG. 5, two or more databases (of the same type, different types, or a combination of types) can be used according to the specific needs, requirements or specific embodiments of computer 602, as well as the described functionality. Database 606 is illustrated as an internal component of computer 602, but in alternative embodiments, database 606 can be external to computer 602.

[0073] Computer 602 also includes a memory 607 that can hold data for computer 602, or a combination of components connected to network 630 (whether shown or not). Memory 607 can store any data consistent with the present disclosure. In some embodiments, memory 607 can be a combination of two or more different types of memory (e.g., a combination of semiconductor and magnetic storage) according to the specific needs, requirements or specific embodiments of computer 602, as well as the described functionality. Although a single memory 607 is illustrated in FIG. 5, two or more memories 607 (of the same type, different types, or a combination of types) can be used according to the specific needs, requirements or specific embodiments of computer 602, as well as the described functionality. Memory 607 is illustrated as an internal component of computer 602, but in alternative embodiments, memory 607 can be external to computer 602.

[0074] Application 608 can be an algorithm software engine that provides functions according to the specific needs, requirements, or specific embodiments of computer 602, as well as the described functions. For example, application 608 can serve as one or more components, modules, or applications. Further, although illustrated as a single application 608, application 608 can be implemented as multiple applications 608 on computer 602. Additionally, although illustrated inside computer 602, in alternative embodiments, application 608 can be external to computer 602.

[0075] Computer 602 can also include a power source 614. The power source 614 can include a rechargeable or non-rechargeable battery configured to be either user-exchangeable or non-user-exchangeable. In some embodiments, the power source 614 can include a power conversion and management circuit that includes recharge, standby, and power management functions. In some embodiments, the power source 614 can include a power plug that enables computer 602 to be plugged into an outlet or power source, for example, to supply power to computer 602 or recharge a rechargeable battery.

[0076] Any number of computers 602 can exist that are associated with or external to the computer system including computer 602, and each computer 602 communicates via network 603. Further, the terms "client", "user", and other appropriate terminology can be used interchangeably as appropriate without departing from the scope of this disclosure. Furthermore, this disclosure contemplates that multiple users can use one computer 602 and one user can use multiple computers 602.

[0077] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed in this specification and their structural equivalents, or in one or more combinations of them. The software embodiments of the described subject matter can be implemented as one or more computer programs. Each computer program can include one or more modules of computer program instructions encoded in a tangible, non-transitory computer-readable storage medium for execution by, or to control the operation of, a data processing apparatus. Alternatively or additionally, the program instructions can be encoded in an artificially generated propagated signal. For example, the signal can be a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to a suitable receiver device configured to be executed by a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.

[0078] The terms "data processing apparatus", "computer", and "electronic computer device" (or equivalents as understood by those skilled in the art) refer to data processing hardware. For example, a data processing apparatus can include, by way of example, programmable processors, computers, or any kind of apparatus, device, and machine that processes data, including multiple processors or computers. The apparatus can also include dedicated logic circuits, such as, for example, a central processing unit (CPU), a field programmable gate array (FPGA), or an application specific integrated circuit (ASIC). In some embodiments, the data processing apparatus or the dedicated logic circuit (or a combination of the data processing apparatus or the dedicated logic circuit) can be hardware-based, software-based (or a combination of both hardware-based and software-based). The apparatus can optionally include code that creates an execution environment for a computer program, such as code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of execution environments. The present disclosure contemplates the use of data processing apparatuses that include or do not include conventional operating systems, such as LINUX, UNIX, WINDOWS, MAC OS, ANDROID, or IOS.

[0079] A computer program, which may also be referred to as or described as a program, software, software application, module, software module, script, or code, can be written in any form of programming language. Examples of programming languages include, for example, compiled languages, interpreted languages, declarative languages, or procedural languages. The program can be deployed in any form, including stand-alone programs, modules, components, subroutines, or units used in a computing environment. A computer program can correspond to a file in a file system, but this is not essential. The program can be stored in a file that holds one or more scripts stored in a markup language document, in a single file dedicated to the program, or in multiple coordinated files that store one or more modules, subprograms, or portions of code. A computer program can be deployed to be executed on one computer or multiple computers, for example, located at one site or distributed across multiple sites interconnected by a communication network. The portions of the program shown in the various figures may be shown as individual modules that implement various configurations and functions through various objects, methods, or processes, but the program can alternatively include a number of sub-modules, third-party services, components, and libraries. Conversely, the configurations and functions of various components can be combined into a single component as appropriate. Thresholds used to make computational decisions can be determined statically, dynamically, or both statically and dynamically.

[0080] The methods, processes, or logical flows described herein can be implemented by one or more programmable computers executing one or more computer programs to manipulate input data and generate output so as to perform the functions. The methods, processes, or logical flows can also be implemented by dedicated logic circuitry, such as a CPU, FPGA, or ASIC, and the apparatus can be implemented as such.

[0081] Computers suitable for the execution of a computer program can be based on one or more of general and special purpose microprocessors, and other types of CPUs. The elements of a computer are a CPU for executing or performing instructions, and one or more memory devices for storing the instructions and data. In general, a CPU can receive (and write data to) instructions and data from memory. A computer can also include, or be operatively coupled to, one or more mass storage devices for storing data. In some embodiments, a computer can receive data from, and transfer data to, a mass storage device including, for example, a magnetic disk, magneto-optical disk, or optical disk. Further, a computer can be incorporated in another device, such as a mobile phone, a personal digital assistant (PDA), a portable audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive.

[0082] A computer-readable medium (either temporary or non-temporary as appropriate) suitable for storing computer program instructions and data can include all forms of permanent / non-permanent and volatile / non-volatile memory, media, and memory devices. The computer-readable medium can include, for example, semiconductor memory devices such as random access memory (RAM), read-only memory (ROM), phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices. The computer-readable medium can also include magnetic devices such as tapes, cartridges, cassettes, and internal / removable disks. The computer-readable medium can further include magneto-optical disks and optical memory devices and technologies such as digital video disks (DVDs), CD ROMs, DVD+ / -Rs, DVD-RAMs, DVD-ROMs, HD-DVDs, and BLURAYs. The memory can store various objects or data including caches, classes, frameworks, applications, modules, backup data, jobs, web pages, web page templates, data structures, database tables, repositories, and dynamic information. Examples of the types of objects and data stored in the memory can include parameters, variables, algorithms, instructions, rules, constraints, and references. Additionally, the memory can include logs, policies, security or access data, and report files. The processor and memory can be complemented by or incorporated into dedicated logic circuitry.

[0083] Embodiments of the subject matter described in this disclosure provide an interaction with a user that includes presenting information to the user (and receiving input from the user) on a display device It can be implemented on a computer having a chair. Examples of the type of display device include a cathode ray tube (CRT), a liquid crystal display (LCD), a light emitting diode (LED), and a plasma monitor. The display device can include a keyboard and a pointing device such as a mouse, a trackball, or a trackpad. User input can also be provided to the computer by using a touch screen, such as the surface of a tablet computer having pressure sensitivity, or a multi-touch screen using capacitance or electrical sensing. Other types of devices can be used to provide user interaction, including receiving user feedback including perceptual feedback such as visual feedback, auditory feedback, or tactile feedback. Input from the user can be received in the form of acoustic, voice, or tactile input. In addition, the computer can interact with the user by sending documents to and receiving documents from the devices used by the user. For example, the computer can send a web page to a web browser in response to receiving a request from the web browser of the user's client device.

[0084] The term "graphical user interface", i.e., "GUI", can be used in the singular or plural to refer to one or more graphical user interfaces and to each display of a particular graphical user interface. Thus, a GUI can represent any graphical user interface that processes information and efficiently presents the information results to the user, including, but not limited to, a web browser, a touch screen, or a command line interface (CLI). Generally, a GUI can include a plurality of user interface (UI) elements, some or all of which are related to a web browser, such as interactive fields, drop-down lists, and buttons. These and other UI elements can be related to or represent the functions of a web browser.

[0085] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes back-end components (such as, for example, a data server), or a computing system that includes middleware components (such as, for example, an application server). Further, the computing system can include a front-end component, for example, a client computer having one or both of a graphical user interface or a web browser through which a user can interact with the computer. The components of the system can be interconnected by any form or medium of wired or wireless digital data communication (or a combination of data communications) in a communication network. Examples of communication networks include local area network (LAN), radio access network (RAN), metropolitan area network (MAN), wide area network (WAN), worldwide interoperability for microwave access (WIMAX), wireless local area network (WLAN) (using, for example, 802.11a / b / g / n or 802.20 or a combination of protocols), all or a portion of the Internet, or any other one communication system or a plurality of communication systems (or a combination of communication networks) in one or more locations. The network can communicate, for example, in a combination of Internet protocol (IP) packets, frame relay frames, asynchronous transfer mode (AMT) cells, voice, video, data, or communication types between network addresses.

[0086] The computing system can include a client and a server. The client and server can generally be remote from each other and typically communicate over a communication ne Interactions can be made through the workpiece. The relationship between the client and the server can be brought about by computer programs that are executed on respective computers and have a client-server relationship.

[0087] A clustered file system can be of any file system type that is accessible from multiple servers for reading and updating. Since the locking of the file exchange system can be done at the application layer, locking or consistency tracking may not be necessary. Further, Unicode data files may be different from non-Unicode data files.

[0088] This specification includes details of many specific embodiments, but these should not be construed as limitations on the scope of what can be claimed, but rather as descriptions of configurations that may be specific to particular embodiments. Some of the configurations described in the context of separate embodiments herein can also be implemented in combination in a single embodiment. Conversely, the various configurations described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. Further, the configurations described above are described as acting in some combinations and may even be claimed as such initially, but one or more of the configurations from the claimed combination can, in some cases, be deleted from that combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.

[0089] In the foregoing description, embodiments of the present invention have been described with reference to numerous specific details that may vary from embodiment to embodiment. Accordingly, the description and drawings are to be regarded in an illustrative rather than a limiting sense. The sole and exclusive indicator of the scope of the present invention, and what the applicant intends to be the scope of the present invention, is the literal and equivalent scope of the set of claims from this application, including any subsequent amendments (in the specific form in which such claims become a patent). Any definitions explicitly set forth herein for terms contained in such claims shall apply to the meaning of such terms as used in the claims. Additionally, when the terms "comprising" or "including" are used in the foregoing description or the following claims, what precedes this phrasing can be additional steps or entities, or sub-steps / sub-entities of the previously recited steps or entities.

[0090] Specific embodiments of the subject matter have been described. Other embodiments, modifications, and alternative forms of the described embodiments will be apparent to those skilled in the art and are within the scope of the following claims. Although the methods of operation are shown in a specific order in the drawings or the claims, this should not be understood as requiring that such methods of operation be performed in the specific order shown or sequentially, or that all of the illustrated methods of operation be performed (some methods of operation may be considered optional). Depending on the circumstances, multitasking or parallel processing (or a combination of multitasking and parallel processing) may be advantageous and can be implemented when considered appropriate.

[0091] Furthermore, the separation or integration of the various system modules and components in the embodiments described above should not be construed as requiring such separation or integration in all embodiments, and it should be understood that the described program components and systems can generally be integrated together into a single software product or packaged into multiple software products.

[0092] Therefore, the above-described example embodiments do not define or constrain the present disclosure. Other changes, substitutions, and modifications are possible without departing from the spirit and scope of the present disclosure.

[0093] Furthermore, any claimed embodiment is considered applicable to a computer system including at least a computer-implemented method; a non-transitory computer-readable medium storing computer-readable instructions for implementing the computer-implemented method; and a computer memory operatively coupled to a hardware processor configured to execute the instructions stored in the computer-implemented method or the non-transitory computer-readable medium.

[0094] Numerous embodiments of these systems and methods have been described. Nevertheless, it will be understood that various changes may be made without departing from the spirit and scope of the present disclosure.

Claims

1. A computer-implemented method for redeveloping an existing drug, comprising: receiving, by a computer system, data representing the medical records of a plurality of patients; based on the medical records: determining at least one target signaling pathway associated with the drug; and determining one or more indicators based on one or more factors corresponding to a diagnosis associated with the target signaling pathway; thereby selecting a set of patients; determining a plurality of patient characteristics of the set of patients, each patient in the set of patients exhibiting at least one of the plurality of patient characteristics; grouping, by the computer system, the set of patients to generate a plurality of distinct groups according to the plurality of patient characteristics, each of the distinct groups including at least one patient from the set of patients; selecting, based on one or more group selection criteria, a set of distinct groups from the plurality of distinct groups; identifying one or more associated patient characteristics by analyzing each distinct group of the set of distinct groups; the method as described above.

2. Grouping a set of patients comprises: executing a machine learning system configured to execute one or more unsupervised clustering techniques the method according to claim 1.

3. The method according to claim 2, wherein the one or more unsupervised clustering techniques include a binary k-means clustering technique.

4. Grouping a set of patients includes performing multiple correspondence analysis to reduce the dimensions of the plurality of patient characteristics, the method according to any one of claims 1 to 3.

5. Selecting a set of distinct groups comprises: determining, for each distinct group of the plurality of distinct groups, a feature score for each patient characteristic exhibited by that distinct group; comparing the feature scores of each distinct group of the plurality of distinct groups with a feature score threshold; the method according to any one of claims 1 to 4.

6. Selecting a set of distinct groups comprises at least one of: determining a stability measure for each distinct group of a plurality of distinct groups; determining a purity measure for each distinct group of a plurality of distinct groups; and determining the number of patients in a set of patients included in each distinct group of a plurality of distinct groups, the method of claim 1.

7. Identifying one or more relevant patient characteristics comprises: for each distinct group of a set of distinct groups, generating a plurality of potentially relevant patient characteristics by selecting, from among a plurality of patient characteristics, those patient characteristics indicated by that distinct group and corresponding to the theme of that group; and ranking each of the potentially relevant patient characteristics of the plurality of potentially relevant patient characteristics, the method of claim 1.

8. Ranking each of the potentially relevant patient characteristics comprises, for each potentially relevant patient characteristic, assigning a rank value based on the frequency of co-occurrence of that potentially relevant patient characteristic and at least one reference indication related to the drug, the method of claim 7.

9. Co-occurrence is measured by determining, for each of the potentially relevant patient characteristics, the ratio of a set of distinct groups that includes both that potentially relevant patient characteristic and at least one reference indication, the method of claim 8.

10. Identifying one or more relevant patient characteristics comprises, for each of the potentially relevant patient characteristics: determining at least one of clinical feasibility and commercial feasibility, the method of claim 7.

11. A data processing system for redeveloping an existing drug, comprising a computer-readable memory including computer-executable instructions; and at least one processor configured to execute executable logic including the computer-executable instructions and at least one machine learning model to perform one or more operational methods, wherein the one or more operational methods comprise: receiving data representing the medical records of a plurality of patients; and based on the medical records: determining at least one target signaling pathway related to the drug; and ​ Determining one or more indicators based on one or more factors corresponding to a diagnosis associated with a target signal transduction pathway; Thereby selecting a set of patients; Determining a plurality of patient characteristics of a set of patients, wherein each patient in the set of patients exhibits at least one of the plurality of patient characteristics; Grouping a set of patients to generate a plurality of distinct groups according to a plurality of patient characteristics while using a machine learning model, wherein each of the distinct groups includes at least one patient among the set of patients; Selecting a set of distinct groups among the plurality of distinct groups based on one or more group selection criteria; Identifying one or more relevant patient characteristics by analyzing each distinct group of the set of distinct groups; The data processing system including the above.

12. The data processing system according to claim 11, wherein the machine learning model is trained to group a set of patients using one or more unsupervised clustering techniques.

13. The data processing system according to claim 12, wherein the one or more unsupervised clustering techniques include a binary k-means clustering technique.

14. The data processing system according to any one of claims 11 to 13, wherein grouping a set of patients includes performing multiple correspondence analysis to reduce the dimensions of a plurality of patient characteristics.

15. Selecting a set of distinct groups includes: For each distinct group of the plurality of distinct groups, determining a feature score for each patient characteristic indicated by the distinct group; Comparing the feature scores of each distinct group of the plurality of distinct groups with a feature score threshold; The data processing system according to any one of claims 11 to 14 including the above.

16. Selecting a set of distinct groups includes at least one of: determining a stability measure for each distinct group of the plurality of distinct groups, determining a purity measure for each distinct group of the plurality of distinct groups, and determining the number of patients in the set of patients included in each distinct group of the plurality of distinct groups. The data processing system according to any one of claims 11 to 15 including the above.

17. Identifying one or more relevant patient characteristics: For each distinct group of a set of distinct groups, generating a plurality of potentially relevant patient characteristics by selecting, for each distinct group, the patient characteristics among the plurality of patient characteristics that are indicated by that distinct group and that correspond to the theme of that group; Ranking each of the potentially relevant patient characteristics of the plurality of potentially relevant patient characteristics; The data processing system according to claim 11, comprising:

18. Ranking each of the potentially relevant patient characteristics comprises assigning a rank value to each potentially relevant patient characteristic based on the frequency of co-occurrence of that potentially relevant patient characteristic and at least one reference indication related to the drug. The data processing system according to claim 17.

19. Co-occurrence is measured by determining, for each of the potentially relevant patient characteristics, the ratio of a set of distinct groups that includes both that potentially relevant patient characteristic and at least one reference indication. The data processing system according to claim 18.

20. A computer-implemented method for redeveloping an existing drug, comprising: Receiving, by a computer system, data representing the medical records of a plurality of patients; Based on the medical records: Determining at least one target signaling pathway related to the drug; and Determining one or more indicators based on one or more factors corresponding to a diagnosis associated with the target signaling pathway; Thereby selecting a set of patients; Determining a plurality of patient characteristics of a set of patients, wherein each patient of the set of patients exhibits at least one of the plurality of patient characteristics; Grouping, by a computer system, a set of patients to generate a plurality of distinct groups according to a plurality of patient characteristics, wherein each of the distinct groups includes at least one patient of the set of patients; Selecting a set of distinct groups among the plurality of distinct groups based on one or more group selection criteria; Identifying one or more relevant patient characteristics by analyzing each of the set of distinct groups; Identifying at least one of one or more associated patient characteristics as a target indication for drug redevelopment, and The method comprising.

Citation Information

Patent Citations

  • A Bayesian Causal Network Model for Health Diagnosis and Treatment Based on Patient Data

    JP2017537365A

  • Relevance feedback to improve the performance of classification models that co-classify patients with similar profiles

    JP2019512795A

  • Method and System for Extracting Data From a Plurality of Electronic Data Stores of Patient Data to Provide Provider and Patient Data Similarity Scoring

    US20180121604A1