Data processing system and method for redeveloping existing drugs
Patent Information
- Application Number
- JP2025064018
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-09
- Filing Date
- 2025-04-09
- Publication Date
- 2026-08-27
- Estimated Expiration
- 2040-12-09
AI Technical Summary
【0010】 本開示の実施態様は、以下の利点のうちの1つまたはそれ以上を提供することができる。従来の技術と比較すると、本明細書に記載するシステムおよび方法は、既存薬物の再開発に使用するために、関連する可能性があるインジケーション(本明細書では、患者特性と称する場合がある)を特定するための異なるレベルの複雑性を有する大量のデータを処理する計算時間を短縮することにより、計算効率を向上させることができる。特定の機械学習技法を使用して、従来の技術によって特定することができない新たなインジケーションを発見することができる。従来の技術と比較すると、本明細書に記載するシステムおよび方法は、人間の入力および技能への依存を低減させることができる。
Smart Images

Figure 0007912108000002 
Figure 0007912108000003 
Figure 0007912108000004
Abstract
Description
Technical Field
[0001] Claims of Priority This application claims the benefit of European Patent Application No. 20315299.6, filed on June 9, 2020, and U.S. Provisional Patent Application No. 62 / 945,814, filed on December 9, 2019. The entire contents of the above are incorporated herein by reference.
[0002] The present disclosure generally relates to data processing systems and methods for redeveloping existing drugs.
Background Art
[0003] The redevelopment of existing clinical drugs (drug repurposing, sometimes referred to as repositioning) may refer to a strategy for drug discovery that can be relatively low-cost and provide high efficiency. The redevelopment of existing drugs usually involves analyzing whether a drug approved to treat one type of medical condition (e.g., a disease) can be used to treat other types of medical conditions (e.g., common and / or rare diseases). There are several therapeutic areas that show high potential for the redevelopment of existing drugs, including oncology, immunology, infectious diseases, and orphan diseases.
Summary of the Invention
Means for Solving the Problems
[0004] In at least one aspect of this disclosure, a data processing system is provided. The data processing system includes a computer-readable memory containing computer-executable instructions, and at least one processor configured to execute executable logic, which includes computer-executable instructions and at least one machine learning model. When at least one processor is executing computer-executable instructions, the at least one processor is configured to perform one or more acts. One or more acts include receiving data representing the medical records of multiple patients. One or more acts include selecting a set of patients based on the medical records by: determining at least one target signaling pathway associated with a drug; and determining one or more indicators based on one or more factors corresponding to a diagnosis associated with the target signaling pathway. One or more acts include determining multiple patient characteristics of the set of patients, each patient in the set exhibiting at least one of the multiple patient characteristics. One or more acts include grouping the set of patients according to the multiple patient characteristics and using a machine learning model to generate multiple distinct groups, each distinct group including at least one patient from the set of patients. One or more methods of operation include selecting a set of distinct groups from a number of distinct groups based on one or more group selection criteria. One or more methods of operation include identifying one or more relevant patient characteristics by analyzing each of the distinct groups of the set of distinct groups.
[0005] Machine learning models can be trained to group sets of patients using one or more unsupervised clustering techniques. These one or more techniques may include bisecting k-means clustering. Grouping sets of patients may involve performing multiple correspondence analysis to reduce the dimensionality of multiple patient characteristics.
[0006] Selecting a set of distinct groups may include determining a feature score for each patient characteristic represented by a distinct group within a set of distinct groups. Selecting a set of distinct groups may include comparing the feature score of each distinct group within a set of distinct groups to a feature score threshold. Selecting a set of distinct groups may include at least one of the following: determining a stability measure for each distinct group within a set of distinct groups; determining a purity measure for each distinct group within a set of distinct groups; and determining the number of patients in a set of patients included in each distinct group within a set of distinct groups.
[0007] Identifying one or more relevant patient characteristics may include generating multiple potentially relevant patient characteristics by selecting, for each distinct group of a set of distinct groups, the patient characteristics represented by that distinct group and corresponding to the theme of that group from among multiple patient characteristics. Identifying one or more relevant patient characteristics may include ranking each of the multiple potentially relevant patient characteristics. Ranking each potentially relevant patient characteristic may include assigning a rank value to each potentially relevant patient characteristic based on the frequency of co-occurrence between that potentially relevant patient characteristic and at least one reference indication related to the drug. Co-occurrence can be measured by determining, for each potentially relevant patient characteristic, the proportion of a set of distinct groups that include both that potentially relevant patient characteristic and at least one reference indication. Identifying one or more relevant patient characteristics may include determining, for each potentially relevant patient characteristic, at least one of clinical feasibility and commercial feasibility.
[0008] One or more methods of action may include identifying at least one of one or more relevant patient characteristics as a targeted indication for redeveloping an existing drug.
[0009] These and other aspects, configurations and embodiments may be expressed in other ways as methods, apparatus, systems, components, program products, means or processes for carrying out functions.
[0010] Embodiments of this disclosure may offer one or more of the following advantages: Compared to prior art, the systems and methods described herein can improve computational efficiency by reducing the computation time required to process large amounts of data with varying levels of complexity for identifying potentially relevant indications (which may be referred to herein as patient characteristics) for use in the redevelopment of existing drugs. Certain machine learning techniques can be used to discover novel indications that cannot be identified by prior art. Compared to prior art, the systems and methods described herein can reduce reliance on human input and skill.
[0011] These and other aspects, configurations and embodiments will become apparent from the following description, including the claims. [Brief explanation of the drawing]
[0012] [Figure 1] This figure shows an example of a data processing system for the redevelopment of existing drugs. [Figure 2] This flowchart shows examples of methods for redeveloping existing drugs. [Figure 3] This figure shows an experiment using the system and method described herein. [Figure 4] This figure shows experimental results obtained from using the systems and methods described herein. [Figure 5]This is a block diagram of an example computer system used to provide computational capabilities related to the algorithms, methods, functions, processes, flows, and procedures described in this disclosure. [Modes for carrying out the invention]
[0013] Redeveloping existing drugs can be used to find new clinical indications (e.g., reasons for using a particular drug) for clinically approved drugs. When hundreds of new indications are explored, key opinion leaders (KOLs) may not be able to cope with the high level of complexity associated with these new indications. Processing large amounts of data and using advanced analytics can support KOL-centric approaches. Driven by relevant clinical questions, powerful advanced analytical techniques can extract clinically relevant information hidden within large amounts of data, which can later aid in clinical decision-making.
[0014] Computational methods for redeveloping existing drugs can detect new drug-disease relationships using similarity measures (chemical similarity, molecular activity similarity, gene expression similarity, or side effect similarity), molecular docking, or shared molecular pathology. These existing drug redevelopment methods can be classified into network-based, text mining (literature search), and semantic methods.
[0015] Network-based methods can include creating integrated networks by combining multiple data sources such as drugs, proteins, genes, and diseases. For example, the Connectivity Map (C-Map) method can leverage the transcriptome by using gene expression profiling to relate biological, chemical, and clinical symptoms, thereby facilitating the discovery of new disease-gene-drug connections. Network-based clustering methods can be used to find biological modules. This method is inspired by the fact that biological entities (diseases, drugs, proteins, etc.) in the same module of a biological network typically share similar characteristics. Clustering can use the network's topological structure to find drug-disease, drug-drug, disease-disease, or drug-target relationships. Implementing clustering methods can present several challenges. For example, the edges linking drugs and diseases may depend on incomplete collected drug-disease relationships, which may require the integration of multiple databases to improve predictive accuracy. In addition, potential safety signals can be collected by incorporating different data sources that provide information on drug side effects. Furthermore, distinguishing between negative and positive associations can be difficult, and there is no "absolute" standard method for examining associations between biological modules.
[0016] Text mining-based methods can utilize keyword co-occurrence and semantic inference for novel drug-disease relationships. Some methods can be based on Swanson's analogy-based inference method (ABC model), which assumes that if "B" is one of the features of disease "C" and substance "A" affects "B", then an implicit link "A:C" can be inferred through the B connection. However, literature mining-based methods alone may have limitations due to the ambiguous nature of language, the limited scope of biomedical relationships, and the limited precision of text mining techniques.
[0017] Embodiments of the data processing systems and methods described herein utilize real-world data-driven protocols for the redevelopment of existing drugs to identify indications. This can mitigate the aforementioned disadvantages. In some embodiments, the data processing systems and methods described herein combine analytical and real-world data with KOL clinical output. In some embodiments, the data processing systems and methods described herein use real-world data and analytical methods in an unsupervised manner. In some embodiments, patients are clustered using machine learning techniques to identify groups of conditions appearing across multiple clusters. In some embodiments, the identified groups of conditions may correspond to common biological pathways. The data processing systems and methods described herein can be used to identify potential indications that were not previously identified using conventional techniques.
[0018] The following description provides numerous specific details for illustrative purposes to ensure that this disclosure is fully understood. However, it will become clear that this disclosure can be implemented without these specific details. Where else, well-known structures and devices are shown in block diagram form to avoid unnecessarily obscuring this disclosure.
[0019] In drawings, the specific arrangement or ordering of schematic elements, such as those representing devices, modules, instruction blocks, and data elements, is shown for the purpose of facilitating explanation. However, those skilled in the art will understand that the specific arrangement or ordering of schematic elements in drawings is not intended to imply that a particular order or sequence of operations or separation of processes is required. Furthermore, the inclusion of schematic elements in drawings is not intended to imply that such elements are essential in all embodiments, or that the configurations represented by such elements cannot be included in or combined with other elements in some embodiments.
[0020] Furthermore, in the drawings, when using connecting elements such as solid or dashed lines or arrows to indicate the connection, relationship or association between two or more other schematic elements, the absence of any of these connecting elements is not intended to mean that there is no possibility of a connection, relationship or association. In other words, in the drawings, some connections, relationships or associations between elements are not shown so as not to obscure the present disclosure. In addition, for ease of illustration, a single connecting element is used to represent multiple connections, relationships or associations between elements. For example, when a connecting element represents the communication of signals, data or instructions, those skilled in the art should understand that, if necessary, such elements represent one or more signal paths (e.g., buses) in order to affect the communication.
[0021] Here, reference will be made in detail to the embodiments illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various described embodiments. However, it will be apparent to those skilled in the art that the various described embodiments can be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.
[0022] Several configurations will be described below that can each be used independently of one another or in any combination with other configurations. However, any individual configuration may not address any of the above-described problems, or may only address one of the above-described problems. Some of the above-described problems may not be fully addressed by any of the configurations described herein. Although headings are provided, data related to a particular heading but not found in the section having that heading may also be found elsewhere in this specification.
[0023] Examples of Data Processing Systems and Methods Figure 1 shows an example of a data processing system 100. In some embodiments, the data processing system 100 is configured to process data that can represent the medical records of multiple patients to identify new indications for drugs (existing drugs to be redevelopment). The system 100 includes a computer processor 110. The computer processor 110 includes a computer-readable memory 111 and computer-readable instructions 112. The system 100 also includes a machine learning system 150. The machine learning system 150 includes a machine learning model 120. The machine learning model 120 can be separated from the computer processor 110 or integrated with the computer processor 110.
[0024] The computer-readable medium 111 (or computer-readable memory) can include any type of data storage technology suitable for the local technical environment, including but not limited to semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, removable memory, disk memory, flash memory, dynamic random access memory (DRAM), static random access memory (SRAM), electronically erasable programmable read-only memory (EEPROM), etc. In some embodiments, the computer-readable medium 111 includes a code segment having executable instructions.
[0025] In some embodiments, the computer processor 110 includes a general-purpose processor. In some embodiments, the computer processor 110 includes a central processing unit (CPU). In some embodiments, the computer processor 110 includes at least one application-specific integrated circuit (ASIC). The computer processor 110 may also include a general-purpose programmable microprocessor, a graphics processing unit, a dedicated programmable microprocessor, a digital signal processor (DSP), a programmable logic array (PLA), a field-programmable gate array (FPGA), dedicated electronic circuits, etc., or a combination thereof. The computer processor 110 is configured to execute program code such as computer executable instructions 112, and is also configured to execute executable logic including a machine learning model 120.
[0026] The computer processor 110 is configured to receive data representing the medical records of multiple patients. For example, the computer processor 110 can receive data from a database containing electronic medical records (EMRs) for approximately 94 million (or more) patients, identifiable by key identifiers (IDs) that enable patient matching across different data tables. In some embodiments, the data may include diagnoses, clinical tests, treatments, medications, patient events, insurance, biomarkers, measurements, clinical status, lifestyle parameters, microbiology, prescriptions, etc. In some embodiments, the data may include natural language processing-driven data. The data can be received through any of a variety of techniques, such as wireless communication, fiber optic communication, USB, CD-ROM, etc.
[0027] The machine learning system 150 can apply machine learning techniques to train the machine learning model 120. As part of training the machine learning model 120, the machine learning system 150 can form a training set of input data by identifying a positive training set of input data items that are determined to have the target characteristics, and in some embodiments, it can form a negative training set of input data items that do not have the target characteristics.
[0028] The machine learning system 150 extracts feature values from the input data of the training set, and these features are variables that are considered to be related to whether or not an input data item has one or more related characteristics. In this specification, an ordered list of features for the input data is referred to as the feature vector of the input data. In some embodiments, the machine learning system 150 applies dimensionality reduction to reduce the amount of data in the feature vector of the input data. Next, the data set is reduced to a more representative set. For example, the machine learning system 150 can apply multiple correspondence analysis (MCA), linear discriminant analysis (LDA), principal component analysis (PCA), etc.
[0029] In some embodiments, the machine learning system 150 uses unsupervised machine learning to train the machine learning model 120. Typically, unsupervised machine learning techniques make inferences from a dataset using input vectors without referring to known, i.e., labeled results. In some embodiments, the machine learning system 150 can perform clustering to divide the data points into multiple groups such that data points within the same group are more similar to other data points within the same group and less similar to data points in other groups. In some embodiments, the clustering includes performing K-means clustering, where a one-level non-nested partition of data points is created by iteratively partitioning the dataset. That is, if K is the desired number of clusters, in each iteration the dataset is partitioned into K independent clusters. This process can continue until a specified clustering criterion function value is optimized. In some embodiments, the machine learning system 150 is configured to perform bipartite K-means clustering. Bipartite K-means clustering typically involves partitioning one cluster into two subclusters (e.g., by using k-means) at each bipartite step until k clusters are obtained. Bipartite K-means clustering may be more beneficial than K-means clustering because it can reduce computation time when K is a relatively large value, generate clusters of similar size, and generate clusters with lower entropy.
[0030] The computer processor 110 is configured to execute computer executable instructions 112 to perform one or more methods of operation. In some embodiments, one or more methods of operation include receiving data representing the medical records of multiple patients. For example, the computer processor 110 can receive data from a database containing electronic medical records (EMRs) of approximately 94 million (or more) patients, identifiable by key identifiers (IDs) that enable patient matching across different data tables. In some embodiments, the data may include diagnoses, clinical tests, treatments, medications, patient events, insurance, biomarkers, measurements, clinical status, lifestyle parameters, microbiology, prescriptions, etc. In some embodiments, the data may include natural language processing-driven data. The data may be received through any of a variety of techniques, such as wireless communication, fiber optic communication, USB, CD-ROM, etc.
[0031] In some embodiments, one or more operating methods include selecting a set of patients based on medical records. The selection of a set of patients includes determining at least one target signaling pathway associated with the existing drug to be redeveloped. For example, if the existing drug to be redeveloped is dupilumab, the computer processor 110 may determine, based on the drug's known function, that the drug modulates the interleukin-4 (IL-4) and interleukin-13 (IL-13) signaling pathways. The selection of a set of patients also includes determining one or more indicators based on one or more factors corresponding to a diagnosis associated with the target signaling pathway. For example, using factors such as pathway mechanisms, associated clinical symptoms, therapeutic analogues, data and epidemiology, and the coherence of drug lifecycle management, sources including medical databases and medical evidence software can be searched to identify diseases associated with the determined signaling pathway. These diseases can be categorized based on the strength of the link to the determined signaling pathway. Categories include focused, medium, and broad groups. This can be done. For example, returning to the IL-4 / IL-13 example, the intensive group of diseases may include diseases directly related to the IL-4 / IL-13 mechanism of action on the Th2 pathway, the intermediate group of diseases may include diseases indirectly related to the IL-4 / IL-13 mechanism of action on the Th2 pathway, and the broad group of diseases may include diseases associated with a broader inflammatory response. By moving from the intensive group to the broad group, the number of indicators to consider when selecting a set of patients can be increased and the potential influence of molecules can be reduced. Therefore, in some embodiments, only the intensive group, or the intensive and intermediate groups, are used to select a set of patients. In some embodiments, only patients with at least one diagnosis, drug, clinical test and / or treatment related to the determined signaling pathway are selected so as to be included in a set of patients. Detailed examples of factors and indicators will be provided later with reference to Table 1.
[0032] In some embodiments, one or more methods of operation include determining multiple patient characteristics (which may be referred to herein as characteristics) of a set of patients, where each patient in the set exhibits at least one of the multiple patient characteristics. Determining multiple patient characteristics may include analyzing initially received data to identify broader patient characteristics for incorporating all or substantial portion of the received data. For example, broader patient characteristics may correspond to diagnoses (e.g., immune disorders, diabetes), prescriptions (e.g., immunotherapy, other drug classes), treatments (e.g., human leukocyte antigen typing), and laboratory results (e.g., abnormally high / low IgE). In some embodiments, determining multiple patient characteristics may include receiving user input (e.g., via a user interface communicating with a computer processor 110). For example, a user may input patient characteristics based on clinical input, demographics, medications, comorbidities, treatments, and immunologically specific clinical laboratory data. Customized characteristic classes may also be added to improve data integrity and representativeness, and to collect more information about disease and drug response. In some embodiments, determining multiple patient characteristics involves validating multiple patient characteristics. Validation may include calculating the proportion of selected patients having at least one characteristic family (e.g., the proportion of patients with prescription records) and comparing this proportion to the proportion of patients in the initially received data having at least one characteristic family, thereby determining whether the patient characteristics in the initially received data accurately map to a selected set of patients. Closer values for the two numbers indicate a more accurate mapping. Validation may also include identifying multiple patients present in both the initially received data and the selected set of patients and verifying identical mappings of patient characteristics between the patients in the initially received data and the selected set of patients, thereby determining whether the patient characteristics accurately map to the patients.
[0033] In some embodiments, one or more methods of operation include grouping a set of patients to generate multiple distinct groups, each containing at least one patient from a set, according to multiple patient characteristics (for example, defined by features related to a determined signaling pathway). For example, one or more computer processors 110 can run a machine learning model 120 to implement a clustering technique such as the bipartite k-means clustering technique described above. Clustering can result in multiple clusters (e.g., distinct groups) of patients, where patients in one cluster are more similar to each other with respect to their corresponding patient characteristics than patients in other clusters. In some embodiments, the generated clusters may show correlations between patient characteristics, even if they do not exist in the same patient. Clinical input can be received and used at various stages of the clustering process to ensure the clinical relevance of the resulting clusters. For example, clinical input from a disease expert can be used in the inclusion and grouping of clinically relevant features, and in the validation of the clusters. In evaluation, it is possible to facilitate the creation of clinically relevant cohorts. Patient characteristics can be identified as distinctive in a cluster if they occur more frequently in the general population (e.g., overall in a selected group of patients) than in the general population.
[0034] In some embodiments, multiple correspondence analysis (MCA) is used to reduce the dimensionality of patient characteristics. Bipartite K-means facilitate appropriate and effective separation of patients into sufficiently "tight" but stable clusters, and the numerous clusters showing immune relevance can be used to score patient characteristics (described in more detail later). The resulting clusters can be presented to users (e.g., clinical professionals) (e.g., via a user interface) for validation and evaluation. This reduces the risk of cluster noninterpretation and ensures that there are no overlapping features between different clusters.
[0035] In some embodiments, one or more methods of operation include selecting a set of distinct groups from a plurality of distinct groups based on one or more group selection criteria. In some embodiments, selecting a set of distinct groups includes ranking the groups and selecting a number of the highest-ranked groups (e.g., the top 60 ranked groups). Groups may be ranked based on immunological richness, stability, purity, and size. In some embodiments, one or more measures (which may be referred to herein as feature scores) are calculated for each patient trait to rank clusters. These one or more measures may include, for example, specificity (which may be referred to herein as a “lift score”), the number of patients in the cluster that present the patient trait, and an immunological score. The specificity score measures how distinctive the patient trait is to the rest of the population within the cluster (for example, if males make up 50% of the population and 75% of the cluster, the “lift score” may be equal to 1.5). In some embodiments, only patient characteristics that have a lift score exceeding a threshold lift score (e.g., 1) and appear in a proportion of patients exceeding a threshold proportion (e.g., 10%) are considered to define a cluster and correspond to the theme of the cluster. Patient characteristics considered to define a cluster may be referred to herein as potentially relevant patient characteristics. Next, an immunological score can be assigned to the patient characteristics (e.g., any of the patient characteristics considered to define a cluster, or any of the patient characteristics), which scores the patient characteristics according to their type (e.g., disease, drug, clinical test, procedure, etc.) and immunological relevance. The patient characteristic scores within each cluster can then be aggregated (e.g., summed) and normalized. Clusters that satisfy a threshold cluster score (e.g., 50%) can then be considered immunologically specific.
[0036] Selecting a set of distinct groups may involve evaluating one or more of the following: stability, purity, and the number of patients in each cluster. Stability can be evaluated using one or more of the following methods: (1) replicating clusters with different sized data; (2) changing the cluster initialization seed; (3) changing the number of clusters generated; and (4) applying a training-test method. For each cluster in the training set, stability can be defined as the maximum proportion of patients who are also grouped together in the test set. Purity can be measured by the intra-cluster variance of the MCA component of patients within a cluster, which can result in homogeneous and dense clusters. In some embodiments, clusters are selected if they exceed both a threshold stability percentage (e.g., 50%) and a threshold purity percentage (e.g., the cluster is at most 20% purity of all clusters).
[0037] In some embodiments, one or more operating methods are a set of separate groups This involves identifying one or more relevant patient characteristics (e.g., indications) by analyzing each distinct group. Identifying one or more relevant patient characteristics may include ranking the patient characteristics presented by each selected cluster (e.g., all patient characteristics, or those considered to define the cluster). Ranking may be based on the frequency of co-occurrence with each of the numerous established (reference) characteristics of the existing drug being redeveloped (e.g., if the drug is dupilumab, the reference characteristics may include asthma, atopic dermatitis, IgE allergy, and composite immunological score). Co-occurrence can be measured by calculating the proportion of patient-weighted clusters that include both patient characteristics and references. In some embodiments, one or more patient characteristics that a content domain expert has determined to be relevant to a core cluster theme (e.g., indicated by user input received via a user interface) may also be considered for evaluation, regardless of the number of patients in which these characteristics appear (which may be less than 10%).
[0038] Identifying one or more patient characteristics may include evaluating the clinical and commercial feasibility of those characteristics. For example, patient characteristics that indicate a distinct clinical diagnosis may be identified. The commercial evaluation may be based on available sales forecasts and data showing competitor assets, determined links to target signaling pathways (whether or not they are documented in publications), global morbidity of the patient characteristics, and disability-adjusted life years (DALYs) of the patient characteristics (e.g., per 100,000 life years). In some embodiments, at least one of the one or more patient characteristics may be identified as a targeted indication for redeveloping an existing drug. As a result, in some embodiments, one or more methods of action generally produce one or more new indications for the existing drug being redeveloped.
[0039] In general, this specification describes patients as human patients, but embodiments are not limited thereto. For example, patients may refer to non-human animals, plants, or human replica systems.
[0040] For illustrative purposes, this specification states that data will be received for approximately 94 million patients, but it should be understood that the data may be for fewer or more patients.
[0041] For illustrative purposes, this specification describes dupilumab as an existing drug to be redeveloped, but it should be understood that the existing drug to be redeveloped could be any drug.
[0042] Figure 2 is a flowchart illustrating an example method 200 for redeveloping an existing drug. Method 200 can be carried out by the data processing system 100 described above with reference to Figure 1. Method 200 includes receiving data representing a medical record (block 210), selecting a set of patients (block 220), determining several patient characteristics (block 230), grouping the set of patients to generate several distinct groups (block 240), selecting a set of distinct groups (block 250), and identifying one or more relevant patient characteristics (block 260).
[0043] In block 210, data representing the medical records of multiple patients is received. The data can be received from a database containing EMRs of approximately 94 million (or more) patients, identifiable by key IDs that enable patient matching across different data tables. In some embodiments, the data may include diagnoses, clinical tests, procedures, medications, patient events, insurance, biomarkers, measurements, clinical status, lifestyle parameters, microbiology, prescriptions, medical images, etc. In some embodiments, the data may include: It includes data driven by natural language processing. The data can be received via any of the following techniques: wireless communication, fiber optic communication, USB, CD-ROM, etc.
[0044] In block 220, at least one target signaling pathway associated with the existing drug to be redeveloped is determined. For example, if the drug is dupilumab, based on the drug's known function, it may be determined that the drug modulates the interleukin-4 (IL-4) and interleukin-13 (IL-13) signaling pathways. In some embodiments, one or more indicators can be determined based on one or more factors corresponding to a diagnosis associated with the target signaling pathway. For example, sources, including medical databases and medical evidence software, can be searched using factors such as pathway mechanism, associated clinical symptoms, therapeutic analogues, data and epidemiology, and the coherence of drug lifecycle management to identify diseases associated with the determined signaling pathway. These diseases can be categorized based on the strength of the link to the determined signaling pathway. Categories may include intensive, intermediate, and broad groups. For example, returning to the IL-4 / IL-13 example, the intensive group of diseases may include diseases directly related to the IL-4 / IL-13 mechanism of action on the Th2 pathway, the intermediate group of diseases may include diseases indirectly related to the IL-4 / IL-13 mechanism of action on the Th2 pathway, and the broad group of diseases may include diseases associated with a broader inflammatory response. By moving from the intensive group to the broad group, the number of indicators to consider when selecting a set of patients can be increased and the potential influence of molecules can be reduced. Therefore, in some embodiments, either only the intensive group or both the intensive and intermediate groups are used to select a set of patients. In some embodiments, only patients with at least one diagnosis, drug, clinical test and / or treatment related to the determined signaling pathway are selected so as to be included in a set of patients.
[0045] In block 230, multiple patient characteristics are determined for a set of patients, where each patient in the set exhibits at least one of the multiple patient characteristics. Determining multiple patient characteristics may involve analyzing the initially received data to identify broader patient characteristics to incorporate all or substantial portions of the received data. For example, broader patient characteristics may correspond to diagnoses (e.g., immune disorders, diabetes), prescriptions (e.g., immunotherapy, other drug classes), treatments (e.g., human leukocyte antigen typing), and laboratory results (e.g., abnormally high / low IgE). In some embodiments, determining multiple patient characteristics may involve receiving user input (e.g., via a user interface). For example, a user may input patient characteristics based on clinical input and demographic, medication, comorbidities, treatments, and immunology-specific clinical laboratory data. Customized characteristic classes may also be added to improve data integrity and representativeness, and to collect more information about disease and drug response. In some embodiments, determining multiple patient characteristics may involve validating the multiple patient characteristics. Validation may include determining whether the patient characteristics in the initial data accurately map to a selected set of patients by calculating the proportion of selected patients who have at least one characteristic family (e.g., the proportion of patients with prescription records) and comparing this proportion to the proportion of patients in the initial data who have at least one characteristic family. The closer the two numbers are, the more accurate the mapping. Validation may also include determining whether the patient characteristics accurately map to patients by identifying multiple patients who are included in both the initial data and the selected set of patients and verifying the identical mapping of patient characteristics between the patients in the initial data and the selected set of patients.
[0046] In block 240, (for example, by features related to the determined signaling pathway) As defined, a set of patients is grouped according to multiple patient characteristics, generating multiple distinct groups, each containing at least one patient from the set. For example, a clustering technique such as the bipartite k-means clustering technique described above can be performed using multiple patient characteristics for a set of patients. Clustering can result in multiple clusters (e.g., distinct groups) of patients where patients in one cluster are more similar to each other with respect to their corresponding patient characteristics than patients in other clusters. In some embodiments, the generated clusters may show correlations between patient characteristics even if they do not exist in the same patient. Clinical input can be received and used at various stages of the clustering process to ensure the clinical relevance of the resulting clusters. For example, clinical input from disease experts can facilitate the creation of clinically relevant cohorts in the inclusion and grouping of clinically relevant features and in the validation and evaluation of clusters. A patient characteristic can be identified as distinctive in a cluster if it occurs more frequently in the general population (e.g., overall in a selected set of patients) than in the general population.
[0047] In some embodiments, multiple correspondence analysis (MCA) is used to reduce the dimensionality of patient characteristics. Bipartite K-means facilitate appropriate and effective separation of patients into sufficiently "tight" but stable clusters, and the numerous clusters showing immune relevance can be used to score patient characteristics (described in more detail later). The resulting clusters can be presented to users (e.g., clinical professionals) (e.g., via a user interface) for validation and evaluation. This reduces the risk of cluster noninterpretation and ensures that there are no overlapping features between different clusters.
[0048] In block 250, a set of distinct groups is selected from a plurality of distinct groups based on one or more group selection criteria. In some embodiments, selecting a set of distinct groups involves ranking the groups and selecting a number of the highest-ranked groups (e.g., the top 60 ranked groups). Groups can be ranked based on immunological richness, stability, purity, and size. In some embodiments, one or more measures are calculated for each patient trait to rank clusters. These one or more measures may include, for example, specificity (which may be referred to herein as the “lift score”), the number of patients in the cluster that present the patient trait, and the immunological score. The specificity score measures how distinctive the patient trait is to the rest of the population within the cluster (for example, if males make up 50% of the population and 75% of the cluster, the “lift score” may be equal to 1.5). In some embodiments, only patient characteristics that have a lift score exceeding a threshold lift score (e.g., 1) and appear in a proportion of patients exceeding a threshold proportion (e.g., 10%) are considered to define a cluster and correspond to the theme of the cluster. Patient characteristics considered to define a cluster may be referred to herein as potentially relevant patient characteristics. Next, an immunological score can be assigned to the patient characteristics (e.g., any of the patient characteristics considered for defining a cluster, or any of the patient characteristics), which scores the patient characteristics according to their type (e.g., disease, drug, clinical test, procedure, etc.) and immunological relevance. The patient characteristic scores within each cluster can then be aggregated (e.g., summed) and normalized. Clusters that satisfy a threshold cluster score (e.g., 50%) can then be considered immunologically specific.
[0049] Selecting a set of distinct groups may involve evaluating one or more of the following: stability, purity, and the number of patients in each cluster. Stability can be assessed by: (1) replicating the clusters with different data sizes; and (2) initializing the clusters. The results can be evaluated by one or more of the following: (3) changing the set; (4) changing the number of clusters generated; and (5) applying the training-test method. For each cluster in the training set, stability can be defined as the maximum proportion of patients who are also grouped together in the test set. Purity can be measured by the intra-cluster variance of the MCA component of patients within a cluster, which can result in homogeneous and dense clusters. In some embodiments, a cluster is selected if it exceeds both a threshold stability percentage (e.g., 50%) and a threshold purity percentage (e.g., the cluster is at most 20% purity of all clusters).
[0050] In block 260, one or more relevant patient characteristics are identified by analyzing each distinct group of a set of distinct groups. Identifying one or more relevant patient characteristics may include ranking the patient characteristics presented by each selected cluster (e.g., all patient characteristics, or patient characteristics considered to define the cluster). Ranking may be based on the frequency of co-occurrence with each of the numerous established (reference) characteristics of the existing drug being redeveloped (e.g., if the drug is dupilumab, the reference characteristics may include asthma, atopic dermatitis, IgE allergy, and composite immunological score). Co-occurrence can be measured by calculating the proportion of patient-weighted clusters that include both patient characteristics and references. In some embodiments, one or more patient characteristics that a content domain expert has determined to be relevant to a core cluster theme (e.g., as indicated by user input received via a user interface) may also be considered for evaluation, regardless of the number of patients in which these characteristics appear (which may be less than 10%).
[0051] Identifying one or more patient characteristics may include evaluating the clinical and commercial feasibility of those characteristics. For example, patient characteristics that indicate a distinct clinical diagnosis may be identified. The commercial evaluation may be based on available sales forecasts and data showing competitor assets, determined links to target signaling pathways (whether or not they are documented in publications), global morbidity of the patient characteristics, and disability-adjusted life years (DALYs) of the patient characteristics (e.g., per 100,000 life years). In some embodiments, at least one of the one or more patient characteristics may be identified as a targeted indication for redeveloping existing drugs.
[0052] Experimental results Figure 3 shows an experiment using the system and methods described herein. This experiment was conducted to identify novel indications for dupilumab, an anti-IL4 / IL13 drug, and to validate a real-world development (RWD) driven protocol for the redevelopment of existing drugs. One objective of this experiment was to reduce drug development costs and shorten time to market while minimizing dropouts and risks. A hybrid approach was applied, utilizing scientific and clinical capabilities through analysis combined with KOL expertise, commercial evaluation, and real-world data.
[0053] Data Source: The Optum Humedica dataset from 2014 to 2018 was used. This database contained electronic medical records for 94 million patients, identifiable by key identifiers, enabling patient matching across different data tables. The database collected information on EMR data, including diagnosis, clinical tests, procedures, medications, patient events, insurance, biomarkers, measurements, clinical status and lifestyle parameters, microbiology, and prescriptions. Natural language processing (NLP) driven tables were not included due to their limited scope and clinical relevance. Furthermore, data tables containing incomplete or irrelevant information were excluded. A total of five data tables were included, reducing the data source to 40 million patients.
[0054] Patient selection: Indicators for patient selection were based on a clinical framework related to the underlying immunological pathways, as shown in Table 1.
[0055] [Table 1]
[0056] We selected only adult patients (18 years of age or older) who had received a diagnosis related to the IL4 / 13 pathway (i.e., a signaling pathway) and who had undergone at least one diagnosis, medication, clinical test, and procedure. Using the criteria and data completeness of these immunological conditions, we obtained a cohort consisting of 17 million unique patients.
[0057] To identify indicators, five factors (as shown in Table 1) were considered. These factors consisted of several platforms (DOC Library, DOC Data, Doctor Evidence, DOC Label, DOC Search) and were searched through sources including PubMed, ClinicalTrials.gov, WHO, and the Doctor Evidence engine database, a medical evidence software and service company. The information within these factors was then classified according to three lenses based on Th2 responses: intensive, intermediate, and broad. For example, diseases were assigned to the intensive or intermediate lens based on their direct and indirect involvement in the mechanism of action of IL4 / IL13 in the Th2 pathway, respectively, and to the broad lens if they were associated with a broader inflammatory response. Moving from the intensive lens to the broad lens broadened the range of indicators to consider and reduced the potential for molecular influence. Dicaters were ultimately excluded from the analysis because their mechanistic links were not specific enough to meet the criteria for identifying patient populations with similar characteristics. Therefore, only intensive and intermediate lens indicators were included in the experiment. A final list of 208 indicators was included for analysis across 17 broad systems.
[0058] Feature Selection: A broad range of features (patient characteristics) were selected to capture the information available in the Optum dataset; these features were then prioritized and validated by clinical experts to ensure that all essential variables were included and that the variable values were clinically meaningful. Where appropriate, some features were retained, while others (certain demographics) were newly created (as shown by V1, as described later in Figure 4). New features were added based on clinical input and demographics, medications, comorbidities, treatments, and immunology-specific clinical laboratory data. Customized feature classes were generated and iteratively added to the analysis (as shown by V2 and V3, as described later in Figure 4) to improve data completeness and representativeness, and to collect more information about the severity of disease and drug responses.
[0059] Robust methods were used to ensure the completeness of features in the final database. Two verification steps were performed based on the mapping of patients and features across the Optum database and the generated tables to verify that the features were accurately generated. First, to confirm whether Optum patient features were accurately mapped to the data tables of the present invention, the proportion of patients having at least one feature family was calculated and it was determined whether this number was the same in Optum Humedica and the data tables of the present invention. Second, to confirm whether the features were mapped to the correct patients, 10 random patients were tracked from the raw Optum Humedica data to the generated dataset to ensure identical mapping of features in the two datasets. After accurate mapping was demonstrated by feature verification, the algorithm was run on 17 million patients containing 2700 features.
[0060] Clustering: Using clustering techniques, patients sharing similar characteristics, such as those defined by features related to the IL4 / 13 pathway, were grouped together. Clustering sought similarities between patients based on their characteristics. The generated clusters found correlations between symptoms, even if they were not present in the same patient. To ensure the clinical relevance of the results, clinical input was incorporated at various stages of the process. Thus, clinical input from disease experts helped in creating clinically relevant cohorts, inclusion and grouping of clinically relevant features, and finally, validation and evaluation of the clusters. Features were identified as characteristic within a cluster if they occurred more frequently than in the general population.
[0061] Multiple correspondence analysis (MCA) was used to reduce the dimensionality of the features. Then, using bipartite K-means, the data was divided into 500 clusters, providing appropriate and effective separation of patients in sufficiently "tight" but stable clusters, allowing a large number of clusters demonstrating immune relevance to be used for indication scoring. Clinical experts validated and evaluated the clusters identified throughout this process. This step facilitated the reduction of the risk of cluster non-interpretation and ensured that there were no overlapping features between different clusters. This clustering method was performed on data representing 2700 features (e.g., MCA components) and 17 million patients. The final number of clusters generated by this algorithm was 500.
[0062] Identifying new indications (i.e., relevant patient characteristics): For further evaluation, conduct clinical and commercial judgments to obtain a short list of priority signals, and class Based on the stan output, the most clinically relevant indications across clusters were identified. Four methodological steps were used. In the first step, the top 60 clusters were selected from 500 ranked clusters based on immunological richness, stability, purity, and size. Three measures were calculated for the features contained in each cluster: specificity, the number of patients exhibiting that feature within each cluster, and an immunological score to determine the selection. Specificity, also referred to as the “lift score,” measured how distinctive a feature is within the cluster compared to the rest of the population (for example, if males make up 50% of the population and 75% of the cluster, the “lift score” is equal to 1.5). Only features with a lift score greater than 1 (meaning they occur within the cluster at a surprisingly high rate compared to the broader population) and appearing in more than 10% of patients were considered to define (and name) the cluster and develop the cluster's theme. In addition, each selected feature was assigned a score according to its type (disease, drug, clinical test, or procedure) and immunological relevance. The feature scores within each cluster were summed and normalized. Next, a cluster was considered immunologically unique if it met a predefined threshold of a 50% score. In the second step, clusters were selected according to stability, purity, and the number of patients. Stability was assessed using four methods: 1. replicating clusters with different data sizes, 2. changing the cluster initialization seed, 3. changing the number of generated clusters, and 4. applying a training-test method. For each cluster in the training set, stability was defined as the maximum proportion of patients who were also grouped together in the test set. Purity could be measured by the intra-cluster variance of the MCA component of patients within that cluster, which could result in homogeneous and high-density clusters. Clusters were included in the analysis if they had stability above 50% and purity at a maximum of 20%. In addition, all indications that content area experts deemed relevant to the core cluster theme were considered for evaluation, regardless of the number of patients in which these features appeared (which may be less than 10%).In the third step, these new indications were ranked based on their frequency of co-occurrence with each of the four established (reference) indications for dupilumab (asthma, atopic dermatitis, IgE allergy, and composite immunological score). Co-occurrence was measured by calculating the proportion of patient-weighted clusters that included both the indication and the reference. In the final step, the final list of indications was further characterized by clinical and commercial feasibility. Clinical evaluation retained indications that indicated distinct clinical diagnoses. Based on ranking with IL4 / IL13 reference, conditions that did not appear in the top 30 were removed as they were considered to have a low relevance to IL4 / IL13 regulation. Commercial evaluation was possible only for a subset of indications that appeared clinically relevant, where data on sales forecasts and competitor assets were available. In addition, for commercial evaluation, several factors were also considered: links to the IL4 / 13 pathway, whether or not they are documented in the literature; the global incidence of indications; and disability-adjusted life years (DALYs) of indications (per 100,000 survival years).
[0063] Results: Figure 4 shows the experimental results obtained from the use of the systems and methods described herein. The final cohort of 17 million patients extracted from Optum Humedica was analyzed to assess the completeness and representativeness of the data across selected features. In the median lens population, the patient cohort consisted of 59% females, the majority Caucasian (77%), the mean age at last activity was 53 years (SD=7 years), and the mean follow-up period was 7.1 years. Patients most frequently presented with ICD-10 codes for acute sinusitis (25.2%), allergic rhinitis of unknown origin (20.6%), and other asthma of unknown origin (19.4%) as immunological conditions. The most frequently used immunological-related drugs were prednisone (28.0%), fluticasone furoate (22.2%), and methylprednisolone (13.3%), and 0.4% of patients received allergen immunotherapy injections and β2 glycoprotein antibody measurements. The majority of patients had a high white blood cell count (70.8%) and were ill. Patients underwent testing for neutrophil count (ANC) (64.3%) and absolute lymphocyte count (ALC) (63.8%). A clustering procedure generated 500 clusters, of which 125 were classified as robust and stable for immune disorders. Of these 125, 110 were also classified as high purity. Of these, 60 clusters containing the largest number of patients per cluster were retained and analyzed for clinically relevant signals. Following a confirmation process, a training-test method was used to consider 84% of the clusters as highly stable, and reproduced 90% and 99%, respectively, of the top 20 clusters, regardless of seed position in the data table and the number of patients. Based on the features contained in the clusters, six cluster themes were identified: multi-organ immunological effects, neoplasms, asthma and other hypersensitivity, musculoskeletal dysfunction, cardiac metabolic spectrum, and gynecological and obstetric conditions, which were also partially reproduced in V2 and V3 replicates. Of these, 250 indications were selected for cluster evaluation and ranked by co-occurrence with each reference. Clinical and commercial feasibility further characterized the list of approximately 85 indications: about 20 of these did not represent a distinct clinical diagnosis or lacked sufficient clinical evidence for IL4 / 13 regulation, while others lacked readily available commercial evaluation information. The final list of indications from the hybrid approach identified approximately 90% already in lifecycle management, along with approximately 60% of additional potential new indications.
[0064] Figure 5 is a block diagram of Example Computer System 600 used to provide computational functions related to the algorithms, methods, functions, processes, flows, and procedures described in this disclosure (such as Method 200 described above with reference to Figure 2), according to several embodiments of this disclosure. The illustrated computer 602 is intended to encompass any computing device, including a server, desktop computer, laptop / notebook computer, wireless data port, smartphone, personal digital assistant (PDA), tablet computing device, or one or more processors within such devices, including physical instances, virtual instances, or both. Computer 602 may include input devices such as keypads, keyboards, and touchscreens that can accept user information. Computer 602 may also include output devices that can transmit information related to how computer 602 operates. The information may include digital data, visual data, audio information, or a combination of information. The information may be presented in a graphical user interface (UI) (or GUI).
[0065] Computer 602 can function as a client, network component, server, database, persistence, or component of a computer system that implements the subject matter described herein. The illustrated computer 602 is connected to network 630 in a communicative manner. In some embodiments, one or more components of computer 602 can be configured to operate in different environments, including cloud computing-based environments, local environments, global environments, and combinations of environments.
[0066] Broadly speaking, computer 602 is an electronic computing device capable of receiving, transmitting, processing, storing, and managing data and information related to the subject matter described. According to some embodiments, computer 602 may also include, or be communicably linked to, an application server, an email server, a web server, a cache server, a streaming data server, or a combination of such servers.
[0067] Computer 602 can receive requests via network 630 from a client application (for example, running on another computer 602). Computer 602 can respond to incoming requests by processing them using software applications. Requests can also be sent to computer 602 from internal users (e.g., from a command console), external (or third parties), automation applications, entities, individuals, systems, and computers.
[0068] Each component of computer 602 can communicate using system bus 603. In some embodiments, any or all of the components of computer 602, including hardware or software components, can interface with each other or with interface 604 (or a combination of both) via system bus 603. The interface can be an application programming interface (API) 612, a service layer 613, or a combination of API 612 and service layer 613. API 612 can include specifications for routines, data structures, and object classes. API 612 may or may not be dependent on a computer language. API 612 can refer to a complete interface, a single function, or a set of APIs.
[0069] The service layer 613 can provide software services to computer 602 and other components (whether or not shown) that are communicatively connected to computer 602. The functionality of computer 602 may be accessible to all service consumers using this service layer. Software services, such as those provided by service layer 613, can provide reusable, defined functionality through a defined interface. For example, the interface may be software written in a language that provides data in Java, C++, or Extensible Markup Language (XML) format. Although illustrated as an integrated component of computer 602, in alternative embodiments, API 612 or service layer 613 may be a standalone component in relation to other components of computer 602 and other components communicatively connected to computer 602. Furthermore, any or all parts of API 612 or service layer 613 may be implemented as a child or submodule of another software module, enterprise application, or hardware module without departing from the scope of this disclosure.
[0070] Computer 602 includes interface 604. While illustrated as a single interface 604 in Figure 5, two or more interfaces 604 may be used depending on the specific needs, requirements, or embodiments of computer 602, as well as the functions described. Interface 604 can be used by computer 602 to communicate with other systems connected to network 603 in a distributed environment (whether or not illustrated). Generally, interface 604 may include, or be implemented using, software or hardware (or a combination of software and hardware) encoded logic capable of communicating with network 630. More specifically, interface 604 may include software supporting one or more communication protocols related to communication. Thus, network 630 or interface hardware may be capable of communicating physical signals inside or outside computer 602 as illustrated.
[0071] Computer 602 includes a processor 605. Although it is illustrated as a single processor 605 in Figure 5, two or more processors 605 can be used depending on the specific needs, requirements or specific embodiments of computer 602 and the functions described. Generally, a processor 605 can execute instructions and manipulate data to the present disclosure. The computer 602 can be operated in a manner that includes methods of operation using algorithms, methods, functions, processes, flows, and procedures as described.
[0072] Computer 602 also includes a database 606 that can hold data for computer 602 and other components connected to the network 630 (whether or not shown). For example, database 606 may be an in-memory, conventional, or data-storing database consistent with this disclosure. In some embodiments, database 606 may be a combination of two or more different database types (e.g., a hybrid in-memory data database and a conventional database) depending on the specific needs, requirements or particular embodiment of computer 602 and the functions described. Although shown as a single database in Figure 5, two or more databases (of the same type, different types, or a combination of types) may be used depending on the specific needs, requirements or particular embodiment of computer 602 and the functions described. Although database 606 is shown as an internal component of computer 602, in alternative embodiments, database 606 may be external to computer 602.
[0073] Computer 602 also includes memory 607, which can hold data for computer 602 or (whether illustrated or not) a combination of components connected to network 630. Memory 607 can store any data consistent with this disclosure. In some embodiments, memory 607 may be a combination of two or more different types of memory (e.g., a combination of semiconductor and magnetic storage) depending on the specific needs, requirements or a particular embodiment of computer 602 and the functions described. Although Figure 5 illustrates a single memory 607, two or more memories 607 (of the same type, different types, or a combination of types) may be used depending on the specific needs, requirements or a particular embodiment of computer 602 and the functions described. Although memory 607 is illustrated as an internal component of computer 602, in alternative embodiments, memory 607 may be external to computer 602.
[0074] Application 608 may be an algorithmic software engine that provides functionality according to the specific needs, requirements, or specific embodiments of computer 602, as well as the functions described. For example, application 608 may function as one or more components, modules, or applications. Furthermore, although illustrated as a single application 608, application 608 may be implemented as multiple applications 608 on computer 602. In addition, although illustrated as being inside computer 602, in alternative embodiments, application 608 may be outside computer 602.
[0075] The computer 602 may also include a power supply 614. The power supply 614 may include a rechargeable or non-rechargeable battery that can be configured to be either user-replaceable or non-user-replaceable. In some embodiments, the power supply 614 may include power conversion and management circuits that include recharge, standby, and power management functions. In some embodiments, the power supply 614 may include a power plug that allows the computer 602 to be plugged into an outlet or power source, for example, to supply power to the computer 602 or to recharge its rechargeable battery.
[0076] Any number of computers 602 may exist that are associated with or outside of the computer system including computer 602, and each computer 602 communicates via network 603. Furthermore, the terms “client,” “user,” and other appropriate terms may be used interchangeably as appropriate without departing from the scope of this disclosure. Furthermore, this disclosure is intended to enable many users to use one computer 602, and for one user to use multiple computers 602.
[0077] Embodiments of the subject matter and functional operating methods described herein can be implemented in digital electronic circuits, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed herein and their structural equivalents, or in one or more combinations thereof. Software embodiments of the subject matter described herein can be implemented as one or more computer programs. Each computer program may include one or more modules of computer program instructions encoded in a tangible, non-temporary computer-readable storage medium to be executed by a data processing device or to control the operating method of a data processing device. Alternatively or further, program instructions may be encoded in / on an artificially generated propagating signal. For example, this signal may be a machine-generated electrical, optical, or electromagnetic signal generated to encode information to be transmitted to a receiving device suitable for execution by a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage board, a random or serial access memory device, or a combination of computer storage media.
[0078] The terms “data processing device,” “computer,” and “electronic computer device” (or equivalents as understood by those skilled in the art) refer to data processing hardware. For example, a data processing device can encompass all types of devices, machines, and equipment that process data, including, for example, a programmable processor, a computer, or multiple processors or computers. The device may also include dedicated logic circuits, including, for example, a central processing unit (CPU), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). In some embodiments, a data processing device or dedicated logic circuit (or a combination of a data processing device or dedicated logic circuit) may be hardware-based or software-based (or a combination of both). The device may optionally include code that creates an execution environment for computer programs, such as processor firmware, a protocol stack, a database management system, an operating system, or code that constitutes a combination of execution environments. This disclosure intends to describe the use of data processing devices with or without conventional operating systems, such as LINUX, UNIX, WINDOWS, MAC OS, ANDROID, or IOS.
[0079] Computer programs, which may be called, or also referred to as, programs, software, software applications, modules, software modules, scripts, or code, can be written in any form of programming language. Examples of programming languages include compiled languages, interpreted languages, declarative languages, or procedural languages. Programs can be deployed in any form, including standalone programs, modules, components, subroutines, or units used in a computing environment. Computer programs can, but are not required, correspond to files in a file system. A program can be stored in a single file dedicated to it, in part of a file holding one or more scripts stored in a markup language document, or in multiple coordinated files storing one or more modules, subprograms, or parts of code, along with other programs or data. Computer programs can be deployed to run on one or more computers, for example, located at one site or distributed across multiple sites interconnected by a communication network. (As illustrated in various diagrams) While parts of a program may be presented as individual modules implementing various configurations and functions through different objects, methods, or processes, a program can instead contain numerous submodules, third-party services, components, and libraries. Conversely, the configurations and functions of various components can be combined into a single component as appropriate. Thresholds used to make computational decisions can be determined statically, dynamically, or both statically and dynamically.
[0080] The methods, processes, or logic flows described herein can be implemented by one or more programmable computers running one or more computer programs, such that they perform their functions by manipulating input data to produce outputs. The methods, processes, or logic flows can also be implemented by dedicated logic circuits, such as a CPU, FPGA, or ASIC, and the device may be implemented in such a manner.
[0081] A computer suitable for running computer programs can be based on one or more general-purpose and dedicated microprocessors and other types of CPUs. The elements of a computer are a CPU that executes or runs instructions and one or more memory devices that store instructions and data. Generally, a CPU can receive instructions and data from (and write data to) memory. A computer also includes, or can be operably coupled to, one or more mass storage devices that store data. In some embodiments, a computer can receive data from and transfer data to mass storage devices, including, for example, magnetic disks, magneto-optical disks, or optical disks. Furthermore, a computer can be incorporated into another device, such as a mobile phone, personal digital assistant (PDA), portable audio or video player, game console, Global Positioning System (GPS) receiver, or portable storage device such as a Universal Serial Bus (USB) flash drive.
[0082] Computer-readable media (temporary or non-temporary, as appropriate) suitable for storing computer program instructions and data may include all forms of permanent / non-permanent and volatile / non-volatile memory, media, and memory devices. Computer-readable media may include semiconductor memory devices, such as random access memory (RAM), read-only memory (ROM), phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices. Computer-readable media may also include magnetic devices, such as tapes, cartridges, cassettes, and internal / removable disks. Computer-readable media may also include magneto-optical disks and optical memory devices and technologies, such as digital video discs (DVDs), CD-ROMs, DVD+ / -R, DVD-RAM, DVD-ROMs, HD-DVDs, and Blu-rays. Memory can store a variety of objects or data, including caches, classes, frameworks, applications, modules, backup data, jobs, web pages, web page templates, data structures, database tables, repositories, and dynamic information. Types of objects and data stored in memory include parameters, variables, algorithms, instructions, rules, constraints, and references. Furthermore, memory can include logs, policies, security or access data, and reporting files. The processor and memory can be complemented by or integrated into dedicated logic circuits.
[0083] Embodiments of the subject matter described herein include display devices that provide user interaction, including displaying information to a user (and receiving input from the user). It can be implemented on a computer with a chair. Examples of display device types include cathode ray tubes (CRTs), liquid crystal displays (LCDs), light-emitting diodes (LEDs), and plasma monitors. The display device may include a keyboard and a pointing device, such as a mouse, trackball, or trackpad. User input can also be provided to the computer by using a touchscreen, such as a pressure-sensitive tablet computer surface or a multi-touch screen using capacitive or electrical sensing. Interaction with the user can be provided using other types of devices, including receiving user feedback, such as perceptual feedback including visual, auditory, or tactile feedback. Input from the user can be received in the form of acoustic, voice, or tactile input. In addition, the computer can interact with the user by sending documents to and receiving documents from devices used by the user. For example, the computer can send a web page to a web browser in response to a request received from the web browser of the user's client device.
[0084] The term "graphical user interface," or "GUI," can be used in the singular or plural form to describe one or more graphical user interfaces, and each of the displays of a particular graphical user interface. Therefore, a GUI can represent any graphical user interface, including but not limited to web browsers, touchscreens, or command-line interfaces (CLIs), that processes information and efficiently presents the results to the user. Generally, a GUI can include multiple user interface (UI) elements, some or all of which are related to a web browser, such as interactive fields, pull-down lists, and buttons. These and other UI elements may be related to or represent the functionality of a web browser.
[0085] Embodiments of the subject matter described herein can be implemented in a computing system including backend components (e.g., as a data server) or in a computing system including middleware components (e.g., an application server). Furthermore, the computing system may include a frontend component, such as a client computer having either or both a graphical user interface or a web browser that allows a user to interact with the computer. The components of the system can be interconnected by any form or medium of wired or wireless digital data communication (or a combination of data communication) in a communication network. Examples of communication networks include local area networks (LANs), wireless access networks (RANs), metropolitan area networks (MANs), wide area networks (WANs), Worldwide Interoperability for Microwave Access (WiMAX), wireless local area networks (WLANs) (e.g., using 802.11a / b / g / n or 802.20 or a combination of protocols), all or part of the Internet, or any other one or more communication systems (or combinations of communication networks) in one or more locations. Networks can communicate using, for example, Internet Protocol (IP) packets, Frame Relay frames, Asynchronous Transfer Mode (AMT) cells, voice, video, data, or a combination of communication types between network addresses.
[0086] A computing system can include clients and servers. Clients and servers can generally be remote from each other, typically via a communication network. Interaction can occur through a network. The client-server relationship can arise from computer programs running on each computer that have a client-server relationship.
[0087] A cluster file system can be any file system type that is accessible from multiple servers for reading and updating. Locking of the file exchange system can be done at the application layer, so locking or consistency tracking may not be necessary. Furthermore, Unicode data files may differ from non-Unicode data files.
[0088] This specification includes details of many specific embodiments, but these should not be construed as limitations on the scope of what can be claimed, but rather as descriptions of configurations that may be specific to a particular embodiment. Some configurations described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various configurations described in the context of a single embodiment may also be implemented individually in multiple embodiments or in any preferred subcombination. Furthermore, some of the configurations described above may be described as working in several combinations, and may even be initially claimed as such, but one or more configurations from a claimed combination may, in some cases, be removed from that combination, and the claimed combination may be directed towards a subcombination or a variation of a subcombination.
[0089] In the above description, embodiments of the present invention have been described with reference to numerous specific details, which may differ from embodiment to embodiment. Therefore, the description and drawings should be considered illustrative rather than restrictive. The sole and exclusive indicator of the scope of the present invention, and what the applicant intends to be the scope of the present invention, is the set of claims that constitute the literal and equivalent scope patentable from this application (in the particular form in which such claims become patentable), including any subsequent amendments. Any definitions expressly provided herein for terms contained in such claims shall apply to the meaning of such terms as used in the claims. In addition, where the terms “comprising” or “including” are used in the above description or in the following claims, what precedes this phrase may be an additional step or entity, or a substep / sub-entity of a previously enumerated step or entity.
[0090] Specific embodiments of the subject matter have been described. Other embodiments, modifications, and substitutions of the described embodiments are within the scope of the following claims, as will be apparent to those skilled in the art. Although the drawings or claims show the operating methods in a specific order, this should not be understood as requiring that such operating methods be performed in a specific illustrated order or sequentially, or that all the illustrated operating methods be performed (some operating methods may be considered optional), in order to achieve the desired result. In some circumstances, multitasking or parallel processing (or a combination of multitasking and parallel processing) may be advantageous and may be implemented where deemed appropriate.
[0091] Furthermore, the separation or integration of various system modules and components in the embodiments described above should not be interpreted as requiring such separation or integration in all embodiments. It should be understood that the described program components and systems can generally be integrated together in a single software product or packaged in multiple software products.
[0092] Therefore, the embodiments described above do not define or restrict this disclosure. Other modifications, substitutions, and alterations are possible without departing from the spirit and scope of this disclosure.
[0093] Furthermore, any claimed embodiment is deemed applicable to a computer system comprising at least a computer implementation method; a non-temporary computer-readable medium for storing computer-readable instructions for implementing the computer implementation method; and computer memory interconnected to a hardware processor configured to implement the instructions stored in the computer implementation method or the non-temporary computer-readable medium.
[0094] Numerous embodiments of these systems and methods have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of this disclosure.
Claims
1. A method performed by one or more computers: This involves processing medical record data for a patient population to generate a feature sequence that characterizes each patient within that population; To generate data that identifies a set of patient clusters by applying clustering operations to a feature sequence that characterizes patients in a patient population; Selecting a subset of patient clusters for use in identifying new indications about drugs; This includes processing data characterizing patient clusters in a selected subset of patient clusters in order to identify one or more new indications about a drug, To process: (i) ranking multiple candidate drug indications based on a subset of selected patient clusters, and (ii) one or more reference drug indications associated with the drug; This includes identifying one or more candidate drug indicators as a new indication for a drug, based on ranking multiple candidate drug indicators. The aforementioned method.
2. Selecting a subset of patient clusters to use when identifying new indications about drugs is: For each patient cluster, determine the stability of the patient cluster under perturbations of the clustering operation parameters; For each patient cluster, the decision of whether to select the patient cluster for use in identifying new indications for a drug is based at least in part on the stability of the patient cluster under perturbations of the clustering operation parameters. The method according to claim 1, including the method described in claim 1.
3. A set of patient clusters to be used when identifying new indications about drugs Selecting a subset of the standard is: For each patient cluster, the purity of the patient cluster is determined based on a measure of the variance between the feature sequences of the patients included in the patient cluster; For each patient cluster, the decision of whether to select the patient cluster for use in identifying new indications for a drug is made, at least in part, based on the purity of the patient cluster, which is determined based on a measure of variance between the feature sequences of patients included in the patient cluster. The method according to claim 1 or 2, including the method described in claim 1 or 2.
4. Selecting a subset of patient clusters to use when identifying new indications about drugs is: For each patient cluster, determine the number of patients included in that cluster; For each patient cluster, a decision is made, based at least partly on the number of patients included in the cluster, whether to select the patient cluster for use in identifying new indications for the drug. The method according to any one of claims 1 to 3, including the method described in any one of claims 1 to 3.
5. Ranking multiple candidate drug indications based on (i) a subset of selected patient clusters and (ii) one or more reference drug indications associated with the drug: For each candidate drug indication, determine a co-occurrence score that defines the frequency of co-occurrence of (i) the candidate drug indication and (ii) one or more reference drug indications within a selected subset of patient clusters; To determine the ranking of multiple candidate drug indications based at least part on their co-occurrence scores, The method according to claim 1, including the method described in claim 1.
6. For each candidate drug injection, within a selected subset of patient clusters, (i) determine a co-occurrence score that defines the frequency of co-occurrence of candidate drug indicators and (ii) one or more reference drug indicators: (i) Determine the number of patient clusters associated with both candidate drug indicators and (ii) one or more reference drug indicators: The method according to claim 5, including the method described in claim 5.
7. The method according to any one of claims 1 to 6, wherein the clustering operation includes a bipartite k-means clustering operation.
8. The method according to any one of claims 1 to 7, wherein the application of a clustering operation includes performing a multiple correspondence analysis to reduce the dimensionality of a feature array that characterizes a group of patients.
9. The method according to any one of claims 1 to 8, further comprising determining that each patient in the patient population has features related to a drug-targeted signaling pathway.
10. The method according to claim 9, wherein the signal transduction pathway is the IL4 / IL13 pathway.
11. To determine that each patient in the patient population possesses characteristics related to the signaling pathway targeted by the drug: This includes determining that one or more patients in a patient population are associated with one or more clinical conditions related to the IL4 / IL13 pathway; The method according to claim 10, wherein the clinical conditions associated with the IL4 / IL13 pathway include one or more of eosinophilic esophagitis, eosinophilic granulomatosis with polyangiitis (Churg-Strauss syndrome), anaphylaxis, allergic conjunctivitis, urticaria, thyroiditis, pancreatitis, amyloidosis, or basal cell carcinoma.
12. To determine that each patient in the patient population possesses characteristics related to the signaling pathway targeted by the drug: The method according to claim 10 or 11, comprising determining that one or more patients in a patient population are associated with one or more of the following: a diagnosis related to the IL4 / IL13 pathway, medication related to the IL4 / IL13 pathway, a clinical test related to the IL4 / IL13 pathway, or a treatment related to the IL4 / IL13 pathway.
13. The method according to any one of claims 1 to 12, wherein the drug comprises an anti-interleukin-4 receptor α (anti-IL-4Rα) antibody.
14. The method according to claim 13, wherein the drug comprises dupilumab.
15. The method according to any one of claims 1 to 14, wherein the patient population's medical record data includes data characterizing the diagnosis, clinical tests, treatments, drug prescriptions, and biomarker measurements of patients in the patient population.
16. The method according to any one of claims 1 to 15, wherein the patient population includes at least 94 million patients.
17. The method according to any one of claims 1 to 16, wherein for multiple patients in a patient population, the feature sequence characterizing the patients includes at least 2,700 features.
18. One or more computers, Including one or more storage devices communicatively coupled to one or more computers, A system in which one or more storage devices, when executed by one or more computers, store instructions causing one or more computers to perform the operating methods of any one of the methods described in claims 1 to 17.
19. One or more non-temporary computer storage media, which, when executed by one or more computers, store instructions causing one or more computers to perform the operating method of any one of the methods described in claims 1 to 17.
Citation Information
Patent Citations
A Bayesian Causal Network Model for Health Diagnosis and Treatment Based on Patient Data
JP2017537365A
Relevance feedback to improve the performance of classification models that co-classify patients with similar profiles
JP2019512795A
Method and System for Extracting Data From a Plurality of Electronic Data Stores of Patient Data to Provide Provider and Patient Data Similarity Scoring
US20180121604A1