Generalized biomarker model

A generic biomarker model addresses the inefficiency of individualized biomarker identification by accurately identifying patients with specific biomarkers, enhancing the precision and applicability of cohort selection in medical data analysis.

JP2025160302APending Publication Date: 2025-10-22FLATIRON HEALTH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025123114
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-10-29
Filing Date
2025-07-23
Publication Date
2025-10-22

AI Technical Summary

Technical Problem

Existing methods for identifying patients with specific biomarkers are inefficient and require individualized models for each biomarker, which is not feasible due to the wide range of biomarkers and limited data availability, especially when dealing with large amounts of medical data containing handwritten notes and varied medical records.

Method used

A generic biomarker model is developed using a processor to access population data, trained on one or more second biomarkers, and identifies a first group exceeding a likelihood threshold for being tested for a first biomarker, determining cohort candidates based on the model's output.

Benefits of technology

The generic biomarker model accurately identifies patients associated with specific biomarkers, improving efficiency and accuracy over traditional text searches, and can be applied to various characteristics beyond the trained biomarkers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025160302000001_ABST
    Figure 2025160302000001_ABST
Patent Text Reader

Abstract

To provide a model support system that identifies a patient related to a specific biomarker regardless of availability of medical data related to the specific biomarker.SOLUTION: A method for identifying candidates for a cohort on the basis of a biomarker includes: accessing a database from which information associated with a population of individuals can be derived; training a generalized biomarker model on the basis of a second biomarker using information, and providing, to the generalized biomarker model, a first biomarker different from the second biomarker and associated with a cohort; obtaining, from the generalized biomarker model, a first output indicating a first group of the population of individuals exceeding a first likelihood threshold having been tested for the first biomarker; and determining, on the basis of the first output, whether an individual in the first group of the population of individuals is a candidate for the cohort.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 751,990, filed October 29, 2018. The contents of the above application are incorporated herein by reference in their entirety.

[0002] background Technical Field FIELD OF THE DISCLOSURE

[0002] The present disclosure relates to cohort selection, and more particularly to using one or more generic models for automated cohort selection. [Background technology]

[0003] Background information

[0003] There is a growing trend to provide personalized treatments for patients when treating cancer and various other diseases. As an example, to provide more effective treatment, patients with certain forms of cancer (e.g., lung cancer, breast cancer, etc.) can be provided with personalized treatment plans based on genomic markers in the individual's tumor cells. Each tumor cell may have a specific genetic profile that determines how it interacts with other cells in the body and defines the types of biological pathways that may enable the most effective treatment.

[0004]

[0004] Thus, as the medical industry moves toward more personalized treatment plans, being able to identify patients with specific treatment histories and / or characteristics may become increasingly important. Returning to the example of cancer patients, it may be desirable to identify patients who exhibit specific biomarkers. For example, patients may be identified as candidates for a particular treatment, specific laboratory test, or other similar grouping based on whether they have been tested for a particular biomarker and the results of the treatment. However, identifying patients with specific biomarkers can be difficult when examining large amounts of medical data. For example, such identification may require combing through thousands of medical records for indications of whether the patient has been tested for the biomarker and to find the results of the test. Further complicating the problem is that individual patients often have been tested for hundreds of different biomarkers, many of which are not used as the basis for the patient's treatment. In addition, medical records often contain handwritten notes or other text, which may make automating this process more difficult. Some solutions may include developing machine learning models to determine whether a patient has been tested for a particular biomarker. For example, if it is known whether a patient has been tested for a particular biomarker, a model can be trained based on a set of medical records. However, such a solution requires an individualized model for each biomarker, which may not be feasible due to the wide range of biomarkers that may be tested and the limited data available for certain biomarkers.

[0005]

[0005] Therefore, there is a need for improved techniques for identifying patients with specific therapeutic characteristics. A solution should enable the development of machine learning models that are independent of the specific biomarkers (or other characteristics) used to train the model. Thus, a generic biomarker model can be used to identify patients associated with a specific biomarker regardless of the availability of medical data related to that specific biomarker. Summary of the Invention [Means for solving the problem]

[0006] overview

[0006] Embodiments consistent with the present disclosure include systems and methods for identifying candidates associated with specific biomarkers. In one embodiment, a model-assisted system may include at least one processor. The processor may be programmed to: access a database from which information associated with a population of individuals can be derived; provide a first biomarker associated with the cohort to a generic biomarker model, the generic biomarker model being trained based on one or more second biomarkers using the information, the first biomarker being different from the one or more second biomarkers; obtain a first output from the generic biomarker model indicative of a first group of the population of individuals that exceed a first likelihood threshold for being tested for the first biomarker; and determine whether an individual in the first group of the population of individuals is a candidate for the cohort based on the first output.

[0007] In another embodiment, a computer-implemented method can identify cohort candidates based on biomarkers. The method can include accessing a database from which information related to a population of individuals can be derived, providing a first biomarker related to the cohort to a generic biomarker model, the generic biomarker model being trained based on one or more second biomarkers using the information, the first biomarker being different from the one or more second biomarkers, obtaining a first output from the generic biomarker model indicative of a first group of the population of individuals that exceed a first likelihood threshold for being tested for the first biomarker, and determining whether individuals in the first group of the population of individuals are candidates for the cohort based on the first output.

[0008] In another embodiment, a model-assisted system may include at least one processor. The processor may be programmed to: access a database from which information related to a population of individuals can be derived; provide a first characteristic associated with the cohort to a generic model, the generic model being trained using the information based on one or more second characteristics, the first characteristic being different from the one or more second characteristics; obtain a first output from the generic model indicative of a first group of the population of individuals that exceeds a first likelihood threshold being associated with the first characteristic; and determine whether individuals in the first group of the population of individuals are candidates for the cohort based on the first output.

[0009]

[0009] Consistent with other disclosed embodiments, a non-transitory computer-readable storage medium may include program instructions that are executed by at least one processing device to perform any of the methods described herein.

[0010] BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings, which are incorporated in and constitute a part of this specification, and together with the description, serve to illustrate and explain the principles of various exemplary embodiments. [Brief explanation of the drawings]

[0011] [Figure 1]

[0011] FIG. 1 is a block diagram illustrating an exemplary system environment for implementing embodiments consistent with the present disclosure. [Figure 2]

[0012] FIG. 1 is a block diagram illustrating an exemplary medical record for a patient, consistent with the present disclosure. [Figure 3]

[0013] FIG. 1 is a block diagram illustrating an example machine learning process for implementing embodiments consistent with the present disclosure. [Figure 4A]

[0014] FIG. 1 is a block diagram illustrating an example of a process for building a generic biomarker model consistent with the present disclosure. [Figure 4B]

[0015] FIG. 1 is a block diagram illustrating an example of a technique for extracting features for building a generic biomarker model consistent with the present disclosure. [Figure 5]

[0016] 1 is a flow chart illustrating an exemplary process for identifying cohort candidates based on biomarkers consistent with the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0012] Detailed Description

[0017] The following detailed description refers to the accompanying drawings. Wherever possible, the same reference numbers are used in the drawings and the following description to refer to the same or similar parts. While several exemplary embodiments are described herein, modifications, adaptations, and other implementations are possible. For example, components shown in the figures may be substituted, added, or modified, and the exemplary methods described herein may be modified by substituting, rearranging, removing, or adding steps to the disclosed methods. Therefore, the following detailed description is not limited to the disclosed embodiments and examples. Rather, the appropriate scope is defined by the appended claims.

[0013]

[0018] Embodiments herein include computer-implemented methods, tangible non-transitory computer-readable media, and systems. Computer-implemented methods may be executed by at least one processor (e.g., processing unit) that receives instructions, for example, from a non-transitory computer-readable storage medium. Similarly, systems consistent with the present disclosure may include at least one processor (e.g., processing unit) and memory, where the memory may be a non-transitory computer-readable storage medium. As used herein, a non-transitory computer-readable storage medium refers to any type of physical memory in which information or data readable by at least one processor may be stored. Examples include random access memory (RAM), read-only memory (ROM), volatile memory, non-volatile memory, hard drives, CD-ROMs, DVDs, flash drives, disks, and any other known physical storage medium. Singular terms such as "memory" and "computer-readable storage medium" may also refer to multiple structures, such as multiple memories and / or computer-readable storage media. As referred to herein, "memory" may include any type of computer-readable storage medium unless otherwise specified. A computer-readable storage medium can store instructions for execution by at least one processor, including instructions for causing the processor to perform steps or stages consistent with embodiments herein. Additionally, one or more computer-readable storage media can be utilized in implementing a computer-implemented method. The term "computer-readable storage medium" should be understood to include tangible items and to exclude carrier waves and transient signals.

[0014]

[0019] Embodiments of the present disclosure provide systems and methods for identifying patients based on a generic model. Users of the disclosed systems and methods may include any individual who may wish to access and / or analyze patient data and / or conduct experiments using selected patient cohorts. Thus, throughout this disclosure, references to "users" of the disclosed systems and methods may include physicians, researchers, quality assurance departments at health care organizations, and / or any other individuals.

[0015]

[0020] Figure 1 illustrates an exemplary system environment 100 for implementing embodiments consistent with the present disclosure, described in detail below. As shown in Figure 1, system environment 100 includes several components, including client devices 110, data sources 120, systems 130, and / or networks 140. It will be understood from this disclosure that the number and arrangement of these components is exemplary and is shown for illustrative purposes. Other arrangements and numbers of components may be used without departing from the teachings and embodiments of the present disclosure.

[0016]

[0021] As shown in FIG. 1 , exemplary system environment 100 includes system 130. System 130 may include one or more server systems, databases, and / or computing systems configured to receive information from entities over a network, process the information, store the information, and display / transmit the information to other entities over the network. Thus, in some embodiments, the network may facilitate cloud sharing, storage, and / or computing. In one embodiment, system 130 may include processing engine 131 and one or more databases 132, which are illustrated within the area bounded by the dashed lines representing system 130 in FIG. 1 . Processing engine 140 may include at least one processing device, such as one or more general-purpose processors, e.g., central processing units (CPUs), graphics processing units (GPUs), etc., and / or one or more special-purpose processors, e.g., application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc.

[0017]

[0022] Components of environment 100 (including systems 130, client devices 110, and data sources 120) can communicate with each other and with other components via network 140. Network 140 may include the Internet, a wide area network (WAN), a wired local area network (LAN), a wireless WAN (e.g., WiMAX), a wireless LAN (e.g., IEEE 802.11, etc.), a mesh network, a mobile / cellular network, an enterprise or private data network, a storage area network, a virtual private network using a public network, a short-range wireless communication technique (e.g., Bluetooth, infrared, etc.), or various other types of network communications. In some embodiments, communications may occur across two or more of these types of networks and protocols.

[0018]

[0023] System 130 may be configured to identify patients based on particular qualities or characteristics associated with the patient and / or the treatment the patient receives. In some embodiments, the characteristics may be based on particular biomarkers. For example, system 130 may be configured to identify patients based on whether they have been tested for particular biomarkers, specific test results associated with the biomarkers (e.g., positive or negative), or various other characteristics. While patient selection based on biomarkers or biomarker status is used throughout this disclosure, it will be understood that the disclosed systems, methods, and / or techniques can be used for other patient identification means as well (e.g., whether the patient has been prescribed a particular medication, whether the patient has undergone a particular treatment, etc.). Likewise, in other embodiments, it will be understood that the disclosed systems, methods, and / or techniques can be similarly used to identify other individuals, subjects, entities, etc. based on generic models.

[0019]

[0024] System 130 may be configured to receive patient medical and other information from data sources 120 or other sources within network 140. In some embodiments, medical information may be stored in the form of one or more medical records, each associated with a patient. More specifically, system 130 may be configured to receive and store data transmitted over network 140 from various data sources, including data sources 120, process the received data, and transmit data and results based on the processing to client device 110. Data sources 120 may include a variety of sources of medical information regarding a patient. For example, data sources 120 may include the patient's healthcare provider, such as a doctor, nurse, specialist, consultant, hospital, clinic, etc. Data sources 120 may also include laboratories, such as radiology or other imaging laboratories, hematology laboratories, pathology laboratories, etc. Data sources 120 may also include insurance companies or any other source of patient data.

[0020]

[0025] The system 130 may be configured to develop and use one or more models for identifying patients with specific characteristics based on medical records. For example, the system 130 may use machine learning techniques to develop the models based on training data. In some embodiments, the system 130 may develop a generic model that may be trained based on a set of specific characteristics or properties, but may be used more broadly to identify patients with other characteristics within the patient medical record that can be treated similarly. For example, if the system 130 is used to identify patients associated with specific biomarkers, the system 130 may develop or implement a generic biomarker model. While it may be desirable to develop a separate model for each biomarker, this may not be feasible. For example, some biomarkers may be commonly tested in a broad patient population, while other biomarkers may be tested relatively infrequently for a small number of patient samples. Thus, while it may be possible to develop specific biomarker models for more common biomarkers for which sample data are readily available, it may be too difficult or expensive to develop specific biomarker models for all biomarkers due to the sheer number of biomarkers that may be tested and the limited data sets that may be available for some biomarkers.

[0021]

[0026] Thus, a generic biomarker model can be developed that can be trained using one or more biomarkers included in the first set. The first set of biomarkers can be biomarkers for which sufficient information is available in medical records or other data to develop an accurate or reliable machine learning model. Because medical records can describe and / or discuss many biomarkers in a similar manner (e.g., with a similar structure, using common terminology, etc.), the generic biomarker model can be used for biomarkers other than those included in the first set. For example, a physician describing test results for common biomarkers (e.g., those included in the first set) can describe test results related to other biomarkers in a similar manner. As a result, the generic biomarker model can be configured to identify not only patients being tested for the first set of biomarkers, but also patients being tested for biomarkers other than those in the first set. System 130 can apply one or more generic models to received medical results to identify patients associated with particular characteristics (e.g., being tested for a particular biomarker, testing positive for a particular biomarker, etc.). Using a generic biomarker model may yield more accurate results than simply performing a text search for a given biomarker identifier. For example, a doctor's note containing "withhold EGFR testing" may indicate that the patient has not been tested for the EGFR biomarker, yet a text search will still display the result. It will be understood that this is an example and that more complex relationships may arise between biomarkers and the surrounding text. While the generic model is described with respect to biomarkers, it will be understood that this is by way of example and that generic models can be developed to identify patients based on other characteristics (e.g., prescribed medications, treatments administered, other forms of testing, etc.).

[0022]

[0027] System 130 may further communicate with one or more client devices 110 via network 140. For example, system 130 may provide client device 110 with results based on analyzing information from data source 120. Client device 110 may include any entity or device capable of sending and receiving data via network 140. For example, client device 110 may include a computing device such as a server or a desktop or laptop computer. Client device 110 may also include other devices such as mobile devices, tablets, wearable devices (i.e., smart watches, implantable devices, fitness trackers, etc.), virtual machines, IoT devices, or various other technologies. In some embodiments, client device 110 may transmit queries for information about one or more patients to system 130 via network 140, such as queries about patients being tested for specific biomarkers or queries about various other information about patients.

[0023]

[0028] In some embodiments, the system 130 may be configured to select one or more cohorts. As used herein, a cohort may include any group of information (e.g., people, objects, etc.) that share at least one common characteristic or exhibit attributes that meet a set of predetermined criteria. In some embodiments, a cohort may include individuals that exhibit at least one common characteristic from a medical perspective (e.g., demographic or clinical characteristic). An individual may include any member of one or more groups (e.g., subjects, people, objects, etc.). For example, individuals from a population determined to have a particular type of disease, or more specifically individuals from a population being tested for a particular biomarker associated with that disease, may be identified and placed into a common cohort. Cohorts can be constructed for a variety of purposes. In some examples, a cohort may be constructed to form a group used to analyze characteristics of a particular disease, such as the epidemiology of the disease, treatment, or how outcomes such as disease mortality or progression depend on certain variables.

[0024]

[0029] The various components of the system environment 100 may include hardware, software, and / or firmware assemblies, including memory, a central processing unit (CPU), and / or a user interface. Memory may include any type of RAM or ROM implemented by a physical storage medium, such as magnetic storage, including floppy disks, hard disks, or magnetic tape; semiconductor storage, such as solid-state disks (SSDs) or flash memory; optical disk storage; or magneto-optical disk storage. The CPU may include one or more processors for processing data according to a set of programmable instructions or software stored in memory. The functions of each processor may be provided by a single dedicated processor or by multiple processors. Furthermore, the processor may include, without limitation, digital signal processor (DSP) hardware or any other hardware capable of executing software. The optional user interface may include any type or combination of input / output devices, such as a display monitor, keyboard, and / or mouse.

[0025]

[0030] Data transmitted and / or exchanged within system environment 100 may occur across a data interface. As used herein, a data interface may include any boundary at which two or more components of system environment 100 exchange data. For example, environment 100 may exchange data between software, hardware, databases, devices, people, or any combination of the above. Furthermore, it will be understood that any suitable configuration of software, processors, data storage, and networks may be selected to implement the components of system environment 100 and related embodiment features.

[0026]

[0031] FIG. 2 illustrates an exemplary medical record 200 for a patient. The medical record 200 may be received from a data source 120 and processed by the system 130 to identify the patient, as described above. As shown in FIG. 2 , the record received from the data source 120 (or elsewhere) may include both structured data 210 and unstructured data 220. The structured data 210 may include quantifiable or categoriseable data about the patient, such as gender, age, race, weight, vital signs, test results, date of diagnosis, type of diagnosis, stage of disease (e.g., billing code), timing of treatment, procedure performed, date of visit, type of care, insurance provider and start date, medication instructions, medication management, or any other measurable data about the patient. The unstructured data may include information about the patient that is not quantifiable or easily categoriseable, such as a doctor's notes or patient lab reports. The unstructured data 220 may include information such as a doctor's description of the treatment plan, notes describing what happened at the visit, a description of the patient's condition, radiation therapy reports, pathology reports, etc. In some embodiments, the unstructured data may include data related to one or more biomarkers. For example, the unstructured data may include notes (e.g., from a doctor, nurse, lab assistant, etc.) discussing test results related to a particular biomarker (e.g., whether the patient underwent the test, the results of the test, an analysis of the results, etc.).

[0027]

[0032] Within the data received from data source 120, each patient may be represented by one or more records generated by one or more medical professionals or patients. For example, a doctor associated with the patient, a nurse associated with the patient, a physical therapist associated with the patient, etc. may each generate a patient's medical record. In some embodiments, one or more records may be collated and / or stored within the same database. In other embodiments, one or more records may be distributed across multiple databases. In some embodiments, a record may be stored and / or given multiple electronic data representations. For example, a patient record may be represented as one or more electronic files, such as a text file, a Portable Document Format (PDF) file, an Extensible Markup Language (XML) file, etc. If a document is stored as a PDF file, an image, or other file without text, the electronic data representation may also include text associated with the document derived from an optical character recognition process. In some embodiments, unstructured data may be captured by an extraction process, while structured data may be entered by a medical professional or calculated using an algorithm.

[0028]

[0033] FIG. 3 illustrates an exemplary machine learning system 300 for implementing embodiments consistent with the present disclosure. The machine learning system 300 can be implemented as part of system 130 (FIG. 1). For example, the machine learning system 300 can be a component of processing engine 131 or a process executed using processing engine 131. According to disclosed embodiments, the machine learning system 300 can generate a generic model (e.g., a supervised machine learning system) based on a set of training data related to patients, and can use the model to identify patients associated with a particular characteristic. For example, as shown in FIG. 3, the machine learning system 300 can build a generic biomarker model 330 for identifying patients associated with a test biomarker 315. The machine learning system 300 can develop the model 330 through a training process, e.g., using a training algorithm 320.

[0029]

[0034] Training the model 330 may include using a training dataset 310, which may be input into a training algorithm 320 to develop the model. The training data 310 may include multiple patient medical records 312 (e.g., "medical record 1," "medical record 2," etc.) for which results associated with various training biomarkers 311 may already be known. For example, a training biomarker 311 may be associated with one or more medical records 312 in which the patient was tested for the training biomarker 311. In some embodiments, each training biomarker 311 may be associated with one or more medical records 312. For example, as shown in FIG. 3 , training biomarker A may be associated with multiple medical records 312 (e.g., medical record 1 and medical record 2). The training biomarkers 311 may represent biomarkers for which sufficient data is available to accurately construct the generic biomarker model 330.

[0030]

[0035] In some embodiments, the training data 310 may also be cleaned, conditioned, and / or manipulated before being input into the training algorithm 320 to facilitate the training process. The machine learning system 300 may extract one or more features (or feature vectors) from the recordings and apply the training algorithm 320 to determine correlations between text discussing a particular biomarker, whether a patient has been tested for that biomarker, and what the test results may indicate. These features may be extracted from structured and / or unstructured data, as described above with respect to FIG. 2. For example, the training process may correlate words or word combinations around a biomarker identifier in the unstructured data with whether a patient has been tested for the biomarker, the results of the test, etc. The process for building the generic model 330 is described in more detail below with respect to FIG. 4A.

[0031]

[0036] Once the model 330 is constructed, test data, such as test biomarkers 331, and medical records 332 may be input into the generic biomarker model 330. The medical records 440 may correspond to the medical records 200 described above. For example, the medical records 440 may include structured and unstructured data associated with multiple patients, such that each patient has one or more medical records associated with them. The generic model 330 may extract features from the medical records 440 to generate output 350. The output 350 may identify medical records 332 associated with the patient that are also associated with the test biomarkers 331. For example, the output 350 may identify patients undergoing testing for the test biomarkers 311. In some embodiments, the output 350 may indicate other patient groups associated with the test biomarkers 311. For example, the output 350 may indicate that the patient tested positive for the test biomarker 331, tested negative for the test biomarker 331, was diagnosed with a particular medical condition based on the biomarker 331, was administered a particular treatment based on the test biomarker 331, etc. Each of the different groups 351 may be determined by a separate generic biomarker model 330, or one generic biomarker model 330 may be configured to provide multiple outputs 350 and / or patient groups 351.

[0032]

[0037] In some embodiments, patients may be selected for one or more groups based on the patient exceeding a particular likelihood threshold. For example, the generic biomarker model 330 may generate a likelihood or confidence value for each patient who has been tested for a biomarker, tested positive for a biomarker, etc. The generic biomarker model 330 may select patients for inclusion in one or more of the groups 351 based on whether the patient exceeds a particular likelihood threshold (e.g., 50%, 60%, 70%, 80%, 90%, 99%, etc.) or confidence value threshold. In some embodiments, the threshold may be adjustable based on a desired level of efficiency and performance. For example, as described above, the model may be retrained based on test data (which may include records from databases not used to develop the model). One or more loss functions may be used to adjust the threshold.

[0033]

[0038] In some embodiments, the output 350 can be used to identify patients for inclusion in a cohort, as described above. For example, the generic biomarker model 330 can be used to identify patients who have been tested for a test biomarker 331, who have tested positive for a test biomarker 331, etc. Further analysis can then determine whether the patient is a candidate for the cohort. In some embodiments, such analysis can include confirming that an individual has been tested for a biomarker, tested positive for a biomarker, etc., based on medical records associated with the individual, depending on the cohort. In some embodiments, confirmation can be a manual process (e.g., performed by a trained medical professional).

[0034]

[0039] In some embodiments, the remaining portion of the training data 310 can be used to test the trained model 330 to evaluate its performance. For example, for each individual in the remaining portion of the training data set 310, a feature vector can be extracted from the medical records associated with that patient. The feature vector can be provided to the model 330, and the output for that individual can be compared to known results for that individual (e.g., whether the individual tested positive for a particular training biomarker 311). As shown in FIG. 3 , the deviation between the output of the model 330 and the known biomarkers being tested for any individual in the training data set 310 can be used to generate a performance measure 360. The performance measure 360 ​​can be used to update the model 330 (e.g., retrain the model) to reduce the deviation between the output 350 and known patient outcomes. For example, one or more functions of the model can be added, removed, or modified (e.g., a quadratic function can be modified to a cubic function, an exponential function can be modified to a polynomial function, etc.). Thus, the deviation may be used to inform decisions to modify how the features included in the model 330 are constructed or what type of model is used. Alternatively, in some embodiments, one or more weights of the regression (or one or more weights of the nodes if the model includes a neural network) may be adjusted to reduce the deviation. If the level of deviation is within desired limits (e.g., 10%, 5%, or less), the one or more models 330 may be deemed suitable for operating on datasets where patient outcomes are unknown. Although described above in terms of "deviation," one or more loss functions may also be used to measure the accuracy of the model. For example, a squared loss function, a hinge loss function, a logistic loss function, a cross-entropy loss function, or any other loss function may be used. In such embodiments, model updates may be configured to reduce (or even at least locally minimize) one or more loss functions.

[0035]

[0040] The accuracy of the generic biomarker model 330 can be assessed in a variety of other ways. In some embodiments, the accuracy of the generic biomarker model 330 can be assessed based on one or more biomarker-specific models. For example, a specific biomarker model can be generated for a particular training biomarker 311. The biomarker-specific model can be developed using the techniques described above, but can be trained based on medical records that indicate whether the patient was tested for that particular biomarker. The generic biomarker model 330 should be able to identify patients who have been tested for a particular biomarker as accurately as, or with similar accuracy to, the biomarker-specific model. Thus, the processing engine 131 can be configured to compare the output from the biomarker-specific model with the output 350 to assess the accuracy of the generic biomarker model 330.

[0036]

[0041] In other embodiments, the accuracy of the generic biomarker model 330 can be evaluated based on a text search for the biomarkers. For example, the processing engine 131 can perform a basic text search on test biomarkers 331 in medical records to identify a group of patients undergoing testing for the generic biomarker. The generic biomarker model 330 should outperform a basic text search because it can glean additional information from the piece of information. Therefore, a comparison between the results of the text search and the output 350 can be used to evaluate the accuracy of the generic biomarker model 330. Additionally, various other diagnostic queries can be performed that may indicate inaccuracy of the generic biomarker model 330, such as determining whether the generic biomarker model 330 identified medical records that were not identified in the text search.

[0037]

[0042] 4A is a block diagram illustrating an example of a process 400 for building a generic biomarker model consistent with the present disclosure. For example, the process 400 can be used to build the generic biomarker model 330 using the training dataset 330 as discussed above with respect to FIG.

[0038]

[0043] As shown in FIG. 4A , relevant training biomarkers 410 can be selected for use in building the model. For example, the training biomarkers 410 can be selected by medical professionals trained to make manual, subjective judgments about whether a patient is associated with a particular biomarker. While the "EGFR" and "ALK" biomarkers are shown as examples, it will be understood that the generic biomarker model 330 can be built using any suitable biomarker or other data. The training biomarkers 410 can represent biomarkers for which sufficient data is available to accurately build the generic biomarker model 330. The training biomarkers 410 can correspond to the training biomarkers 311 discussed above.

[0039]

[0044] The training biomarkers 410 can be input to information fragment extraction 412, where text associated with the biomarkers 410 is extracted from the patient medical record. While some or a portion of the documentation in a patient's medical record may be available electronically, typed, handwritten, or printed text in the record can be converted to machine-coded text (e.g., by optical character recognition (OCR)). The electronic text can then be searched for specific keywords or phrases associated with the particular biomarker. In some embodiments, information fragments of text near identified training biomarkers 410 can be examined to gather additional information about the context of the word or phrase. By evaluating information fragments surrounding the training biomarkers 410 rather than the biomarkers alone, a model can be trained to distinguish "ALK" from terms such as "ALK not tested," which may have significantly different meanings.

[0040]

[0045] After information piece extraction 412, feature vectorization 414 can be performed on the extracted information pieces to identify a set of feature vectors. In some embodiments, structured data contained within the medical records from which the information pieces were extracted can also be evaluated along with the information pieces. For example, the extracted phrases and any structured data considered can be converted into multidimensional vectors that correlate scores to the phrases and other structured data. The score for each phrase and / or portion of structured data can represent a magnitude along the dimension associated with the corresponding phrase and / or portion. In some embodiments, the scores can be binary, such that the presence of the phrase results in a magnitude of 1 along the dimension associated with the phrase, while the absence of the phrase results in a magnitude of 0 along the dimension associated with the phrase. For example, if the extracted information piece includes the phrase "EGFR tested," the vector can have a component magnitude of 1 along the "EGFR" dimension; if the extracted information piece includes only the phrase "EGFR not tested," and does not include the phrase "EGFR" apart from the "not" modifier, the vector can have a component magnitude of 0 along the "EGFR" dimension. In other embodiments, the scores can be non-binary, indicating, for example, an incidence rate associated with the phrase. For example, if the extracted piece of information contains five instances of the phrase "EGFR," the vector may have a component magnitude of 5 along the "EGFR" dimension, and if the extracted piece of information contains only two instances of the phrase "ALK," the vector may have a component magnitude of 2 along the "ALK" dimension. The rate of occurrence may represent a normalized measure of instances, such as total instances per a particular number of characters, a particular number of words, a particular number of sentences, a particular number of paragraphs, a particular number of pages, etc.

[0041]

[0046] The machine learning system 300 can use any suitable machine learning algorithm to develop the model 330 based on the feature vector. For example, the training algorithm 320 can include a logistic regression 416 to determine a score based on the feature vector. The score can correlate to or indicate whether a patient associated with the medical record has been tested for biomarkers, etc. Additionally or alternatively, the training algorithm 320 can include one or more neural networks that adjust the weights of one or more nodes so that an input layer of features passes through one or more hidden layers and then through an output layer of patient outcomes (with associated probabilities). Other types of machine learning techniques, such as linear regression models, lasso regression analysis, random forest models, K-nearest neighbor (KNN) models, K-means models, decision trees, Cox proportional hazards regression models, naive Bayes models, support vector machine (SVM) models, or gradient boosting algorithms, can also be used in combination with or apart from the logistic regression 416. Models can also be developed using unsupervised or reinforcement machine learning processes, which do not require manual training. Based on the application of logistic regression 416, a resulting model can be developed in step 418. For example, a generic biomarker model 330 can be constructed based on the training biomarkers 311, as described above.

[0042]

[0047] 4B is a block diagram illustrating an example of a technique for extracting features for building a generic biomarker model consistent with the present disclosure. The blocks shown in FIG. 4B may correspond to process 400.

[0043]

[0048] As described above, training biomarkers 410 are input into information piece extraction 412. As indicated by block 420, system 130 can identify training biomarkers 410 (e.g., "EGFR") from within the patient medical record. In some embodiments, this block can include converting typed, handwritten, or printed text within the unstructured data of the patient medical record to machine-coded text (e.g., via optical character recognition (OCR)). In some embodiments, as indicated by block 430, the text of the biomarkers can be replaced with tokens 431 (e.g., "[biomarker]") that represent the training biomarkers 410 within the text. By substituting tokens 431 for one or more training biomarkers 410, a generic model can be built based on how biomarkers are treated within the text of the medical record, rather than models based on individual biomarkers. Information pieces 432 of text near the identified tokens 431 can be examined to gather additional information about the context of the word or phrase. For example, the piece of information 431 may be based on a predetermined number of characters or words before or after the phrase 431, all text in the same paragraph as the phrase 431, or a variety of other techniques.

[0044]

[0049] A plurality of feature vectors 440 can be extracted based on the pieces of information 431. For example, features can be extracted based on Term-Frequency Inverse-Document-Frequency (TFIDF) vectorization or other means. As shown in FIG. 4B, the features can be individual words or bigrams (e.g., "lung [biomarker]"). Various other forms of features (e.g., trigrams, n-grams, etc.) can also be used. The system 130 can then select the features and perform logistic regression (or various other algorithms mentioned above) to construct the generic biomarkers 330.

[0045]

[0050] 5 illustrates an exemplary process 500 for identifying cohort candidates based on biomarkers, consistent with disclosed embodiments. Method 500 may be implemented by at least one processor of processing engine 131 of system 100 shown in FIG. 1, for example. In some embodiments, process 500 may be performed by another device within system 100, such as client device 110 or another device with access to system 130.

[0046]

[0051] At step 510, method 500 may include accessing a database from which information related to the population of individuals can be derived. In some embodiments, the information may include medical records related to the population of individuals. For example, processing engine 131 may access medical records from data source 120 or various other sources over network 140. As discussed above, data source 120 may include various sources of patient medical data, including, for example, medical professionals, laboratories, insurance companies, etc. Alternatively, or in addition, the processing engine may access a local database, such as database 132, to access patient medical records.

[0047]

[0052] A medical record may include one or more electronic files, such as a text file, an image file, a PDF file, an XML file, a YAML file, etc. In some embodiments, a medical record (e.g., medical record 200) may include structured information (e.g., structured data 212) and unstructured information (e.g., unstructured data 211) associated with a population of individuals, as described above. For example, structured information may include gender, birth date, race, weight, test results, vital signs, diagnosis date, visit date, medication instructions, diagnosis codes, procedure codes, drug codes, previous treatments, or medication regimens. Unstructured information may include text written by a medical professional, radiation therapy reports, pathology reports, or various other forms of text associated with a patient. In some embodiments, at least a portion of the unstructured information has been subjected to an optical character recognition process, as discussed above. Each medical record may be associated with a particular patient, and in some embodiments, multiple medical records may be associated with a particular patient. Medical records may not be limited to data from medical institutions, but may also include other relevant data types, such as insurance assessment data (e.g., from insurance companies), patient-reported data, or other information related to the patient's treatment or health.

[0048]

[0053] At step 520, method 500 includes providing first biomarkers associated with the cohort to a generic biomarker model, where the generic biomarker model is trained based on one or more second biomarkers using the information, where the first biomarkers are different from the one or more second biomarkers. For example, the one or more second biomarkers may correspond to training biomarkers 311 discussed above with respect to FIG. 3 , and the first biomarkers may correspond to test biomarkers 331. Thus, the one or more second biomarkers may be used to construct the generic biomarker model 330. In some embodiments, the one or more second biomarkers may represent biomarkers for which sufficient data is available to construct the generic biomarker model 330. For example, the one or more second biomarkers may be more prevalent in the information than the first biomarkers. In some embodiments, the generic biomarker model may be trained based on unstructured information, as discussed above. In some embodiments, the generic biomarker model may be developed at least in part based on a feature vector extracted from information based on one or more second biomarkers. For example, the generic biomarker model 330 may be developed based on the feature vector 440 depicted in FIG. 4B. Further, in some embodiments, the feature vector may include at least one biomarker lexicon (e.g., lexicon 431) representing text associated with at least one second biomarker.

[0049]

[0054] Step 520 may include additional substeps to facilitate analysis of the medical record, such as adjusting or modifying information in the record. The processing engine 131 may use various techniques to interpret structured or unstructured information. For example, typed, handwritten, or printed text in the medical record may be converted to machine-coded text (e.g., by optical character recognition (OCR)).

[0050]

[0055] At step 530, method 500 may include obtaining a first output from the biomarker model indicating a first group of a population of individuals above a first likelihood threshold being tested for the first biomarker. For example, the generic biomarker model 330 may generate output 350, which may include group 351 indicating patients being tested for the first biomarker. In some embodiments, the likelihood threshold may be adjusted based on the model's efficiency and performance level. In some embodiments, the biomarker model may generate the first output using a binary classification algorithm. For example, the binary classification algorithm may include at least one of logistic regression, random forest, gradient boosting tree, support vector machine, or neural network. In some embodiments, the classification algorithm may include various other algorithms described above (e.g., Cox proportional hazards regression, Lasso regression network, etc.). In some embodiments, step 530 may include further steps, such as storing the first output for access by a user of the generic biomarker model. In some embodiments, step 530 may include transmitting the first output to one or more users or one or more devices. For example, the system 120 can transmit the first output to the client device 100 over the network 140 .

[0051]

[0056] In some embodiments, process 500 may further include obtaining a second output from the generic biomarker model indicating a second group of individuals above a second likelihood threshold for testing positive for the first biomarker, the individuals being included in the second group. In some embodiments, a second group of individuals may be identified in the first output along with the first group of individuals. For example, the generic biomarker model may be configured to determine both the first group of individuals being tested for the biomarker and the second group of individuals testing positive for the biomarker. In other embodiments, a separate generic biomarker model may be used to identify the second group of individuals.

[0052]

[0057] At step 540, method 500 may include determining whether an individual in the first group of individuals is a candidate for the cohort based on the first output. For example, determining whether an individual is a candidate for the cohort may include confirming that the individual has been tested for the biomarker based on medical records associated with the individual. As discussed above, this may be a manual process (e.g., by a trained medical professional) to determine whether the individual has actually been tested for the first biomarker. In embodiments in which the generic biomarker model is configured to determine whether a patient is associated with a particular test result (e.g., the patient tests positive for the first biomarker), determining whether the individual is a candidate for the cohort may include confirming that the individual has tested positive for the biomarker based on medical records associated with the individual.

[0053]

[0058] In some embodiments, process 500 may further include additional steps. For example, process 500 may be configured to confirm the accuracy of the generic biomarker model. In some embodiments, the accuracy of the generic biomarker model may be evaluated based on a biomarker model specific to a first biomarker. Thus, process 500 may include providing the first biomarker to a biomarker-specific model, where the biomarker-specific model is trained based on the first biomarker using medical records. Process 500 may further include obtaining a third output from the biomarker-specific model indicative of a third group of individuals above a likelihood threshold who have been tested for at least one biomarker. Process 500 may further include confirming the accuracy of the generic biomarker model by comparing the first output with the third output. For example, a difference between the results of the generic biomarker model and the results of the biomarker-specific model may indicate whether the generic biomarker model is effective in identifying patients being tested for a wide variety of different biomarkers.

[0054]

[0059] In other embodiments, the accuracy of the generic biomarker model can be confirmed by comparing the results to a search function. Accordingly, process 500 may include searching the medical records for the first biomarker to generate a fourth output indicating a fourth group of individuals undergoing testing for at least one biomarker. For example, system 130 may use a plain text search function to search for words related to the first biomarker in the medical records. Process 500 may further include confirming the accuracy of the generic biomarker model by comparing the first output to the fourth output. Ideally, the generic biomarker model performs better than a basic text search for the first biomarker in identifying patients for inclusion in the cohort. Various other means for testing the accuracy of the generic biomarker model may also be used. Process 500 may further include additional steps, such as updating the generic biomarker model based on the determined accuracy.

[0055]

[0060] In some embodiments, process 500 may be applied to other traits in addition to biomarkers. Thus, in some embodiments, process 500 may include accessing a database from which information related to a population of individuals can be derived (step 520); providing a first trait associated with the cohort to a generic model, the generic model being trained using the information based on one or more second traits, the first trait being distinct from the one or more second traits (step 540); obtaining a first output from the generic model indicative of a first group of individuals in the population that exceed a first likelihood threshold for being associated with the first trait (step 560); and determining whether individuals in the first group of individuals are candidates for the cohort based on the first output (step 580). In some implementations, traits may correspond to biomarkers, as discussed above. Thus, the first trait may include a first biomarker, the one or more second traits may include one or more second biomarkers, and the first output may be indicative of a first group of individuals being tested for the first biomarker. In other embodiments, the first characteristic can include a first drug, the one or more second characteristics can include one or more second drugs, and the first output can indicate a first group of individuals being treated using the first drug.

[0056]

[0061] The above description has been presented for illustrative purposes. It is not exhaustive or limited to the precise form or embodiment disclosed. Modifications and adaptations will be apparent to those skilled in the art from consideration of the specification and practice of the disclosed embodiments. In addition, while aspects of the disclosed embodiments are described as being stored in memory, those skilled in the art will appreciate that these aspects can also be stored on other types of computer-readable media, such as secondary storage, e.g., a hard disk, or CD ROM, or other forms of RAM or ROM, USB media, DVD, Blu-ray, 4K Ultra HD Blu-ray, or other optical drive media.

[0057]

[0062] Computer programs based on the written descriptions and disclosed methods are within the skill of an experienced developer. The various programs or program modules can be created using any of the techniques known to those skilled in the art, or can be designed in conjunction with existing software. For example, program sections or program modules can be designed in or with the .Net Framework, .Net Compact Framework (and related languages ​​such as Visual Basic, C, etc.), Java, Python, R, C++, Objective-C, HTML, a combination of HTML / AJAX, XML, or HTML with included Java applets.

[0058]

[0063] Furthermore, while exemplary embodiments have been described herein, the scope of any and all embodiments includes equivalent elements, modifications, omissions, combinations (e.g., of aspects across various embodiments), adaptations, and / or variations, as will be understood by those skilled in the art based on this disclosure. Limitations in the claims should be interpreted broadly based on the language used in the claims and not limited to the examples described herein or examples in the prosecution of this application. Such examples should be construed as non-exclusive. Furthermore, the steps of the disclosed methods can be modified in any manner, including rearranging steps and / or inserting or deleting steps. Accordingly, it is intended that the specification and examples be considered as exemplary only, with the true scope and spirit being indicated by the appended claims and their full scope of equivalents.

Claims

1. accessing a database from which information relating to a group of individuals can be derived; providing a first biomarker associated with the cohort to a generic biomarker model, said generic biomarker model being trained based on one or more second biomarkers using said information, said first biomarker being different from said one or more second biomarkers; obtaining a first output from the generic biomarker model indicative of a first group of the population of individuals above a first likelihood threshold being tested for the first biomarker; and determining whether an individual in the first group of the population of individuals is a candidate for the cohort based on the first output; at least one processor programmed to perform A model support system including:

2. The model-assisted system of claim 1 , wherein the information comprises medical records associated with the population of individuals.

3. The model-assisted system of claim 2 , wherein the medical records include structured and unstructured information related to the population of individuals.

4. The model-aided system of claim 3 , wherein the unstructured information comprises text written by a medical professional, a radiation treatment report, or a pathology report.

5. The model-aided system of claim 4 , wherein the generic biomarker model is trained based on the unstructured information.

6. The model-aided system of claim 5 , wherein at least a portion of the unstructured information has been subjected to an optical character recognition process.

7. 2. The model-assisted system of claim 1, wherein determining whether the individual is a candidate for the cohort comprises verifying that the individual has been tested for the biomarker based on medical records associated with the individual.

8. the at least one processor: obtaining a second output from the generic biomarker model indicative of a second group of the population of individuals above a second likelihood threshold of testing positive for the first biomarker, the individual being included in the second group; 2. The model support system of claim 1, further programmed to:

9. 9. The model-assisted system of claim 8, wherein determining whether the individual is a candidate for the cohort comprises confirming that the individual has tested positive for the biomarker based on medical records associated with the individual.

10. The model-assisted system of claim 1 , wherein the at least one processor is further programmed to store the first output for access by a user of the generic biomarker model.

11. The model-assisted selection system of claim 1 , wherein the generic biomarker model uses a binary classification algorithm to generate the first output.

12. 12. The model-assisted selection system of claim 11, wherein the binary classification algorithm comprises at least one of a logistic regression, a random forest, a gradient boosting tree, a support vector machine, or a neural network.

13. The model-aided system of claim 1 , wherein the generic biomarker model is developed at least in part based on feature vectors extracted from the information based on the one or more second biomarkers.

14. The model-assisted system of claim 13 , wherein the feature vector includes at least one biomarker lexicon representing text associated with the at least one second biomarker.

15. The model-assisted selection system of claim 1 , wherein the one or more second biomarkers appear more frequently in the information than the first biomarker.

16. the at least one processor: providing the first biomarker to a biomarker-specific model, the biomarker-specific model being trained based on the first biomarker using the information; obtaining a third output from the biomarker-specific model indicative of a third group of the population of individuals above a likelihood threshold undergoing testing for the at least one biomarker; and verifying the accuracy of the generic biomarker model by comparing the first output with the third output.

2. The model support system of claim 1, further programmed to:

17. the at least one processor: searching the information for the first biomarker to generate a fourth output indicative of a fourth group of the population of individuals being tested for the at least one biomarker; and verifying the accuracy of the generic biomarker model by comparing the first output with the fourth output.

2. The model support system of claim 1, further programmed to:

18. 1. A computer-implemented method for identifying cohort candidates based on biomarkers, comprising: accessing a database from which information relating to a group of individuals can be derived; providing a first biomarker associated with the cohort to a generic biomarker model, said generic biomarker model being trained based on one or more second biomarkers using said information, said first biomarker being different from said one or more second biomarkers; obtaining a first output from the generic biomarker model indicative of a first group of the population of individuals above a first likelihood threshold being tested for the first biomarker; and determining whether an individual in the first group of the population of individuals is a candidate for the cohort based on the first output; A computer-implemented method comprising:

19. 20. The computer-implemented method of claim 18, wherein the information comprises medical records associated with the population of individuals.

20. 20. The computer-implemented method of claim 19, wherein the medical records include structured and unstructured information related to the population of individuals.

21. 21. The computer-implemented method of claim 20, wherein the unstructured information comprises text written by a medical professional, a radiation treatment report, or a pathology report.

22. 22. The computer-implemented method of claim 21, wherein the generic biomarker model is trained based on the unstructured information.

23. 20. The computer-implemented method of claim 18, wherein determining whether the individual is a candidate for the cohort comprises confirming that the individual has been tested for the biomarker based on medical records associated with the individual.

24. 20. The computer-implemented method of claim 18, wherein the likelihood threshold is adjustable based on the efficiency and performance level of the model.

25. accessing a database from which information relating to a group of individuals can be derived; providing a first characteristic associated with the cohort to a generic model, the generic model being trained using the information based on one or more second characteristics, the first characteristic being different from the one or more second characteristics; obtaining a first output from the generic model indicative of a first group of the population of individuals that exceed a first likelihood threshold associated with the first trait; and determining whether an individual in the first group of the population of individuals is a candidate for the cohort based on the first output; at least one processor programmed to perform A model support system including:

26. the first characteristic comprises a first biomarker; the one or more second characteristics include one or more second biomarkers; the first output indicating the first group of individuals being tested for the first biomarker; 26. The model-aided system of claim 25.

27. the first property comprises a first drug; the one or more second properties include one or more second drugs; the first output indicating the first group of individuals being treated with the first drug; 26. The model-aided system of claim 25.