Systems and methods for susceptibility modeling and screening based on multi-domain data analysis
By employing machine learning algorithms to analyze multi-domain data, the method optimizes cancer screening plans for individual patients, enhancing diagnostic accuracy and resource efficiency through personalized susceptibility modeling and screening.
Patent Information
- Application Number
- JP2025515398
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-14
- Filing Date
- 2023-09-12
- Publication Date
- 2025-09-25
AI Technical Summary
Existing healthcare diagnosis and management systems face challenges in efficiently utilizing vast amounts of data from multiple sources to provide personalized and accurate cancer susceptibility modeling and screening, often resulting in time- and resource-intensive processes that yield insufficient generalized guidance.
A computer-readable medium executes a method using machine learning algorithms to generate susceptibility and screening models based on multi-domain data, including proteomic, genetic, and environmental data, iteratively refining these models to optimize screening schedules and modalities for individual patients.
This approach enables personalized cancer screening plans that improve diagnostic accuracy, reduce false positives, and optimize resource utilization by tailoring screening frequencies and methods to patient-specific risk profiles.
Smart Images

Figure 2025531891000001_ABST
Abstract
Description
[Background technology]
[0001] In the field of healthcare diagnosis and management, ever-increasing amounts of data and data sources are now available to researchers, analysts, organizational entities, and others. This influx of information enables advanced analytics but also presents many new challenges in sifting through the available data and data sources to find the most relevant and useful information. As the use of technology continues to increase, so too does the availability of new data sources and information.
[0002] Due to the abundance of data available from numerous data sources, determining the optimal values and sources to use presents complex problems that are difficult to overcome. Accurately and fully utilizing available data across multiple sources can require both a team of individuals with extensive domain expertise and months or years of work to evaluate outcomes. This process can involve exhaustively searching large amounts of raw data to identify and study relevant data sources. Often, applying these types of analytical techniques to areas requiring accurate results obtainable only through time- and resource-intensive research is at odds with the demands of modern applications. For example, a process developed to evaluate outcomes may not match the specific situation or individual considerations. In this scenario, applying the process requires extrapolation to fit the specific situation, diluting the process's effectiveness, or requiring the expenditure of valuable time and resources to modify the process. As a result, processes developed in this manner typically provide only generalized guidance that is insufficient for reuse in other settings or by other users. As more detailed and personalized data becomes available, there is an increasing demand for the ability to accurately identify relevant data points from the ocean of information available across multiple data sources and efficiently apply that data across a myriad of personalized scenarios.
[0003] Multidomain data processing and analysis play a key role in the diagnosis and management of diseases such as cancer. Cancer is a complex and heterogeneous disease regulated by multiple factors across many domains, including genetic, molecular, cellular, tissue, population, environmental, and socioeconomic factors. For example, early cancer diagnosis can help improve outcomes by providing care at the earliest possible stage, providing patients with multiple benefits, including less invasive treatments, improved quality of life, and improved overall survival. As a result, public health programs are placing particular emphasis on effective screening strategies for early cancer detection. Summary of the Invention
[0004] Certain embodiments of the present disclosure relate to a non-transitory computer-readable medium including instructions executable by one or more processors to cause a system to perform a method for modeling and screening for cancer susceptibility. The method may include receiving input data associated with one or more cancers and a patient, determining a susceptibility model and a data enrichment rate based on the input data, and obtaining data associated with the patient from a plurality of data domains based on the susceptibility model and the data enrichment rate, where the patient data includes at least proteomic data. The method may also include using one or more machine learning algorithms to generate a set of susceptibility data associated with the patient based on the susceptibility model and the patient data. The method may also include using one or more machine learning algorithms to determine a screening model for the patient based on the susceptibility data, and generating one or more screening models based on the screening model. and screening patients for multiple types of cancer.
[0005] According to some disclosed embodiments, the input data may include at least cancer type, cancer prevalence, cancer prognosis, cancer stage, time of cancer diagnosis, or any combination thereof.
[0006] According to some disclosed embodiments, the patient data may further include patient characteristic data, medical history data, genetic data, immunological data, insurance data, health coverage data, environmental data, or biological sampling data from the patient.
[0007] According to some disclosed embodiments, the proteomic data may be based on a biological sample of the patient.
[0008] According to some disclosed embodiments, the method may further include determining, from the patient data, a set of features associated with the patient's susceptibility to cancer.
[0009] According to some disclosed embodiments, determining the susceptibility model may include selecting the susceptibility model from a plurality of data models in a model databank.
[0010] According to some disclosed embodiments, a patient-specific cancer screening model may include one or more screening methods, each associated with one or more screening schedules.
[0011] According to some disclosed embodiments, the method may further include iteratively refining the screening model using one or more machine learning algorithms by adjusting a screening schedule of the screening method based on the one or more outcome metrics until the one or more outcome metrics reach a threshold value.
[0012] According to some disclosed embodiments, the outcome metric may include a positive predictive value, a screening burden measure, an estimated risk measure, or any combination thereof.
[0013] According to some disclosed embodiments, the method may further include iteratively refining the susceptibility model using one or more machine learning algorithms by adjusting a data enrichment rate based on the one or more outcome metrics until the one or more outcome metrics reach a threshold value, and generating an improved set of susceptibility data based on the adjusted enrichment rate.
[0014] Other systems, methods, and computer-readable media are also described herein. [Brief explanation of the drawings]
[0015] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments and, together with the description, serve to explain the disclosed principles.
[0016] [Figure 1] FIG. 1 is a block diagram illustrating an exemplary system for determining and refining susceptibility and screening models based on data input from multiple data sources, according to some embodiments of the present disclosure. [Figure 2] FIG. 1 is a block diagram illustrating an example susceptibility model engine for generating and refining susceptibility data based on multi-domain data, in accordance with some embodiments of the present disclosure. [Figure 3] FIG. 1 is a block diagram illustrating an exemplary screening model engine for generating, validating, and refining screening adjustments and predictive outputs, according to some embodiments of the present disclosure. [Figure 4] FIG. 1 is a block diagram illustrating an exemplary machine learning platform, in accordance with some embodiments of the present disclosure. [Figure 5] FIG. 2 is a schematic diagram of an example server of a distributed system, in accordance with some embodiments of the present disclosure. [Figure 6]FIG. 1 is a flow diagram illustrating an exemplary process for receiving input based on potential outcomes, performing multi-domain data acquisition, generating individualized data, and performing screening and data and model refinement based on measured performance, according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0017] In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the disclosed exemplary embodiments. However, those skilled in the art will understand that the principles of the exemplary embodiments may be practiced without all the specific details. Well-known methods, procedures, and components have not been described in detail so as not to obscure the principles of the exemplary embodiments. Unless explicitly stated, the exemplary methods and processes described herein are not constrained to a particular order or sequence, or to a particular system configuration. Furthermore, some of the described embodiments or elements thereof may occur or be performed simultaneously, contemporaneously, or in parallel. Reference will now be made in detail to the disclosed embodiments, examples of which are illustrated in the accompanying drawings.
[0018] Considering the emerging paradigm for personalized cancer screening and diagnosis based on each patient's individual susceptibility to one or more specific cancer types, traditional analytical approaches focused solely on single-factor analysis (e.g., single-gene studies) may prove to be of limited value. Utilizing comprehensive patient proteomic data (e.g., by exploiting the myriad protein immune responses in the form of seropositives to tumor antigens) has the potential to complement these types of approaches for cancer detection. Furthermore, personalized cancer susceptibility modeling may be achieved by analyzing proteomic patterns in synergistic combination with data from multiple other domains, such as patient characteristics and environmental data. This type of data enrichment can help further tailor susceptibility modeling for each individual patient with respect to specific cancer types, effectively compensating for rare cancer types with low prevalence in a given population that are often difficult to detect and screen for.
[0019] Effective cancer susceptibility modeling then enables the optimization of cancer screening plans for patients. For example, traditional screening plans that do not consider a patient's individual susceptibility are often associated with low positive predictive value in cancer diagnosis, leading to high rates of false-positive diagnoses. This, in turn, requires additional confirmatory testing through expensive imaging tests (e.g., CT, MRI, PET scans, etc.) or invasive techniques such as tissue biopsies, resulting in unnecessary screening costs and potentially harming patients without reaching a reliable diagnosis. At the same time, screening schedules (e.g., the frequency with which a patient is administered a particular screening method) may be inappropriately adjusted without considering the patient's unique susceptibility to disease. For example, some screening techniques (e.g., colonoscopy, fine-needle aspiration biopsy) may not be cost-effective if they only result in a negligible amount of lead time in diagnosis (e.g., they have no impact on disease management) or may even be harmful if administered at high intervals to patients with a certain level of disease susceptibility.
[0020] In this regard, artificial intelligence systems and machine learning algorithms can be used to identify patients' susceptibility data. This can be a valuable tool for building a comprehensive yet personalized screening model for a given patient based on machine learning techniques. For example, a screening model based on machine learning techniques can help identify high-risk patients for a particular type of cancer according to their individual susceptibility, allowing for the calibration of screening plans based on susceptibility and risk stratification. For example, an exemplary system can increase screening frequency for high-risk patients or decrease it for low-risk patients. An exemplary system can also change screening modalities (e.g., deploying a costly and invasive procedure such as a colonoscopy for patients at high risk of colorectal cancer, or using a cost-effective and less invasive procedure such as a fecal occult blood test for low-risk patients). Furthermore, a machine learning-based screening model can maintain a high level of robustness through iterative refinement based on one or more performance measures (e.g., based on its predictive accuracy or screening burden) to determine the optimal screening schedule and the most reliable diagnosis, while also having the ability to output further patient-specific recommendations, such as lifestyle modifications.
[0021] The embodiments described herein provide techniques and techniques for evaluating numerous data sources and vast amounts of data used to create machine learning models. These techniques can use information related to the specific domain and application of the machine learning model to prioritize potential data sources. Furthermore, the techniques and techniques herein can interpret available data sources and data to extract probabilities and outcomes associated with the specific domain and application of the machine learning model. The described techniques can synthesize data into a coherent machine learning model, which can be used to analyze and compare various paths or courses of action.
[0022] These technologies can efficiently evaluate data sources and data, prioritize their importance based on domain- and situation-specific needs, and provide effective and accurate predictions that can be used to evaluate potential courses of action. The technologies and methods allow data models to be applied to individualized situations. These methods and technologies enable detailed evaluations that can potentially improve decision-making. Furthermore, these technologies can evaluate systems in which the process for evaluating data outcomes is easily configured and can be reused by other uses of the technology.
[0023] The technology can utilize machine learning models to automate the process and predict responses without human intervention. The performance of such machine learning models is typically improved by providing more training data. The predictive quality of the machine learning model is manually evaluated to determine whether the machine learning model requires further training. Embodiments of these described technologies can help improve the predictions of the machine learning model using prediction quality metrics requested by the user.
[0024] FIG. 1 is a block diagram illustrating various exemplary components of a system 100 for determining and refining susceptibility and screening models based on data input from multiple data sources, consistent with embodiments of the present disclosure. The system 100 can include a data input engine 110, which can further include a data extractor 111, a data converter 112, and a data loader 113. The data input engine 110 can process data from the data sources 101-104. In some embodiments, the data input engine 110 can be implemented using a computing device. For example, the data from the data sources 101-104 can be obtained via an I / O device or a network interface. Furthermore, the data can be stored in appropriate storage or system memory during processing. The data input engine 110 can also interact with a data storage 115, which can further include a data processor 116. , may be implemented on a computing device that stores data in storage or system memory. System 100 may include a characterization engine 120. Characterization engine 120 may comprise an annotator 121, a data censor 122, a summarizer 123, and a Boolean operator 124. System 100 may also include an analysis engine 130 and a feedback engine 140. Like data input engine 110, characterization engine 120 may be implemented on a computing device. Similarly, characterization engine 120 may utilize storage or system memory to store data and may utilize one or more I / O devices or network interfaces to transmit or receive data. Each of the data input engine 110, data extractor 111, data transformer 112, data loader 113, characterization engine 120, annotator 121, data censor 122, summarizer 123, Boolean operator 124, analysis engine 130, and feedback engine 140 may be a module, which is a packaged functional hardware unit designed for use with other components or portions of programs that perform specific functions of a related function. Each of these modules may be implemented using a computing device. Each of these components is described in more detail below. In some embodiments, the functionality of system 100 may be divided across multiple computing devices to enable distributed processing of data. In these embodiments, the different components may communicate via one or more I / O devices or network interfaces.
[0025] System 100 can be relevant to many different areas or fields of use. Description of embodiments relevant to particular areas, such as disease diagnosis and management (e.g., related to cancer), is not intended to limit the disclosed embodiments to those particular areas; embodiments consistent with the present disclosure can be applied to any area that utilizes predictive modeling based on available data.
[0026] The data input engine 110 is a module that can retrieve data from various data sources (e.g., data sources 101, 102, 103, and 104) and process the data for use by the rest of the system 100. The data input engine 110 can further include a data extractor 111, a data transformer 112, and a data loader 113.
[0027] Data extractor 111 retrieves data from data sources 101, 102, 103, and 104. Each of these data sources can represent a different type of data source. For example, data source 101 can be a database. Data source 102 can represent structured data. Data sources 103 and 104 can be flat files. Furthermore, data sources 101-104 can include overlapping or entirely different datasets. In some embodiments, data sources 101-104 can include input data associated with potential outcomes. For example, in a cancer screening or diagnostic setting, the input data can include data associated with cancer type, cancer prevalence, cancer prognosis, time of cancer diagnosis, or cancer stage. In some embodiments, data sources 101-104 can include values including data enrichment ratios. In some embodiments, the data enrichment ratio values are based on user input. In some embodiments, data source 101 can include proteomic data, while data sources 102, 103, and 104 include various data from other domains or sources. For example, data source 102 may include patient characteristic data such as the patient's age, sex, race / ethnicity, height / weight, etc. In another example, data source 103 may include environmental data including smoking history, diet type, etc. In another example, data source 104 may include data from biological sampling from the patient, such as data regarding blood, plasma, serum, or urine samples. In another example, the data sources 101-104 may include data obtained from a variety of sources. In another example, the data sources 101-104 may include demographic data, medical history data, clinical visit data, or surgical history data. In another example, the data sources 101-104 may include family history data. In another example, the data sources 101-104 may include genetic data or immunological data. In another example, the data sources 101-104 may include insurance data or health coverage data. The data extractor 111 can interact with the various data sources to extract relevant data and provide the data to the data converter 112.
[0028] Data converter 112 can receive data from data extractor 111 and process the data into a standard format. In some embodiments, data converter 112 can normalize data, such as dates or numeric values, based on specific units of measurement. For example, data source 101 can store dates in day-month-year format, data source 102 can store dates in year-month-day format, or data source 103 can store weight measured in kilograms or pounds. In this example, data converter 112 can correct the data provided via data extractor 111 into a consistent date format or standardized unit format, respectively. Thus, data converter 112 can effectively clean the data provided via data extractor 111 so that all of the data, although originating from various sources, has a consistent format.
[0029] Additionally, the data converter 112 can extract additional data points from the data. For example, the data converter can process dates in year-month-day format by extracting separate data fields for the year, month, and day. The data converter can also perform other linear and non-linear transformations and extractions on categorical and numerical data, such as normalization and discrimination. The data converter 112 can provide the transformed or extracted data to the data loader 113.
[0030] Data loader 113 can receive the normalized data from data transformer 112. Data loader 113 can merge the data into various formats depending on the specific requirements of system 100 and store the data in an appropriate storage mechanism, such as data storage 115. In some embodiments, data storage 115 can be data storage for a distributed data processing system (e.g., Hadoop Distributed File System, Google File System, ClusterFS, or OneFS). In some embodiments, data storage 115 can be a relational database (described in more detail below). Depending on the particular embodiment, data loader 113 can optimize the data for storage and processing in data storage 115. In some embodiments, various types of data structures, such as databases, flat files, data stored in memory (e.g., system memory), or data stored in any other suitable storage mechanism, can be stored in data storage 115 by data loader 113.
[0031] System 100 may optionally include a characterization engine 120. The characterization 120 may process data prepared by data input engine 110 and stored in data storage 115. The characterization engine 120 may include an annotator 121, a data censor 122, a summarizer 123, and a Boolean operator 124. The characterization may retrieve data from data storage 115 prepared by data input engine 110. Various types of data structures may be suitable inputs to the characterization engine 120, such as, for example, a database, a flat file, data stored in memory (e.g., system memory), or data stored in any other suitable storage mechanism that may be stored in data storage 115 by data loader 113.
[0032] The featurization engine 120 can then convert the data into features that can be used for further analysis. Features can be data that represent other data. Features can be determined based on domain, categorical data type, or many other factors associated with the data stored in the data structure. Furthermore, features can represent information about multiple data records in a dataset or about a single category within a data record. Moreover, multiple features can be generated to represent the same data.
[0033] The featurization engine 120 can include an annotator 121. The annotator 121 can provide context to the data structures from the data storage 115. The annotator 121 can further determine which additional data records are associated with the event of interest and should be used in the predictive model. After the annotator 121 processes and identifies relevant restrictions on the data, the data censor 122 can filter out data that does not meet the established criteria. After the data is censored, the summarizer 123 can analyze the remaining data structures and data to generate features for the dataset. In some embodiments, the features can be based on the particular type of data under consideration, and many features can be generated from a single data point or a set of data points. The summarizer 123 can further consider data points that occur across multiple data records for an individual, or can consider data points related to multiple individuals.
[0034] After features are established for a particular data set, the established features can be stored in data storage 115, provided directly to analysis engine 130, or provided to Boolean operator 124 for further processing before analysis.
[0035] The Boolean operator 124 can process the determined features from the summarizer 123 and establish corresponding Boolean or binary data for the features. Using a binary representation of the features can allow the dataset to be analyzed using statistical analysis techniques optimized for binary data. The Boolean operator 124 can generate a Boolean or binary value based on whether a particular feature or attribute is present or not. For example, a data feature that establishes whether a particular type of request is present or not for a user can be easily represented by a "1" for "true" and a "0" for "false." In this example, the feature could be whether an individual has a particular protein level above a particular numeric threshold, or whether an individual is a smoker, or whether they have a family history of cancer.
[0036] After processing the data, the featurization engine 120 can generate feature data directly from the summarizer 123 or binary feature data from the Boolean operator 124. This data can be stored in the data storage 115 for later analysis or passed directly to the analysis engine 130.
[0037] The analysis engine 130 can analyze data stored by the data loader 113 or in the data storage 115. In some embodiments, the analysis engine 130 can analyze normalized data based on multiple data sources, as exemplified by the data sources 101-104. The analysis engine 130 may include a data model selector 131, a susceptibility model engine 132, or a screening model engine 134.
[0038] The data model selector 131 can select a data model from the model databank 108. The selected data model can be a type of susceptibility model or a type of screening model. The selected model can be input to the data input engine 110 or the data In some embodiments, the data model selector 131 may select a data model based on input data from the data entry engine 110 or data storage 115, or based on information requested by the user interface 150. In some embodiments, the exemplary input data may include data associated with cancer type or cancer prevalence. In some embodiments, the exemplary input data may include patient-related data such as age, sex, race / ethnicity, smoking history, or personal or family history of cancer.
[0039] In some embodiments, the data model selector 131 can send the selected data model to the susceptibility model engine 132. The susceptibility model engine 132 can receive the data model from the data model selector 131 or the model databank 108. The susceptibility model engine 132 can receive data from the data loader 113 or from data stored in the data storage 115. In some embodiments, the susceptibility model engine 132 can receive normalized data based on multiple data sources, as exemplified by the data sources 101-104. The susceptibility model engine 132 can receive values including data enrichment rates from the data input engine 110 or the data storage 115. In a cancer screening or diagnostic setting, the susceptibility model engine 132 can also receive input data associated with cancer type, cancer prevalence, cancer prognosis, time of cancer diagnosis, or cancer stage. The susceptibility model engine 132 can also receive input data. In some embodiments, the susceptibility model engine 132 can perform data enrichment on the received data. In some embodiments, the susceptibility model engine 132 can optionally perform feature selection based on the received data. The susceptibility model engine 132 can also apply the received data to a selected model. The susceptibility model engine 132 can also estimate susceptibility. The susceptibility model engine 132 can also generate susceptibility data based on the estimated susceptibility. In some embodiments, the susceptibility model engine 132 can also validate the generated data based on input from the feedback engine 140. In some embodiments, the susceptibility model engine 132 can also refine the generated data based on data validation input. In some embodiments, the data refinement is performed by adjusting a data enhancement rate. In some embodiments, the data validation input is based on the feedback engine 140. In some embodiments, the susceptibility model engine 132 can send the susceptibility data to the screening model engine 134.
[0040] In some embodiments, the data model selector 131 can send the selected data model to the screening model engine 134. The screening model engine 134 can receive the data model from the data model selector 131. The screening model engine 134 can receive data from the data loader 113 or from data stored in the data storage 115. In some embodiments, the screening model engine 134 can receive normalized data based on multiple data sources, as exemplified by the data sources 101-104. The screening model engine 134 can receive susceptibility data from the susceptibility model engine 132. In some embodiments, the screening model engine 134 can select a screening method. The screening model engine 134 can also perform adjustment of the screening frequency of the selected screening method. In some embodiments, the screening model engine 134 can apply the screening method at the adjusted screening frequency to generate predicted output values. The screening model engine 134 can also validate the predicted output values based on input from the feedback engine 140. In some embodiments, the susceptibility model engine 134 can select a screening method. The screening model engine 134 can also perform adjustment of the screening frequency of the selected screening method. In some embodiments, the screening model engine 134 can apply the screening method at the adjusted screening frequency to generate predicted output values. The screening model engine 134 can also validate the predicted output values based on input from the feedback engine 140. The rule engine 134 may also perform screening refinement. In some embodiments, screening refinement is performed by selecting new screening methods or adjusting screening frequencies.
[0041] The analysis engine 130 can optionally analyze the features or binary data generated by the featurization engine 120 to determine which features are most indicative of the occurrence of an event of interest. The analysis engine 130 can use a variety of methods to analyze the thousands, millions, or billions of features that may be generated by the featurization engine 120. Example analysis techniques include feature subset selection, stepwise regression testing, chi-squared (χ) testing, or other regularization methods that promote sparsity (e.g., coefficient shrinkage). The analysis engine 130 can use this output to generate a model to apply to existing and future data to identify individuals who are likely to experience an event of interest (e.g., a positive diagnosis of a type of cancer).
[0042] The analysis engine 130 can store the data model in data storage 115 for future use. Additionally, the data model can be provided to the feedback engine 140 for refinement. The feedback engine 140 can apply the data model to a wider dataset to determine the accuracy of the model. The feedback engine 140 can utilize one or more outcome metrics, such as 240 in FIG. 2 and 340 in FIG. 3, to measure the performance of the model. Based on those results, the feedback engine 140 can report the results to the analysis engine 130. The feedback engine 140 can report to the susceptibility model engine 132 to refine the generated susceptibility data (as shown in FIG. 2). The feedback engine 140 can also report to the screening model engine 134 to refine the screening model (as shown in FIG. 3). In some embodiments, the feedback engine 140 can optionally report to the featurization engine 120 to iteratively update certain inputs used by the annotator 121, the data censor 122, and the summarizer 123 to tune the models. In this way, the featurization engine 120 can train as more and more data is analyzed.
[0043] In some embodiments, the analysis engine 130 can use various statistical analysis techniques to test the accuracy and usefulness of a particular model or multiple models generated for a target event. Models can be evaluated using evaluation metrics such as precision, refinement, accuracy, area under the receiver operating characteristic (ROC) curve, area under the precision-recall (PR) curve, lift, or accuracy in rank, among others. The feedback engine 140 can provide feedback aimed at optimizing the model based on the model's particular domain and use case. For example, in the context of cancer diagnosis, if a model is being used to identify individuals with a high susceptibility to a particular type of cancer, the feedback engine 140 can provide feedback and adjustments to the data model selector 131, the susceptibility model engine 132, the screening model engine 134, or optionally the characterization engine 220 to optimize the model for greater accuracy to ensure diagnostic accuracy by minimizing false positives, understanding that false positives can lead to unwarranted additional testing that may be costly or invasive. Additionally, the feedback engine 140 can test the data model using techniques such as cross-validation to optimize the number of features selected for the model by the analysis engine 130.
[0044] The system 100 may further include a user interface 150. The user interface 150 is a graphical user interface implemented on a computing device that utilizes a graphics memory, a GPU, and a display device. The user interface 150 may be a graphical user interface (GUI). The user interface 150 may provide a representation of data from the analysis engine 130 or the feedback engine 140, or optionally the characterization engine 120. The user interface 150 may be a read-only interface that does not accept user input. In some embodiments, the user interface 150 may accept user input to control the representation. In other embodiments, the user interface 150 may accept user input to control or modify components of the system 100. The user interface 150 may be text-based or may include graphical components that represent the displayed data.
[0045] In some embodiments, user interface 150 can provide a user with the ability to make recommendations based on the predictive models generated by system 100. For example, system 100 can be used to generate a predictive model for the diagnosis of a particular cancer type. The results of this model can be presented to patients whose data may indicate they have a certain level of susceptibility to that particular cancer type. Individual users do not have insight into the particular data model itself, but benefit from the ability to seek preventative care based on the diagnosis.
[0046] In some embodiments, user interface 150 may provide a representation of the functionality of characterization engine 120, analysis engine 130, or feedback engine 140. In some embodiments, user interface 150 may display feedback information from feedback engine 140. In these embodiments, a domain expert may use user interface 150 to validate the generated model, provide feedback on the generated model, and / or modify inputs or data used by analysis engine 130 or characterization engine 120 to generate the model.
[0047] System 100 can be used as described to quickly and accurately generate effective predictive models across many different domains. Instead of requiring labor- and time-intensive methods to generate narrow predictive models, system 100 can be used to quickly generate and iterate predictive models that are general enough to be applied to a wide range of future data, while utilizing statistically significant features to optimally predict events of interest.
[0048] FIG. 2 is a block diagram illustrating an example susceptibility model engine 210 for generating and refining susceptibility data based on multi-domain patient data 201, according to some embodiments of the present disclosure.
[0049] The susceptibility model engine 210 may include a data enrichment engine 212, a model application engine 216, a susceptibility estimation engine 218, a data generation engine 220, a data validation engine 222, and a data refinement engine 224. The susceptibility model engine 210 may optionally include a feature selector engine 214. In some embodiments, the susceptibility model engine 210 may be exemplified by the susceptibility model engine 132 shown in FIG. 1 .
[0050] The multi-domain patient data 201 may include proteomic data 202, patient characteristic data 203, environmental data 204, biological sampling data 205, medical history data 206, or genetic data 207. In some embodiments, the multi-domain patient data 201 may be exemplified by data stored by the data loader 113 or in the data storage 115. In some embodiments, the multi-domain patient data 201 may include proteomic data 202, patient characteristic data 203, environmental data 204, biological sampling data 205, medical history data 206, or genetic data 207. In some embodiments, the multi-domain patient data 201 may be exemplified by data stored by the data loader 113 or in the data storage 115. Data 201 may be exemplified by normalized data based on multiple data sources, such as data sources 101 to 104 in FIG.
[0051] The susceptibility model engine 210 can receive multi-domain patient data 201 as input from multiple data sources. In some embodiments, the multiple data sources may include, but are not limited to, data from multiple domains or modalities. For example, in a disease diagnosis setting (e.g., for cancer detection), the data sources may include, but are not limited to, proteomic data 202, patient characteristic data 203, environmental data 204, biological sampling data 205, medical history data 206, family history data, genetic data 207, or immunological data. The data sources may also include insurance or health coverage data.
[0052] The proteomic data may include data associated with one or more sets of proteins obtained from a patient's biological sampling. In some embodiments, the biological sampling may include obtaining a blood, body fluid, or tissue sample from the patient. The proteomic data may include individual abundance values for the set of proteins obtained from the biological sampling, or individual immune responses for the set of proteins obtained from the biological sampling in the form of seropositivity to one or more tumor antigens. In some embodiments, the proteomic data may include mass spectrometry data, data associated with identified peptides, post-translational modification data, fluorescence microscopy data, protein subcellular localization data, fluorescence energy transfer experimental data, protein region and three-dimensional structure prediction data, protein-protein interaction data, or protein-protein interaction prediction data associated with one or a set of proteins obtained from the biological sampling.
[0053] In some embodiments, the patient characteristic data may include data associated with the patient's age, sex, race / ethnicity, and height / weight. In some embodiments, the patient characteristic data may further include demographic data. In some embodiments, the environmental data may include data associated with smoking history or diet type (e.g., a primarily red meat diet or a plant-based diet). In some embodiments, the biological sampling data may include data regarding blood, plasma, serum, or urine samples obtained from the patient. In some embodiments, the medical history data may include clinical visit data or surgical history data. In some embodiments, the family history data may include a genetic history of a positive cancer diagnosis or treatment associated with the patient's family. In some embodiments, the genetic data may include data associated with tumor-associated genetic markers for the patient. In some embodiments, the immunological data may include data associated with an immune or autoimmune response associated with one or more tumor markers, or an individualized response to immunotherapy.
[0054] In some embodiments, the multi-domain patient data 201 may be processed, transformed, normalized, or stored in a data storage by a data entry engine. In some embodiments, the data entry engine may be exemplified by data entry engine 110 of FIG. 1. In some embodiments, a characterization engine may extract a set of data features from the multi-domain patient data 201. In some embodiments, the characterization engine may be exemplified by characterization engine 120 of FIG. 1.
[0055] The susceptibility model engine 210 may include a data enrichment engine 212. The data enrichment engine 212 may receive values including a data enrichment rate from the data input engine 110 or the data storage 115. In some embodiments, the data enrichment rate value may be a factor (e.g., 2x, 5x, 10x, etc.) or a percentage (e.g., 10%, 20%, etc.). ) The data enrichment engine 212 may enrich the multi-domain patient data 201 based on a data enrichment rate. For example, based on a 10% data enrichment rate, the data enrichment engine 212 may accordingly increase the amount of data received from the data input engine 110 or data storage 115 by 10%. In another example, the data enrichment engine 212 may double the amount of data received based on a 2x data enrichment rate.
[0056] The susceptibility model engine 210 may optionally include a feature selector engine 214. The feature selector engine 214 may receive a set of features from the featurization engine 120, as in FIG. 1. In some embodiments, the set of features may be based on input data from the data input engine 110, data from the data storage 115 (as shown in FIG. 1). In some embodiments, the set of features may be based on multi-domain patient data 201.
[0057] The susceptibility model engine 210 may include a model application engine 216. In some embodiments, the model application engine 216 may receive a selected data model from the data model selector 130 and the analysis engine 130, as shown in FIG. 1 . The model application engine 216 may apply a subset of the multi-domain patient data 201 to the selected data model by selecting data that conforms to one or more data parameters of the selected data model. In some embodiments, the selected subset of the multi-domain patient data 201 may be enriched by the data enrichment engine 212. For example, the model application engine 216 may select a data model based on proteomic data or environmental data. The model application engine 216 may select a subset of the multi-domain patient data 201 that includes proteomic data 202 or environmental data 204. The data enrichment engine 212 may enrich the selected data based on a data enrichment rate. The model application engine 216 may apply the data to the corresponding data parameters in the selected data model. In some embodiments, the model application engine 216 may optionally apply the set of features received from the feature selector engine 214 to the data parameters in the selected model.
[0058] The susceptibility model engine 210 may include a susceptibility estimation engine 218. The susceptibility estimation engine 218 may calculate a susceptibility score based on the application of data to a selected data model or a selected model from the model application engine 216. In some embodiments, for example in a cancer diagnostic setting, the susceptibility score may include a numerical percentage representing the patient's likelihood of having a particular type of cancer. In some embodiments, the susceptibility score may include a category (e.g., high risk, medium risk, or low risk) representing the patient's risk stratification level for a particular type of cancer. In some embodiments, the susceptibility score may include a multiplicative coefficient (e.g., a 2-fold or 10-fold risk of developing cancer) representing the patient's risk of developing a particular type of cancer compared to a reference population.
[0059] The susceptibility model engine 210 may include a data generation engine 220. The data generation engine 220 may generate a set of susceptibility data based on a data model and application of data (e.g., a subset of the multi-domain patient data 201) to the data model from the model application engine 216. The data generation engine 220 may generate a set of susceptibility data based on susceptibility scores from the susceptibility estimation engine 218. In a cancer diagnostic setting, the data generation engine 220 may generate a set of susceptibility data for an individual patient for a particular cancer type. The generated susceptibility data may include the susceptibility scores from the susceptibility estimation engine 218 or further data associated with the susceptibility estimation performed by the engine 218. For example, For example, the generated susceptibility data may include a susceptibility score (e.g., in the form of a risk stratification ratio or category), a subset of the multi-domain patient data 201 (e.g., proteomic data 202, environmental data 204), or data parameters from a data model selected from the data model selector 131 (as in FIG. 1).
[0060] The susceptibility model engine 210 may include a data validation engine 222. The data validation engine 222 may receive the susceptibility data generated by the data generation engine 220. In some embodiments, the data validation engine 222 may output the generated susceptibility data, as exemplified by output susceptibility data 260. In some embodiments, the output susceptibility data 260 may be stored in a database, as exemplified by data storage 115 or data sources 101-104 shown in FIG. 1. The data validation engine 222 may perform validation of the generated susceptibility data by interacting with a feedback engine 226. In some embodiments, the feedback engine 226 may be exemplified by feedback engine 140 of FIG. 1. In some embodiments, the feedback engine 226 may apply the data model to a broader dataset to determine the accuracy of the model. In some embodiments, the feedback engine 226 may utilize one or more outcome metrics, such as outcome metrics 240, to generate a set of performance measurement data based on the generated susceptibility data or the selected data model. In some embodiments, the outcome metrics 240 may include a positive predictive value, a screening burden measure, or an estimated risk measure. For example, in a cancer diagnosis setting, the outcome metric 240 may include a positive predictive value as an indicator of the proportion of correct prediction outputs in the case of a cancer diagnosis by the prediction output generation engine 308 in the screening model engine 310 (as shown in FIG. 3 ). The screening burden measure may be a measure based on a cost-effectiveness analysis of applying a cancer screening methodology at a particular screening frequency by the screening application engine 306 in the screening model engine 310 (as shown in FIG. 3 ). The screening burden may be a measure of cost-effectiveness based on patient insurance or Medicare data. The estimated risk measure may be a measure based on any reported harm to patients (e.g., from invasive screening techniques such as needle biopsy) resulting from applying a screening methodology at a particular screening frequency.In some embodiments, the set of performance measure data may include a numerical analysis based on positive predictive values from outcome metrics 240. In some embodiments, the set of performance measure data may include a numerical cost-effectiveness analysis of applying a screening method (e.g., a cancer screening method in a cancer diagnostic setting) at a given screening frequency. In some embodiments, the set of performance measure data may include a probabilistic risk analysis based on reported harms (e.g., from invasive screening techniques to patients in a cancer diagnostic setting) resulting from applying a screening method at a given screening frequency.
[0061] The feedback engine 226 can send the performance measurement data to the data validation engine 222. The data validation engine 222 can generate a data validation score for data refinement based on the performance measurement data. In some embodiments, the data validation score for data refinement may include a binary value (e.g., "data refinement needed," "data refinement not needed") or a gradient of a value on a numerical scale (e.g., on a scale of 1 to 10, the need for data refinement for a particular data set is "6 out of 10"). The data validation engine 222 can send the data validation score for data refinement to the data refinement engine 224.
[0062] The susceptibility model engine 210 may include a data refinement engine 224. The data refinement engine 224 may receive data validation scores for data refinement from the data validation engine 222 or performance measurement data from the feedback engine 226. The data refinement engine 224 can perform data refinement based on input from the data validation engine 222 or performance measurement data from the feedback engine 226. The data refinement engine 224 can perform data refinement by calculating an adjusted data refinement rate based on the performance measurement data. The data refinement engine 224 can send the adjusted data refinement rate to the data refinement engine 212 for iterative data refinement.
[0063] In some embodiments, the susceptibility model engine 210 can enable iterative cycles of data refinement. The iterative cycles of data refinement may include the data refinement engine 224 sending an adjusted data enrichment rate to the data enrichment engine 212. The iterative cycles may include the data enrichment engine 212 performing data enrichment based on the adjusted data enrichment rate. For example, if the adjusted data enrichment rate is determined to be 20% by the data refinement engine 224 based on the output of the data validation engine 222 or the feedback engine 226, and the original unadjusted data enrichment rate was 10%, the data enrichment engine 212 may increase the data enrichment rate by 10% to generate an improved dataset. In some embodiments, the improved dataset is a subset of the multi-domain patient data 201 calibrated by the adjusted data enrichment rate. The iterative cycles of data refinement may optionally include the feature selector engine 214 selecting a set of features based on the improved dataset. The iterative cycles of data refinement may also include the model application engine 216 applying the improved dataset to the selected data model. The iterative cycle of data refinement may include the susceptibility estimation engine 218 performing susceptibility estimation based on the improved dataset and data model. The iterative cycle of data refinement may also include the data generation engine 220 generating a set of improved susceptibility data based on output from the susceptibility estimation engine 218. The iterative cycle of data refinement may also include the data validation engine 222 interacting with the feedback engine 226 to validate the improved susceptibility data. In some embodiments, the iterative cycle of data refinement may continue until the data validation engine 222 or the feedback engine 226 determines that one or more outcome metrics 240 have reached a threshold. In some embodiments, the data validation engine 222 may perform output of the set of improved susceptibility data, as exemplified by output susceptibility data 260.In some embodiments, the output sensitivity data 260 may be stored in a database, such as exemplified by the data storage 115 or data sources 101-104 shown in FIG.
[0064] The susceptibility model engine 210 can interact with the machine learning engine 230. According to some embodiments, the machine learning engine 230 can use one or more machine learning models to analyze one or more datasets or data subsets utilized by the susceptibility model engine 210. The machine learning engine 230 may be trained using output from the data validation engine 222 or performance measurement data from the feedback engine 226 based on one or more outcome metrics 240. The machine learning engine 230 may be configured to predict an optimal set of datasets or data subsets based on the training data. In some embodiments, the machine learning engine 230 may be configured to predict an optimal adjusted data enrichment rate based on the training data. In some embodiments, the optimal dataset or data subset may include data associated with a set of susceptibility data generated by the data generation engine 220 that has an optimal data validation score from the data validation engine 222. For example, in a cancer diagnostic setting, the machine learning engine 230 may determine whether a subset of multi-domain patient data 201 (e.g., including a particular combination of multi-domain data such as proteomic data, patient characteristic data, or environmental data) is associated with a particular It can be determined that in combination with the data model, it can be used to generate a set of susceptibility data associated with an optimal data validation score based on the output from the data validation engine 222.
[0065] The machine learning engine 230 can measure the effectiveness of one or more outcomes based on one or more outcome metrics, such as those exemplified by outcome metrics 240. The machine learning engine 230 can measure the effectiveness of an outcome based on output from the data validation engine 222 or the feedback engine 226. The machine learning engine 230 can perform iterative cycles of training and hypothesis refinement by automatically generating alternative hypotheses based on different subsets of data in the multi-domain patient data 201 or different data enrichment rates. In some embodiments, the machine learning engine 230 can also perform hypothesis refinement based on user-defined data selections or user-defined data enrichment rates as part of the input data from the data input engine 110 of FIG. 1. The machine learning engine 230 can perform iterative cycles of hypothesis generation, validation, and refinement until an effectiveness outcome measure, such as that from outcome metrics 240, reaches a threshold value. In some embodiments, the machine learning engine 230 may be exemplified by a machine learning platform 402 (shown in FIG. 4).
[0066] FIG. 3 is a block diagram illustrating an exemplary screening model engine for generating, validating, and refining screening adjustments and predictive outputs, according to some embodiments of the present disclosure.
[0067] The screening model engine 310 may include a screening method selector engine 302, a screening frequency adjustment engine 304, a screening application engine 306, a predicted output generation engine 308, an output validation engine 309, and a screening refinement engine 312. In some embodiments, the screening model engine 310 may be exemplified by the screening model engine 134 shown in FIG.
[0068] The screening method selector engine 302 can receive as input a set of susceptibility data 301. In some embodiments, the susceptibility data 301 may be generated by a susceptibility model engine, exemplified by the susceptibility model engine 210 in FIG. 2. In some embodiments, the susceptibility data 301 may be exemplified by the output susceptibility data 260 in FIG. 2. In some embodiments, the multi-domain patient data 201 shown in FIG. 2 can also be input to the screening model engine 310. The screening method selector engine 302 can select one or more screening methods from a database of screening methods. In some embodiments, the database of screening methods may be exemplified by the data storage 115 or data sources 101-104 shown in FIG. 1. For example, in a cancer diagnostic setting, the database of screening methods may include screening methods corresponding to a particular type of cancer (e.g., mammography for breast cancer or colonoscopy for colon cancer). The screening method selector engine 302 can select one or more screening methods based on the susceptibility data 301. In some embodiments, the screening method selector engine 302 may select a screening method using one or more machine learning algorithms based on the machine learning engine 330. The screening method selector engine 302 may send the selected screening method to the screening frequency adjustment engine 304.
[0069] The screening frequency adjustment engine 304 can receive one or more screening methods from the screening method selector engine 302. The screening frequency adjustment engine 304 can select one or more screening methods based on the susceptibility data 301. The frequency of a method can be set or adjusted. For example, in a cancer screening setting, the screening frequency adjustment engine 304 can set high-frequency interval screening for patients with a high susceptibility to a particular cancer type or low-frequency interval screening for patients with a low susceptibility. The screening frequency adjustment engine 304 can also set the screening frequency based on a cost-effectiveness analysis of the screening method. The screening frequency adjustment engine 304 can also set the screening frequency based on an analysis of the estimated harm of applying the screening method to an individual patient. The screening frequency adjustment engine 304 can also set the screening frequency based on screening guidelines from an external data source, such as exemplified by the data storage 115 or data sources 101-104 in FIG. 1. The screening frequency adjustment engine 304 can send the selected screening method associated with the screening frequency to the screening application engine 306.
[0070] The screening application engine 306 can perform application of one or more selected screening methods based on the associated screening frequency, as determined by the screening frequency adjustment engine 304. The screening application model 306 can generate a personalized screening model based on one or more screening methods, the associated screening frequency, or the susceptibility data 301. In some embodiments, such as in a cancer screening setting, the screening application engine 306 can simulate the application of a screening method to an individual patient over a time interval based on the personalized screening model. In some embodiments, the screening application engine 306 can simulate the application of a screening method based on the multi-domain patient data 201 or the susceptibility data 301 generated by the susceptibility model engine 210. In some embodiments, the screening application engine 306 can output a set of personalized application data based on the personalized screening model. In some embodiments, the set of personalized application data may include results associated with the application of the screening method. For example, if the selected screening method is mammography, the application data may include a BI-RADS score or a detailed description of mammographic findings. The set of application data may also include periodic or progressive data associated with the screening method over multiple time intervals according to the screening frequency. For example, if the screening frequency for a mammography screening method is set to annual (i.e., every year), the set of application data may include a data series of BI-RADS scores or other consecutive mammographic findings for an individual patient over a predetermined number of years. In some embodiments, the set of personalized application data may also include cost-effectiveness data associated with the screening method at the screening frequency over a predetermined time period. In some embodiments, the cost-effectiveness data may include insurance or Medicare data associated with an individual patient.In some embodiments, the set of personalized application data may also include estimated risk data associated with any adverse events or harms from applying the screening method to an individual patient. The screening application engine 306 can send the set of personalized application data or personalized screening models to the predictive output generation engine 308.
[0071] The predictive output generation engine 308 can perform data analysis based on the personalized application data or a set of personalized screening models from the screening application engine 306. For example, in a cancer screening and diagnostic setting, the predictive output generation engine 308 can perform data analysis based on personalized application data associated with one or more cancer screening methods administered at a particular screening frequency within a personalized screening model. The predictive output generation engine 308 generates a set of predictive output data associated with the diagnosis of one or more types of cancer. In some embodiments, the set of predictive output data may include a binary value (e.g., a positive cancer diagnosis versus a negative cancer diagnosis). In some embodiments, the set of predictive outputs may include a numerical scale of likelihood (e.g., on a scale of 1 to 10, where the likelihood of having a particular type of cancer is 6 for an individual patient). In some embodiments, the set of predictive outputs may include risk stratification categories (e.g., for an individual patient, there is a high risk, intermediate risk, or low risk for a particular cancer type). In some embodiments, the set of predictive output data may include data associated with the risk of a particular cancer type within a predetermined time frame of obtaining the patient data (e.g., within a certain number of years of biological sampling of the patient).
[0072] The screening model engine 310 may include an output validation engine 309. The output validation engine 309 may receive the set of predicted output data generated by the predicted output generation engine 308. In some embodiments, the output validation engine 309 may perform output of the set of predicted output data, as exemplified by output data 360. In some embodiments, the output data 360 may be stored in a database, as exemplified by data storage 115 or data sources 101-104 shown in FIG. 1. The output validation engine 309 may perform validation of the generated predicted output data by interacting with a feedback engine 320. In some embodiments, the feedback engine 320 may be exemplified by feedback engine 140 of FIG. 1. In some embodiments, the feedback engine 320 may apply the personalized screening model from the screening application engine 306 to a broader dataset to determine the accuracy of the model. In some embodiments, the feedback engine 320 may utilize one or more outcome metrics, such as outcome metrics 340, to generate a set of performance measurement data based on the generated predicted output data or personalized screening model. In some embodiments, the outcome metric 340 may include a positive predictive value, a screening burden measure, or an estimated risk measure. For example, in a cancer diagnosis setting, the outcome metric 340 may include a positive predictive value as an indicator of the proportion of correct predicted outputs of a cancer diagnosis by the prediction output generation engine 308 within the screening model engine 310. The screening burden measure may be a measure based on a cost-effectiveness analysis based on an individualized screening model (e.g., by applying a cancer screening methodology at a particular screening frequency by the screening application engine 306). The screening burden measure may be based on a set of cost-effectiveness data. In some embodiments, the set of cost-effectiveness data may be based on patient insurance or Medicare data.The estimated risk measure may be a measure based on any reported harm to patients (e.g., from invasive screening techniques such as needle biopsy) resulting from application of the screening method at a particular screening frequency for each individualized screening model. In some embodiments, the set of performance measure data may include a numerical analysis based on positive predictive values from outcome metrics 340. In some embodiments, the set of performance measure data may include a numerical cost-effectiveness analysis of applying a screening method (e.g., a cancer screening method in a cancer diagnostic setting) at a given screening frequency. In some embodiments, the set of performance measure data may include an estimated risk analysis based on reported harm (e.g., from invasive screening techniques to patients in a cancer diagnostic setting) resulting from application of the screening method at a given screening frequency.
[0073] The feedback engine 320 can send the performance measurement data to the output validation engine 309. The output validation engine 309 can generate a screening validation score for screening refinement based on the performance measurement data. In some embodiments, the screening validation score for screening refinement is a binary value (e.g., "needs screening refinement," "no screening refinement required") or a positive or negative value on a numerical scale. It may include a gradient of the value (e.g., on a scale of 1 to 10, the need for screening refinement for a particular dataset is "6 out of 10"). The output validation engine 222 may send the screening validation scores to the screening refinement engine 224 for screening refinement.
[0074] The screening model engine 310 may include a screening improvement engine 312. The screening improvement engine 312 may receive screening validation scores for screening improvement from the output validation engine 309 or performance measurement data from the feedback engine 320. The screening improvement engine 312 may perform screening improvement based on the input from the output validation engine 309 or the performance measurement data from the feedback engine 320. The screening improvement engine 312 may perform screening improvement by generating alternative screening methods from a database of screening methods. The screening improvement engine 312 may perform screening improvement by adjusting a screening frequency associated with the screening method. The screening improvement engine 312 may perform improvement based on the performance measurement data. The screening improvement engine 312 may interact with the screening method selector engine 302 for iterative data improvement.
[0075] In some embodiments, the screening model engine 310 can enable iterative cycles of screening refinement. The iterative cycles of screening refinement may include the screening refinement engine 312 sending alternative screening methods or adjusted screening frequencies to the screening method selector engine 302. The iterative cycles may include the screening method selector engine 302 selecting an alternative screening method based on output from the screening refinement engine 312. For example, in breast cancer screening, the screening refinement engine 312 may generate an alternative method to mammography, such as a breast ultrasound scan or a breast MRI scan. The iterative cycles of screening refinement may also include the screening frequency adjustment engine 304 adjusting the screening frequency associated with the alternative screening method selected by the screening method selector engine 302. For example, the screening frequency adjustment engine 304 may increase or decrease the time interval between screenings based on the adjusted frequency from the screening refinement engine 312. The iterative cycles of screening refinement may also include the screening application engine 306 generating one or more alternative screening methods, adjusted screening frequencies, or an improved screening model based on the susceptibility data 301. The iterative cycle of screening refinement may include the predicted output generation engine 308 generating an improved set of predicted output data based on the improved screening model. The iterative cycle of screening refinement may also include the output validation engine 309 interacting with the feedback engine 320 to validate the improved predicted output data. In some embodiments, the iterative cycle of screening refinement may continue until the output validation engine 309 or the feedback engine 320 determines that one or more outcome metrics 340 have reached a threshold. In some embodiments, the output validation engine 309 may perform output of the improved set of predicted output data, as exemplified by output data 360.In some embodiments, the output data 360 may be stored in a database, such as exemplified by the data storage 115 or data sources 101-104 shown in FIG.
[0076] The screening model engine 310 may interact with a machine learning engine 330. According to some embodiments, the machine learning engine 330 may use one or more machine learning models to generate one or more of the models utilized by the screening model engine 210. may analyze multiple datasets or data subsets. The machine learning engine 330 may be trained using output from the output validation engine 309 or performance measurement data from the feedback engine 320 based on one or more outcome metrics 340. The machine learning engine 330 may be configured to predict an optimal set of predicted output data based on the training dataset. For example, in a cancer screening and diagnostic setting, the machine learning engine 330 may be configured to generate a cancer diagnosis optimized for precision or accuracy. In some embodiments, the machine learning engine 330 may be configured to predict an optimal screening method or an optimal screening frequency based on the training dataset. In some embodiments, the training dataset may include data from the susceptibility data 301. In some embodiments, the training dataset may include data from an external data source, such as the data storage 115 or data sources 101-104, as in FIG. 1 . In some embodiments, the optimal screening method or optimal screening frequency may be associated with an optimal screening validation score from the output validation engine 309. For example, in a breast cancer diagnostic setting, the machine learning engine 330 may determine, based on output from the output validation engine 309 or the feedback engine 320, that the optimal screening method for an individual patient with a particular susceptibility profile based on the susceptibility data 301 (e.g., age over 50 years, positive family history of breast cancer, negative smoking history, etc.), is mammography with an adjusted screening frequency of once a year based on data from the screening refinement engine 312.
[0077] The machine learning engine 330 can measure the effectiveness of one or more outcomes based on one or more outcome metrics, such as those exemplified by outcome metrics 340. The machine learning engine 330 can measure the effectiveness of an outcome based on output from the output validation engine 309 or the feedback engine 320. The machine learning engine 330 can perform iterative cycles of training and hypothesis refinement by automatically generating alternative hypotheses based on different subsets of data in the susceptibility data 301, different screening methods, or screening frequencies. In some embodiments, the machine learning engine 330 can also perform hypothesis refinement based on user-defined data selections or user-defined screening methods as part of the input data from the data input engine 110 of FIG. 1. The machine learning engine 330 can perform iterative cycles of hypothesis generation, validation, and refinement until an effectiveness outcome measure, such as that from outcome metrics 340, reaches a threshold value. In some embodiments, the machine learning engine 330 may be exemplified by a machine learning platform 402 (shown in FIG. 4).
[0078] FIG. 4 is a block diagram illustrating various example components of a machine learning platform, according to some embodiments of the present disclosure.
[0079] As shown in Figure 4, machine learning (ML) platform 402 can generate performance measures using machine learning (ML) models in machine learning (ML) model repository 460 as input. In some embodiments, machine learning platform 402 may be exemplified by machine learning engine 230 shown in Figure 2. In some embodiments, machine learning platform 402 may be exemplified by machine learning engine 330 shown in Figure 3. Performance measures 470 generated by ML platform 402 may include adjusted measures 472, predicted measures 474, and performance metrics 476.
[0080] The ML platform 402 can generate the performance measures by taking as input the further measures generated by the aggregation module 462.
[0081] In some embodiments, the ML platform 402 can also obtain additional measures from the measures 464. In some embodiments, the measures 464 are stored as input in an external data source. In some embodiments, the aggregation module 462 can directly provide measures generated by querying a database.
[0082] The ML platform 402 can use input measures 466 from measures 464 and ML models 465 from the ML model repository 460 to generate performance measures 470, including adjusted measures 472, predicted measures 474, and performance metrics 476. The adjusted measures 472 can be adjustments to the input measures 466 adjusted for a patient's susceptibility to a particular cancer type based on susceptibility data generated by a susceptibility model (as shown in FIG. 2). The ML platform 402 can make multiple adjustments to various data parameters associated with the input measures 466. In some embodiments, the ML platform 402 can also generate multiple adjusted measures for each data type or data source associated with the input measures. The ML platform 402 can predict the performance of an ML model used by the susceptibility model engine 210 (as shown in FIG. 2) in generating susceptibility data. The ML platform 402 can predict the performance of an ML model used by the screening model engine 310 (as shown in FIG. 3) in adjusting a screening or predicting an output.
[0083] The ML platform 402 can use different layers of the ML model 465 to generate different types of measures in the performance measure 470. In some embodiments, the ML platform 402 can use a different ML model for each type of performance measure. In some embodiments, the ML platform 402 can generate the ML model as part of the performance measure generation. The ML platform 402 can store the generated ML model in the ML model repository 460. The ML platform 402 can generate a new ML model by adjusting the ML model 465 based on the generated performance measure 470.
[0084] The ML platform 402 can link the performance measures 470 with the input measures 466. The relationship between the adjusted measures 472 or predicted measures 474 and the input measures 466 may be stored in external data storage.
[0085] In some embodiments, the ML platform 402 can also store a relationship between the input measures 466 and performance metrics 476 of the machine learning models in the ML model repository 460. The performance metrics 476 can indicate a discrepancy between the predicted output of the ML model used by the screening model engine 310 and an outcome measure, as illustrated by the feedback engine 320 in FIG. 3 . The relationship between the performance metrics 476 and the input measures 466 may only exist when the discrepancy exceeds a threshold. The ML platform 402 can request the measurement system 100 to store the performance measures 470 in data storage.
[0086] The ML models in the ML model repository 460 may be based on one or more ML algorithms. In some embodiments, the ML algorithm may include, for example, a Viterbi algorithm, a Naive Bayes algorithm, a neural network, an elastic net regression model, or the like, or joint dimensionality reduction techniques (e.g., cluster canonical correlation analysis, partial least squares, bilinear models, cross-modal factor analysis). In some embodiments, the ML model may include an accelerated time to failure model or a proportional hazards model. In some embodiments, the ML model may be based on a Weibull distribution, a log-logistic distribution, an exponential distribution, or a gamma distribution. The type of model selected from the model repository 460 may be The group may influence the inputs required for both the susceptibility model engine 210 and the screening model engine 310.
[0087] In some embodiments, the at least one ML model 465 may be trained, for example, using a supervised learning method (e.g., gradient descent or stochastic gradient descent optimization). In some embodiments, the ML model may be trained based on user-generated training data or automatically generated markup data.
[0088] FIG. 5 illustrates a schematic diagram of an exemplary server of a distributed system according to some embodiments of the present disclosure. In some embodiments, system 100 (shown in FIG. 1) and its components may be implemented in a distributed computing system such as that illustrated by distributed computing system 500. According to FIG. 5, server 510 of distributed computing system 500 comprises a bus 512 or other communication mechanism for communicating information, one or more processors 516 communicatively coupled with bus 512 for processing information, and one or more main processors 517 communicatively coupled with bus 512 for processing information. Processor 516 may be, for example, one or more microprocessors. In some embodiments, one or more processors 516 comprise processor 565 and processor 566, where processor 565 and processor 566 are connected via an inter-chip interconnect of an interconnection topology. Main processor 517 may be, for example, a central processing unit (“CPU”).
[0089] The server 510 further comprises a storage device 514, which may include memory 561 and physical storage 564 (e.g., a hard drive, a solid-state drive, etc.). The memory 561 may include random access memory (RAM) 562 and read-only memory (ROM) 563. The storage device 514 may be communicatively coupled to the processor 516 and the main processor 517 via the bus 512. The storage device 514 may include a main memory that may be used for storing temporary variables or other intermediate information during execution of instructions executed by the processor 516 and the main processor 517. Such instructions, once stored in a non-transitory storage medium accessible to the processor 516 and the main processor 517, render the server 510 a specialized machine customized to perform the operations specified in the instructions. As used herein, the term “non-transitory medium” refers to any non-transitory medium that stores data or instructions that cause a machine to operate in a specific manner. Such non-transitory medium may include non-volatile or volatile media. Non-transitory media include, for example, optical or magnetic disks, dynamic memory, floppy disks, flexible disks, hard disks, solid state drives, magnetic tape or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROMs and EPROMs, FLASH-EPROMs, NVRAMs, flash memories, registers, caches, any other memory chips or cartridges, and networked versions thereof.
[0090] The server 510 can transmit data to or communicate with another server 530 via a network 522. The network 522 can be a local network, an internet service provider, the internet, or any combination thereof. The communication interface 518 of the server 510 is connected to the network 522, which can enable communication with the server 530. In addition, the server 510 can be coupled to peripheral devices 540, including a display (e.g., a cathode ray tube (CRT), a liquid crystal display (LCD), a touch screen, etc.) and input devices (e.g., a keyboard, a mouse, a soft keypad, etc.), via the bus 512.
[0091] The server 510 may be implemented using customized hardwired logic, one or more ASICs or FPGAs, firmware, or program logic that, in combination with the server, makes the server 510 a dedicated machine.
[0092] Various forms of media may be involved in carrying one or more sequences of one or more instructions to the processor 516 or main processor 517 for execution. For example, the instructions may initially be carried on a magnetic disk or solid state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to the server 510 can receive the data on the telephone line and use an infrared transmitter to convert the data to an infrared signal. An infrared detector can receive the data carried in the infrared signal and appropriate circuitry can place the data on the bus 512. The bus 512 carries the data to main memory in storage device 514, from which the processor 516 or main processor 517 retrieves and executes the instructions.
[0093] Multi-domain data-driven system 100 or one or more of its components may reside on either server 510 or 530 and be executed by processor 516 or 517. In some embodiments, components of system 100 may be distributed across multiple servers 510 and 530. For example, data entry engine 110 or characterization engine 120 may be executed on multiple servers. Similarly, analysis engine 130 or feedback engine 140 may be maintained by multiple servers 510 and 530.
[0094] FIG. 6 is a flow diagram illustrating an exemplary process for receiving input based on potential outcomes, performing multi-domain data acquisition, generating individualized data, and performing screening and data and model refinement based on measured performance, according to some embodiments of the present disclosure.
[0095] Process 600 may be performed by a system such as system 100 of Figure 1. In some embodiments, process 600 may be implemented using one or more instructions that may be stored on a computer-readable medium (e.g., storage device 514 of Figure 5).
[0096] In some embodiments, process 600 begins at step 603. In step 603, the system can receive data input across multiple data sources or data domains based on potential outcomes. For example, in a cancer diagnosis or screening setting, the input data may include data associated with cancer type, cancer prevalence, cancer prognosis, time of cancer diagnosis, or cancer stage, while the potential outcome may include a diagnosis of a specific cancer type. In some embodiments, the input data may include proteomic data, patient characteristic data (e.g., patient age, sex, race / ethnicity, height / weight), demographic data, environmental data (e.g., smoking history, diet type, etc.), data obtained from biological sampling (e.g., regarding blood, plasma, serum, or urine samples), medical history data, clinical visit data, surgical history data, family history data, genetic data, or immunological data. In some embodiments, the input data may include cost-effectiveness data, including insurance data or Medicare data.
[0097] In step 604, using the input data, the system can select a susceptibility data model to be used to determine the user's susceptibility to potential outcomes. For example, in a cancer diagnostic setting, the system may select a model data model (such as that illustrated by 108 in FIG. 1) to be used to determine a patient's individual susceptibility to a particular type of cancer. A data model can be selected from multiple data models stored in the data bank. In some embodiments, the system can select a data model based on a subset of input data. For example, the system can select a model based on a subset of input data including cancer type, prevalence, stage, or severity. In some embodiments, the system can select a data model to compensate for specific characteristics in the target outcome. For example, the system can select a different data model for cancers with low prevalence than for commonly occurring cancer types. In some embodiments, some of the input data may be more or less relevant to which susceptibility data model is selected. For example, a patient's smoking history may be more relevant for lung cancer compared to colon cancer. Therefore, less relevant input data may be discarded.
[0098] In step 605, the system may select or calibrate a value including a data enhancement rate from the input data. In some embodiments, the data enhancement rate value may include a factor (e.g., 2x, 5x, 10x, etc.) or a percentage (e.g., 10%, 20%, etc.). Based on the data enhancement rate, the system may enhance the input data or data acquired via multi-domain data acquisition (e.g., as in step 606). For example, based on a data enhancement rate of 10%, the system may increase the amount of data it receives or acquires by 10% accordingly.
[0099] In step 606, the system can perform data acquisition from multiple data domains or data sources based on the data parameters of the selected data model. The system can enhance the acquired data based on the data enhancement rate selected in step 605. In some embodiments, the system can perform data acquisition by performing biological data sampling 607, in which a biological sample related to a blood, plasma, serum, or urine sample is acquired. In some embodiments, the system can perform data acquisition via proteomic assay 608 based on the acquired biological sample.
[0100] In step 620, the system can generate a set of susceptibility data for an individual user related to potential outcomes. In some embodiments, the set of susceptibility data may be exemplified by the susceptibility model engine 210 shown in FIG. 2. In some embodiments, for example, in a cancer diagnosis setting, the system can generate a set of susceptibility data for a particular patient related to a particular cancer type diagnosis based on the input data or acquired data and a selected susceptibility data model. The generated susceptibility data may include an individualized susceptibility score based on a susceptibility estimation analysis. In some embodiments, the susceptibility score may include a likelihood percentage for the occurrence of a potential outcome or risk stratification category (e.g., in cancer diagnosis). In some embodiments, the susceptibility data or susceptibility score may include a multiplicative coefficient representing the likelihood of the occurrence of a potential outcome for an individual (e.g., a patient) relative to a reference population. For example, a patient with a particular set of susceptibilities may have a two-fold or ten-fold increased risk of developing cancer compared to the general population. In some embodiments, the susceptibility data may include data associated with the likelihood of the occurrence of a potential outcome within a predetermined time frame. For example, in a cancer diagnosis setting, the susceptibility data may include a patient's susceptibility score for a particular type of cancer over a particular time frame (e.g., from the time the patient's biological sampling data is obtained).
[0101] In step 630, the system may determine a screening model for the potential outcome of interest. In some embodiments, the system may determine the screening model based on personalized susceptibility data associated with the potential outcome of interest. In some embodiments, the screening model may include one or more screening methods, each associated with a screening frequency. be.
[0102] In step 633, the system may adjust screening parameters associated with a screening model or screening method. In some embodiments, the system may select from among various screening methods. In some embodiments, the system may adjust the frequency of screening for one or more screening methods. In some embodiments, the screening method may be selected by the screening model engine 310 shown in FIG. 3. In some embodiments, adjusting the frequency of the screening method may be performed by the screening model engine 310. In some embodiments, for example, in a cancer screening setting, the system may adjust the frequency of screening by setting more frequent interval screening for patients with a high susceptibility to a particular cancer type or less frequent interval screening for patients with a low susceptibility. The system may also adjust the screening frequency based on a cost-effectiveness analysis of the screening method or an analysis of the estimated harms of applying the screening method to an individual patient. The system may also adjust the screening frequency based on screening guidelines from an external data source.
[0103] Process 600 then proceeds to step 635. In step 635, the system can use the screening model to perform screening for potential outcomes of interest. In some embodiments, such as in a cancer screening setting, the system can simulate the application of a screening method to an individual patient over a time interval based on the personalized screening model or input or acquired data. In some embodiments, the system can output results associated with the application of the screening. For example, if the selected screening method is mammography, the system can output a BI-RADS score or a detailed description of the mammographic findings. In some embodiments, the system can also output cost-effectiveness data associated with the screening method at a screening frequency over a predetermined time period. In some embodiments, the cost-effectiveness data is based on the patient's insurance or Medicare data. In some embodiments, the system can also output estimated risk data associated with adverse events or harms associated with the application of the screening method.
[0104] Also, in step 635, the system may generate a predictive output associated with a potential outcome of interest based on the results associated with the application of the screening. For example, in a cancer screening and diagnostic setting, the system may generate a set of predictive output data associated with the diagnosis of one or more types of cancer. In some embodiments, the set of predictive output data may include a binary value (e.g., a positive cancer diagnosis vs. a negative cancer diagnosis). In some embodiments, the set of predictive outputs may include a numerical scale of likelihood (e.g., on a scale of 1 to 10, the likelihood of having a particular type of cancer is 6 for an individual patient). In some embodiments, the set of predictive outputs may include risk stratification categories (e.g., there is a high / medium / low risk for a particular type of cancer for an individual patient).
[0105] The system may also optionally output data associated with the screening results or generated predicted output data to a user or operator in step 635. In some embodiments, the system outputs the data via an external display device, exemplified by 150 in FIG.
[0106] In step 639, the system evaluates the performance using one or more outcome metrics. The performance of the cleaning model can be measured. In some embodiments, the outcome metric may include a positive predictive value, a screening burden measure, or an estimated risk measure. In some embodiments, the system can determine a set of performance measurement data based on one or more outcome metrics. In some embodiments, the system can validate the screening model based on the performance measurement data.
[0107] In step 640, the system can make a decision for further cycles of iterative data refinement or data model refinement based on validation of the screening model, performance measurement data, or one or more outcome metrics. The system can perform data refinement by adjusting a data enrichment rate and generate an improved set of susceptibility data based on the adjusted data enrichment rate. The system can also perform screening refinement by selecting an alternative screening method and adjusting a screening frequency and generate an improved screening model based on the alternative screening method and the adjusted screening frequency. In some embodiments, the system can perform iterative cycles of data refinement or screening refinement based on one or more outcome metrics until a determination is reached that measured performance associated with a potential outcome of interest has reached a threshold.
[0108] As used herein, unless otherwise stated, the term "or" includes all possible combinations unless impracticable. For example, if it is stated that a component may include A or B, it may include A, or B, or A and B, unless otherwise stated or impracticable. As a second example, if it is stated that a component may include A, B, or C, it may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C, unless otherwise stated or impracticable.
[0109] Exemplary embodiments are described above with reference to flowchart diagrams or block diagrams of methods, apparatus (systems), and computer program products. It will be understood that each block of the flowchart diagrams or block diagrams, and combinations of blocks in the flowchart diagrams or block diagrams, can be implemented by a computer program product or instructions on a computer program product. These computer program instructions may be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, such that the instructions, executed by a processor of the computer or other programmable data processing apparatus, create means for performing the functions / acts specified in one or more blocks of the flowchart or block diagrams.
[0110] These computer program instructions may be stored on a computer-readable storage medium that can direct one or more hardware processors of a computer, other programmable data processing apparatus, or other device to function in a number of ways, such that the instructions stored on the computer-readable storage medium form an article of manufacture including instructions that implement the functions / acts specified in one or more blocks of the flowchart or block diagram.
[0111] The computer program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device to cause the computer, other programmable apparatus, or other device to perform a series of operational steps to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide a process for implementing the functions / acts specified in one or more blocks of the flowchart or block diagram.
[0112] Any combination of one or more computer-readable mediums may be utilized. The computer-readable medium may be a non-transitory computer-readable storage medium. In the context of this specification, a computer-readable storage medium may be any tangible medium that contains or can store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0113] The program code embodied on the computer readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, fiber optic cable, RF, IR, etc., or any suitable combination thereof.
[0114] Computer program code for carrying out operations, e.g., embodiments, may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider).
[0115] The flowcharts and block diagrams in the figures illustrate examples of the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of code, including one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams or flowchart diagrams, and combinations of blocks in the block diagrams or flowchart diagrams, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or a combination of dedicated hardware and computer instructions.
[0116] It is understood that the described embodiments are not mutually exclusive, and that elements, components, materials, or steps described in connection with one exemplary embodiment may be combined with or excluded from other embodiments in any suitable manner to achieve desired design objectives.
[0117] The disclosed embodiments may be further described using the following clauses.
[0118] 1. A non-transitory computer-readable medium containing instructions executable by one or more processors to cause a system to perform a method, the method comprising: receiving input data associated with one or more types of cancer and a patient; determining a susceptibility model and a data enrichment rate based on the input data; obtaining data associated with the patient from a plurality of data domains based on the susceptibility model and the data enrichment rate, the patient data including at least proteomic data; and using a machine learning algorithm to determine a cancer susceptibility associated with the patient using the susceptibility model and the patient data. A non-transitory computer-readable medium comprising: generating data; determining a screening model for one or more types of cancer for the patient based on the cancer susceptibility data; and providing information for screening the patient for one or more types of cancer based on the screening model and the patient data.
[0119] 2. The non-transitory computer-readable medium of clause 1, wherein the input data includes at least one of a type of cancer, a prevalence of cancer, a time of cancer diagnosis, a prognosis of cancer, or a stage of cancer.
[0120] 3. The non-transitory computer-readable medium of any one of clauses 1-2, wherein the patient data further includes patient characteristic data, insurance data, health coverage data, medical history data, genetic data, immunological data, environmental data, or biological sampling data from the patient.
[0121] 4. The non-transitory computer-readable medium of any one of clauses 1-3, wherein the proteomic data is based on a patient biological sample.
[0122] 5. The non-transitory computer-readable medium of any one of clauses 1-5, wherein instructions executable by one or more processors further cause the system to determine, from the patient data, a set of features associated with the patient's susceptibility to one or more types of cancer.
[0123] 6. The non-transitory computer-readable medium of any one of clauses 1 to 5, wherein determining the susceptibility model includes selecting the susceptibility model from a plurality of data models in a model data bank.
[0124] 7. The non-transitory computer-readable medium of any one of clauses 1-6, wherein the patient-specific screening model includes one or more screening methods, each associated with one or more screening schedules.
[0125] 8. The non-transitory computer-readable medium of any one of clauses 1-7, wherein the instructions executable by the one or more processors further cause the system to iteratively refine the screening model using one or more machine learning algorithms by adjusting a screening schedule of the screening method based on one or more outcome metrics until the one or more outcome metrics reach a threshold.
[0126] 9. The non-transitory computer-readable medium of clause 8, wherein the outcome metric comprises at least one of a positive predictive value, a screening burden measure, or an estimated risk measure.
[0127] 10. The non-transitory computer-readable medium of clause 8, wherein the instructions executable by the one or more processors further cause the system to iteratively refine the susceptibility model using one or more machine learning algorithms by adjusting a data enrichment rate based on one or more outcome metrics until the one or more outcome metrics reach a threshold, and generating an improved set of susceptibility data based on the adjusted enrichment rate.
[0128] 11. Receiving input data associated with one or more types of cancer and a patient; determining a susceptibility model and a data enrichment rate based on the input data; obtaining data associated with the patient from a plurality of data domains based on the susceptibility model and the data enrichment rate, wherein the patient data includes at least proteomic data; and generating a set of cancer susceptibility data associated with the patient based on the susceptibility model and the patient data using one or more machine learning algorithms; A method for data modeling and analysis that includes using one or more machine learning algorithms to determine a screening model for one or more types of cancer for a patient based on cancer susceptibility data, and screening the patient for the one or more types of cancer based on the screening model.
[0129] 12. The method of clause 11, wherein the input data includes at least the type of cancer, the prevalence of cancer, the time of cancer diagnosis, the prognosis of cancer, the stage of cancer, or any combination thereof.
[0130] 13. The method of any one of clauses 11-12, wherein the patient data further comprises patient characteristic data, medical history data, insurance data, health coverage data, genetic data, immunological data, environmental data, or biological sampling data from the patient.
[0131] 14. The method of any one of clauses 11 to 13, wherein the proteomic data is based on a biological sample of the patient.
[0132] 15. The method of any one of clauses 11 to 14, wherein determining the susceptibility model includes selecting the susceptibility model from a plurality of data models in a model databank.
[0133] 16. The method of any one of clauses 11-15, further comprising determining from the patient data a set of features associated with the patient's susceptibility to one or more types of cancer.
[0134] 17. The method of any one of clauses 11 to 16, wherein the patient-specific screening model comprises one or more screening methods, each associated with one or more screening schedules.
[0135] 18. The method of any one of clauses 11-17, further comprising iteratively refining the screening model using one or more machine learning algorithms by adjusting the screening schedule of the screening method based on one or more outcome metrics until the one or more outcome metrics reach a threshold value.
[0136] 19. The method of clause 18, wherein the outcome metric comprises a positive predictive value, a screening burden measure, an estimated risk measure, or any combination thereof.
[0137] 20. The method of clause 18, further comprising iteratively refining the susceptibility model using one or more machine learning algorithms by adjusting the data enrichment rate based on one or more outcome metrics and generating an improved set of susceptibility data based on the adjusted enrichment rate.
[0138] 21. A computer-implemented system for data modeling and analysis, the system comprising: a memory for storing instructions; and at least one processor configured to execute the instructions to cause the computer-implemented system to perform the operations described in any one of clauses 11 to 20.
[0139] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Certain adaptations and modifications of the described embodiments may be made. Other embodiments will be apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. The specification and examples are intended to be exemplary only. Also, the sequence of steps depicted in the figures is intended to be illustrative only and is not intended to be limited to any particular sequence of steps. As such, Those skilled in the art will appreciate that these steps may be performed in different orders while performing the same method.
Claims
1. 1. A non-transitory computer-readable medium containing instructions executable by one or more processors to cause a system to perform a method, the method comprising: receiving input data associated with one or more types of cancer and patients; determining a susceptibility model and a data enhancement rate based on the input data; obtaining data associated with the patient from a plurality of data domains based on the susceptibility model and the data enrichment rate, the patient data including at least proteomic data; using a machine learning algorithm to generate cancer susceptibility data associated with the patient using the susceptibility model and the patient data; determining a screening model for the one or more types of cancer for the patient based on the cancer susceptibility data; providing information for screening said patient for said one or more types of cancer based on said screening model and patient data; 1. A non-transitory computer-readable medium comprising:
2. 10. The non-transitory computer-readable medium of claim 1, wherein the input data comprises at least one of a type of cancer, a prevalence of cancer, a time of cancer diagnosis, a prognosis of cancer, or a stage of cancer.
3. 10. The non-transitory computer-readable medium of claim 1, wherein the patient data further comprises patient characteristic data, insurance data, health coverage data, medical history data, genetic data, immunological data, environmental data, or biological sampling data from the patient.
4. The non-transitory computer-readable medium of claim 1 , wherein the proteomic data is based on a biological sample of the patient.
5. the instructions executable by the one or more processors: determining from the patient data a set of features associated with the patient's susceptibility to the one or more types of cancer; The non-transitory computer-readable medium of claim 1 , further causing the system to execute:
6. The non-transitory computer-readable medium of claim 1 , wherein determining the susceptibility model comprises selecting the susceptibility model from a plurality of data models in a model databank.
7. 10. The non-transitory computer-readable medium of claim 1, wherein the screening model for the patient includes one or more screening methods each associated with one or more screening schedules.
8. the instructions executable by the one or more processors: iteratively improving the screening model using one or more machine learning algorithms by adjusting a screening schedule of the screening method based on one or more outcome metrics until the one or more outcome metrics reach a threshold value. The non-transitory computer-readable medium of claim 1 , further causing the system to execute:
9. The outcome metric may be a positive predictive value, a screening burden measure, or an estimated risk measure. The non-transitory computer-readable medium of claim 8 , comprising at least one of a constant value.
10. the instructions executable by the one or more processors: iteratively improving the susceptibility model using one or more machine learning algorithms by adjusting the data enrichment rate based on the one or more outcome metrics until the one or more outcome metrics reach a threshold value, and generating an improved set of susceptibility data based on the adjusted enrichment rate. The non-transitory computer-readable medium of claim 8 , further causing the system to perform the following:
11. receiving input data associated with one or more types of cancer and patients; determining a susceptibility model and a data enhancement rate based on the input data; obtaining data associated with the patient from a plurality of data domains based on the susceptibility model and the data enrichment rate, the patient data including at least proteomic data; generating a set of cancer susceptibility data associated with the patient based on the susceptibility model and the patient data using one or more machine learning algorithms; determining a screening model for the one or more types of cancer for the patient based on the cancer susceptibility data using one or more machine learning algorithms; screening said patient for said one or more types of cancer based on said screening model; A method for data modeling and analysis, including:
12. 12. The method of claim 11, wherein the input data comprises at least a type of cancer, a prevalence of cancer, a time of cancer diagnosis, a prognosis of cancer, a stage of cancer, or any combination thereof.
13. 12. The method of claim 11, wherein the patient data further comprises patient characteristic data, medical history data, insurance data, Medicare data, genetic data, immunological data, environmental data, or biological sampling data from the patient.
14. The method of claim 11 , wherein the proteomic data is based on a biological sample of the patient.
15. The method of claim 11 , wherein determining the susceptibility model comprises selecting the susceptibility model from a plurality of data models in a model databank.
16. determining from the patient data a set of features associated with the patient's susceptibility to the one or more types of cancer; The method of claim 11 further comprising:
17. The method of claim 11 , wherein the patient-specific screening model includes one or more screening methods each associated with one or more screening schedules.
18. iterating the screening model using one or more machine learning algorithms by adjusting a screening schedule of the screening method based on the one or more outcome metrics until the one or more outcome metrics reach a threshold value. To improve The method of claim 11 further comprising:
19. 20. The method of claim 18, wherein the outcome metric comprises a positive predictive value, a screening burden measure, an estimated risk measure, or any combination thereof.
20. iteratively improving the susceptibility model using one or more machine learning algorithms by adjusting the data enrichment rate based on the one or more outcome metrics and generating an improved set of susceptibility data based on the adjusted enrichment rate.
20. The method of claim 18, further comprising:
21. 1. A computer-implemented system for data modeling and analysis, said system comprising: a memory for storing instructions; Execute the instructions, receiving input data associated with one or more types of cancer and patients; determining a susceptibility model and a data enhancement rate based on the input data; obtaining data associated with the patient from a plurality of data domains based on the susceptibility model and the data enrichment rate, the patient data including at least proteomic data; using a machine learning algorithm to generate cancer susceptibility data associated with the patient using the susceptibility model and the patient data; determining a screening model for the one or more types of cancer for the patient based on the cancer susceptibility data; providing information for screening said patient for said one or more types of cancer based on said screening model and patient data; at least one processor configured to cause the computer-implemented system to perform operations including A system comprising: