Systems and methods for predicting immunologically active peptides with machine learning models
By employing machine learning models to predict immunologically active peptides by filtering out human genome peptides and ranking potential candidates, the method enhances the accuracy and efficiency of peptide prediction, addressing the limitations of existing algorithms.
Patent Information
- Application Number
- PCT/US2024/060419
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-12-16
- Publication Date
- 2025-06-19
AI Technical Summary
Existing predictive algorithms for peptide interactions often fail to identify immunologically active peptides due to lack of consideration for self-tolerance and individual variability among humans, resulting in low accuracy and increased costs for further testing.
The development of systems and methods using machine learning models that receive a relevant target associated with an immunologically active peptide, determine human genome peptides, compare them to target peptides, remove human genome peptides, generate a ranked list of immunologically active peptides, and test them to refine the list.
This approach improves the prediction of immunologically active peptides by reducing the number of peptides that need to be tested, conserving computational resources, and accelerating the time to market while reducing costs.
Smart Images

Figure 00000032_0000 
Figure 00000033_0000 
Figure 00000034_0000
Abstract
Description
SYSTEMS AND METHODS FOR PREDICTING IMMUNOLOGICALLY ACTIVE PEPTIDES WITH MACHINE LEARNING MODELSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 611,019, titled “AN ALGORITHM TO PREDICT IMMUNOLOGICALLY ACTIVE HLA-PEPTIDE COMPLEXES,” filed December 15, 2023, the full disclosure of which is hereby incorporated by reference in its entirety for all purposes.BACKGROUND1. Field of Disclosure
[0002] Embodiments of the present disclosure relate to systems and methods to predict peptide binding interactions. Specifically, one or more embodiments are directed toward identifying unique peptides based on their immunological activity and strength of interaction.2. Description of Related Art
[0003] Predictive algorithms for peptide interactions often fail to identify peptides that are immunologically active when subjected to laboratory testing. For example, the algorithms may not account for self-tolerance with binding peptides and may not account for the large variability with human beings, such as those due to age, sex, race and ethnicity, and / or the like. As a result, a small percentage of these predicted peptides may show sufficient immunological responses for further testing, which may lead to increased costs and longer times to market.SUMMARY
[0004] Applicant recognized the problems noted above herein and conceived and developed embodiments of systems and methods, according to the present disclosure, for predicting immunologically active peptides.
[0005] In an embodiment, a computer-implemented method includes receiving a relevant target associated with an immunologically activate peptide for treating one or more conditions. The method also includes determining one or more human genome peptides associated with a human genome. The method further includes comparing the one or more human genome peptides to a list of target peptides generated, based at least in part, on the relevant target. The method includes removing the one or more human genome peptides from the list of target peptides. The method further includes generating, based at least in part on the relevant target and the list of target peptides, a ranked list of immunologically active peptides. The method also includes testing the ranked list of immunologically active peptides. The method includes generating an active list based, at least in part, on the testing.
[0006] In another embodiment, a processor includes one or more circuits to generate a set of ranked results for an immunologically active peptide. The one or more circuits are further to discard one or more results from the set of ranked results based, at least in part, on testing to generate a refined set of ranked results. The one or more circuits are also to add the refined set of ranked results to a training dataset. The one or more circuits are further to train one or more machine learning models using at least a portion of the training dataset. The one or more circuits are to execute one or more validation tests on the one or more machine learning models. The one or more circuits are further to determine the one or more validation tests exceed one or more metrics. The one or more circuits are to deploy the one or more machine learning models. The one or more circuits are further to collect, over a period of time, results from execution of the one or more machine learning models. The one or more circuits are further to add the results to the training dataset
[0007] In another embodiment, a computer-implemented method includes generating a set of ranked results for an immunologically active peptide. The method also includes discarding one ormore results from the set of ranked results based, at least in part, on testing to generate a refined set of ranked results. The method further includes adding the refined set of ranked results to a training dataset. The method includes training one or more machine learning models using at least a portion of the training dataset. The method also includes executing one or more validation tests on the one or more machine learning models. The method further includes determining the one or more validation tests exceed one or more metrics. The method also includes deploying the one or more machine learning models. The method includes collecting, over a period of time, results from execution of the one or more machine learning models. The method also includes adding the results to the training dataset.BRIEF DESCRIPTION OF DRAWINGS
[0008] The present technology will be better understood on reading the following detailed description of non-limiting embodiments thereof, and on examining the accompanying drawings, in which:
[0009] FIG. 1 illustrates an example environment for predicting immunologically active peptides, in accordance with embodiments of the present disclosure;
[0010] FIG. 2 illustrates an example environment for predicting immunologically active peptides, in accordance with embodiments of the present disclosure;
[0011] FIG. 3 A illustrates an example environment for testing and retraining one or more machine learning system, in accordance with embodiments of the present disclosure;
[0012] FIG. 3B illustrates an example flow chart for developing one or more machine learning systems, in accordance with embodiments of the present disclosure;
[0013] FIG. 4 illustrates an example environment for refining a target list upstream of one or more machine learning systems, in accordance with embodiments of the present disclosure;
[0014] FIG. 5 A is a flow chart of a process for predicting one or more immunologically active peptides, in accordance with embodiments of the present disclosure;
[0015] FIG. 5B is a flow chart of a process for training a machine learning system to identify immunologically active peptides, in accordance with embodiments of the present disclosure;
[0016] FIG. 6 is an example configuration for a computing device, in accordance with embodiments of the present disclosure.DETAILED DESCRIPTION
[0017] The foregoing aspects, features, and advantages of the present disclosure will be further appreciated when considered with reference to the following description of embodiments and accompanying drawings. In describing the embodiments of the disclosure illustrated in the appended drawings, specific terminology will be used for the sake of clarity. However, the disclosure is not intended to be limited to the specific terms used, and it is to be understood that each specific term includes equivalents that operate in a similar manner to accomplish a similar purpose. Additionally, like reference numerals may be used for like components, but such use should not be interpreted as limiting the disclosure.
[0018] When introducing elements of various embodiments of the present disclosure, the articles "a", "an", "the", and "said" are intended to mean that there are one or more of the elements. The terms "comprising", "including", and "having" are intended to be inclusive and mean that there may be additional elements other than the listed elements. Any examples of operating parameters and / or environmental conditions are not exclusive of other parameters / conditions of the disclosed embodiments. Additionally, it should be understood that references to "one embodiment", "an embodiment", “certain embodiments”, or “other embodiments” of the present disclosure are not intended to be interpreted as excluding the existence of additional embodiments that also incorporate the recited features. Furthermore, reference to terms such as “above”, “below”,“upper”, “lower”, “side”, “front”, “back”, or other terms regarding orientation or direction are made with reference to the illustrated embodiments and are not intended to be limiting or exclude other orientations or directions. Like numbers may be used to refer to like elements throughout, but it should be appreciated that using like numbers is for convenience and clarity and not intended to limit embodiments of the present disclosure. Moreover, references to “substantially” or “approximately” or “about” may refer to differences within ranges of + / - 10 percent.
[0019] Embodiments of the present disclosure may be directed toward systems and methods to predict immunologically active peptide complexes using one or more machine learning (ML) systems. Systems and methods may be directed toward one or more ML systems that may be trained on peptide binding data, which may include human leukocyte antigen (HLA) alleles and / or functional peptides for one or more immunological targets. In operation, an input may be received for a desired target (e.g., a relevant target), the desired target may be processed by one or more data processing systems, ranked, and then one or more sets of ranked results may be provided for potential peptides for the given desired target. The ranked list may then be used for laboratory testing, which may provide information to refme / update the ranked list and / or model. The model may be continuously updated (e.g., at intervals, as a threshold quantity of data is obtained, etc.) to provide updated immunological information. In at least one embodiment, one or more lists generated by the one or more ML systems may be modified based, at least in part, on a similarity with peptides that are found in the human genome and / or have already been experimentally determined to be insufficient for one or more purposes. In this manner, a reduced number of peptides may be generated for laboratory review. Furthermore, by culling the data at an upstream location in the pipeline, compute resources may be reduced. Accordingly, systems and methods of the present disclosure may be used to generate lists or peptides for one or more desired targets.
[0020] In at least one embodiment, systems and methods of the present disclosure may be used to address and overcome problems with existing peptide prediction algorithms. Many peptide prediction algorithms have low accuracy, where clinically testing may find that identified algorithms are immunologically active in less than approximately 20% of predicted peptides. Systems and methods address and overcome these problems by introducing both an interrogation against the human genome for presence of predicted binding peptides along with one or more ML systems to identify and rank peptides not only on immunological activity, but also at least strength of interaction. In this manner, embodiments provide for improved peptide prediction systems and methods with reduced use of compute resources, thereby reducing costs associated with prediction, costs for evaluation of predicted peptides, and reducing time to market. In one or more embodiments, systems and methods may interrogate the human genome to identify predicted binding peptides for one or more desired targets (e.g., targets of interest). Because peptides present within the human genome and shared with the target of interest would be anticipated to produce immunologic self-tolerance, embodiments may preemptively identify and remove such peptides from a generated list, thereby reducing downstream testing and / or reducing downstream compute to rank and / or determine likely interaction strengths for peptides that are associated with the human genome. Furthermore, one or more embodiments may use one or more ML systems, which may include a variety of trained algorithms and / or training processes, to assess the differences between peptides which are not found in the human genome which do or do not function effectively as anticipated. Algorithms may be developed by using information, which may be curated, and may be extracted from one or more data lakes and / or databases. Information used to train and / or tune ML models may include HLA alleles and functional peptides for one or more immunologic targets. One or more models may include neural networks consisting of multiple nodes trained on existing content and subsequently retrained, as target identification increases, where certainty of theoutcome exists and can provide declarative direction for the algorithm. Accordingly, result information, such as information from laboratory testing, may be evaluated for declarative direction, thereby tuning training data based on experimental validations. Furthermore, embodiments may also be used to assess strength of immunological interactions. For example, because not all immunologically targeted peptides perform identically with regard to the numbers of CD8+ Cytotoxic T Lymphocytes (CD8+CTLs) they induce, (i.e. some may be immunodominant), peptides may be assessed both for whether they are recognized immunologically as well as the strength of that interaction. One or more metrics may be established to rank or otherwise assess strength of interaction, such as on non-limiting example of, using the number of tetramer positive CD8+CTLs induced by a particular peptide after a defined series of stimulations.
[0021] FIG. 1 illustrates an example system 100 that may be used with embodiments of the present disclosure. In this example, a computing device 102 (e.g., user device, compute device, client device, etc.) can submit a request over at least one network 104 to be received by a provider environment 106. The prediction environment 106 (e.g., environment) may be an online platform provided by a service provider and / or for an affiliate, for example the environment 106 may be hosted or otherwise provided via one or more cloud resource providers on behalf of a service provider. The client computing device 102 may be a representative and / or act as a proxy for one or more users that may be submitting requests. For example, a user may navigate to one or more dashboards, web applications, landing pages, or access points using the device to submit a request, among other options. Additionally, in at least one embodiment, the client computing device 102 may act as a proxy to execute stored instructions to make and receive requests. For example, the client computing device 102 may send a request responsive to receiving one or more inputs and / or the like. As another example, a request may be transmitted as part of an automated or semi-automated workflow, which may or may not receive user interaction. For example, upon submitting laboratory testing results and / or selecting a target identification, one or more workflows may be initiated to select one or more ML models, process one or more datasets and / or inputs associated with the target identification, and then provide a curated set of results. Accordingly, the client computing device 102 may be used with direct input from one or more users, from stored software instructions, from executions of various workflows, or combinations thereof.
[0022] In at least one embodiment, the request can include a request to execute one or more workflows associated with peptide prediction, which may be based, at least in part, on a target of interest. It should be appreciated that peptide prediction is provided by way of non-limiting example and systems and methods of the present disclosure may be used in a variety of different types of prediction tasks, which may include biological prediction tasks for the development and / or identification of one or more biological compounds that may be used to treat, prevent, or otherwise address one or more diseases or illnesses associated with humans, animals, plant life, and / or the like.
[0023] In many cases, the analysis and / or prediction tasks may include a request to access data (e.g., stored data, streaming data, etc.) and then to process the data using one or more workflows associated with the environment 106. In at least one embodiment, a selected workflow may be based, at least in part, on information provided by the computing device 102, such as a command, or based on data received by the environment 106. The network(s) 104 can include any appropriate network, such as the Internet, a local area network (LAN), a cellular network, an Ethernet, or other such wired and / or wireless network. The prediction environment 106 can include any appropriate resources for accessing data or information, such as laboratory results, prior training data, medical information, and / or the like, as may include various servers, data stores, and other such components known or used for accessing data and / or processing data from across a network (orfrom the “cloud”). Moreover, the client computing device 102 can be any appropriate computing or processing device, as may include a desktop or notebook computer, smartphone, tablet, wearable computer (e.g., smart watch, glasses, contacts, headset, etc.), server, or other such system or device.
[0024] An interface layer 108, when receiving a request or call, can determine the type of call or request and cause information to be forwarded to the appropriate component or sub-system. For example, the interface 108 may be associated with one or more landing pages, as an example, to guide a user toward a workflow or action. In at least one embodiment, the interface layer 108 may include other functionality and implementations, such as load balancing and the like.
[0025] Various embodiment of the present disclosure are directed toward predictive systems to identify one or more peptides to treat an immunological target. In at lest one embodiment, a manger 110 may be associated with the prediction environment 106 to receive and route incoming requests, for example to determine whether a requestor has authority to make the request by querying a user datastore 112, by accessing one or more third party resources 114 and / or databases 116, and / or by directing the request toward systems and sub-systems of the prediction environment 106. For example, the manager 110 may incorporate an authentication service to verify credentials provided by the client device 102, for example against the user datastore 112, to verify and permit access to the environment 106. Furthermore, in at least one embodiment, verification may also determine a level of accessibility within the environment 106, which may be on an application-basis, a userbasis, or some combination thereof. For example, a first user may have access to the environment, but only have a limited set of applications that are accessible, while a second user may have access to more applications, and a third user may be entirely barred from the environment.
[0026] Systems and methods may include a web-based or application-based portal that permits receipt of immunological targets and / or data to refine one or more ML models. At least oneembodiment discussed herein may be related to generating predictions for peptides that may be used to treat one or more target conditions, diseases, and / or the like.
[0027] In this example, a data evaluation engine 118 may be used to process data generated by one or more ML models and / or prepare data for execution with one or more ML models. For example, systems and methods of the present disclosure may use the data evaluation engine 118 for one or more extra, transform, and load techniques to combine data from multiple sources into a centralized repository, such as the data lake 120. The information may be sorted or otherwise prepared according to one or more specifications, which may include formatting, removing portions of the data, compressing the data, and / or the like. Data may include clinical data, such as metrics associated with immunological recognition, strength of interaction, and / or the like. Furthermore, systems and methods may process multi-modal data, which may include text data, image data, video data, and / or combinations thereof.
[0028] The illustrated embodiment may be used for an initial analysis of data and / or to refine or otherwise filter data that is generated responsive to execution of one or more ML systems. By way of example, a human genome evaluator 122 may be used to evaluate the human genome in order to determine the presence of predicted binding peptides. As discussed herein, a peptide present within the human genome is likely to produce immunologic self-tolerance, and as a result, would be unlikely to be useful in treating one or more diseases or conditions. Accordingly, by evaluating the human genome (or other genomes for embodiment where peptides are predicted for nonhumans), embodiments may be used to reduce a total number of peptides that are evaluated, either by computational methods or in a laboratory settings, which may reduce costs, time to market, and / or compute. The results of the evaluator 122 may be used to filter or otherwise reduce a number of predicted peptides and / or as training data with one or more ML models to block or prevent identification of certain peptides for a given target identification process.
[0029] Systems and methods may also include a processing engine 124 and a ranking engine 126. The processing engine 124 may be used to prepare data, for example prior to assembly within the data lake 120, and / or may be used to evaluate an output of one or more ML systems and / or to evaluate and process laboratory and / or simulation data, among other types of data. The ranking engine 126 may also be used to apply one or more metrics to rank different ML outputs to identify peptides that have a greater likelihood of providing a strong immunological interaction. For example, even if a peptide is recognized, it may induce a response is that is too low to be clinically effective. By ranking based on multiple metrics, systems and methods may be used to prioritize further testing and analysis.
[0030] One or more embodiments may further implement an ML system 128 that may include one or more ML or artificial intelligence (Al) models, such as transformer-based models, convolutional neural networks, recurrent neural networks, and / or combinations thereof. Various embodiments associated with the ML system 128 may include execution of different software instructions based, at least in part, on a request received from the user device. In this example, an ML model may be selected from one or more models of a model datastore 130, which may include a set of models that may be trained for a domain and / or are general or foundational models that may be used with embodiments of the present disclosure. These models may be trained and or execute using one or more different datastores, which may include training data, model parameters, model settings, rules, and / or the like. The models in the model datastore 130 may undergo training using a training engine 132, which may use training data from a training datastore 134, which may include outputs or tagged infomiation. The training data, which may be labeled or unlabeled, and also may be augmented or otherwise influenced by one or more human reviewers, but it should be appreciated that raw training data may be used with one or more self-supervised learning processes.Accordingly, models may be trained for specific use cases and / or a general model may be trainedfor a specific domain, such as for a specific type of ailment (e.g., cancer, viral, etc.). In operation, an ML service 136 may be used to execute and run the model selected from the model datastore 130 and may output one or more predictions 138. The predictions 138 may include a set of potential peptides that may be candidates for bonding for the immunologic target. In certain embodiments, predictions 138 may be further processed using the data evaluation engine 118, such as to provide additional rankings or other processing steps.
[0031] Various types of architectures may be implemented in various embodiments, and in certain embodiments, architecture may be technique-specific. As one example, architectures may include recurrent neural networks (RNNs) or long short-term memory networks (LSTMs), transformer architectures (e.g., self-attention mechanisms, encoder and / or decoder blocks, etc.), convolutional neural networks (CNNs), and / or the like.
[0032] In various embodiments, the models may be trained using unsupervised learning, in which models learn patterns from large amounts of unlabeled training data (e.g., text, audio, video, image, etc.). Furthermore, one or more models may be task-specific or domain-specific, which may be based on the type of training data used. Additionally, foundational models may be used and then tuned for specific tasks or domains. Some types of foundational models may include questionanswering, summarization, filling in missing information, and translation. Additionally, specific models may also be used and / or augmented for certain tasks, using techniques like prompt tuning, fine-tuning, retrieval augmented generation (RAG), adding adapters, and / or, the like. Systems and methods may incorporate a number of different models to execute one or more operations, including, but not limited to, NetMHC, PSSM, PickPocket, MHCFlurry, AI-MHC, SYFPEFITHI, and ACME.
[0033] FIG. 2 illustrates an example environment 200 that may be used with embodiments of the present disclosure. In this example, an input 202 is provided to the processing engine 124 alongwith information from the data like 120. In at least one embodiment, the input 202 may correspond to an input for a relevant target, which may include a request to generate a prediction or be an input of a known relevant target to be used during training or proving of a model, among other examples. In the illustrated embodiment, the input 302 and at least a portion of the data associated with the data lake 120 are processed using one or more techniques, which may include conforming formatting and the like in accordance with a desired specification and / or in accordance with an output analysis. The ranking engine 126 may the generate one or more rankings, which may include relevancy rankings and / or similarity scoring, among other options. The output ranking may be provided for analysis by one or more ML systems 128, as shown by the numeral 1. As discussed herein, one or more models may execute, using one or more computer resources, as part of the ML service 136 to generate one or more predictions 138. The one or more predictions 138 may include set of predicted immunologically active peptides, as one example. In at least one embodiment, the one or more predictions 138 may be further processed for different metrics or readability, such as ranking by binding score, using the ranking engine 126 and / or one or more additional ranking engines. A ranked result may be generated, as shown by the numeral 2, and stored in one or more datasets 304 for the given target. The one or more datasets 204 may then be used in a variety of downstream tasks, including laboratory testing, refinement, training, validation, and / or combinations thereof.
[0034] FIG. 3A illustrates an example environment that may be used with embodiments of the present disclosure to generate a set of immunologically active peptides using one or more machine learning systems. In this example, a data evaluation engine 118 may receive an input 302, which may correspond to a relevant target for a given evaluation, and / or data that may be acquired from the one or more data lakes 120. As discussed herein, the data evaluation engine 118 may include one or more pre and / or post-processing engines in order to receive input data, format the data fora given machine learning task, and / or manipulate or modify the data to generate updated data, such as data that has been ranked or compared to various additional databases or defined parameters.
[0035] The input 302 may be directed toward the data evaluation engine 118 and, based at least in part on the information from the input 302, the data evaluation engine 118 may query or otherwise pull information from the data lake 120. However, in various embodiments, information from the data lake 120 may be identified and / or provided as part of the input 302. As discussed herein, one or more processing tasks may be performed on the information from the data lake 120 and / or to the input 302 to conform to one or more specifications or preferences for analysis. Thereafter, additional processing tasks such as ranking, similarity scoring, and / or the like may be executed. In certain embodiments, additional tasks such as comparison with peptides in the human genome may also be performed in order to eliminate peptides and / or reduce a quantity of data for processing by the one or more ML systems. In operation, information is provided to the ML system 128 for further processing and analysis, which may include using compute resources to analyze data using one or more target ML algorithms. For example, peptides with a likelihood of binding that exceed a threshold may be identified and then further ranked, for example using one or more metrics, such as a binding score. Accordingly, the ML system 128 may be used to generate a list 304, which may be a ranked result set. In certain embodiments, the list 304 may be committed to a database, such as one associated with the data lake 120.
[0036] One or more embodiments may also provide the list 304 for testing 306, which may include laboratory testing, simulations, and / or combinations thereof. During testing, irrelevant and / or weak bindings may be discarded to further cull or otherwise reduce a size of the list 304. Accordingly, a set of results 308 may be generated from the testing 306, which may be saved or otherwise committed to one or more datastores, which may be associated with the data lake 120.
[0037] In at least one embodiment, the results 308 may also be used by the one or more training engines 132 to retrain the ML systems 128, or components therefor. For example, weights may be tuned for different models and / or ranking or sorting rules may be adjusted based on the results 308. In this manner, the ML system 128 may undergo continued refinement and improvement as new peptides are discovered and / or as new binding information is developed. Additionally, the training engine 132 may delay training until additional results 308 are generated. For example, multiple different input 302 may be provided to analysis different HLA complexes. By collecting information before retraining, compute resources may be conserved. After training the model(s) associated with the ML system 128 may be validated, tested, optimized as necessary, and then deployed for use with different inputs 302. Accordingly, systems and methods of the present disclosure may be associated with training and / or developing different models to identify immunologically active peptides.
[0038] FIG. 3B illustrates an example flow chart of an example process 320 for developing, training, and maintaining one or more ML models. In the process beings 322 and an input is received for a relevant target 324. The input may be provided by one or more users, which may be human users, or may be provided as part of a workflow, for example code that is executed to process a set or subset of data in an automated or semi-automated fashion. The input may be associated with a target peptide or HLA complex. Responsive to the input, additional data may be imported or otherwise acquired by the process, such as data that has been processed or prepared for one or more ML operations. In this example, further data processing is performed to modify and format the data it one or more desired specifications 326. The data, which may have been modified or passed through, may then be ranked tuned or otherwise analyzed by one or more systems 328. For example, relevancy ranking, similarity scoring, and / or the like may be performed.In certain embodiments, the scoring may be associated with one or more metrics. The scoring maybe used to determine which data is processed by a downstream ML system. For example, scoring may be used to remove likely irrelevant information in order to reduce compute resource usage and / or to provide faster downstream results. The documents may be prepared for ML analysis 330 and then compute resources maybe used to analyze the data using one or more ML algorithms 332, as discussed herein. For example, one or more neural networks may be used to identify a group or cluster of documents that may be associated with known documents for a given HLA complex that has shown immunological activity for a target condition, such as a virus or cancer.
[0039] The ML analysis may create a ranked result set of documents, which may be committed to a database 334. For example, the database may be part of the data lake, and / or may be a database that is formatted specifically for the ranked result set of documents. The ranked set may then be provided to a lab for testing 336, which may include physical testing and / or simulations. Testing may be used to discard irrelevant / weak bindings 338 and generate a refined dataset, which may be received 340 and then committed to one or more databases 342. The process may be repeated for each HLA complex 344, and / or until one or more stop conditions are reached, and then a set of results may be committed for storage 346 and importation as a finalized dataset 348. The finalized dataset may then be used to train and / or retrain one or more neural networks 350, which may be validated after training 352, and then tested 354. In certain embodiments, testing may yield additional improvements that may be incorporated to optimize the model 356. Upon validation and verification of the model, one or more systems may be deployed which may interact with one or more users 358. The one or more systems may incorporate various features, such as natural language processing (NLP) or other interface options for ease of use and accessibility. As the model is used, results may be continuously monitored and refined 360, thereby providing a continuously improving system and ending the development and training process 362.
[0040] FIG. 4 illustrates an example environment 400 that may be used with embodiments of the present disclosure. This example illustrates a pre-ML processing step that may be used to remove peptides identified as being part of the human genome, which may be unlikely to produce a sufficient immune response to treat one or more conditions. As a result, by identifying and removing a subset of potential target peptides, embodiments may reduce compute use and processing times. Furthermore, downstream testing, either physical or simulation, may be conserved because testing could be targeted toward those peptides with a higher likelihood of producing a desired result. In this example, a target set of peptides 402 and a human genome binding set of peptides 404 are processed using one or more evaluation engines 406. The evaluation engine 406 may be used to identify overlapping sets of peptides between the target 402 and the human genome binding peptides 404 to remove or otherwise reduce the total number of processed peptides to generate a ranked list 408. The ranked list 408 may include less information than the target 402, and therefore, may be less compute intensive when processed by the ML system 128, which may then generate a list of peptides 410 for further processing and testing, as discussed herein.
[0041] FIG. 5A illustrates an example flow chart for a process 500 for determining a set of immunologically active peptides. It should be appreciated that steps for the method may be performed in any order, or in parallel, unless otherwise specifically stated. Moreover, the method may include more or fewer steps. In this example, a relevant target is received associated with an immunologically active peptide 502. For example, as part of an evaluation process, an input may be provided with a target condition (e.g., a virus, a cancer, etc.) and / or may include a target peptide for evaluation, among other options. One or more human genome peptides may be determined 504 and may be compared to a list associated with the relevant target 506. For example, the relevant target may be provided as an input and a first list may be generated of potential peptides that maybe associated with that target. The comparison may determine whether or not there are overlapping peptides between the target list and the human genome list 508. If so, then the one or more human genome peptides may be removed from the target list 510. As discussed herein, the one or more human genome peptides may be unlikely to develop an immunological response, and therefore, removing them from an initial list prior to further processing or evaluation may reduce resource use. One or more ranked lists of likely immunologically active peptides may then be generated 512. The ranked list may be tested 514, which may include simulation testing and / or lab testing, to then generated a refined list 516. In this manner, ML systems may be used with lab testing to reduce a total number of peptides tested to reduce time to market and overall costs associated with peptide identification and development.
[0042] FIG. 5B illustrates an example flow chart for a process 520 for developing a peptide identification system. In this example, a set of ranked results for an immunologically active peptide is generated 522. The set of ranked results may be generated by one or more ML systems, as discussed herein. The set of results may then be evaluated, for example using laboratory analysis, and one or more results may be discarded to generate a refined set of ranked results 524. The refined set of ranked results may then be added to a training dataset 526, which may be used, at least in part, to train one or more ML systems 528. As discussed herein, the one or more ML systems may include one or more models used to generate the initial set of ranked results.
[0043] In at least one embodiment, one or more validation tests may be executed on the one or more ML systems 530 and it may be determined whether or not to retrain the model 532. For example, an error rate may be evaluated, or some other metric may be used, such as one or more loss functions, to detemiine whether or not to retrain models associated with the one or more ML systems. If it is determined not to retrain the one or more ML systems, then the one or more MLsystems may be deployed 534. In certain embodiments, active use of the one or more ML systems may be used to collect additional information that may be used to retrain or refine the model 536.
[0044] FIG. 6 illustrates a set of general components of an example computing device 600. In this example, the device includes a processor 602 for executing instructions that can be stored in a memory 604. The device can include many types of memory, data storage, or non-transitory computer-readable storage media, such as a first data storage for program instructions for execution by the processor 602, a separate storage for images or data, a removable memory for sharing information with other devices, etc. The device may optionally include a display element 606, such as a touch screen or liquid crystal display (LCD), although devices such as portable media players might convey information via other means, such as through audio speakers, and other devices may not include displays, such as server components executing within data centers, among other options. As discussed, the device in many embodiments will include at least one interaction component 608 able to receive input from a user. This input can include, for example, a push button, touch pad, touch screen, wheel, joystick, keyboard, mouse, keypad, or any other such device or element whereby a user can input a command to the device. In some embodiments, however, such a device might not include any buttons at all and might be controlled only through a combination of visual and audio commands, such that a user can control the device without having to be in contact with the device. In some embodiments, the computing device 600 of FIG. 6 can include one or more network interface or communication components 610 for communicating over various networks, such as a Wi-Fi, Bluetooth, RF, wired, or wireless communication systems. The device may be configured to communicate with a network, such as the Internet, and may be able to communicate with other such devices. The device will also include one or more power components 612, such as power cords, power ports, batteries, wirelessly powered or rechargeable receivers, and the like.
[0045] Storage media and other non-transitory computer readable media for containing code, or portions of code, can include any appropriate media known or used in the art, including storage media and communication media, such as but not limited to volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules, or other data, including RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by the a system device. Based on the disclosure and teachings provided herein, a person of ordinary skill in the art will appreciate other ways and / or methods to implement the various embodiments.
[0046] Embodiments may also be described in view of the following clauses:1. A computer-implemented method, comprising: receiving a relevant target associated with an immunologically activate peptide for treating one or more conditions; determining one or more human genome peptides associated with a human genome; comparing the one or more human genome peptides to a list of target peptides generated, based at least in part, on the relevant target; removing the one or more human genome peptides from the list of target peptides; generating, based at least in part on the relevant target and the list of target peptides, a ranked list of immunologically active peptides; testing the ranked list of immunologically active peptides; and generating an active list based, at least in part, on the testing.2. The computer-implemented method of clause 1, wherein the testing includes at least one of laboratory testing or simulations.3. The computer-implemented model of clause 1, further comprising: retrieving one or more datasets based, at least in part, on the relevant target.4. The computer-implemented method of clause 1, further comprising: determining one or more models associated with the ranked list of immunologically active peptides exceeds one or more validation metrics; and retraining the one or more models.5. The computer-implemented method of clause 1, wherein the relevant target is provided via a natural language processing interface.6. The computer-implemented method of clause 1, wherein the ranked list of immunologically active peptides is ranked based on at least one of immunological activity or strength of interaction.7. The computer-implemented method of clause 1, wherein the list of target peptides is generated by a first trained model and the ranked list is generated by a second trained model.8. The computer-implemented method of clause 1, wherein each of the list of target peptides and the ranked list is generated by a trained neural network.9. A processor, comprising: one or more circuits to: generate a set of ranked results for an immunologically active peptide; discard one or more results from the set of ranked results based, at least in part, on testing to generate a refined set of ranked results; add the refined set of ranked results to a training dataset;train one or more machine learning models using at least a portion of the training dataset; execute one or more validation tests on the one or more machine learning models; determine the one or more validation tests exceed one or more metrics; deploy the one or more machine learning models; collect, over a period of time, results from execution of the one or more machine learning models; and add the results to the training dataset.10. The processor of clause 9, wherein the one or more circuits are further to: receive a relevant target for the immunologically active peptide; generate a first set of potential peptides; and remove one or more peptides from the first set of potential peptides based, at least in part, on a comparison to one or more genome peptides.11. The processor of clause 10, wherein the one or more genome peptides share one or more binding sites with the relevant target.12. The processor of clause 9, wherein the one or more processors are further to: retrieve one or more datasets based, at least in part, on the relevant target; and extract a portion of the one or more datasets based, at least in part, on at least one of a similarity score or a relevance ranking.13. The processor of clause 9, wherein the deployed one or more machine learning models include a natural language processing interface.14. The processor of clause 9, wherein the one or more validation tests are executed against one or more pre-existing datasets.15. The processor of clause 9, wherein the testing includes at least one of physical laboratory test or a computer-aided simulation.16. A computer-implemented method, comprising: generating a set of ranked results for an immunologically active peptide; discarding one or more results from the set of ranked results based, at least in part, on testing to generate a refined set of ranked results; adding the refined set of ranked results to a training dataset; training one or more machine learning models using at least a portion of the training dataset; executing one or more validation tests on the one or more machine learning models; determining the one or more validation tests exceed one or more metrics; deploying the one or more machine learning models; collecting, over a period of time, results from execution of the one or more machine learning models; and adding the results to the training dataset.17. The computer-implemented method of clause 16, further comprising: receiving a relevant target for the immunologically active peptide; generating a first set of potential peptides; and removing one or more peptides from the first set of potential peptides based, at least in part, on a comparison to one or more genome peptides.18. The computer-implemented method of clause 17, wherein the one or more genome peptides share one or more binding sites with the relevant target.19. The computer-implemented method of clause 16, further comprising: retrieving one or more datasets based, at least in part, on the relevant target; andextracting a portion of the one or more datasets based, at least in part, on at least one of a similarity score or a relevance ranking.20. The computer-implemented method of clause 16, wherein the testing includes at least one of physical laboratory test or a computer-aided simulation.21. A computer-implemented method, comprising: receiving a relevant target associated with an immunologically activate peptide for treating one or more conditions; determining one or more human genome peptides associated with a human genome; comparing the one or more human genome peptides to a list of target peptides generated, based at least in part, on the relevant target; removing the one or more human genome peptides from the list of target peptides; generating, based at least in part on the relevant target and the list of target peptides, a ranked list of immunologically active peptides; testing the ranked list of immunologically active peptides; and generating an active list based, at least in part, on the testing.22. The computer-implemented method of clause 21, wherein the testing includes at least one of laboratory testing or simulations.23. The computer-implemented model of any of clauses 21 or 22, further comprising: retrieving one or more datasets based, at least in part, on the relevant target.24. The computer-implemented method of any of clauses 21-23, further comprising: determining one or more models associated with the ranked list of immunologically active peptides exceeds one or more validation metrics; and retraining the one or more models.25. The computer-implemented method of any of clauses 21-24, wherein the relevant target is provided via a natural language processing interface.26. The computer-implemented method of any of clauses 21-25, wherein the ranked list of immunologically active peptides is ranked based on at least one of immunological activity or strength of interaction.27. The computer-implemented method of any of clauses 21-26, wherein the list of target peptides is generated by a first trained model and the ranked list is generated by a second trained model.28. The computer-implemented method of any of clauses 21-27, wherein each of the list of target peptides and the ranked list is generated by a trained neural network.29. A processor, comprising: one or more circuits to: generate a set of ranked results for an immunologically active peptide; discard one or more results from the set of ranked results based, at least in part, on testing to generate a refined set of ranked results; add the refined set of ranked results to a training dataset; train one or more machine learning models using at least a portion of the training dataset; execute one or more validation tests on the one or more machine learning models; determine the one or more validation tests exceed one or more metrics; deploy the one or more machine learning models; collect, over a period of time, results from execution of the one or more machine learning models; and add the results to the training dataset.30. The processor of clause 29, wherein the one or more circuits are further to: receive a relevant target for the immunologically active peptide; generate a first set of potential peptides; and remove one or more peptides from the first set of potential peptides based, at least in part, on a comparison to one or more genome peptides.31. The processor of clause 30, wherein the one or more genome peptides share one or more binding sites with the relevant target.32. The processor of any of clauses 29-31, wherein the one or more processors are further to: retrieve one or more datasets based, at least in part, on the relevant target; and extract a portion of the one or more datasets based, at least in part, on at least one of a similarity score or a relevance ranking.33. The processor of any of clauses 29-32, wherein the deployed one or more machine learning models include a natural language processing interface.34. The processor of any of clauses 29-33, wherein the one or more validation tests are executed against one or more pre-existing datasets.35. The processor of any of clauses 29-34, wherein the testing includes at least one of physical laboratory test or a computer-aided simulation.36. A computer-implemented method, comprising: generating a set of ranked results for an immunologically active peptide; discarding one or more results from the set of ranked results based, at least in part, on testing to generate a refined set of ranked results; adding the refined set of ranked results to a training dataset;training one or more machine learning models using at least a portion of the training dataset; executing one or more validation tests on the one or more machine learning models; determining the one or more validation tests exceed one or more metrics; deploying the one or more machine learning models; collecting, over a period of time, results from execution of the one or more machine learning models; and adding the results to the training dataset.37. The computer-implemented method of clause 36, further comprising: receiving a relevant target for the immunologically active peptide; generating a first set of potential peptides; and removing one or more peptides from the first set of potential peptides based, at least in part, on a comparison to one or more genome peptides.38. The computer-implemented method of clause 37, wherein the one or more genome peptides share one or more binding sites with the relevant target.39. The computer-implemented method of any of clauses 36-38, further comprising: retrieving one or more datasets based, at least in part, on the relevant target; and extracting a portion of the one or more datasets based, at least in part, on at least one of a similarity score or a relevance ranking.40. The computer-implemented method of any of clauses 36-39, wherein the testing includes at least one of physical laboratory test or a computer-aided simulation.
[0047] Although the technology herein has been described with reference to particular embodiments, it is to be understood that these embodiments are merely illustrative of the principles and applications of the present technology. It is therefore to be understood that numerousmodifications may be made to the illustrative embodiments and that other arrangements may be devised without departing from the spirit and scope of the present technology as defined by the appended claims.
Claims
CLAIMS1. A computer-implemented method, comprising: receiving a relevant target associated with an immunologically activate peptide for treating one or more conditions; determining one or more human genome peptides associated with a human genome; comparing the one or more human genome peptides to a list of target peptides generated, based at least in part, on the relevant target; removing the one or more human genome peptides from the list of target peptides; generating, based at least in part on the relevant target and the list of target peptides, a ranked list of immunologically active peptides; testing the ranked list of immunologically active peptides; and generating an active list based, at least in part, on the testing.
2. The computer-implemented method of claim 1, wherein the testing includes at least one of laboratory testing or simulations.
3. The computer-implemented model of claim 1, further comprising: retrieving one or more datasets based, at least in part, on the relevant target.
4. The computer-implemented method of claim 1, further comprising: determining one or more models associated with the ranked list of immunologically active peptides exceeds one or more validation metrics; and retraining the one or more models.
5. The computer-implemented method of claim 1, wherein the relevant target is provided via a natural language processing interface.
6. The computer-implemented method of claim 1, wherein the ranked list of immunologically active peptides is ranked based on at least one of immunological activity or strength of interaction.
7. The computer-implemented method of claim 1, wherein the list of target peptides is generated by a first trained model and the ranked list is generated by a second trained model.
8. The computer-implemented method of claim 1, wherein each of the list of target peptides and the ranked list is generated by a trained neural network.
9. A processor, comprising: one or more circuits to: generate a set of ranked results for an immunologically active peptide; discard one or more results from the set of ranked results based, at least in part, on testing to generate a refined set of ranked results; add the refined set of ranked results to a training dataset; train one or more machine learning models using at least a portion of the training dataset; execute one or more validation tests on the one or more machine learning models; determine the one or more validation tests exceed one or more metrics; deploy the one or more machine learning models; collect, over a period of time, results from execution of the one or more machine learning models; and add the results to the training dataset.
10. The processor of claim 9, wherein the one or more circuits are further to: receive a relevant target for the immunologically active peptide; generate a first set of potential peptides; and remove one or more peptides from the first set of potential peptides based, at least in part, on a comparison to one or more genome peptides.
11. The processor of claim 10, wherein the one or more genome peptides share one or more binding sites with the relevant target.
12. The processor of claim 9, wherein the one or more processors are further to:retrieve one or more datasets based, at least in part, on the relevant target; and extract a portion of the one or more datasets based, at least in part, on at least one of a similarity score or a relevance ranking.
13. The processor of claim 9, wherein the deployed one or more machine learning models include a natural language processing interface.
14. The processor of claim 9, wherein the one or more validation tests are executed against one or more pre-existing datasets.
15. The processor of claim 9, wherein the testing includes at least one of physical laboratory test or a computer-aided simulation.
Citation Information
Patent Citations
Immunogenic peptide presentation prediction method, system and device and storage medium
CN116994643A
Methods and Systems for Identification of Human Leukocyte Antigen Peptide Presentation and Applications Thereof
US20210033608A1
Method, system and computer program product for determining peptide immunogenicity
WO2022079255A1
Systems and methods for evaluating immunological peptide sequences
WO2023086999A1
Cited By
Beta-lactoglobulin ACE inhibitory peptide multi-attribute evaluation method based on machine learning
CN121075444A