Machine learning-assisted pipeline for predicting properties of atomic systems in the absence of training data
The machine learning-assisted pipeline leverages UDAL and self-tuning HMC to predict atomic system properties without initial training data, overcoming existing inefficiencies and limitations, and achieving accelerated and automated discovery of desirable atomic systems.
Patent Information
- Application Number
- PCT/IB2024/054072
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-04-26
- Publication Date
- 2025-06-26
AI Technical Summary
Current methods for predicting properties of atomic systems are hindered by the need for extensive training data, inefficiencies in simulation methods, and limitations in exploring complex configurational and compositional spaces.
A machine learning-assisted pipeline that utilizes uncertainty-driven active learning (UDAL) to train candidate-specific machine learning interatomic potential (MLIP) models, combined with a self-tuning Hamiltonian Monte Carlo (HMC) simulator, to generate predictive data without initial training data, and a probabilistic graph machine learning model to refine predictions over time.
This approach accelerates the discovery of atomic systems with desirable properties by generating accurate and representative training data efficiently, reducing computational resources, and enabling fully automated candidate selection and validation.
Smart Images

Figure IB2024054072_26062025_PF_FP_ABST
Abstract
Description
MACHINE LEARNING-ASSISTED PIPELINE FOR PREDICTING PROPERTIES OFATOMIC SYSTEMS IN THE ABSENCE OF TRAINING DATACROSS-REFERENCE TO PRIOR APPLICATION
[0001] Priority is claimed to U.S. Provisional Application No. 63 / 611,795, filed on December 19, 2023, the entire contents of which is hereby incorporated by reference herein. FIELD
[0002] The present invention relates to artificial intelligence (Al) and machine learning (ML), and in particular to a method, system, data structure, computer program product, and computer-readable medium for predicting properties of atomic systems.BACKGROUND
[0003] The classical discovery of new drugs and materials is a process that requires in-depth domain knowledge and is extremely time-consuming and resource intensive. Often, the selection of a potentially good system to test in the wet lab from a very large pool of candidates relies on the so-called “chemical intuition” of the expert, which is unfortunately difficult to replicate in an automated system. On the other hand, state-of-the-art drug and materials discovery technologies require searching through exponentially more complex configurational (atom positions) and compositional (atom types) spaces. These vast spaces are inaccessible to an expert a priori but are of interest for designing new products.
[0004] For these reasons, machine learning has been identified as a promising tool to assist the experimentalist in quickly identifying systems with desirable properties (see Yoo, Jiho, et al., “Industrializing AI / ML during the end-to-end drug discovery process,” Current Opinion in Structural Biology 79 (2023): 102528, which is hereby incorporated by reference herein). This has resulted in a few different material / drug discovery pipelines, which for instance, filter down a set of candidate compounds through a combination of classical (i.e., not necessarily accurate) molecular dynamics (MD) simulations to generate training data and ML models that, after training, act as a surrogate model in place of MD (see Gupta, Aayush, and Huan-Xiang Zhou, “Machine learning-enabled pipeline for large-scale virtual drug screening,” Journal of Chemical Information and Modeling 61.9 (2021): 4236-4244, which is hereby incorporated by reference herein). Other works rely on a closed-loop active learning (AL) scheme (Park, Mijung, Marcel Nassar, and Haris Vikalo, “Bayesian active learning for drug combinations,” IEEE transactions on biomedical engineering 60. 11 (2013): 3248-3255, hereinafter “Park et al.”; Reker, Daniel, “Practical considerations for active machine learning in drug discovery,” Drug Discovery Today: Technologies 32 (2019): 73-79; and Kavalsky, Lance, et al., “By how much can closed- loop frameworks accelerate computational materials discovery?” Digital Discovery 2.4 (2023):1112-1125, each of which is hereby incorporated by reference herein), where the goal is to build good surrogate models for predicting properties of interest, for instance interatomic potentials. Park et al. describe to iteratively select the next candidate using a Hamiltonian Monte Carlo (HMC) algorithm that relies on a surrogate model (trained incrementally as new candidates are selected) to maximize an expected improvement criterion.SUMMARY
[0005] In an embodiment, the present disclosure provides a computer-implemented machine learning method for using one or more predicted properties of one or more atomic systems to train a graph machine learning model. The method includes obtaining a candidates pool comprising a plurality of candidates associated with a plurality of atomic systems and using an uncertainty-driven active learning (UDAL) to obtain a candidate specific machine learning interatomic potential (MLIP) model for a first batch of candidates from the plurality of candidates. The method further includes incorporating the candidate specific MLIP model into a self-tuning Hamiltonian Monte Carlo (HMC) simulator to generate an HMC output indicating the one or more predicted properties of the one or more atomic systems associated with the first batch of candidates. The method also includes training the graph machine learning model based on a dataset comprising the HMC output. The method has applications including, but not limited to, use cases in medicine / healthcare, e.g., for drug design or treatment, and discovery of new materials, to optimize predictions or support decision making.BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Embodiments of the present invention will be described in even greater detail below based on the exemplary figures. The present invention is not limited to the exemplary embodiments. All features described and / or illustrated herein can be used alone or combined in different combinations in embodiments of the present invention. The features and advantages of various embodiments of the present invention will become apparent by reading the following detailed description with reference to the attached drawings which illustrate the following:
[0007] FIG. 1 illustrates a simplified block diagram depicting an exemplary computing environment according to one or more embodiments of the present invention;
[0008] FIG. 2 schematically illustrates a method and machine learning pipeline for predicting candidates to test in a wet lab according to one or more embodiments of the present invention; and
[0009] FIG. 3 is a block diagram of an exemplary processing system, which can be configured to perform any and all operations disclosed herein.DETAILED DESCRIPTION
[0010] The Al-assisted discovery of atomic systems with desired properties often relies on the availability of system-specific and task-specific training data sets. Embodiments of the present invention introduce an Al-assisted pipeline to determine (e.g., find) promising candidates for wet lab experiments in a data-driven fashion without requiring any a priori data set. Additionally, and / or alternatively, the Al-assisted pipeline according to embodiments of the present invention is self-adaptive (e.g., its performance can improve as it runs).
[0011] Embodiments of the present invention overcome a number of technical challenges and provide a number of improvements over existing technology. In particular, some of the problems of existing approaches overcome by embodiments of the present invention include:It is technically challenging to identify good simulation data for training surrogate models as it is hard to acquire for a specific atomic system of interest. Transfer learning strategies built on top of public datasets can only marginally help since each atomic system has unique properties.Active learning (AL) approaches are typically not guaranteed to sample from the correct distribution to study the properties of an atomic system, as the end goal there is only to build a good heuristic surrogate. Computer simulations such as HMC, on the other hand, sample configurations that explore the space of the property of interest, but when used alone, suffer from other deficiencies in performance.Obtaining ground truth data from wet labs can be costly and slow, and can consume a great deal of resources. Similarly, using density functional theory (DFT) to obtain very accurate interatomic forces for an atomic system and predict its properties by running atomistic simulations is unfeasible for even medium-sized (e.g., a few hundred of atoms) atomic systems, whereas classical force fields might not be sufficiently accurate.
[0012] Embodiments of the present invention extend the AL closed loop and provide solutions to the above-described technical problem(s). In particular, embodiments of the present invention provide numerous improvements including, but not limited to, the following improvements:Using an uncertainty driven active learning (UDAL) method to accelerate the training of a candidate-specific machine learning interatomic potential (MLIP) model, which is more accurate than classical force fields and faster than DFT. This advantageously requires no a priori training data and produces an MLIP faster than other solutions, thereby being flexible, computationally efficient, and enabling to conserve computational resources.Obtaining a valid dataset by running a self-tuning Hamiltonian Monte Carlo (HMC) algorithm to sample interesting configurations of the atomic system of interest more efficiently than classical approaches. The MLIP model trained using UDAL is used to compute interatomicpotentials which are needed in the atomic system’s simulation. The algorithm is used to generate training data, and the dataset is incrementally built as more candidates are simulated.Using the current dataset to train a probabilistic graph machine learning model that returns the probability of a candidate atomic system to have the desired property. The model is refined and improves in performance as the dataset grows, and it is used to suggest the batch of candidates to consider in the next iteration.
[0013] After running (e.g., executing) a self-tuning HMC simulation for the batch of selected candidates, if a candidate with desirable property is found, it is sent to the wet lab for further validation. This pipeline effectively accelerates and fully automatizes the discovery of suitable, molecular, drug or material candidates.
[0014] One or more embodiments of the present invention advantageously combine three strategies (e.g., UDAL, self-tuning HMC, and probabilistic machine learning) that leverage the strengths of each other to move beyond the classical active learning setup and allow to refine a pool of candidates more efficiently without the need for any human intervention or “chemical intuition”.
[0015] According to a first aspect, the present invention provides a computer-implemented machine learning method for using one or more predicted properties of one or more atomic systems to train a graph machine learning model. For instance, the method includes obtaining a candidates pool comprising a plurality of candidates associated with a plurality of atomic systems and using an uncertainty-driven active learning (UDAL) to obtain a candidate specific machine learning interatomic potential (MLIP) model for a first batch of candidates from the plurality of candidates. The method further includes incorporating the candidate specific MLIP model into a self-tuning Hamiltonian Monte Carlo (HMC) simulator to generate an HMC output indicating the one or more predicted properties of the one or more atomic systems associated with the first batch of candidates. The method also includes training the graph machine learning model based on a dataset comprising the HMC output .
[0016] According to a second aspect, the method according to the first aspect further comprises selecting a new batch of candidates from the candidates pool based on using the trained graph machine learning model.
[0017] According to a third aspect, the method according to the first or second aspect further comprises extracting the first batch of candidates from the candidates pool, wherein using the UDAL to obtain the candidate specific MLIP model comprises training the candidate specific MLIP model for the first batch of candidates based on using the UDAL.
[0018] According to a fourth aspect, the method according to any of the first to third aspects further comprises that extracting the first batch of candidates is based on using a randomized process or user input indicating domain-specific pre-screening information
[0019] According to a fifth aspect, the method according to any of the first to fourth aspects further comprises that extracting the first batch of candidates comprises: obtaining operator input indicating to enable the graph machine learning model; and using the graph machine learning model to select the first batch of candidates.
[0020] According to a sixth aspect, the method according to any of the first to fifth aspects further comprises that training the candidate specific MLIP model for the first batch of candidates comprises: obtaining atomic system information and molecular dynamics (MD) associated with the first batch of candidates; and inputting the atomic system information and the molecular dynamics (MD) into the UDAL to train the candidate specific MLIP model for the first batch of candidates, wherein the candidate specific MLIP model is used to determine a prediction energy of different configurations of a candidate system associated with the first batch of candidates.
[0021] According to a seventh aspect, the method according to any of the first to sixth aspects further comprises that the self-tuning HMC simulator comprises a first tunable parameter and a second tunable parameter, wherein the first tunable parameter is a time-step for a velocity Verlet integrator, and wherein the second tunable parameter is a number of integration steps.
[0022] According to an eighth aspect, the method according to any of the first to seventh aspects further comprises using the UDAL to obtain one or more additional candidate specific MLIPs models for one or more additional batches of candidates from the plurality of candidates; incorporating the one or more additional candidate specific MLIPs models into the self-tuning HMC simulator to generate one or more additional HMC outputs; and generating the dataset, wherein the dataset comprises the HMC output and the one or more additional HMC outputs.
[0023] According to an ninth aspect, the method according to any of the first through eighth aspects further comprises that generating the dataset comprises: labelling the HMC output and the one or more additional HMC outputs based on using a tuple indicating a specific configuration of a candidate and a property value, and wherein training the graph machine learning model is based on the labelling of the HMC output and the one or more additional HMC outputs.
[0024] According to an tenth aspect, the method according to any of the first through ninth aspects further comprises that the graph machine learning model maps a configuration of acandidate system associated with the first batch of candidates into a value of interest, wherein the configuration is mapped together with a measure of uncertainty over a predicted value.
[0025] According to an eleventh aspect, the method according to any of the first through tenth aspects further comprises that training the graph machine learning model is further based on using probabilistic machine learning, and wherein the graph machine learning model is a deep graph network (DGN).
[0026] According to a twelfth aspect, the method according to any of the first through eleventh aspects further comprises using the UDAL to obtain a new candidate specific MLIP model for the new batch of candidates; incorporating the new candidate specific MLIP model into the self-tuning HMC simulator to generate new HMC output; and performing further training of the graph machine learning model based on using an updated dataset comprising the HMC output and the new HMC output.
[0027] According to a thirteenth aspect, the method according to any of the first through twelfth aspects further comprises, based on the HMC output satisfying one or more properties of interest, providing, to a wet lab computing system, candidate information indicating the first batch of candidates and the HMC output, wherein the candidates pool indicates a pool of T-cell receptor (TCR)-binding candidates or a pool of candidate materials, wherein the first batch of candidates is a first batch of TCR candidates, wherein the HMC output indicates a simulation of an interaction between the first batch of TCR candidates and a peptide, wherein the candidate information comprises information indicating binding of the first batch of TCR candidates to the peptide, wherein selecting the new batch of candidates comprises selecting a new batch of TCR candidates based on the simulation of the interaction between the first batch of TCR candidates and the peptide, and wherein the new batch of candidates is used for drug vaccine development or for performing medical treatments.
[0028] According to a fourteenth aspect of the present disclosure, a computer system is provided for using one or more predicted properties of one or more atomic systems to train a graph machine learning model, the system comprising one or more hardware processors, which, alone or in combination, are configured to provide for execution of the method according to any of the first to thirteenth aspects.
[0029] A fifteenth aspect of the present disclosure provides a tangible, non-transitory computer-readable medium having instructions thereon, which, upon being executed by one or more processors, provides for execution of the method according to any of the first to the thirteenth aspects.
[0030] FIG. 1 illustrates a simplified block diagram depicting an exemplary computing environment 100 according to one or more embodiments of the present invention. Theenvironment 100 includes one or more computing systems 104. Although certain entities within environment 100 can be described below and / or depicted in the FIGs. as being singular entities, it will be appreciated that the entities and functionalities discussed herein can be implemented by and / or include one or more entities.
[0031] In operation, the one or more computing systems 104 can perform one or more processes, methods, and / or algorithms such as, but not limited to, the method and machine learning pipeline 200 for predicting candidates to test in a wet lab, which is described in FIG. 2 below. For example, the one or more computing systems 104 can obtain information 102, perform the method and machine learning pipeline 200, and output information 106. For instance, the one or more computing systems 104 can obtain information 102 from one or more data sources. The information 102 obtained from the data sources can include, but is not limited to, a pool of candidate materials and / or T-cell receptors (e.g., the candidates’ pool 212 such as the drug candidates that is shown in FIG. 2). Then, using the method and machine learning pipeline 200 that is described in FIG. 2, the one or more computing systems 104 can determine and output information 106. For example, the one or more computing systems 104 can output information 106 to a computing system, device, apparatus, and / or entity associated with a wet lab. The information 106 can include, but is not limited to a stream of candidates that is sent for validation to a wet lab. This will be described in more detail below.
[0032] In some variations, the environment 100 can include one computing system 104 that performs the method and machine learning pipeline 200. In other variations, the environment 100 can include a plurality of computing systems 104. For example, each step of the method and machine learning pipeline 200 can be performed by one or more computing systems 104. Then, using a network, the one or more computing systems 104 can provide an output of the step to the next computing system(s) 104 (e.g., a first computing system 104 can provide the candidate selection of promising candidates to a second computing system 104). The next computing system(s) 104 can perform the next step of the method and machine learning pipeline 200 (e.g., the second computing system 104 can perform the uncertainty driven active learning from the second step 204 of FIG. 2, and provide the output to a third computing system 104), and so on. In such variations, the computing systems 104 can communicate via a network. The network can be a global area network (GAN) such as the Internet, a wide area network (WAN), a local area network (LAN), or any other type of network or combination of networks. The network may provide a wireline, wireless, or a combination of wireline and wireless communication between the entities within the environment 100 (e.g., the plurality of computing systems 104).
[0033] The computing systems 104 can be and / or include, but are not limited to, one or more computing devices, computing platforms, cloud computing platforms, systems, servers, desktops,laptops, tablets, mobile devices, internet of things (IOT) devices, and / or other apparatuses / devices. In some examples, the computing system 104 can include and / or comprise one or more communication components, one or more processing components, and one or more memory components.
[0034] It will be appreciated that the exemplary environment depicted in FIG. 1 is merely an example, and that the principles discussed herein can also be applicable to other situations and / or examples. For instance, in some examples, the environment 100 can include the data sources that provide the information 102 to the one or more computing systems 104. Additionally, and / or alternatively, the environment 100 can include the computing systems, devices, and / or apparatuses associated with the wet labs. These computing systems, devices, and / or apparatuses can obtain the information 106 from the one or more computing systems 104.
[0035] FIG. 2 schematically illustrates a method and machine learning pipeline 200 for predicting candidates to test in a wet lab according to one or more embodiments of the present invention. For example, FIG. 2 visually depicts a machine learning pipeline according to one or more embodiments of the present invention. The method and machine learning pipeline 200 can be performed by the one or more computing systems 104 of environment 100 shown in FIG. 1. It will be recognized that the method and machine learning pipeline 200 can be performed in any suitable environment. The descriptions, illustrations, methods and processes of FIG. 2 are merely exemplary and the method and machine learning pipeline 200 may use other descriptions, illustrations, and processes.
[0036] For example, embodiments of the present invention can address a problem where the goal is to select suitable candidates (e.g., possibly selected from publicly available repositories) from a pool P = {s1;s2, ... , sw} of N atomic systems according to one or more specific properties y that are desired to either minimize or maximize. Each system (e.g., Sj such asor s2) is described as SjG IRn‘xFdenotes the matrix of atoms’ attributes. For instance, the system st, which includes ntatoms, can be described by a matrix that includes ntrows and “F” columns. This can be a way to describe the information uniquely to each of the atoms in the system. For instance, each row can indicate an atom and its continuous attributes, which is represented by “F”. For example, row j of this matrix encodes the F continuous attributes, such as atom type, atom mass, etc., that are specific to atom j in the system.
[0037] Additionally, and / or alternatively, embodiments of the present invention can also consider an additional target system that is paired with each candidate. For instance, in an embodiment, the target is a virus and candidate drugs can be evaluated against the target virus to determine whether the candidate drugs can be effective on that virus. For example, in some embodiments, the target virus and a candidate drug can be considered as a unique system, s, (asdescribed above), adding the information that some atoms belong to the virus and others to the drug. Thus, the drug system is virtually augmented with the target virus and can be processed by the pipeline 200. In other words, the unique system can include information (e.g., atoms) for both the target virus and the target candidate drug. The computing system 104 can use the unique system with the atoms for both the target virus and the target candidate drug for the pipeline 200.
[0038] The first step 202 of the pipeline 200 is to select one or more promising candidates that needs to be evaluated (e.g., a batch of promising candidates to be evaluated). This process can be random, guided by a first domain-specific pre-screening carried out by humans (e.g., domain-specific pre-screening information), or carried out by a graph machine learning model described below.
[0039] For example, the computing system 104 (e.g., one or more of the computing systems 104 from environment 100) can perform the first through fifth steps 202-210 shown in FIG. 2. In the first step 202, the computing system 104 can select (e.g., extract) one or more promising candidates (e.g., selected candidates) such as a batch of promising candidates to be evaluated. In some examples, the computing system 104 can select the one or more promising candidates using a randomized process (e.g., by using a random number generator (RNG) and / or other processes). In other examples, the computing system 104 can select the one or more promising candidates based on user information (e.g., user information and / or user input indicating a first domain-specific pre-screening carried out by humans) and / or based on one or more machine learning (ML) models such as the graph ML model described below.
[0040] At first, it can be expected that the random and / or human-guided strategies will be the most reasonable ones as the graph machine learning model is not yet trained, but the performance of the model starts to improve and become advantageous as more data is collected by the pipeline. The policy of when to enable the ML model can be left to the expert user.
[0041] For example, initially (e.g., the first iteration of steps 202-210 and / or the first few iterations of steps 202-210), the computing system 104 can select the one or more promising candidates based on using a randomized process and / or based on the user information (e.g., user input). Afterwards, based on operator input (e.g., input from the expert user indicating a policy to enable the graph ML model), the computing system 104 can select the one or more promising candidates based on using the ML model (e.g., the graph ML model described below).
[0042] The batch of selected candidates (e.g., one or more candidates), together with the additional target system (if required), are passed as input to the uncertainty-driven active learning module (e.g., uncertainty-driven active learning processor), which is described further below. The candidates are also removed from the pool.
[0043] For instance, for the first step 202, the computing system 104 can obtain information 102. The information 102 can include the candidates’ pool 212 (e.g., the drug candidates).Then, the computing system 104 can perform candidate selection based on the above (e.g., select the batch of promising candidates based on a randomized process, user input, and / or using a graph ML model). The computing system 104 then passes the promising candidates (e.g., the batch of selected candidates) to the next step 204 (e.g., an uncertainty -driven active learning module). Additionally, and / or alternatively, the computing system 104 can determine one or more additional target systems and provide the additional target systems to the uncertainty- driven active learning module. For instance, the computing system 104 can obtain additional input 214. In some instance, the additional input 214 can be user and / or operator input indicating one or more target systems (e.g., a target virus). Additionally, and / or alternatively, the computing system 104 can remove the candidates from the pool (e.g., from the candidates’ pool 212).
[0044] At a second step 204, the computing system 104 can use an uncertainty driven active learning module to perform uncertainty driven active learning. The uncertainty driven active learning module can be and / or include one or more processors, engines, software instructions, controllers, hardware devices, and / or computing apparatuses. In other words, in some variations, the uncertainty driven active learning module can be and / or include one or more hardware devices that can be used by the computing system 104. In other variations, the uncertainty driven active learning module can be software-instructions (e.g., a batch of softwareinstructions) stored in memory. The computing system 104 (e.g., one or more hardware processors of the computing system 104) can execute the uncertainty driven active learning module to perform uncertainty driven active learning.
[0045] The uncertainty-driven active learning (UDAL) is a method for generating comprehensive (e.g., covers the relevant configurational and compositional spaces) yet minimal (e.g., reduces the number of DFT calculations) training data sets for atomic systems. UDAL uses uncertainty-biased atomistic simulations to guarantee uniform root-mean-squared errors (RMSEs) reduction in predicted energies and atomic forces. Thus, it guarantees that developed MLIPs 216 can be used for large-scale atomistic simulations for computing thermodynamic and other properties of atomic systems without loss of accuracy (a property required for running, e.g., self-tuning HMC with surrogate models). Using gradient-based uncertainties and batch selection algorithms improves the computational efficiency of UDAL and results in maximally diverse training data sets.
[0046] Inputs to UDAL are the atomic system, including atom positions and atom types, physically inspired atomistic simulations (e.g., MD) for generating the candidate pool, and aquantum mechanical solver that provides reference energies and atomic forces. UDAL is an iterative approach for training MLIP models, and outputs trained MLIP models and generated training data sets.
[0047] For instance, based on being provided the set of candidates (e.g., drug candidates) that are selected from the pool of candidates, UDAL is used to train the MLIPs from scratch (i.e., without any prior training data). Each candidate can be described by its atomic positions (e.g., 3-dimensional (3-D) coordinates of atoms) and atomic numbers (e.g., types of atoms such as 1 for hydrogen atom and 6 for carbon atom).
[0048] For each candidate, the computing system 104 can perform the following steps:1) Create a small initial training data set by randomly perturbing atomic positions, e.g., draw displacements from a random uniform distribution and add them to atomic positions.2) Train an MLIP using this initial training data set.3) Run atomistic simulations, e.g., molecular dynamics (MD) simulations, using the trained MLIP to generate a pool of candidates. It is noted that these candidates are different from those used to select the drug candidates (e.g., the drug candidates from step 202). These candidates are generated and used for UDAL only and represent different sets of atomic positions for the same drug molecule. In some instances, the computing system 104 can additionally bias the respective atomistic simulation with the model's prediction uncertainty, realized by adding this uncertainty to the predicted energy. This bias effectively drives the respective atomistic simulation to regions where the current MLIP is uncertain about its predictions, e.g., regions where the current MLIP has the greatest likelihood to make a wrong energy and force prediction. In some variations, the computing system 104 can drive the atomistic simulation by biased forces, computed as the sum of actual (true) forces and forces obtained as a gradient of uncertainty with respect to atomic positions. Including these regions can be crucial for training uniformly accurate MLIPs, e.g., an MLIP that can predict energy and force for any candidate with similar accuracy.4) Select a batch from the above candidates and compute labels (e.g., energies and forces) using the quantum mechanical solver.5) Update the training data set.6) Train a new MLIP with the updated data set and repeat step (3) until the maximal training data set size is acquired.
[0049] In some embodiments, gradient-based uncertainties allow ensemble-free uncertainty quantification. That is, they allow to quantify uncertainties without training an ensemble (e.g., more than three MLIPs at the same time). Typically, the model's prediction uncertainty is estimated as the variance between separate predictions of models in the ensemble. Thus,gradient-based uncertainties improve UDAL by reducing the computational cost by a factor proportional to the number of models used in ensemble. To compute gradient-based uncertainties, the computing system 104 employs gradients of the output of the MLIP with respect to its parameters, e.g., weights and biases. The computing system 104 can compute the gradient-based uncertainty itself as a distance between those gradients (referred to as gradient features) or as a posterior uncertainty (e.g., Gaussian process regression).
[0050] Batch selection algorithms allow the computing system 104 to select more than one candidate at once such that they are informative (e.g., high uncertainty) and maximally diverse. Particularly, the computing system 104 can use two algorithms:1) Maximize the distance between selected candidates in the space spanned by the gradient features.2) Maximize the determinant of the covariance matrix (e.g., Gaussian process regression) built using gradient features.
[0051] Selecting a maximally diverse batch can be crucial. If it is not enforced, similar candidates can be chosen, worsening the performance of next-generation MLIP. Also, using batch selection allows running multiple atomistic simulations for generating the pool of candidates in parallel, which can accelerate the generation of corresponding pools.
[0052] For example, at the second step 204, the computing system 104 can obtain information (e.g., from step 202) such as the one or more promising candidates, the additional input 214 (e.g., the target system), and / or other information. As such, the computing system 104 can obtain the atomic system information (e.g., atom positions and / or atom types), physically inspired atomistic simulations (e.g., molecular dynamics) for generating the candidate pool, and / or a quantum mechanical solver that provides reference energies and atomic forces. Based on the obtained information, the computing system 104 can use the UDAL module to train one or more MLIP models 216 for the selected candidate(s).
[0053] For instance, using UDAL, the computing system 104 uses uncertainty-biased atomistic simulations to guarantee uniform RMSEs reduction in predicted energies and atomic forces. For example, as mentioned above, in the third step of performing UDAL, the computing system 104 can additionally bias the respective atomistic simulation with the model's prediction uncertainty, realized by adding this uncertainty to the predicted energy. This bias effectively drives the respective atomistic simulation to regions where the current MLIP is uncertain about its predictions. Additionally, and / or alternatively, the computing system 104 can use gradientbased uncertainties and batch selection algorithms to improve the computational efficiency of UDAL and results in maximally diverse training data sets. For example, as also mentioned above in the third step, the computing system 104 can drive the atomistic simulation by biasedforces, computed as the sum of actual (true) forces and forces obtained as a gradient of uncertainty with respect to atomic positions. Further, in the fourth step, the computing system 104 can select a batch from the candidates of the third step and compute labels using the quantum mechanical solver. To perform this, the computing system 104 can use a batch selection algorithm such as, but not limited to, a batch selection algorithm that maximizes the distance between selected candidates in the space spanned by the gradient features, a batch selection algorithm that maximizes the determinant of the covariance matrix (e.g., Gaussian process regression) built using gradient features, and / or other batch selection algorithms.
[0054] As such, based on using UDAL, the computing system 104 determines (e.g., generates) one or more MLIP models 216 for the selected candidates (e.g., the MLIP models 216 can be candidate specific). The MLIP models 216 can be used to predict energies (e.g., determine prediction energies) of the different configurations of the candidate system, atomic forces, and / or stress tensors. For example, the atomistic simulations can force atomic forces to evolve the system in time. A stress tensor can be used to set up specific environmental conditions defined by the external pressure. Further, the MLIP models 216 can be machine learning models for determining interatomic potentials (e.g., the interaction between a pair of atoms and / or the interaction of an atom with a group of atoms in a condensed phase). In an example, the MLIP models 216 can utilize high-dimensional mathematical regression to interpolate between reference energies associated with the interatomic potentials. The computing system 104 provides the generated MLIP models 216 to the next step 206.
[0055] At a third step 206, the computing system 104 obtains the generated MLIP models 216 and uses a Hamiltonian Monte Carlo (HMC) method (e.g., a self-tuning HMC simulator) to generate one or more HMC outputs. For example, the HMC method is a method that allows the sampling from a given target distribution. For many phenomena in nature, it is desirable to find representative configurations of the atomistic system at a fixed temperature which implies to sample from the Boltzmann distribution. The input to this method is thus a description of the system, mainly given by MLIP and a target distribution. The output are approximations of the expected values of atom properties in the given distribution, representing the natural environment.
[0056] HMC uses a local transition kernel between states, given often by the velocity Verlet integrator. There are two major tunable parameters, which the self-tuning HMC optimizes for: 1) the time-step of the integrator and 2) the number of integration steps. Further, the functional form of the velocity Verlet integrator is not fixed (e.g., it is possible to extend it with more parameters, such as, for instance, plugging in neural networks). Special care is taken in defining the loss (regarding invariances of the system), as well as making sure that the loss acts as a goodproxy for the effectiveness of the method (e.g., that it reduces the correlation between subsequent states using only local information).
[0057] For example, the computing system 104 can construct the HMC such that the resulting samples of the Markov chain are picked from the target distribution. The computing system 104 can then estimate the thermodynamic expectation values by simply evaluating the measure prescription on the obtained samples and averaging over them. If the expected squared jump distance as loss is used, it is guaranteed to minimize the one-jump autocorrelation in real coordinates (e.g., based on using the expected squared jump distance as a loss, the computing system 104 can guarantee to minimize the one-jump autocorrelation in real coordinates). Other losses are not guaranteed to minimize some specific measure of the Markov chain, but effectively lead to lower autocorrelations in the measurements. In addition, the parameters are self-tuning in the sense that their value is optimized without prescribing explicit target values (e.g. for the acceptance rate). Thus, the simulations tunes “itself’ to be optimal.
[0058] For instance, at the third step 206, the computing system 104 can obtain information (e.g., from step 204) such as the trained MLIPs 216 (e.g., MLIPs 216 for a specific candidate) and / or other information. The computing system 104 can determine a description of the system (e.g., based on using the trained MLIPs 216) and / or a target distribution. Using the determinations, the computing system 104 can use the HMC to determine an HMC output indicating properties of the atomic system associated with the batch of candidates. For instance, the HMC output can indicate the expected values of atom properties in a given target distribution, which can represent the natural environment. For example, when using the HMC, the computing system 104 can use the local transition kernel between states (e.g., velocity Verlet integrator), the two tunable parameters that are optimized by the HMC (e.g., the time-step of the integrator and the number of integration steps), and / or other factors to determine the output. In some embodiments, the computing system 104 can use and / or optimize a plurality of tunable parameters (e.g., the two tunable parameters described above as well as other tunable parameters), which can also be integration step dependent. The computing system 104 can provide the output (e.g., the expected values of atom properties in the given distribution) to the next step 208. For instance, the computing system 104 can generate a dataset 218 based on the output. The dataset 218 can include the outputs from the HMC step 206, and can grow based on the number of simulations performed and the candidates selected. The computing system 104 can provide the dataset to the fourth step 208.
[0059] At a fourth step 208, embodiments of the present invention utilize probabilistic machine learning for efficient candidate selection. For example, the new set of data samples generated by the HMC (third step 206) is combined with the existing ones obtained throughmultiple iterations of the pipeline 200 according to one or more embodiments of the present invention. This data is labelled as each sample can be a tuple (e.g., specific configuration of the candidate, [optional target system], desired property value). With this data, a (probabilistic) graph machine learning model (see Bacciu, Davide, et al., “A gentle introduction to deep learning for graphs,” Neural Networks 129 (2020): 203-221, which is hereby incorporated by reference herein) is trained to predict the desired value from the input configuration together with a measure of uncertainty about the prediction itself (see Errica, Federico, Davide Bacciu, and Alessio Micheli, “Graph mixture density networks,” International Conference on Machine Learning. PMLR (2021), which is hereby incorporated by reference herein).
[0060] The trained machine learning model can be used to make very efficient predictions about the desired property values of the candidates in the pool P. As candidates are selected and new data is generated through the repetitive execution of steps 202-206, the machine learning model is refined and becomes more accurate and less uncertain.
[0061] For example, at the fourth step 208, the computing system 104 obtains the dataset 218 (e.g., the dataset 218 indicating the expected values of atom properties in the given distribution). The dataset 218 can include data from multiple iterations of performing method 200 and / or portions of method 200 (e.g., steps 202-206). For instance, after each iteration, the computing system 104 combines new data (e.g., new set of data samples generated by the HMC method from step 206) with existing data already within the dataset 218. The existing data can be obtained through previous iterations of performing the method 200 and / or the steps 202-206. The computing system 104 can label the data (e.g., each sample can be a tuple indicating a specific configuration of the candidate, a target system, and / or a desired property value). The computing system 104 can train a graph machine learning model based on the dataset 218. For instance, using probabilistic machine learning, the computing system 104 can train the graph machine learning model. In some instances, the graph machine learning model can be a deep graph network (DGN) 220.
[0062] The graph machine learning model can map a configuration of a candidate system into the value of interest, together with a measure of uncertainty over the predicted value. For example, machine learning can be applied to data represented as vectors (e.g., logistic regression), images (e.g., convolutional neural networks), but also graphs (e.g., neural networks for graphs). The latter computes a “representation” (e.g., a vector) for each of the entities or for the whole graph, by iteratively exchanging information between the entities of the graph according to their pairwise relations. Once a vector representation (for individual entities / nodes or the entire graph) is produced, a standard machine learning classifier can be applied to produce a value. In embodiments of the present invention, the computing system 104 performs one stepfurther: the values that are produced can be the parameters of a distribution (e.g., mean and variance of a Gaussian, but also mixtures of Gaussians), which can convey information about which value is more likely to be observed in the output given the input graph. The computing system 104 can interpret this output distribution as the uncertainty of the graph ML model with respect to the possible output values.
[0063] In some embodiments, the computing system 104 can output the trained machine learning model (e.g., the DGN 220). The computing system 104 can use the trained machine learning model / DGN 220 to make efficient predictions about the desired property values of the candidates in the candidate pool 212. Further, as new candidate are selected and new data is generated / provided to the dataset 218, the computing system 104 can refine the machine learning model to become more accurate and less uncertain.
[0064] In some instances, predictions (at inference time) that are associated with a low level of uncertainty are considered to screen the candidates. For instance, a candidate associated with a desirable predicted value and a low uncertainty of the machine learning model can be preferred by the computing system 104 to others. At the same time, given different candidates associated with low uncertainty, it is possible to enforce diversity of the selected ones using certain well- known characteristics, for instance different functional groups or number of benzene rings, etc. On the other hand, at training time, embodiments of the present invention can encourage the model to select candidates with higher uncertainty, making sure it explores a meaningful and diverse set of new systems.
[0065] For example, according to the above description of the graph machine learning model, the computing system 104 can be used to predict a Gaussian distribution as the output, where the mean of the Gaussian distribution is understood as the most likely value associated with the input system. While training the pipeline 200, if the variance parameter predicted by the graph machine learning system (e.g., the fifth step 210) is high, then the model is very uncertain about the prediction, which means it might be a new system (never seen before by the model). Consequently, it might be worth choosing this candidate to enforce diversity of the observed candidates (e.g., the computing system 104 can choose this candidate, or the candidate with a variance parameter greater than a certain threshold, to enforce the diversity of the observed candidates). When the fifth step 210 has reached a desirable level of accuracy, after iterating over many candidates in the pipeline 200, embodiments of the present invention can move to the inference phase. In the inference phase, embodiments of the present invention can skip UDAL and HMC and use the graph machine learning model as the surrogate predictor of the desired properties. At inference, embodiments of the present invention are interested in predictions with low uncertainty, therefore high variance if the previous example of the Gaussian is considered.If the model predicts a good value for the property of a candidate with low uncertainty, the candidate is immediately sent to the wet lab for additional screening.
[0066] In some variations, when sufficiently trained, the model can then be used to skip the second and third steps 204 and 206 of FIG. 2 entirely and thus save enormous amounts of compute time and computational resources in selecting the candidates to be sent to the wet lab. For example, after determining that the machine learning model (e.g., DGN 220) is sufficiently trained (e.g., based on user input and / or one or more training thresholds), the computing system 104 can determine to skip steps 204 and 206. As such, in such iterations, the computing system 104 can perform the first step 202 to obtain the candidate selection and move directly towards the four step 208 as well as the decision block 222, which will be described below.
[0067] To determine when the model is sufficiently trained, this can also be performed by evaluating the predictive performance of step 5 on the candidates sent to the wet lab, for which embodiments of the present invention have an accurate assessment of their properties, or by using part of the incrementally built dataset as a "validation" set. The validation set is not used for training but only as a proxy for the generalization performances of the graph ML model, that is, how well it will predict on previously unseen data.
[0068] At a fifth step 210, the computing system 104 can perform ML-assisted candidate filtering. For example, as mentioned above, initially, the computing system 104 can use user input and / or a randomized process for the first step 202 and selecting candidates from the candidates’ pool 212. Then, the computing system 104 can perform the steps described above, and determine an output from the graph ML model. The output from the graph ML model can indicate a variance parameter (e.g., a variance parameter that is predicted by the graph ML model). A high variance parameter (e.g., a variance parameter greater than a certain threshold), can indicate that the graph ML model is very uncertain about the prediction. Initially, when starting out, in some embodiments, the computing system 104 can select the candidate associated with this high variance parameter to enforce the diversity of the observed candidates. For instance, the computing system 104 can use one or more thresholds and based on the variance parameter exceeding the one or more first thresholds, the computing system 104 can select the candidate associated with the variance parameter for the next iteration of pipeline 200. Once the computing system 104 performs sufficient iterations (e.g., based on the variance parameter reaching a desirable level of accuracy such as by reaching one or more second thresholds and / or based on the computing system 104 performing a certain number of iterations), the computing system 104 can move to the inference phase.
[0069] For example, after performing the one or more iterations and based on expert / user input, performing a certain number of iterations of pipeline 200, and / or based on the varianceparameter(s) that are output by the graph ML model, the computing system 104 can move to the inference phase and determine to use the graph ML model (e.g., the graph ML model that was trained at the fourth step 208 such as the DGN 220) for the candidate selection. During the inference phase, the computing system 104 can desire (e.g., select) candidates associated with low uncertainty (e.g., a low variance parameter). As such, at a fifth step 210, the computing system 104 can input information (e.g., the remaining candidates from the candidates’ pool 212) into the trained graph ML model (e.g., the DGN 220) for ML-assisted candidate filtering. Based on this, the computing system 104 can obtain (e.g., select) one or more additional promising candidates. Then, in the next iteration, the computing system 104 can perform steps 204 and 206 to provide additional data into the dataset 218, which is then used to train the graph ML model. Additionally, and / or alternatively, the computing system 104 can determine to skip the steps 204 and 206, and proceed directly from step 202 to steps 208 / 210 (e.g., selecting new candidates from the candidates’ pool 212 using the graph ML model).
[0070] Further, after the third step 206, the computing system 104 can perform decision block 222. For example, as mentioned previously, the output from performing the HMC can be approximations of the expected values of atom properties in the given distribution, which can represent the natural environment. Based on the output, the computing system 104 can perform a decision block 222 that determines whether this output indicates a desirable property (e.g., whether the HMC output satisfies one or more properties of interest). For example, HMC can provide properties of the system. Whether these properties are desirable can be based on the task at hand. In some embodiments, the computing system 104 can obtain user input indicating whether these properties are desirable. In some embodiments, where the goal is to find as large differences as possible, the user input can indicate, for example, how high the free-energy change of the binding of a drug molecule to its target is. In some variations, the computing system 104 can also define threshold standings for “good enough” values. For example, the computing system 104 can obtain one or more threshold standings (e.g., thresholds), and use the threshold standings to perform the decision block 222.
[0071] If no, the computing system 104 can perform result 224, which discards this output.If yes, the computing system 104 can perform result 226, which provides the output to the wet lab. For instance, the computing system 104 can generate information 106 (e.g., candidate information) based on the output from performing the HMC. The candidate information can indicate a batch of candidates and the HMC output. The computing system 104 can provide the information 106 to a computing system, device, apparatus, and / or entity associated with a wet lab (e.g., a wet lab computing system). The wet lab computing system can include and / or beassociated with external partners and / or external databases, which can be used for additional validation if the candidate satisfies the properties of interest.
[0072] Embodiments of the present invention thus provide for general improvements to computers in machine learning systems to more accurately predict candidates while providing improvements in compute time and conserving computational resources. Moreover, embodiments of the present invention can be practically applied to use cases to effect further improvements in a number of technical fields including, but not limited to, medical (e.g., digital medicine, personalized healthcare, Al-assisted drug or vaccine development, etc.) and material design and development (e.g., new material development or material optimization).
[0073] In an embodiment, the present invention can be applied for efficient screening of T- cell receptor (TCR)-binding candidates for vaccine / drug development, which can then be used for medical treatments. A use case addresses that one of the fundamental problems in the successful development of cancer vaccines is understanding whether a TCR, which monitors the health status of cells, identifies (or binds to) specific peptides presented on the surface of a cell. This problem is known as TCR-recognition. The use of Al-assisted tools for automated TCR- recognition relies on the availability of species-dependent datasets, which can be a huge limiting step and technical challenge in the development of Al-based TCR-recognition solutions. The data source for this use case includes a pool of candidate TCRs (or parts of them) and a target peptide. No extra training data is required for the machine learning models. Application of the method according to an embodiment of the present invention will efficiently simulate the interaction between the (batch of) TCR candidates and the peptide using a self-tuning HMC (the third step 206 of FIG. 2), which relies on an accurate MLIP efficiently obtained through UDAL (the second step 204 of FIG. 2). Depending on the result of the simulation, binding TCR candidates are sent to the wet lab for additional validation. The data generated using the selftuning HMC (the third step 206 of FIG. 2) is combined with the existing one and used to re-train a machine learning model that can suggest the next candidates to simulate. The process can be repeated indefinitely, with the accuracy of the filtering system increasing overtime. As output, a stream of candidates is sent for validation to the wet lab. The candidates can be displayed or placed in a queue for testing or experimentation, and / or equipment of the wet lab could be operated in an automated manner to test the candidates.
[0074] In other words, the computing system 104 can select a first batch of TCR candidates, generate an HMC output (based on using the self-tuning HMC) indicating a simulation of an interaction between the first batch of TCR candidates and a peptide, provide to a wet lab computing system candidate information comprising information indicating binding of the first batch of TCR candidates to the peptide, and selecting a new batch of TCR candidates based onthe simulation of the interaction between the first batch of TCR candidates and the peptide. The new batch of candidates can be used for drug vaccine development and for performing medical treatments.
[0075] In another embodiment, the present invention can be applied for efficient discovery of materials with desirable properties (e.g., tensile strength). A use case provides for the discovery of new materials with desirable properties, which heretofore would require to go through a significant time and cost-intensive simulation of the associated atomic systems. Given a large pool of candidate materials, it would be unsustainable to run simulations for all of them, so it is advantageous be as efficient and accurate as possible when identifying suitable materials to avoid wasting time and resources, and to be more competitive in the market. Also, depending on the property of interest, public and private datasets to train Al models might be limited or completely lacking, therefore it is advantageous to provide an intelligent way to select the smallest subset of candidate materials that might be effective in the shortest amount of time. The data source for this use case includes a pool of candidate materials. No extra training data is required for the machine learning models. Application of the method according to an embodiment of the present invention efficiently simulates the materials using a self-tuning HMC (the third step 206 of FIG. 2), which relies on an accurate MLIP efficiently obtained through UDAL (the second step 204 of FIG. 2). Depending on the result of the simulation, candidate materials with good values of a desired property are sent to the wet lab for additional validation. The data generated in the self-tuning HMC (the third step 206 of FIG. 2) is combined with the existing one and used to re-train a machine learning model that can suggest the next candidates to simulate. The process is repeated indefinitely, with the accuracy of the filtering system increasing over time. As output, a stream of candidates is sent for validation to the wet lab. The candidates can be displayed or placed in a queue for testing or experimentation, and / or equipment of the wet lab could be operated in an automated manner to test the candidates.
[0076] In an embodiment, the present invention provides an Al-assisted method for predicting atomic system properties in the absence of training data, the method comprising the steps of:1) Extracting one or a batch of candidates from a given pool (the first step 202 of FIG. 2).2) Using UDAL to efficiently obtain a candidate specific MLIP model that can be used to predict energies of the different configurations of the candidate system (the second step 204 of FIG. 2).3) Incorporating the MLIP into a self-tuning HMC simulator that checks whether the candidate satisfies the properties of interest, generating a dataset of labelled data at the same time (the third step 206 of FIG. 2).4) (In some embodiments) Sending the candidate to external partners / database for additional validation if the candidate satisfies the properties of interest.5) Using the generated dataset (incrementally built as more candidates are simulated) to train a (probabilistic) graph machine learning model that maps a configuration of a candidate system into the value of interest (the fourth step 208 of FIG. 2), together with a measure of uncertainty over the predicted value.6) Employing the trained model of the previous step to propose a new set of candidates, so that the process can be iterated from the first step 202 (the fifth step 210 of FIG. 2)
[0077] Embodiments of the present invention provide for the following improvements and technical advantages over existing technology:1) By powering self-tuning HMC with UDAL’s MLIP, it is provided to more efficiently generate a dataset that is also more representative of the configurational space of the candidate of interest than standard MD and MC methods, (the second and third steps 204, 206 of FIG. 2).2) Utilizing the uncertainty quantification capability of the probabilistic graph machine learning model, trained on the (incrementally built from scratch) dataset, to pick better candidates from the pool, that is, choosing those candidates for which the model is more confident about, (the fourth step 208 of FIG. 2).3) The combination of UDAL with self-tuning HMC and graph machine learning implements a completely automated, self-improving pipeline for discovering atomic systems with desirable properties that extends the classical active learning closed loop and does not require initial training data. When sufficiently trained, the graph machine learning model can entirely replace UDAL and self-tuning HMC by accurately predicting the properties of interest of each candidate and significantly speed up the screening, (the second, third, and fourth steps 204, 206, and 208 of FIG. 2).4) By selecting a diverse batch of candidates based on Al-assisted predictions of candidate properties, we carry out multiple evaluations of promising candidates in parallel, therefore accelerating the discovery of materials or drugs, (the first step 202 of FIG. 2).
[0078] Existing technology for discovery of good candidates is typically limited to the classical active learning closed loop. While this approach can generate models for predicting properties such as the potential of an atomic system, they will never be as accurate as molecular dynamics simulations. In contrast, according to embodiments of the present invention, extending the loop to incorporate a fast and adaptive HMC method, combined with an active learning loop that efficiently generates an accurate MLIP model, has a double advantage: first, it can better explore the configuration space of a system and find if the atomic system has a desirable property; second, the generated data is very well representative of the true distribution andtherefore an excellent source of training samples for the probabilistic graph machine learning model. This model, consequently, is more accurate than other methods relying on data generated by active learning only, and is more effective at fdtering down good candidates from the pool.
[0079] Referring to FIG. 3, a processing system 300 can include one or more processors 302, memory 304, one or more input / output devices 306, one or more sensors 308, one or more user interfaces 310, and one or more actuators 312. Processing system 300 can be representative of each computing system disclosed herein.
[0080] Processors 302 can include one or more distinct processors, each having one or more cores. Each of the distinct processors can have the same or different structure. Processors 302 can include one or more central processing units (CPUs), one or more graphics processing units (GPUs), circuitry (e.g., application specific integrated circuits (ASICs)), digital signal processors (DSPs), and the like. Processors 302 can be mounted to a common substrate or to multiple different substrates.
[0081] Processors 302 are configured to perform a certain function, method, or operation (e.g., are configured to provide for performance of a function, method, or operation) at least when one of the one or more of the distinct processors is capable of performing operations embodying the function, method, or operation. Processors 302 can perform operations embodying the function, method, or operation by, for example, executing code (e.g., interpreting scripts) stored on memory 304 and / or trafficking data through one or more ASICs. Processors 302, and thus processing system 300, can be configured to perform, automatically, any and all functions, methods, and operations disclosed herein. Therefore, processing system 300 can be configured to implement any of (e.g., all of) the protocols, devices, mechanisms, systems, and methods described herein.
[0082] For example, when the present disclosure states that a method or device performs task “X” (or that task “X” is performed), such a statement should be understood to disclose that processing system 300 can be configured to perform task “X”. Processing system 300 is configured to perform a function, method, or operation at least when processors 302 are configured to do the same.
[0083] Memory 304 can include volatile memory, non-volatile memory, and any other medium capable of storing data. Each of the volatile memory, non-volatile memory, and any other type of memory can include multiple different memory devices, located at multiple distinct locations and each having a different structure. Memory 304 can include remotely hosted (e.g., cloud) storage.
[0084] Examples of memory 304 include a non-transitory computer-readable media such as RAM, ROM, flash memory, EEPROM, any kind of optical storage disk such as a DVD, a Blu-Ray® disc, magnetic storage, holographic storage, a HDD, a SSD, any medium that can be used to store program code in the form of instructions or data structures, and the like. Any and all of the methods, functions, and operations described herein can be fully embodied in the form of tangible and / or non-transitory machine-readable code (e.g., interpretable scripts) saved in memory 304.
[0085] Input-output devices 306 can include any component for trafficking data such as ports, antennas (i.e., transceivers), printed conductive paths, and the like. Input-output devices 306 can enable wired communication via USB®, DisplayPort®, HDMI®, Ethernet, and the like. Input-output devices 306 can enable electronic, optical, magnetic, and holographic, communication with suitable memory 306. Input-output devices 306 can enable wireless communication via WiFi®, Bluetooth®, cellular (e.g., LTE®, CDMA®, GSM®, WiMax®, NFC®), GPS, and the like. Input-output devices 306 can include wired and / or wireless communication pathways.
[0086] Sensors 308 can capture physical measurements of environment and report the same to processors 302. User interface 310 can include displays, physical buttons, speakers, microphones, keyboards, and the like. Actuators 312 can enable processors 302 to control mechanical forces.
[0087] Processing system 300 can be distributed. For example, some components of processing system 300 can reside in a remote hosted network service (e.g., a cloud computing environment) while other components of processing system 300 can reside in a local computing system. Processing system 300 can have a modular design where certain modules include a plurality of the features / functions shown in FIG. 3. For example, I / O modules can include volatile memory and one or more processors. As another example, individual processor modules can include read-only-memory and / or local caches.
[0088] While subject matter of the present disclosure has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive. Any statement made herein characterizing the invention is also to be considered illustrative or exemplary and not restrictive as the invention is defined by the claims. It will be understood that changes and modifications may be made, by those of ordinary skill in the art, within the scope of the following claims, which may include any combination of features from different embodiments described above.
[0089] The terms used in the claims should be construed to have the broadest reasonable interpretation consistent with the foregoing description. For example, the use of the article “a” or “the” in introducing an element should not be interpreted as being exclusive of a plurality of elements. Likewise, the recitation of “or” should be interpreted as being inclusive, such that therecitation of “A or B” is not exclusive of “A and B,” unless it is clear from the context or the foregoing description that only one of A and B is intended. Further, the recitation of “at least one of A, B and C” should be interpreted as one or more of a group of elements consisting of A, B and C, and should not be interpreted as requiring at least one of each of the listed elements A, B and C, regardless of whether A, B and C are related as categories or otherwise. Moreover, the recitation of “A, B and / or C” or “at least one of A, B or C” should be interpreted as including any singular entity from the listed elements, e.g., A, any subset from the listed elements, e.g., A and B, or the entire list of elements A, B and C.
Claims
CLAIMSWhat is claimed is:
1. A computer-implemented machine learning method for using one or more predicted properties of one or more atomic systems to train a graph machine learning model, comprising: obtaining a candidates pool comprising a plurality of candidates associated with a plurality of atomic systems; using an uncertainty-driven active learning (UDAL) to obtain a candidate specific machine learning interatomic potential (MLIP) model for a first batch of candidates from the plurality of candidates; incorporating the candidate specific MLIP model into a self-tuning Hamiltonian Monte Carlo (HMC) simulator to generate an HMC output indicating the one or more predicted properties of the one or more atomic systems associated with the first batch of candidates; and training the graph machine learning model based on a dataset comprising the HMC output.
2. The computer-implemented machine learning method of claim 1, further comprising selecting a new batch of candidates from the candidates pool based on using the trained graph machine learning model.
3. The computer-implemented machine learning method of claim 1 or claim 2, further comprising: extracting the first batch of candidates from the candidates pool, wherein using the UDAL to obtain the candidate specific MLIP model comprises training the candidate specific MLIP model for the first batch of candidates based on using the UDAL.
4. The computer-implemented machine learning method of claim 3, wherein extracting the first batch of candidates is based on using a randomized process or user input indicating domainspecific pre-screening information.
5. The computer-implemented machine learning method of claim 3, wherein extracting the first batch of candidates comprises: obtaining operator input indicating to enable the graph machine learning model; and using the graph machine learning model to select the first batch of candidates.
6. The computer-implemented machine learning method of any of the preceding claims, wherein training the candidate specific MLIP model for the first batch of candidates comprises: obtaining atomic system information and molecular dynamics (MD) associated with the first batch of candidates; andinputting the atomic system information and the molecular dynamics (MD) into the UDAL to train the candidate specific MLIP model for the first batch of candidates, wherein the candidate specific MLIP model is used to determine a prediction energy of different configurations of a candidate system associated with the first batch of candidates.
7. The computer-implemented machine learning method of any of the preceding claims, wherein the self-tuning HMC simulator comprises a first tunable parameter and a second tunable parameter, wherein the first tunable parameter is a time-step for a velocity Verlet integrator, and wherein the second tunable parameter is a number of integration steps.
8. The computer-implemented machine learning method of any of the preceding claims, further comprising: using the UDAL to obtain one or more additional candidate specific MLIPs models for one or more additional batches of candidates from the plurality of candidates; incorporating the one or more additional candidate specific MLIPs models into the selftuning HMC simulator to generate one or more additional HMC outputs; and generating the dataset, wherein the dataset comprises the HMC output and the one or more additional HMC outputs.
9. The computer-implemented machine learning method of claim 8, wherein generating the dataset comprises: labelling the HMC output and the one or more additional HMC outputs based on using a tuple indicating a specific configuration of a candidate and a property value, and wherein training the graph machine learning model is based on the labelling of the HMC output and the one or more additional HMC outputs.
10. The computer-implemented machine learning method of any of the preceding claims, wherein the graph machine learning model maps a configuration of a candidate system associated with the first batch of candidates into a value of interest, wherein the configuration is mapped together with a measure of uncertainty over a predicted value.
11. The computer-implemented machine learning method of any of the preceding claims, wherein training the graph machine learning model is further based on using probabilistic machine learning, and wherein the graph machine learning model is a deep graph network (DGN).
12. The computer-implemented machine learning method of any of the preceding claims, further comprising: using the UDAL to obtain a new candidate specific MLIP model for the new batch of candidates;incorporating the new candidate specific MLIP model into the self-tuning HMC simulator to generate new HMC output; and performing further training of the graph machine learning model based on using an updated dataset comprising the HMC output and the new HMC output.
13. The computer-implemented machine learning method of any of the preceding claims, further comprising: based on the HMC output satisfying one or more properties of interest, providing, to a wet lab computing system, candidate information indicating the first batch of candidates and the HMC output, wherein the candidates pool indicates a pool of T-cell receptor (TCR)-binding candidates or a pool of candidate materials, wherein the first batch of candidates is a first batch of TCR candidates, wherein the HMC output indicates a simulation of an interaction between the first batch of TCR candidates and a peptide, wherein the candidate information comprises information indicating binding of the first batch of TCR candidates to the peptide, wherein selecting the new batch of candidates comprises selecting a new batch of TCR candidates based on the simulation of the interaction between the first batch of TCR candidates and the peptide, and wherein the new batch of candidates is used for drug vaccine development or for performing medical treatments.
14. A computer system for using one or more predicted properties of one or more atomic systems to train a graph machine learning model, the system comprising one or more hardware processors, which, alone or in combination, are configured to provide for execution of the following steps: obtaining a candidates pool comprising a plurality of candidates associated with a plurality of atomic systems; using an uncertainty-driven active learning (UDAL) to obtain a candidate specific machine learning interatomic potential (MLIP) model for a first batch of candidates from the plurality of candidates; incorporating the candidate specific MLIP model into a self-tuning Hamiltonian Monte Carlo (HMC) simulator to generate an HMC output indicating the one or more predicted properties of the one or more atomic systems associated with the first batch of candidates; and training the graph machine learning model based on a dataset comprising the HMC output.
15. A tangible, non-transitory computer-readable medium having instructions thereon which, upon being executed by one or more processors, alone or in combination, provide for execution of a method for using one or more predicted properties of one or more atomic systems to train a graph machine learning model comprising the following steps:obtaining a candidates pool comprising a plurality of candidates associated with a plurality of atomic systems; using an uncertainty-driven active learning (UDAL) to obtain a candidate specific machine learning interatomic potential (MLIP) model for a first batch of candidates from the plurality of candidates; incorporating the candidate specific MLIP model into a self-tuning Hamiltonian Monte Carlo (HMC) simulator to generate an HMC output indicating the one or more predicted properties of the one or more atomic systems associated with the first batch of candidates; and training the graph machine learning model based on a dataset comprising the HMC output.