Machine learning-based device, method, and system for engineering mesoscale peptides

Through training and verification using machine learning models, the problem that computational design in the prior art is difficult to explore protein topological space, and efficiently generates polypeptide structures that meet the expectations.

CN114401734BActive Publication Date: 2025-06-03IBIO INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080050301.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-31
Filing Date
2020-05-13
Publication Date
2025-06-03
Estimated Expiration
2040-05-13

AI Technical Summary

Technical Problem

Existing computational design techniques are difficult to fully explore the huge topological space of protein structure, resulting in too much calculation when designing proteins that match the target structure, and it is impossible to effectively generate peptides that meet the expectations.

Method used

By using a machine learning model, a blueprint record set is generated based on a predetermined portion of the reference target structure, the machine learning model is trained to generate a second blueprint record set with the desired score, and the polypeptide structure is verified by computational protein modeling and molecular dynamics simulation.

Benefits of technology

It realizes the exploration of a wider protein topological space in a short time, generates a polypeptide structure that meets the expectations, and improves the efficiency and effectiveness of the computational design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114401734B_ABST
    Figure CN114401734B_ABST
Patent Text Reader

Abstract

The present disclosure provides methods for designing engineered polypeptides that recapitulate molecular structural features of a predetermined portion of a reference protein structure, such as an antibody epitope or a protein binding site. A machine learning (ML) model is trained with blueprint records generated from a reference target structure labeled with scores calculated by computational protein modeling based on the polypeptide structure generated from the blueprint record. The method can include training an ML model based on a first set of blueprint records or a representation thereof and a first set of scores, where each blueprint record from the first set of blueprint records is associated with each score from the first set of scores. After the training, the machine learning model can be executed to generate a second set of blueprint records. Then, a set of engineered polypeptides is generated based on the second set of blueprint records.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit and priority of U.S. Patent Application No. 62 / 855,767, titled "Meso - Scale Engineered Peptides and Methods of Selecting", filed on May 31, 2019, which is incorporated herein by reference in its entirety. Technical Field

[0003] The present disclosure generally relates to the field of artificial intelligence / machine learning, and more particularly to methods and devices for training and using machine - learning models for engineering peptides. Background Art

[0004] Computational design can be used to design novel therapeutic proteins that mimic natural proteins, or to design vaccines that display one or more desired epitopes from pathogenic antigens. Computationally designed proteins can also be used to generate or select binders. For example, an antibody library (such as a phage - display library) can be panned against a designed protein bait to select clones that bind to the bait, or experimental animals can be immunized with a designed immunogen to generate novel antibodies.

[0005] Although there are other platforms, the leading computational design modeling platform is Rosetta (Das and Baker, 2008). This platform can be used to design proteins that match a desired structure. Correia et al., Structure 18:1116 - 26 (2010) disclose a general computational method for designing epitope scaffolds, in which contiguous structural epitopes are grafted into a scaffold protein to achieve conformational stability and immune presentation. Olek et al., PNAS USA 107:17880 - 87 (2010) disclose the grafting of epitopes from the HIV - 1 gp41 protein onto selected receptor scaffolds.

[0006] Conventional computational design techniques typically rely on the grafting of a portion of the target protein structure (e.g., an epitope) onto a pre - existing scaffold. Modeling platforms such as Rosetta are computationally too expensive to fully explore large topological spaces, such as the vast protein topological space that recapitulates a given protein structure. Thus, there is a need for novel and improved devices and methods for the computational design of proteins that mimic a target protein structure. Summary of the Invention

[0007] Typically, in some variations, the device may include a non-transitory processor-readable medium storing code representing instructions to be executed by the processor. The code may include code to cause the processor to train a machine learning model based on a first set of blueprint records or a representation thereof and a first set of scores, where each blueprint record from the first set of blueprint records is associated with each score from the first set of scores. The medium may include code to execute the machine learning model after the training to generate a second set of blueprint records having at least one desired score. The second set of blueprint records may be configured to be received as an input in computational protein modeling to generate an engineered polypeptide based on the second set of blueprint records.

[0008] The medium may include code to cause the processor to receive a reference target structure. The medium may include code to cause the processor to generate the first set of blueprint records from a predetermined portion of the reference target structure, where each blueprint record from the first set of blueprint records includes a target residue position and a scaffold residue position, and each target residue position from the plurality of target residue positions corresponds to one target residue from a plurality of target residues. In some variations, in at least one blueprint record, the target residue positions are discontinuous. In some variations, in at least one blueprint record, the order of the target residue positions is different from the order of the target residue positions in the reference target sequence.

[0009] The medium may include code to cause the processor to label the first set of blueprint records, the labeling being performed by performing computational protein modeling on each blueprint record to generate a polypeptide structure, calculating a score of the polypeptide structure, and associating the score with the blueprint record. In some variations, the computational protein modeling may be based on de novo design without a template matching the reference target structure. In some variations, each score includes an energy term and a structure constraint matching term, and the structure constraint matching term may be determined using one or more structure constraints extracted from a representation of the reference target structure.

[0010] The medium may include code to cause the processor to determine whether the machine learning model needs to be retrained by calculating a second set of scores for the second set of blueprint records. The medium may include additional code to retrain the machine learning model in response to the determination, based on (1) retraining the blueprint records including the second set of blueprint records and (2) retraining the scores including the second set of scores.

[0011] The medium may include code that causes the processor to connect the first blueprint record set and the second blueprint record set after retraining of the machine learning model to generate a retrained blueprint record and to generate a retraining score, with each blueprint record from the retrained blueprint record associated with a score from the retraining score. In some variations, at least one desired score may be a preset value. In some variations, the at least one desired score may be determined dynamically.

[0012] In some variations, the machine learning model may be a supervised machine learning model. The supervised machine learning model may include an ensemble of decision trees, a boosted decision tree algorithm, an extreme gradient boosting (XGBoost) model, or a random forest. In some variations, the supervised machine learning model may include a support vector machine (SVM), a feedforward machine learning model, a recurrent neural network (RNN), a convolutional neural network (CNN), a graph neural network (GNN), or a transformer neural network.

[0013] In some variations, the machine learning model may include an inductive machine learning model. In some variations, the machine learning model may include a generative machine learning model.

[0014] The medium may include code that causes the processor to perform computational protein modeling on the second blueprint record set to generate an engineered polypeptide.

[0015] The medium may include code that causes the processor to filter the engineered polypeptide, the filtering being performed by a static structural comparison with a representation of the reference target structure.

[0016] The medium may include code that causes the processor to filter the engineered polypeptide, the filtering being performed by a dynamic structural comparison of the representation of the reference target structure with the representation of the reference target structure and a molecular dynamics (MD) simulation of each of the engineered polypeptides. In certain variations, the MD simulation is performed in parallel using symmetric multiprocessing (SMP). BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a schematic diagram of an exemplary engineered polypeptide design device.

[0018] Figure 2 is a schematic diagram of an exemplary machine learning model for engineered polypeptide design.

[0019] Figure 3 is a schematic diagram of an exemplary method for engineered polypeptide design.

[0020] Figure 4Schematic diagram of an exemplary method for engineered polypeptide design.

[0021] Figure 5 Schematic diagram of an exemplary method for preparing data for an engineered polypeptide design device.

[0022] Figure 6 Schematic diagram of an exemplary method for engineered polypeptide design.

[0023] Figure 7 Schematic diagram of an exemplary performance of a machine learning model for engineered polypeptide design.

[0024] Figure 8 Schematic diagram of an exemplary method for engineered polypeptide design using a machine learning model.

[0025] Figure 9 Schematic diagram of an exemplary performance of a machine learning model for engineered polypeptide design.

[0026] Figure 10A -D shows an exemplary method of performing molecular dynamics simulations to validate engineered polypeptides.

[0027] Figure 11 Shows an exemplary method of performing molecular dynamics simulations to validate engineered polypeptides.

[0028] Figure 12 Schematic diagram of an exemplary method for parallelizing molecular dynamics simulations.

[0029] Figure 13 Schematic diagram of an exemplary method for validating a machine learning model for engineered polypeptide design. Detailed Description

[0030] Non-limiting examples of various aspects and variations of the present invention are described herein and shown in the drawings.

[0031] The present disclosure provides methods for designing engineered polypeptides, as well as compositions comprising the engineered peptides and methods of using the engineered peptides. For example, methods for using engineered peptides in in vitro antibody selection are provided herein. In some aspects, a user (or program) can select a target protein with a known structure and identify a portion of the target protein as an input for designing an engineered polypeptide. The target protein can be an antigen (or putative antigen) from a pathogenic organism; a protein related to a disease-associated cellular function; an enzyme; a signaling molecule; or any protein for which an engineered polypeptide is desired to recapitulate a portion of the protein. Engineered polypeptides can be used for antibody discovery, vaccination, diagnostics, use in therapeutic methods, biomanufacturing, or other applications. In one variant, the "target protein" can be more than one protein, such as a multimeric protein complex. For the sake of brevity, the present disclosure refers to a target protein, but the methods are also applicable to multimeric structures. In one variant, the target protein is two or more different proteins or protein complexes. For example, the methods disclosed herein can be used to design engineered peptides that mimic common properties of proteins from different species - for example, engineered peptides that target conserved epitopes for antibody selection.

[0032] A computational record of the protein topology is derived, herein referred to as the "reference target structure". The reference target structure can be a conventional protein structure or a structural model, for example represented by the 3D coordinates of all (or most) atoms in the protein or the 3D coordinates of selected atoms (e.g., the Cβ atoms of each protein residue). Optionally, the reference target structure can include dynamic terms derived computationally (e.g., from molecular dynamics simulations) or experimentally (e.g., from spectroscopy, crystallography, or electron microscopy).

[0033] A predetermined portion of the target protein is converted into a blueprint with target residue positions and scaffold residue positions. Each position can be assigned a fixed amino acid residue identity or a variable identity (e.g., any amino acid, or an amino acid with desired physicochemical properties - polar / nonpolar, hydrophobicity, size, etc.). In one variant, each amino acid from the predetermined portion of the target protein is mapped to a target residue position that is assigned to have the same amino acid identity as that present in the target protein. The target residue positions can be contiguous and / or sequential. However, in some variants, an advantage is that the target residue positions can be non-contiguous (interrupted by scaffold residue positions) and non-sequential (different from the order of the target protein). In some variants, unlike transplantation methods, the order of the residues is not restricted. Similarly, the disclosed methods can accommodate discontinuous portions of the target protein (e.g., discontinuous epitopes where different portions of the same protein or even different protein chains contribute to an epitope).

[0034] The positions of the scaffold residues of a blueprint can be specified to have any amino acid at that position (i.e., X represents any amino acid). In a variant, the positions of the scaffold residues are specified by selecting from a subset of possible natural or non-natural amino acids (e.g., small polar amino acid residues, large hydrophobic amino acid residues, etc.). The blueprint can also accommodate optional target and / or scaffold residue positions. In other words, the blueprint can tolerate insertions or deletions of residue positions. For example, a target or scaffold residue position can be specified as present or absent; or the position can be specified as 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more residues.

[0035] Then a subset of blueprints can be used to perform computational modeling to generate corresponding polypeptide structures, and the computational modeling is carried out using, for example, energy terms and topological constraints derived from a reference target structure and scores calculated for each polypeptide structure. A machine learning (ML) model can be trained using the scores and the blueprint or a representation of the blueprint (e.g., a vector representing the blueprint), and the ML model can be executed to generate additional blueprints. One advantage of this approach is that the ML model can explore a larger topological space covered by blueprints compared to what is explored by iterative computational modeling of many blueprints.

[0036] The present disclosure also provides methods and related devices for converting an output blueprint into the sequence and / or structure of an engineered polypeptide, comparing these engineered polypeptides with a target protein - using static comparison, dynamic comparison, or both - and using these comparisons to filter polypeptides.

[0037] Although the methods and devices are described herein as processing data from a set of blueprint records, a set of scores, a set of energy terms, a set of molecular dynamics energies, a set of energy terms, or a set of energy functions, in some cases, such as Figure 1The engineered polypeptide design device 101 shown and described can be used to generate the blueprint record set, the score set, the energy term set, the molecular dynamics energy set, the energy term set, or the energy function set. Thus, the engineered polypeptide design device 101 can be used to generate or process any collection or stream of data, events, and / or objects. For example, the engineered polypeptide design device 101 can process and / or generate any one or more strings, one or more numbers, one or more names, one or more images, one or more videos, one or more executable files, one or more data sets, one or more spreadsheets, one or more data files, one or more blueprint files, and so on. For additional examples, the engineered polypeptide design device 101 can process and / or generate any one or more software codes, one or more web pages, one or more data files, one or more model files, one or more source files, one or more scripts, and so on. As another example, the engineered polypeptide design device 101 can process and / or generate one or more data streams, one or more image data streams, one or more text data streams, one or more numerical data streams, one or more computer-aided design (CAD) file streams, and so on.

[0038] Figure 1 is a schematic diagram of an exemplary engineered polypeptide design device 101. The engineered polypeptide design device can be used to generate an engineered polypeptide design set. The engineered polypeptide design device 101 includes a memory 102, a communication interface 103, and a processor 104. The engineered polypeptide design device 101 can optionally be connected (without intermediate components) or coupled (with or without intermediate components) to a backend service platform 160 via a network 150. The engineered polypeptide design device 101 can be a hardware-based computing device, such as a desktop computer, a server computer, a mainframe computer, a quantum computing device, a parallel computing device, a desktop computer, a laptop computer, a collection of smartphone devices, and so on.

[0039] The memory 102 of the engineered polypeptide design device 101 may include, for example, a memory buffer, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), embedded multi-time programmable (MTP) memory, embedded multimedia card (eMMC), universal flash storage (UFS) device, and so on. The memory 102 may store, for example, one or more software modules and / or code, the software modules and / or code including causing the processor 104 of the engineered polypeptide design device 101 to execute one or more processes or functions (e.g., the data preparation module 105, the computational protein modeling module 106, the machine learning model 107, and / or the molecular dynamics simulation module 108). The memory 102 may store a set of files related to the machine learning model 107 (e.g., generated by execution), the files including data generated by the machine learning model 107 during the operation of the engineered polypeptide design device 101. In some cases, the set of files related to the machine learning model 107 may include temporary variables, return memory addresses, variables, a graph of the machine learning model 107 (e.g., a set of arithmetic operations used by the machine learning model 107 or a representation of the set of arithmetic operations), metadata of the graph, assets (e.g., external files), electronic signatures (e.g., specifying the type of the machine learning model 107 being exported and input / output tensors), and so on, generated during the operation of the engineered polypeptide design device 101.

[0040] The communication interface 103 of the engineered polypeptide design device 101 may be a hardware component of the engineered polypeptide design device 101, which is operably coupled to and used by the processor 104 and / or the memory 102. The communication interface 103 may include, for example, a network interface card (NIC), Wi-Fi TM module, module, an optical communication module, and / or any other suitable wired and / or wireless communication interface. The communication interface 103 may be configured to connect the engineered polypeptide design device 101 to the network 150, as further described in detail herein. In some cases, the communication interface 103 may facilitate receiving or sending data via the network 150. More specifically, in some embodiments, the communication interface 103 may facilitate receiving or sending data, such as receiving a set of blueprint records, a set of scores, a set of energy terms, a set of molecular dynamics energies, a set of energy terms, or a set of energy functions from the backend service platform 160 via the network 150 or sending them to the backend service platform. In some cases, the data received via the communication interface 103 may be processed by the processor 104 or stored in the memory 102, as further described in detail herein.

[0041] Processor 104 may include, for example, a hardware-based integrated circuit (IC) or any other suitable processing device configured to run and / or execute an instruction or set of codes. For example, processor 104 may be a general-purpose processor, a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), an accelerated processing unit (APU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic array (PLA), a complex programmable logic device (CPLD), a programmable logic controller (PLC), and so on. Processor 104 is operatively coupled to memory 102 via a system bus (e.g., an address bus, a data bus, and / or a control bus).

[0042] Processor 104 may include a data preparation module 105, a computational protein modeling module 106, and a machine learning model 107. Processor 104 may optionally include a molecular dynamics simulation module 108. Each of the data preparation module 105, the computational protein modeling module 106, the machine learning model 107, or the molecular dynamics simulation module 108 may be software stored in memory 102 and executed by processor 104. For example, the code that causes the machine learning model 107 to generate a set of blueprint records may be stored in memory 102 and executed by processor 104. Similarly, each of the data preparation module 105, the computational protein modeling module 106, the machine learning model 107, or the molecular dynamics simulation module 108 may be a hardware-based device. For example, the process that causes the machine learning model 107 to generate the set of blueprint records may be implemented on a separate integrated circuit (IC) chip.

[0043] The data preparation module 105 may be configured to receive (e.g., from memory 102 or the backend service platform 160) a data set, including receiving a reference target structure for a reference target. The data preparation module 105 may additionally be configured to generate a set of blueprint records (e.g., a blueprint file encoded in an alphanumeric data table) from a predetermined portion of the reference target structure. In some cases, each blueprint record from the set of blueprint records may include a target residue position and a scaffold residue position, each target residue position corresponding to one of a plurality of target residues.

[0044] In some cases, the data preparation module 105 may additionally be configured to encode a blueprint of a reference target structure into a blueprint record. The data preparation module 105 may additionally convert the blueprint record into a representation of the blueprint record that is generally applicable to a machine learning model. In some cases, the representation may be a one-dimensional digital vector, a two-dimensional alphanumeric data matrix, or a three-dimensional normalized digital tensor. More specifically, in some cases, the representation is a vector of an ordered list of the number of inserted scaffold residue positions. This representation can be used because the order of the target residues can be inferred from the target structure, and thus the representation does not require identification of the amino acid identity of the target residue positions. An example of such a representation is Figure 6 further described.

[0045] In some cases, the data preparation module 105 may generate and / or process a set of blueprint records, a set of scores, a set of energy terms, a set of molecular dynamics energies, a set of energy terms, and / or a set of energy functions. The data preparation module 105 may be configured to extract information from the set of blueprint records, the set of scores, the set of energy terms, the set of molecular dynamics energies, the set of energy terms, or the set of energy functions.

[0046] In some cases, the data preparation module 105 may convert the encoding of the set of blueprint records to have a universal character encoding, such as ASCII, UTF-8, UTF-16, GB, Big5, Unicode, or any other suitable character encoding. In some other cases, the data preparation module 105 may additionally be configured to extract features of the blueprint record and / or the representation of the blueprint record by, for example, identifying a portion of the blueprint record or the representation of the blueprint record that is significant for the engineered polypeptide. In some cases, the data preparation module 105 may convert the units of the set of blueprint records, the set of scores, the set of energy terms, the set of molecular dynamics energies, the set of energy terms, or the set of energy functions from imperial units (such as miles, feet, inches, etc.) to the International System of Units (SI) (such as kilometers, meters, centimeters, etc.).

[0047] The computational protein modeling module 106 can be configured to generate an initial candidate set of blueprint records from a predetermined portion of a reference target structure, and the candidates can be used as starting templates for the computational optimization process described herein. In one instance, the computational protein modeling module 106 can be a Rosetta remodeler. Variations of the method employ other modeling algorithms, including but not limited to molecular dynamics simulations, de novo fragment assembly, Monte Carlo fragment assembly, machine learning structure prediction (such as AlphaFold or trRosetta), protein folding based on a structure knowledge base, neural network protein folding, sequence-based recurrent or transformer network protein folding, generative adversarial network protein structure generation, Markov chain Monte Carlo protein folding, and so on. The initial candidate structures generated using the Rosetta remodeler can be used as a training set for the machine learning model 107. The computational protein modeling module 106 can additionally computationally determine the energy terms for each blueprint from the initial candidates of the blueprint records. Then the data preparation module 105 can be configured to generate scores from the energy terms. In one instance, the scores can be normalized values of the energy terms. The normalized value can be a number from 0 to 1, a number from -1 to -1, a normalized value between 0 and 100, or any other numerical range. In some variations, the computational protein modeling module 106 can be based on de novo design in the absence of a template matching the reference target structure or based on weak distance constraints, where for example the distance between target residues in the target structure is restricted within a target residue distance of 1 angstrom. The weak distance constraints can include constraints that allow a variational noise distribution around the distance constraint (e.g., Gaussian noise with a specific mean and specific variance around the distance constraint). In some variations, the computational protein modeling module 106 can be used by smoothing or adding variational noise to any distance constraints and / or defining the objective function of the computational protein model such that the computational protein model is penalized less severely when long-distance constraints are not satisfied. Additionally, in some cases, the computational protein modeling module 106 can use smoothed labels of the energy terms. The advantage of this method is that by smoothing the energy term labels, the machine learning model 107 can more easily optimize the topological space covered by the blueprints to be explored.

[0048] Compared with the initial candidate set of the blueprint records, the machine learning model 107 can be used to generate an improved blueprint record. The machine learning model 107 can be a supervised machine learning model configured to receive the initial candidate set of the blueprint records and a set of scores calculated by the computational protein modeling module 106. Each score from the set of scores corresponds to a blueprint record from the initial candidate set of the blueprint records. The processor 104 can be configured to associate each corresponding score and blueprint record to generate a labeled training data set.

[0049] In some cases, the machine learning model 107 can include an inductive machine learning model and / or a generative machine learning model. The machine learning model can include boosting decision tree algorithms, decision tree ensembles, extreme gradient boosting (XGBoost) models, random forests, support vector machines (SVMs), feedforward machine learning models, recurrent neural networks (RNNs), convolutional neural networks (CNNs), graph neural networks (GNNs), adversarial network models, instance-based training models, transformer neural networks, and so on. The machine learning model 107 can be configured to include a set of model parameters, including a set of weights, a set of biases, and / or a set of activation functions, which, once trained, can be executed in an inductive mode to generate scores from blueprint records or can be executed in a generative mode to generate blueprint records from scores.

[0050] In one instance, the machine learning model 107 can be a deep learning model that includes an input layer, an output layer, and multiple hidden layers (e.g., 5 layers, 10 layers, 20 layers, 50 layers, 100 layers, 200 layers, etc.). The multiple hidden layers can include normalization layers, fully connected layers, activation layers, convolutional layers, recurrent layers, and / or any other layers suitable for representing the correlation between the set of blueprint records and the set of scores (each score representing an energy term).

[0051] In one instance, the machine learning model 107 can be an XGBoost model that includes a set of hyperparameters, such as the number of boosting rounds or the number of boosting rounds for the trees in the XGBoost model, and the maximum depth that defines the maximum allowed number of nodes from the root to the leaves of the trees in the XGBoost model. The XGBoost model can include a set of trees, a set of nodes, a set of weights, a set of biases, and other parameters that can be used to describe the XGBoost model.

[0052] In some embodiments, the machine learning model 107 (e.g., a deep learning model, an XGBoost model, etc.) can be configured to iteratively receive each blueprint record from the set of blueprint records and generate an output. Each blueprint record from the set of blueprint records is associated with a score from the set of scores. A target function (also referred to as a "cost function") can be used to compare the output and the score to generate a first training loss value. The target function can include, for example, mean squared error, mean absolute error, mean absolute percentage error, logcosh, categorical cross-entropy, and so on. The set of model parameters can be modified in multiple iterations, and the first target function can be executed in each iteration until the first training loss value converges to a first predetermined training threshold (e.g., 80%, 85%, 90%, 97%, etc.).

[0053] In some embodiments, the machine learning model 107 can be configured to iteratively receive each score from the set of scores and generate an output. Each blueprint record from the set of blueprint records is associated with one score from the set of scores. A target function can be used to compare the output and the blueprint record to generate a second training loss value. The set of model parameters can be modified over multiple iterations, and the first target function can be executed in each iteration of the multiple iterations until the second training loss value converges to a second predetermined training threshold.

[0054] Once trained, the machine learning model 107 can be executed to generate an improved set of blueprint records. It can be expected that the improved set of blueprint records has a higher score than the initial candidate set of blueprint records. In some cases, the machine learning model 107 can be a generative machine learning model that is trained (e.g., each score's energy term corresponds to the Rosetta energy of a blueprint record from the set of blueprint records) for a first set of blueprint records corresponding to a first set of scores (e.g., generated using a Rosetta reconstructor) to represent the correlation (e.g., corresponding to the energy term) between the design space of the first set of blueprint records and the first set of scores. Once trained, the machine learning model 107 can generate a second set of blueprint records with a second set of scores associated therewith. In some embodiments, the computational protein modeling module 106 can be used to verify the second set of blueprint records and the second set of scores by calculating a set of energy terms for the second set of blueprint records. The set of energy terms can be used to generate a set of ground truth scores for the second set of blueprint records. A subset of the blueprint records can be selected from the second set of blueprint records such that each blueprint record from the subset of the blueprint records has a ground truth score greater than a threshold. In some cases, the threshold can be a number predetermined by a user of, for example, the engineered polypeptide design device 101. In some other cases, the threshold can be a number dynamically determined based on the set of ground truth scores.

[0055] After the machine learning model 107 is executed to generate a second blueprint record set, the molecular dynamics simulation module 108 can optionally be used to validate the output of the machine learning model 107. The engineered polypeptide design device 101 can filter out a subset of the second blueprint records by: generating engineered polypeptides based on the second blueprint record set, and performing a dynamic structural comparison with the representation of the reference target structure using molecular dynamics (MD) simulations of the representations of each of the reference target structure and the engineered polypeptide structure. For example, the molecular dynamics simulation module 108 can select several (e.g., fewer than 10 hits) engineered polypeptides (based on the second blueprint record set). In some cases, the MD simulation can be performed under boundary conditions, constraints, and / or equilibration. In some cases, the MD simulation can be performed under solution conditions, including the steps of: model preparation, equilibration (e.g., a temperature of 100K to 300K), applying force field parameters and / or solvent model parameters to the representations of each of the reference target structure and the engineered polypeptide structure. In some cases, the MD simulation can perform constrained minimization (e.g., relieve structural clashes), constrained heating (e.g., constrained heating for 100 picoseconds and gradually warming to ambient temperature), relaxation of constraints (e.g., relaxation of constraints for 100 picoseconds and gradual removal of backbone constraints), and so on.

[0056] In some embodiments, the machine learning model 107 is an inductive machine learning model. Once trained, such a machine learning model 107 can, by numerical methods such as calculating scores of blueprints (e.g., calculating scores for a protein modeling module, a density functional theory-based molecular dynamics energy simulator, etc.), predict scores based on blueprint records in a fraction of the time normally taken. Thus, the machine learning model 107 can be used to quickly estimate a set of scores for a set of blueprint records, significantly improving the optimization speed of an optimization algorithm (e.g., 50% faster, 2 times faster, 10 times faster, 100 times faster, 1000 times faster, 1,000,000 times faster, 1,000,000,000 times faster, etc.). In some embodiments, the machine learning model 107 can generate a first set of scores for the first blueprint record set. The processor 104 of the engineered polypeptide design device 101 can execute code representing an instruction set to select the best-performing ones in the first blueprint record set (e.g., the top 10% of the first set of scores, e.g., the top 2% of the first set of scores, etc.). The processor 104 can additionally include code to validate the scores of the best-performing ones in the first blueprint record set. In some variants, if the validation score corresponding to the best-performing one in the first blueprint record set has a value greater than any of those in the first set of scores, it can be generated as an output. In some variants, the machine learning model 107 can be retrained based on a new data set that includes a second blueprint record set and a second set of scores, which include blueprint records and the scores of the best-performing ones.

[0057] Network 150 can be a digital telecommunications network of servers and / or computing devices. The servers and / or computing devices on the network can be connected via one or more wired or wireless communication networks (not shown) to share resources (such as data storage or computing power). The wired or wireless communication network between the servers and / or computing devices of the network can include one or more communication channels, such as one or more radio frequency (RF) communication channels, one or more fiber optic communication channels, and so on. The network can be, for example, the Internet, an intranet, a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a Worldwide Interoperability for Microwave Access network a virtual network, any other suitable communication system, and / or a combination of such networks.

[0058] The backend service platform 160 can be a digital communication network (such as the Internet) operably coupled to servers and / or computing devices and / or a computing device (such as a server) within the digital communication network. In some variations, the backend service platform 160 can include and / or execute cloud-based services, such as software as a service (SaaS), platform as a service (PaaS), infrastructure as a service (IaaS), and so on. In one instance, the backend service platform 160 can provide data storage to store large amounts of data, including protein structures, blueprint records, Rosetta energies, molecular dynamics energies, and so on. In another instance, the backend service platform 160 can provide fast computing to perform a set of computational protein modeling, a set of molecular dynamics simulations, a set of training machine learning models, and so on.

[0059] In some variations, the processes of the computational protein module 106 described herein can be executed in the backend service platform 160 that provides cloud computing services. In such variations, the engineered polypeptide design device 101 can be configured to send signals to the backend service platform 160 using the communication interface 103 to generate a set of blueprint records. The backend service platform 160 can execute the computational protein modeling process for generating the set of blueprint records. Then the backend service platform 160 can send the set of blueprint records to the engineered polypeptide design device 101 via the network 150.

[0060] In some variations, the engineered polypeptide design device 101 can send a file including the machine learning model 107 to a user computing device (not shown) remote from the engineered polypeptide design device 101. The user computing device can be configured to generate a set of blueprint records that meet design criteria (e.g., have a desired score). In some variations, the user computing device receives a reference target structure from the engineered polypeptide design device 101. The user computing device can generate a first set of blueprint records from a predetermined portion of the reference target structure such that each blueprint record includes a target residue position and a scaffold residue position. Each target residue position corresponds to one of a plurality of target residues. The user computing device can additionally train the machine learning model based on the first set of blueprint records or a representation thereof and a first set of scores. After training, the user computing device can execute the machine learning model to generate a second set of blueprint records having at least one desired score (e.g., meeting a specific design criteria). The second set of blueprint records can be received as an input in computational protein modeling to generate an engineered peptide based on the second set of blueprint records.

[0061] Figure 2 is a schematic diagram of an exemplary machine learning model 202 for engineered polypeptide design (similar to the machine learning model 107 as described and shown in Figure 1 ). The machine learning model 202 can be a supervised machine learning model that associates a design space of blueprint records with scores of energy terms corresponding to polypeptides built based on those blueprint records. The machine learning model can have a generative operating mode and / or an inductive operating mode.

[0062] In the generative operating mode, the machine learning model 202 is trained against a first set of blueprint records 201 and a first set of scores 203. Once trained, the machine learning model 202 can generate a second set of blueprint records having a second set of scores that are statistically higher (e.g., have a higher mean) than the first set of scores. In the inductive operating mode, the machine learning model 202 is trained against a first set of blueprint records 201 and a first set of scores 203. Once trained, the machine learning model 202 can generate a second set of scores for the second set of blueprint records. The second set of scores is a predicted set of scores based on historical training data (e.g., the first set of blueprint records and the first set of scores), and is generated faster than using computational protein modeling (similar to the computational protein modeling module 106 as shown and described in Figure 1 ). or molecular dynamics simulation (similar to as shown and described in Figure 1The numerical calculation fractions and / or energy terms of the molecular dynamics module 108) shown and described are significantly faster (e.g., 50% faster, 2 times faster, 10 times faster, 100 times faster, 1000 times faster, 1,000,000 times faster, 1,000,000,000 times faster, etc.).

[0063] Figure 3 is a schematic diagram of an exemplary method 300 for engineered polypeptide design. The method 300 for engineered polypeptide design can be performed, for example, by an engineered polypeptide design device (similar to the engineered polypeptide design device 101) shown and described as Figure 1 The method 300 for engineered polypeptide design optionally includes, in step 301, receiving a reference target structure of a reference target. The method 300 for engineered polypeptide design optionally includes, in step 302, generating a first set of blueprint records from a predetermined portion of the reference target structure, each blueprint record from the first set of blueprint records including a target residue position and a scaffold residue position, each target residue position corresponding to one target residue from a plurality of target residues. In some cases, the target residues are discontinuous. In some cases, the target residues are out of order. The method 300 for engineered polypeptide design can include, in step 303, training a machine learning model (similar to the machine learning model 107) shown and described as Figure 1 based on the first set of blueprint records or a representation thereof and a first set of scores, each blueprint record from the first set of blueprint records being associated with each score from the first set of scores. The representation can be generated using a data preparation module (similar to the data preparation module) based on the first set of blueprint records shown and described as Figure 1 The method 300 for engineered polypeptide design further includes, in step 304, executing the machine learning model after training to generate a second set of blueprint records having at least one desired score (e.g., one score or multiple scores). In some configurations, the machine learning model includes a generative machine learning model, and the at least one desired score is a preset value determined by a user of the engineered polypeptide design device. In some configurations, the machine learning model includes an inductive machine learning model that predicts a set of predicted scores for the second set of blueprint records. A subset of the second set of blueprint records can be selected such that each blueprint record from the subset of blueprint records has a score greater than the at least one desired score. In some configurations, the at least one desired score can be determined dynamically. For example, the at least one desired score can be determined as the 90th percentile of the set of predicted scores.

[0064] The method 300 for engineered polypeptide design optionally includes, at 305, determining whether to retrain a machine learning model by using numerical methods (such as Rosetta re-modeler, ab initio molecular dynamics simulation, machine learning structure prediction (such as AlphaFold or trRosetta), protein folding based on a structural knowledge base, neural network protein folding, sequence-based recurrent or transformer network protein folding, generative adversarial network protein structure generation, Markov chain Monte Carlo protein folding, etc.) to calculate a second set of scores (e.g., a reference true score set). Then the engineered polypeptide design device compares the second set of scores with the predicted score set and determines whether to retrain the machine learning model based on the deviation between the predicted score set and the second set of scores. The method 300 for engineered polypeptide design optionally includes, at 305, in response to the determination, retraining the machine learning model based on: (1) retraining the blueprint records including the second set of blueprint records and (2) retraining the scores including the predicted score set. In some configurations, the engineered polypeptide design device can connect the first set of blueprint records and the second set of blueprint records to generate retrained blueprint records. The engineered polypeptide design device can additionally connect the first set of scores and the second set of scores to generate retrained scores. In some configurations, the retraining of the blueprint records only includes the second set of blueprint records, and the retrained scores only include the second set of scores.

[0065] Figure 4 is a schematic diagram of an exemplary method 400 for engineered polypeptide design. The method 400 for engineered polypeptide design can be performed, for example, by an engineered polypeptide design device (similar to the engineered polypeptide design device 101 as Figure 1 shown and described). The method 400 for engineered polypeptide design includes, at step 401, training a machine learning model (similar to the machine learning model 107 as Figure 1 shown and described) based on a first set of blueprint records or a representation thereof and a first set of scores, where each blueprint record from the first set of blueprint records is associated with each score from the first set of scores. The representation can use a data preparation module (similar to Figure 1The data preparation module shown and described is generated based on the first blueprint record set. The method 400 for engineered polypeptide design further includes, at step 402, executing a machine learning model after training to generate a second blueprint record set having at least one desired score. The method 400 for engineered polypeptide design optionally includes, at step 403, performing computational protein modeling on the second blueprint record set to generate an engineered polypeptide. In some configurations, the method 400 for engineered polypeptide design optionally includes, at step 404, filtering the engineered polypeptide by performing a static structural comparison with a representation of a reference target structure. In some configurations, the method 400 for engineered polypeptide design optionally includes, at step 405, using molecular dynamics (MD) simulations of the representations of each of the reference target structure and the structure of the engineered polypeptide to filter the engineered polypeptide by performing a dynamic structural comparison with a representation of the reference target structure.

[0066] Figure 5 is a schematic diagram of an exemplary method for preparing data for an engineered polypeptide design device. A ribbon diagram of the structure of the target protein is shown on the left. The predetermined portion is shown in a darker color, and the side chains of the amino acid residues of the predetermined portion are shown as stick figures. In this example, the predetermined portion is a part of the target protein that is the desired target epitope of an antibody. By generating an engineered polypeptide that recapitulates the epitope, it is expected that an antibody that specifically binds to this part of the target protein can be obtained.

[0067] Figure 5 The right diagram of shows a schematic diagram of a blueprint set. Each circle represents a residue position. The scaffold residue positions are light gray and the side chains are not shown. The target residue positions are dark gray and the side chains at each position are shown. The side chains are side chains of well-known natural amino acids. In some cases, the target residues and / or scaffold residues are non-natural amino acids. In this example, each target residue position exactly corresponds to a residue of the predetermined portion of the reference target structure of the target protein. The shown blueprint set is "sequential" because in each figure, the order of the target residue positions is the same. The order of the target residues is not necessarily the same as the order of the residues in the target protein sequence. The first and last blueprints have consecutive target residue positions, while the other blueprints are non-consecutive. At least one scaffold residue position is between the first and last target residue positions. The letters N and C represent the amino (N) terminus and carboxyl (C) terminus of the polypeptide that matches a given blueprint.

[0068] Figure 5The five blueprints shown are members of a large number of possible blueprints, represented by the ovals between the lines in the figure. For a blueprint with 35 positions (corresponding to a 35-mer polypeptide), assuming the target residues are in sequence, the total number of potential blueprints is given by the formula: 35! ÷ (11! × (35 - 11)!) = 0.42 trillion. Even using the largest available supercomputing services, it would take years or even a lifetime for the Rosetta reconstructor to compute all possible 35-mers. Thus, using current computing devices and methods, direct computational modeling of each blueprint individually is computationally intractable.

[0069] Figure 6 is a schematic diagram of an exemplary method for engineered polypeptide design. The right side of the schematic shows how a scaffold blueprint (e.g., converted to a blueprint record suitable for use as an input, not shown) is input into a computational protein modeling program (similar to the computational protein modeling module 106 shown and described as Figure 1 including but not limited to the Rosetta reconstructor) to generate a score used as a label. The score generally reflects the energy terms used by the modeling program. In the case of the Rosetta reconstructor, the score includes an energy term reflecting the folding of the designed polypeptide generated from the blueprint and a structural constraint matching term reflecting the structural similarity between the predicted structure of the designed polypeptide and the known structure of a predetermined portion of the reference target structure of the target protein. Other modeling programs and other scoring functions can be used.

[0070] The left side of the schematic shows the conversion of the blueprint-to-blueprint representation. The representation can be any representation suitable for a machine learning model (such as the machine learning model 107 shown and described as Figure 1 here, the representation is a vector. More specifically, the vector is an ordered list of the number of inserted scaffold residues between target residue positions. This representation can be used because the order of the target residue positions is fixed in this representation, and thus the representation does not need to identify the amino acid identity of the target residue positions. This information is implicit. The order of the target residue positions does not necessarily have to be the same as the order in the target structure sequence. The first element 8 of the vector indicates that there are eight scaffold residue positions before the first target residue position. The second element 1 of the vector indicates that there is one scaffold residue position after the first target residue position and before the second target residue position. Subsequent elements 0, 1, 2, or 3 indicate no inserted scaffold residue positions, one, two, or three inserted scaffold residue positions. The last element 4 of the vector indicates that the last four positions in the blueprint are scaffold residue positions.

[0071] One advantage of this change in the representation of the blueprint record is that, except for the first and last elements, the vector is frame-shift constant. That is, the machine learning model has available information about the relative position of the target residue that is independent of the position of the target residue in the blueprint. This allows for the design of similar structures with variable structured / unstructured regions at the N- and C-termini.

[0072] Figure 7 is a schematic diagram of an exemplary performance of a machine learning model for engineered polypeptide design. The scatter plot shows the accuracy with which a machine learning model (such as, for example, the machine learning model 107 shown and described as Figure 1 can generate / predict a set of predicted scores for a set of blueprint records. Each point in the scatter plot represents a blueprint record from the set of blueprint records. The horizontal axis represents the ground truth score for the set of blueprint records that can be calculated by numerical methods such as, for example, the Rosetta reconstructor, ab initio molecular dynamics simulations, and the like. The vertical axis represents the predicted scores for the set of blueprint records generated / predicted by the machine learning model, which runs significantly faster (e.g., 50% faster, 2 times faster, 10 times faster, 100 times faster, 1000 times faster, 1,000,000 times faster, 1,000,000,000 times faster, etc.) than the numerical methods. In an ideal case, the predicted scores correspond to (e.g., equal to, approximate to) the ground truth scores. In cases where the predicted scores do not correspond to the ground truth scores, the machine learning model can be retrained with the set of blueprint records and the ground truth scores until the newly generated predicted scores for the newly generated set of blueprint records correspond to the ground truth scores for the newly generated set of blueprint records. In general, the scores can include energy terms (such as the Rosetta energy function 2015 (REF15)) and structure constraint matching terms (as Figure 6 described). The scores can be defined such that a low score for a blueprint record reflects a low molecular dynamics energy and higher stability of the blueprint record, as shown herein Figure 7 shown. In some variations, the scores can be defined such that a high score for a blueprint record generally reflects a higher stability of the polypeptide constructed based on the blueprint record.

[0073] Figure 8 is a schematic diagram of an exemplary method for engineered polypeptide design using a machine learning model. As Figure 8 shown, an initial data set including a first set of blueprint records and a first set of scores (e.g., representing energy terms such as Rosetta energy or molecular dynamics energy) can be generated and further prepared by a data preparation module (such as the data preparation module 105 shown and described as Figure 1 shown). A machine learning model (similar to the one shown and described as Figure 1The machine learning model 107) shown and described can be trained based on an initial data set. A second blueprint record set can be provided as input to the machine learning model to generate a second set of scores. The second blueprint record set or a portion of the second blueprint record set having a score greater than a predetermined value (e.g., a desired score) can be verified against a ground truth score. If the second set of scores corresponds to the ground truth score accurately enough (e.g., with an accuracy greater than 95%), the second blueprint record set or a portion of the second blueprint record set can be presented to the user. Otherwise, the second blueprint record set or a portion of the second blueprint record set can be used to retrain the machine learning model. In some cases, a third blueprint record set, a fourth blueprint record set, or more blueprint record iterations can be generated to obtain a blueprint with a desired score. In some cases, by iteratively retraining the machine learning model for new blueprint sets and score sets, as many blueprint sets as desired that achieve the desired score can be generated. An exemplary code snippet showing the process of training and using a machine learning model to generate engineered polypeptide designs is as follows:

[0074] training_energies = Rosetta(training_scaffolds) ## Rosetta energies are calculated for an initial training set of scaffolds

[0075] while training_energies has not converged: ## Iterate until Rosetta energies stop improving

[0076] Train xgboost to predict training_energies from training_scaffolds ## Train XGBoost to predict Rosetta energies from the training set of scaffolds

[0077] Predicted_scaffolds = Best predicted scaffolds from xgboost ## Use XGBoost to predict the best scaffolds

[0078] new_energies = Rosetta(predicted_scaffolds) ## Calculate Rosetta energies for the predicted scaffolds

[0079] Add predicted_scaffolds to training_scaffolds ## Add the predicted scaffolds to the training set

[0080] Add new_energies to training_energies ## Add the predicted scaffold energies to the training set

[0081] Figure 9Schematic illustration of exemplary performance of a machine learning model for engineered polypeptide design. As Figure 5 described, for an exemplary blueprint record with 35 positions (consistent with a 35-mer polypeptide), assuming the target residues are sequential, the total number of potential blueprints is given by the formula: 35! ÷ (11! × (35 - 11)!) = 0.42 trillion. Thus, using current computing devices and methods, performing direct computational modeling of each blueprint individually using brute-force discovery / optimization is computationally intractable and may take years or decades. In contrast, using data-driven methods such as the machine learning models described herein can reduce the time for such discovery / optimization (e.g., reduce it to weeks, days, hours, minutes, etc.).

[0082] Figure 10A -D shows an exemplary method for performing molecular dynamics simulations to validate engineered polypeptides. After a machine learning model (such as machine learning model 107 shown and described as Figure 1 shown and described) is trained and executed to generate a set of generated blueprint records that are improved / optimized (e.g., meet design criteria, have a desired score, etc.), an engineered polypeptide design device (as Figure 1 shown and described) can validate the set of generated blueprint records.

[0083] The engineered polypeptide design device can perform computational protein modeling on the set of generated blueprint records (e.g., using computational design modeling module 106 as Figure 1 shown and described) to generate engineered polypeptides. In some embodiments, the engineered polypeptide design device can then filter out a subset of the engineered polypeptides by performing a static structural comparison on a representation of a reference target structure.

[0084] In some embodiments, the engineered polypeptide design device can then use molecular dynamics (MD) simulations of the representations of each of the reference target structure and the structure of the engineered polypeptide to filter out a subset of the engineered polypeptides by performing a dynamic structural comparison with the representation of the reference target structure. For example, the engineered polypeptide design device can select several (e.g., fewer than 10 hits) engineered polypeptides. In some cases, the MD simulation can determine the kinetics of the representations of each of the reference target structure and the structure of the engineered polypeptide under solution conditions, including steps of model preparation, equilibration (e.g., at a temperature of 100K to 300K), and unrestricted MD simulation. In some cases, the MD simulation can include applying force field parameters and solvent model parameters to the representations of each of the reference target structure and the structure of the engineered polypeptide. In some cases, the MD simulation can perform 1000 cycles of constrained minimization (e.g., to relieve structural clashes), constrained heating (e.g., constrained heating for 100 picoseconds and ramping up to ambient temperature), and relaxation of constraints (e.g., relaxation of constraints for 100 picoseconds and gradually removing backbone constraints).

[0085] Figure 11 An exemplary method of performing a molecular dynamics simulation to validate an engineered polypeptide is shown. In some embodiments, in addition to or as an alternative to the method described as in FIG. 10, the MD simulation can be time-limited. For example, the MD simulation can perform 30 ns of unrestricted kinetics. In some embodiments, additionally or alternatively, the MD simulation can be conformationally restricted. For example, the MD simulation can be performed to obtain 80% of the conformational information observed within any time frame to obtain such conformational information. In some embodiments, a metric of the simulation time that determines the throughput and accuracy of the equilibrium MD simulation can be calculated by the cosine similarity score of the simulations of the representations of each of the reference target structure and the structure of the engineered polypeptide.

[0086] Figure 12 is a schematic diagram of an exemplary method of performing molecular dynamics simulations in parallel. In some cases, the engineered polypeptide design can include performing multiple (e.g., 100s, 1000s, 10,000s, etc.) molecular dynamics simulations. In these cases, the processor of the engineered polypeptide design device (such as the processor 104 of the engineered polypeptide design device 101 as Figure 1 shown and described) can include a graphics processing unit (GPU), an accelerated processing unit, and / or any other processing unit that can perform computations in parallel. The GPU can include a set of symmetric multi-processing units (SMP). Thus, the GPU can be configured to parallel-process multiple (e.g., 10s, 100s, etc.) molecular dynamics simulations using the SMP set. In some variants, a cloud computing platform (such as Figure 1The multi-core processing unit on the backend service platform 160) shown and described can be used to perform multiple molecular dynamics simulations in parallel.

[0087] Figure 13 FIG. is a schematic diagram of an exemplary method for validating a machine learning model for engineered polypeptide design. In some embodiments, a scoring method can be used for the molecular dynamics (MD) simulation results of the representation of a reference target structure and the MD simulation results of each of the engineered polypeptides to evaluate each engineered polypeptide. The scoring method can involve using the root mean square deviation (RMSD):

[0088]

[0089] where N is the number of atoms, X i is the reference position vector of the reference target structure, and Y i is the position vector of each engineered polypeptide. Alternatively, the MEM and epitope structure dynamic matching scoring can be performed using the root mean square inner product (RMSIP):

[0090]

[0091] where sorted by the corresponding eigenvalues - from highest to lowest, for N predetermined reference residues, the eigenvectors ψ and are the eigenvector of the reference target structure and the eigenvector of the engineered polypeptide, respectively. Each of the eigenvectors ψ and represents the lowest frequency mode of motion, in this case, the top 10 eigenvectors sorted by the corresponding eigenvalues are used. The eigenvectors of the reference target structure and the eigenvectors of the engineered polypeptide can be calculated, for example, using principal component analysis (PCA).

[0092] For purposes of explanation, the foregoing description uses specific nomenclature to provide a full understanding of the invention. However, it will be apparent to those skilled in the art that specific details are not required to practice the invention. Accordingly, the foregoing description of specific embodiments of the invention is presented for purposes of illustration and description. They are not exhaustive or limit the invention to the precise forms disclosed; obviously, many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described to explain the principles of the invention and its practical application so that others skilled in the art may utilize the invention and various embodiments with various modifications suited to the particular use contemplated. The following claims and their equivalents are intended to define the scope of the invention.

[0093] Enumerated embodiments:

[0094] Embodiment I-1. A method, the method comprising:

[0095] Training a machine learning model based on a first plurality of blueprint records or representations thereof and a first plurality of scores, wherein each blueprint record from the first plurality of blueprint records is associated with each score from the first plurality of scores; and

[0096] After the training, executing the machine learning model to generate a second plurality of blueprint records having at least one desired score,

[0097] The second plurality of blueprint records are configured to be received as inputs in computational protein modeling to generate engineered polypeptides based on the second plurality of blueprint records.

[0098] Embodiment I-2. The method according to Embodiment I-1, comprising:

[0099] Receiving a representation of the reference target structure of a reference target; and

[0100] Generating the first plurality of blueprint records from a predetermined portion of the reference target structure, each blueprint record from the first plurality of blueprint records including a target residue position and a scaffold residue position, and each target residue position corresponding to one target residue from a plurality of target residues.

[0101] Embodiment I-3. The method according to Embodiment I-1 or I-2, wherein in at least one blueprint record, the target residue positions are discontinuous.

[0102] Embodiment I-4. The method according to any one of Embodiments I-1 to I-3, wherein in at least one blueprint record, the order of the target residue positions is different from the order of the target residue positions in the reference target sequence.

[0103] Embodiment I-5. The method according to any one of Embodiments I-1 to I-4, comprising:

[0104] For each blueprint record from the first plurality of blueprint records, marking the first plurality of blueprint records by the following steps:

[0105] Performing computational protein modeling on the blueprint record to generate a polypeptide structure,

[0106] Calculating a score of the polypeptide structure, and

[0107] Associating the score with the blueprint record.

[0108] Embodiment I-6. The method according to any one of Embodiments I-1 to I-5, wherein the computational protein modeling is based on de novo design without a template matching the reference target structure.

[0109] Embodiment I-7. The method according to any one of Embodiments I-1 to I-6, wherein each fraction from the first plurality of fractions includes an energy term and a structure constraint matching term, and the structure constraint matching term is determined using one or more structure constraints extracted from a representation of the reference target structure.

[0110] Embodiment I-8. The method according to any one of Embodiments I-1 to I-7, comprising:

[0111] determining whether the machine learning model needs to be retrained by calculating a second plurality of scores for the second plurality of blueprint records; and

[0112] as a response to the determination, retraining the machine learning model based on: (1) retraining the blueprint records including the second plurality of blueprint records and (2) retraining the scores including the second plurality of scores.

[0113] Embodiment I-9. The method according to Embodiment I-8, comprising:

[0114] after retraining the machine learning model, connecting the first plurality of blueprint records and the second plurality of blueprint records to generate retrained blueprint records and generating retrained scores, and each blueprint record from the retrained blueprint records is associated with a score from the retrained scores.

[0115] Embodiment I-10. The method according to any one of Embodiments I-1 to I-9, wherein the at least one desired score is a preset value.

[0116] Embodiment I-11. The method according to any one of Embodiments I-1 to I-9, wherein the at least one desired score is dynamically determined.

[0117] Embodiment I-12. The method according to any one of Embodiments I-1 to I-10, wherein the machine learning model is a supervised machine learning model.

[0118] Embodiment I-13. The method according to Embodiment I-12, wherein the supervised machine learning model includes an ensemble of decision trees, a boosted decision tree algorithm, an extreme gradient boosting (XGBoost) model, or a random forest.

[0119] Embodiment I-14. The method according to Embodiment I-12, wherein the supervised machine learning model includes a support vector machine (SVM), a feedforward machine learning model, a recurrent neural network (RNN), a convolutional neural network (CNN), a graph neural network (GNN), or a transformer neural network.

[0120] Embodiment I-15. The method according to any one of Embodiments I-1 to I-14, wherein the machine learning model is an inductive machine learning model.

[0121] Embodiment I-16. The method according to any one of Embodiments I-1 to I-14, wherein the machine learning model is a generative machine learning model.

[0122] Embodiment I-17. The method according to any one of Embodiments I-1 to I-16, comprising performing computational protein modeling on the second plurality of blueprint records to generate the engineered polypeptide.

[0123] Embodiment I-18. The method according to any one of Embodiments I-1 to I-17, comprising filtering the engineered polypeptide by performing a static structural comparison with a representation of the reference target structure.

[0124] Embodiment I-19. The method according to any one of Embodiments I-1 to I-18, comprising filtering the engineered polypeptide by performing a dynamic structural comparison with a representation of the reference target structure by using molecular dynamics (MD) simulations of representations of each of the reference target structure and the structure of the engineered polypeptide.

[0125] Embodiment I-20. The method according to Embodiment I-19, wherein the MD simulations are performed in parallel using symmetric multiprocessing (SMP).

[0126] Embodiment I-21. The method according to any one of Embodiments I-1 to I-20, wherein the number of blueprint records in the second plurality of blueprint records is less than the number of blueprint records in the first plurality of blueprint records.

[0127] Embodiment I-22. A non-transitory processor-readable medium storing code representing instructions to be executed by a processor, the code including code to cause the processor to perform the following operations:

[0128] Train a machine learning model based on a first plurality of blueprint records or representations thereof and a first plurality of scores, each blueprint record from the first plurality of blueprint records being associated with each score from the first plurality of scores; and

[0129] After the training, execute the machine learning model to generate a second plurality of blueprint records having at least one desired score,

[0130] The second plurality of blueprint records are configured to be received as inputs in computational protein modeling to generate an engineered polypeptide based on the second plurality of blueprint records.

[0131] Embodiment I-23. The medium as described in Embodiment I-22, comprising code that causes the processor to perform the following operations:

[0132] Receive a representation of a reference target structure; and

[0133] Generate the first plurality of blueprint records from a predetermined portion of the reference target structure, each blueprint record from the first plurality of blueprint records including a target residue position and a scaffold residue position, each target residue position from the plurality of target residue positions corresponding to one target residue from the plurality of target residues.

[0134] Embodiment I-24. The medium as described in Embodiment I-23, wherein in at least one blueprint record, the target residue positions are discontinuous.

[0135] Embodiment I-25. The medium as described in Embodiment I-23 or I-24, wherein in at least one blueprint record, the order of the target residue positions is different from the order of the target residue positions in the reference target sequence.

[0136] Embodiment I-26. The medium as described in any one of Embodiments I-23 to I-25, comprising code that causes the processor to perform the following operations:

[0137] Mark the first plurality of blueprint records by performing the following steps: perform computational protein modeling on each blueprint record to generate a polypeptide structure; calculate a score for the polypeptide structure; and associate the score with the blueprint record.

[0138] Embodiment I-27. The medium as described in Embodiment I-26, wherein the computational protein modeling is based on de novo design in the absence of a template that matches the reference target structure.

[0139] Embodiment I-28. The medium as described in Embodiment I-26 or I-27, wherein each score includes an energy term and a structure constraint matching term, and the structure constraint matching term is determined using one or more structure constraints extracted from the representation of the reference target structure.

[0140] Embodiment I-29. The medium as described in any one of Embodiments I-22 to I-28, comprising code that causes the processor to perform the following operations:

[0141] Determine whether the machine learning model needs to be retrained by calculating a second plurality of scores for the second plurality of blueprint records; and

[0142] In response to the determination, retrain the machine learning model based on: (1) retraining blueprint records including the second plurality of blueprint records and (2) retraining scores including the second plurality of scores.

[0143] Embodiment I-30. The medium as described in Embodiment I-29, including code that causes the processor to perform the following operations:

[0144] After retraining the machine learning model, connect the first plurality of blueprint records and the second plurality of blueprint records to generate retrained blueprint records and generate retrained scores, with each blueprint record from the retrained blueprint records associated with a score from the retrained scores.

[0145] Embodiment I-31. The medium as described in any one of Embodiments I-22 to I-30, wherein the at least one desired score is a preset value.

[0146] Embodiment I-32. The medium as described in any one of Embodiments I-22 to I-31, wherein the at least one desired score is dynamically determined.

[0147] Embodiment I-33. The medium as described in any one of Embodiments I-22 to I-32, wherein the machine learning model is a supervised machine learning model.

[0148] Embodiment I-34. The medium as described in any one of Embodiments I-22 to I-33, wherein the supervised machine learning model includes an ensemble of decision trees, a boosted decision tree algorithm, an extreme gradient boosting (XGBoost) model, or a random forest.

[0149] Embodiment I-35. The medium as described in Embodiment I-33, wherein the supervised machine learning model includes a support vector machine (SVM), a feedforward machine learning model, a recurrent neural network (RNN), a convolutional neural network (CNN), a graph neural network (GNN), or a transformer neural network.

[0150] Embodiment I-36. The medium as described in any one of Embodiments I-22 to I-35, wherein the machine learning model is an inductive machine learning model.

[0151] Embodiment I-37. The medium as described in any one of Embodiments I-22 to I-36, wherein the machine learning model is a generative machine learning model.

[0152] Embodiment I-38. The medium as described in any one of Embodiments I-22 to I-37, including code that causes the processor to perform the following operations:

[0153] Perform computational protein modeling on the second plurality of blueprint records to generate engineered polypeptides.

[0154] Embodiment I-39. The medium as described in Embodiment I-38, comprising code that causes the processor to perform the following operations:

[0155] Filter the engineered polypeptides by performing a static structural comparison with a representation of the reference target structure.

[0156] Embodiment I-40. The medium as described in Embodiment I-38 or I-39, comprising code that causes the processor to perform the following operations:

[0157] Filter the engineered polypeptides by performing a dynamic structural comparison with a representation of the reference target structure using molecular dynamics (MD) simulations of each of the representation of the reference target structure and the engineered polypeptides.

[0158] Embodiment I-41. The medium as described in Embodiment I-40, wherein the MD simulations are performed in parallel using symmetric multiprocessing (SMP).

[0159] Embodiment I-42. The medium as described in any one of Embodiments I-22 to I-41, wherein the number of blueprint records in the second plurality of blueprint records is less than the number of blueprint records in the first plurality of blueprint records.

[0160] Embodiment I-43. An apparatus for selecting an engineered polypeptide, the apparatus comprising:

[0161] A first computing device having a processor and a memory, the memory storing instructions that are executable by the processor to:

[0162] Receive a reference target structure from a second computing device remote from the first computing device;

[0163] Generate a first plurality of blueprint records from a predetermined portion of the reference target structure, each blueprint record from the first plurality of blueprint records including a target residue position and a scaffold residue position, each target residue position corresponding to one of a plurality of target residues;

[0164] Train a machine learning model based on the first plurality of blueprint records or their representations and a first plurality of scores, each blueprint record from the first plurality of blueprint records being associated with each score from the first plurality of scores; and

[0165] After the training, execute the machine learning model to generate a second plurality of blueprint records having at least one desired score,

[0166] The second plurality of blueprint records are configured to be received as inputs in computational protein modeling to generate engineered polypeptides based on the second plurality of blueprint records.

[0167] Embodiment I-44. The apparatus as described in Embodiment I-43, including code that causes the processor to perform the following operations:

[0168] Determine whether the machine learning model needs to be retrained by calculating a second plurality of scores for the second plurality of blueprint records; and

[0169] In response to the determination, retrain the machine learning model based on: (1) retraining the blueprint records including the second plurality of blueprint records and (2) retraining the scores including the second plurality of scores.

[0170] Embodiment I-45. The apparatus as described in Embodiment I-43 or I-44, wherein the desired score is a preset value.

[0171] Embodiment I-46. The apparatus as described in any one of Embodiments I-43 to I-45, wherein the desired score is dynamically determined.

[0172] Embodiment I-47. The apparatus as described in any one of Embodiments I-43 to I-46, wherein the machine learning model is a supervised machine learning model.

[0173] Embodiment I-48. The apparatus as described in Embodiment I-47, wherein the supervised machine learning model includes an ensemble of decision trees, a boosted decision tree algorithm, an extreme gradient boosting (XGBoost) model, or a random forest.

[0174] Embodiment I-49. The apparatus as described in Embodiment I-47 or I-48, wherein the supervised machine learning model includes a support vector machine (SVM), a feedforward machine learning model, a recurrent neural network (RNN), a convolutional neural network (CNN), a graph neural network (GNN), or a transformer neural network.

[0175] Embodiment I-50. The apparatus as described in any one of Embodiments I-43 to I-49, wherein the machine learning model is an inductive machine learning model.

[0176] Embodiment I-51. The apparatus as described in any one of Embodiments I-43 to I-50, wherein the machine learning model is a generative machine learning model.

[0177] Embodiment I-52. The apparatus as described in any one of Embodiments I-43 to I-51, including code that causes the processor to perform the following operations:

[0178] Perform computational protein modeling on the second plurality of blueprint records to generate an engineered polypeptide.

[0179] Embodiment I-53. The apparatus according to embodiment I-52, comprising code that causes the processor to perform the following operations:

[0180] Filter the engineered polypeptide by performing a static structural comparison with a representation of a reference target structure.

[0181] Embodiment I-54. The apparatus according to embodiment I-52 or I-53, comprising code that causes the processor to perform the following operations:

[0182] Filter the engineered polypeptide by performing a dynamic structural comparison with a representation of the reference target structure using molecular dynamics (MD) simulations of each of the representation of the reference target structure and the engineered polypeptide.

[0183] Embodiment I-55. The apparatus according to embodiment I-54, wherein the MD simulations are performed in parallel using symmetric multiprocessing (SMP).

[0184] Embodiment I-56. An engineered polypeptide design generated by a method according to any one of embodiments I-1 to I-21, a medium according to any one of embodiments I-22 to I-42, or an apparatus according to any one of embodiments I-43 to I-55.

[0185] Embodiment I-57. An engineered peptide, wherein the engineered peptide has a molecular mass between 1 kDa and 10 kDa and contains at most 50 amino acids, and wherein the engineered peptide comprises:

[0186] A combination of spatially related topological constraints, wherein one or more of the constraints are constraints derived from a reference target; and

[0187] Wherein 10% to 98% of the amino acids of the engineered peptide satisfy the one or more constraints derived from the reference target,

[0188] Wherein the amino acids that satisfy the one or more constraints derived from the reference target have a backbone root mean square deviation (RSMD) structural homology with the reference target of less than .

[0189] Embodiment I-58. The engineered peptide according to embodiment I-57, wherein the amino acids that satisfy the one or more constraints derived from the reference target have a sequence homology between 10% and 90% with the reference target.

[0190] Embodiment I-59. The engineered peptide as described in Embodiment I-57 or I-58, wherein the combination includes at least two reference target-derived constraints.

[0191] Embodiment I-60. The engineered peptide as described in any one of Embodiments I-57 to I-59, wherein the combination includes an energy term and a structure constraint matching term, and the structure constraint matching term is determined using one or more structure constraints extracted from a representation of the reference target structure.

[0192] Embodiment I-61. The engineered peptide as described in any one of Embodiments I-57 to I-60, wherein the one or more non-reference target-derived constraints describe desired structural features, kinetic features, or any combination thereof.

[0193] Embodiment I-62. The engineered peptide as described in any one of Embodiments I-57 to I-61, wherein the reference target contains one or more atoms associated with a biological reaction or biological function,

[0194] and wherein the atomic fluctuations of the one or more atoms in the engineered peptide associated with the biological reaction or biological function overlap with the atomic fluctuations of the one or more atoms in the reference target associated with the biological reaction or biological function.

[0195] Embodiment I-63. The engineered peptide as described in Embodiment I-62, wherein the root mean square inner product (RMSIP) of the overlap is greater than 0.25.

[0196] Embodiment I-64. The engineered peptide as described in any one of Embodiments I-62 or I-63, wherein the root mean square inner product (RMSIP) of the overlap is greater than 0.75.

[0197] Embodiment I-65. A method for selecting an engineered peptide, the method comprising:

[0198] identifying one or more topological features of a reference target;

[0199] designing spatially related constraints for each topological feature to generate a combination of spatially related topological constraints derived from the reference target;

[0200] comparing the spatially related topological features of a candidate peptide with the combination of spatially related topological constraints derived from the reference target; and

[0201] selecting a candidate peptide having spatially related topological features to generate the engineered peptide, the topological features overlapping with the combination of spatially related topological constraints derived from the reference target.

[0202] Embodiment I-66. The method according to Embodiment I-65, wherein one or more constraints are derived from the energy of each residue and the atomic distances of each residue.

[0203] Embodiment I-67. The method according to any one of Embodiments I-65 or I-66, wherein the characteristics of one or more candidate peptides are determined by computer simulation.

[0204] Embodiment I-68. The method according to Embodiment I-67, wherein the computer simulation includes molecular dynamics simulation, Monte Carlo simulation, coarse-grained simulation, Gaussian network model, machine learning, or any combination thereof.

[0205] Embodiment I-69. The method according to any one of Embodiments I-65 to I-68, wherein the amino acids that satisfy the constraints derived from the one or more reference targets have a sequence homology of between 10% and 90% with the reference target.

[0206] Embodiment I-70. The method according to any one of Embodiments I-65 to I-69, wherein the constraints derived from the one or more non-reference targets describe desired structural features and / or kinetic features.

Claims

1. A method for designing and selecting engineered peptides, the method comprising: receiving a representation of a reference target structure of a reference target; generating a first plurality of blueprint records from a predetermined portion of the reference target structure, each blueprint record from the first plurality of blueprint records including a target residue position and a scaffold residue position, each target residue position corresponding to one target residue from a plurality of target residues; for each blueprint record from the first plurality of blueprint records, marking the first plurality of blueprint records by: performing computational protein modeling on the blueprint record to generate a polypeptide structure, calculating a score of the polypeptide structure, and associating the score with the blueprint record, wherein each score includes an energy term and a structure constraint match term, the structure constraint match term being determined using one or more structure constraints extracted from the representation of the reference target structure; identifying one or more topological features of the reference target and encoding the one or more topological features in a scaffold blueprint; designing spatially related constraints for each topological feature in the scaffold blueprint to generate a combination of spatially related topological constraints derived from the reference target; converting the scaffold blueprint into a vector representation to generate candidate polypeptides in which the spatially related topological features overlap with the combination of spatially related topological constraints derived from the reference target; training a machine learning model using the representation of the scaffold blueprint and the spatially related topological constraints derived from the reference target, wherein the representation is a one-dimensional digital vector, a two-dimensional alphanumeric data matrix, or a three-dimensional normalized digital tensor; executing the machine learning model after the training to generate a second plurality of blueprint records having at least one desired score; the second plurality of blueprint records being configured to be received as an input in computational protein modeling to generate engineered polypeptides based on the second plurality of blueprint records; comparing the spatially related topological features of the candidate peptides with the combination of spatially related topological constraints derived from the reference target; and selecting candidate peptides having spatially related topological features that overlap with the combination of spatially related topological constraints derived from the reference target to produce engineered peptides.

2. The method of claim 1, wherein in at least one blueprint record, the target residue positions are discontinuous.

3. The method of claim 1, wherein in at least one blueprint record, the order of the target residue positions is different from the order of the target residue positions in the reference target sequence.

4. The method of claim 1, wherein the computational protein modeling is based on de novo design in the absence of a template that matches the reference target structure.

5. The method of claim 1, comprising: determining whether the machine learning model needs to be retrained by calculating a second plurality of scores for the second plurality of blueprint records; and as a response to the determination, retraining the machine learning model based on: (1) retraining the blueprint records including the second plurality of blueprint records and (2) retraining the scores including the second plurality of scores.

6. The method of claim 5, comprising: After retraining the machine learning model, connect the first plurality of blueprint records and the second plurality of blueprint records to generate retrained blueprint records and generate a retrained score, where each blueprint record from the retrained blueprint records is associated with a score from the retrained score.

7. The method according to claim 1, wherein the at least one desired score is a preset value.

8. The method according to claim 1, wherein the at least one desired score is dynamically determined.

9. The method according to claim 1, wherein the machine learning model is a supervised machine learning model.

10. The method according to claim 9, wherein the supervised machine learning model includes an ensemble of decision trees, a boosted decision tree algorithm, an extreme gradient boosting (XGBoost) model, or a random forest.

11. The method according to claim 9, wherein the supervised machine learning model includes a support vector machine (SVM), a feedforward machine learning model, a recurrent neural network (RNN), a convolutional neural network (CNN), a graph neural network (GNN), or a transformer neural network.

12. The method according to claim 1, wherein the machine learning model is an inductive machine learning model.

13. The method according to claim 1, wherein the machine learning model is a generative machine learning model.

14. The method according to claim 1, including performing computational protein modeling on the second plurality of blueprint records to generate the engineered polypeptide.

15. The method according to claim 14, including filtering the engineered polypeptide by performing a static structural comparison with a representation of the reference target structure.

16. The method according to claim 14, including filtering the engineered polypeptide by performing a dynamic structural comparison with a representation of the reference target structure using molecular dynamics (MD) simulations of each of the representation of the reference target structure and the structure of the engineered polypeptide.

17. The method according to claim 16, wherein the MD simulations are performed in parallel using symmetric multiprocessing (SMP).

18. The method according to claim 1, wherein the number of blueprint records in the second plurality of blueprint records is less than the number of blueprint records in the first plurality of blueprint records.

19. A non-transitory processor-readable medium storing code representing instructions to be executed by a processor, the code including code to cause the processor to perform the following operations: Receive a representation of a reference target structure of a reference target to: Generate a first plurality of blueprint records from a predetermined portion of the reference target structure, where each blueprint record from the first plurality of blueprint records includes a target residue position and a scaffold residue position, and each target residue position corresponds to one target residue from a plurality of target residues; For each blueprint record from the first plurality of blueprint records, label the first plurality of blueprint records by the following steps: Perform computational protein modeling on the blueprint record to generate a polypeptide structure, Calculate a score of the polypeptide structure, and Associate the score with the blueprint record; Training a machine learning model based on a first plurality of blueprint records or representations thereof and a first plurality of scores, wherein each blueprint record from the first plurality of blueprint records is associated with each score from the first plurality of scores; and After the training, executing the machine learning model to generate a second plurality of blueprint records having at least one desired score, wherein each score from the first plurality of scores includes an energy term and a structural constraint matching term, and the structural constraint matching term is determined using one or more structural constraints extracted from a representation of the reference target structure; The second plurality of blueprint records are configured to be received as inputs in computational protein modeling to generate engineered polypeptides based on the second plurality of blueprint records.

20. The non-transitory processor-readable medium of claim 19, wherein in at least one blueprint record, the target residue positions are discontinuous.

21. The non-transitory processor-readable medium of claim 19, wherein in at least one blueprint record, the order of the target residue positions is different from the order of the target residue positions in the reference target sequence.

22. The non-transitory processor-readable medium of claim 19, wherein the computational protein modeling is based on de novo design in the absence of a template that matches the reference target structure.

23. The non-transitory processor-readable medium of claim 19, comprising code that causes the processor to perform the following operations: Determining whether the machine learning model needs to be retrained by calculating a second plurality of scores for the second plurality of blueprint records; and In response to the determination, retraining the machine learning model based on: (1) retraining the blueprint records including the second plurality of blueprint records and (2) retraining the scores including the second plurality of scores.

24. The non-transitory processor-readable medium of claim 23, comprising code that causes the processor to perform the following operations: After retraining the machine learning model, connecting the first plurality of blueprint records and the second plurality of blueprint records to generate retrained blueprint records and generating retrained scores, wherein each blueprint record from the retrained blueprint records is associated with a score from the retrained scores.

25. The non-transitory processor-readable medium of claim 19, wherein the at least one desired score is a preset value.

26. The non-transitory processor-readable medium of claim 19, wherein the at least one desired score is dynamically determined.

27. The non-transitory processor-readable medium of claim 19, wherein the machine learning model is a supervised machine learning model.

28. The non-transitory processor-readable medium of claim 27, wherein the supervised machine learning model includes an ensemble of decision trees, a boosted decision tree algorithm, an extreme gradient boosting (XGBoost) model, or a random forest.

29. The non-transitory processor-readable medium of claim 27, wherein the supervised machine learning model comprises a support vector machine (SVM), a feedforward machine learning model, a recurrent neural network (RNN), a convolutional neural network (CNN), a graph neural network (GNN), or a transformer neural network.

30. The non-transitory processor-readable medium of claim 19, wherein the machine learning model is an inductive machine learning model.

31. The non-transitory processor-readable medium of claim 19, wherein the machine learning model is a generative machine learning model.

32. The non-transitory processor-readable medium of claim 19, comprising code that causes the processor to perform the following operations: Perform computational protein modeling on the second plurality of blueprint records to generate engineered polypeptides.

33. The non-transitory processor-readable medium of claim 32, comprising code that causes the processor to perform the following operations: Filter the engineered polypeptides by performing a static structural comparison with a representation of the reference target structure.

34. The non-transitory processor-readable medium of claim 32, comprising code that causes the processor to perform the following operations: Filter the engineered polypeptides by performing a dynamic structural comparison with a representation of the reference target structure using molecular dynamics (MD) simulations of the representation of the reference target structure and each of the engineered polypeptides.

35. The non-transitory processor-readable medium of claim 34, wherein the MD simulations are performed in parallel using symmetric multi-processing (SMP).

36. The non-transitory processor-readable medium of claim 19, wherein the number of blueprint records in the second plurality of blueprint records is less than the number of blueprint records in the first plurality of blueprint records.

37. An apparatus for selecting an engineered polypeptide, the apparatus comprising: A first computing device having a processor and a memory, the memory storing instructions executable by the processor to: Receive a representation of a reference target structure of a reference target from a second computing device remote from the first computing device; Generate a first plurality of blueprint records from a predetermined portion of the reference target structure, each blueprint record from the first plurality of blueprint records comprising a target residue position and a scaffold residue position, each target residue position corresponding to one of a plurality of target residues; For each blueprint record from the first plurality of blueprint records, label the first plurality of blueprint records by the following steps: Perform computational protein modeling on the blueprint record to generate a polypeptide structure, Calculate a score of the polypeptide structure, and Associate the score with the blueprint record; Train a machine learning model based on the first plurality of blueprint records or a representation thereof and a first plurality of scores, each blueprint record from the first plurality of blueprint records being associated with each score from the first plurality of scores; And After the training, the machine learning model is executed to generate a second plurality of blueprint records having at least one desired score, wherein each score from the first plurality of scores includes an energy term and a structural constraint matching term, and the structural constraint matching term is determined using one or more structural constraints extracted from a representation of the reference target structure; The second plurality of blueprint records are configured to be received as inputs in computational protein modeling to generate an engineered polypeptide based on the second plurality of blueprint records.

38. The apparatus of claim 37, comprising code that causes the processor to perform the following operations: Determine whether the machine learning model needs to be retrained by calculating a second plurality of scores for the second plurality of blueprint records; and In response to the determination, retrain the machine learning model based on (1) retraining the blueprint records including the second plurality of blueprint records and (2) retraining the scores including the second plurality of scores.

39. The apparatus of claim 37, wherein the desired score is a preset value.

40. The apparatus of claim 37, wherein the desired score is dynamically determined.

41. The apparatus of claim 37, wherein the machine learning model is a supervised machine learning model.

42. The apparatus of claim 41, wherein the supervised machine learning model includes an ensemble of decision trees, a boosted decision tree algorithm, an extreme gradient boosting (XGBoost) model, or a random forest.

43. The apparatus of claim 41, wherein the supervised machine learning model includes a support vector machine (SVM), a feedforward machine learning model, a recurrent neural network (RNN), a convolutional neural network (CNN), a graph neural network (GNN), or a transformer neural network.

44. The apparatus of claim 37, wherein the machine learning model is an inductive machine learning model.

45. The apparatus of claim 37, wherein the machine learning model is a generative machine learning model.

46. The apparatus of claim 37, comprising code that causes the processor to perform the following operations: Perform computational protein modeling on the second plurality of blueprint records to generate an engineered polypeptide.

47. The apparatus of claim 46, comprising code that causes the processor to perform the following operations: Filter the engineered polypeptide by performing a static structural comparison with a representation of the reference target structure.

48. The apparatus of claim 46, comprising code that causes the processor to perform the following operations: Filter the engineered polypeptide by performing a dynamic structural comparison with a representation of the reference target structure using molecular dynamics (MD) simulations of each of the representation of the reference target structure and the engineered polypeptide.

49. The apparatus of claim 48, wherein the MD simulations are performed in parallel using symmetric multi-processing (SMP).

50. A method of designing an engineered polypeptide using a machine learning model, comprising: (a) Receiving a representation of a reference target structure of a reference target by the following steps: Generate a first plurality of blueprint records from a predetermined portion of the reference target structure, each blueprint record from the first plurality of blueprint records including a target residue position and a scaffold residue position, each target residue position corresponding to one of a plurality of target residues; For each blueprint record from the first plurality of blueprint records, label the first plurality of blueprint records by the following steps: Perform computational protein modeling on the blueprint record to generate a polypeptide structure, Calculate a score of the polypeptide structure, and Associate the score with the blueprint record, where each score includes an energy term and a structure constraint matching term, the structure constraint matching term being determined using one or more structure constraints extracted from a representation of the reference target structure; (b) Generate a training set of blueprint records from a predetermined portion of the reference target structure, where each blueprint record includes a target residue position and a scaffold residue position, each target residue position corresponding to one of a plurality of target residues; (c) For each blueprint record of the training set, label each blueprint record in the training set of blueprint records by the following steps: (i) Perform computational protein modeling on the blueprint record to generate a polypeptide structure, (ii) Calculate a score of the polypeptide structure, and (iii) Associate the score with the blueprint record; (d) Train a machine learning model based on the labeled training set; And (e) Apply the trained machine learning model to a desired score set to generate an output set of blueprint records having the desired scores.

51. The method of claim 50, wherein in at least one blueprint record, the target residue positions are discontinuous.

52. The method of claim 50, wherein in at least one blueprint record, the order of the target residue positions is different from the order of the target residue positions in the reference target structure.

53. The method of claim 50, wherein the computational protein modeling is based on de novo design in the absence of a template that matches the reference target structure.

54. The method of claim 50, wherein the desired score set is dynamically determined.

55. The method of claim 50, wherein the machine learning model is a supervised machine learning model.

56. The method of claim 50, including performing computational protein modeling on the output set of blueprint records to generate a predicted structure of an engineered polypeptide.

57. The method of claim 56, including: Filtering the predicted structure of the engineered polypeptide by performing a static structural comparison with a representation of the reference target structure.

58. A non-transitory processor-readable medium storing code representing instructions to be executed by a processor, the processor using a machine learning model to design an engineered polypeptide, the code including code to cause the processor to perform the following operations: (a) Receive a representation of a reference target structure of a reference target; Generate a first plurality of blueprint records from a predetermined portion of the reference target structure, each blueprint record from the first plurality of blueprint records including a target residue position and a scaffold residue position, each target residue position corresponding to one of a plurality of target residues; For each blueprint record from the first plurality of blueprint records, mark the first plurality of blueprint records by: Performing computational protein modeling on the blueprint record to generate a polypeptide structure, Calculating a score for the polypeptide structure, and Associating the score with the blueprint record, where each score includes an energy term and a structural constraint match term, the structural constraint match term being determined using one or more structural constraints extracted from a representation of the reference target structure; (b) Generate a training set of blueprint records from a predetermined portion of the reference target structure, where each blueprint record includes a target residue position and a scaffold residue position, each target residue position corresponding to one of a plurality of target residues, by: (c) For each blueprint record of the training set, mark each blueprint record in the training set of blueprint records by: (i) Performing computational protein modeling on the blueprint record to generate a polypeptide structure, (ii) Calculating a score for the polypeptide structure, and (iii) Associating the score with the blueprint record; (d) Train a machine learning model based on the marked training set; And (e) Apply the trained machine learning model to a desired score set to generate an output set of blueprint records having the desired scores.

59. The non-transitory processor-readable medium of claim 58, wherein the desired score set is dynamically determined.

60. The non-transitory processor-readable medium of claim 58, wherein the machine learning model is a supervised machine learning model.

61. The non-transitory processor-readable medium of claim 58, comprising code that causes the processor to perform the following operations: Perform computational protein modeling on the output set of blueprint records to generate a predicted structure of an engineered polypeptide.

62. The non-transitory processor-readable medium of claim 61, comprising code that causes the processor to perform the following operations: Filter the predicted structure of the engineered polypeptide by performing a static structural comparison with a representation of the reference target structure.

63. The non-transitory processor-readable medium of claim 58, further Comprising: Generate a polypeptide sequence from a blueprint record of the output set of blueprint records by computational protein modeling.

64. The non-transitory processor-readable medium of claim 58, further Comprising: Produce a polypeptide having a polypeptide sequence generated from a blueprint record of the output set of blueprint records by computational protein modeling.

65. The non-transitory processor-readable medium of claim 58, comprising code that causes the processor to perform the following operations: Generate a polypeptide sequence from a blueprint record of the output set of blueprint records by computational protein modeling.

66. The non-transitory processor-readable medium according to claim 58, comprising code that causes the processor to perform the following operations: Generate a polypeptide having a polypeptide sequence that is generated from one blueprint record of an output set recorded from a blueprint by computational protein modeling.

Citation Information

Patent Citations

  • Folded and protease-resistant polypeptides

    WO2018201020A1