Machine learning techniques for predicting thermal stability

JP2025521079A5Pending Publication Date: 2026-05-11AMGEN INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
AMGEN INC
Filing Date
2023-05-09
Publication Date
2026-05-11

AI Technical Summary

Technical Problem

Conventional methods for predicting the thermal stability of single-chain variable fragments (scFvs) are unreliable and inaccurate, as they fail to consider the structure of the scFv, leading to inefficient resource consumption and high costs in experimental screening.

Method used

A method using a trained machine learning model that considers the three-dimensional structure and residue interactions of scFvs to predict thermal stability, incorporating interaction energy metrics and residue sequences, enabling accurate identification of thermally stable scFvs for subsequent manufacturing.

Benefits of technology

This approach allows for efficient computational screening of scFvs, reducing resource consumption and cost by accurately predicting thermal stability, thereby guiding the development of multispecific drugs with improved thermal stability characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Techniques for computationally screening a set of single-chain variable fragments (scFvs). The techniques include using a machine learning model to determine a thermal stability metric for each scFv in the set of scFvs to obtain a plurality of thermal stability metrics, the set of scFvs including a first scFv having a first residue sequence, and determining including using information indicative of the 3D structure of the first scFv to obtain an interaction energy metric for each of a plurality of pairs of residues in the first residue sequence, using the interaction energy metric to generate a first feature set, providing the first feature set as an input to the machine learning model to obtain a corresponding output indicative of the first thermal stability of the first scFv, identifying a subset of the set of scFvs for subsequent manufacture based on the plurality of thermal stability metrics, and manufacturing at least one of the identified scFvs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority under 35 U.S.C. § 119(e) to U.S. Provisional Patent Application No. 63 / 340,332, filed on May 10, 2022, with Attorney Docket No. A1350.70001US00, entitled "MACHINE LEARNING TECHNIQUES FOR PREDICTING SINGLE - CHAIN VARIABLE FRAGMENT (SCFV) THERMOSTABILITY", the entire content of which is incorporated herein by reference.

Background Art

[0002] Monoclonal antibodies (mAbs) represent a major class of therapeutic agents, and more than 100 FDA - approved products are sold in the United States. Multispecific biologics that act on multiple targets or epitopes on the same target are becoming increasingly important for accessing new therapeutic - related pathways and mechanisms of action. Several multispecific biologics have been approved for use, and many more are in clinical and pre - clinical development stages.

[0003] A common component in the construction of multispecific biologics is the single - chain variable fragment (scFv), which consists of a target - binding antibody variable heavy chain (VH) linked to a variable light chain (VL) via a flexible linker. Multispecific format platforms such as BiTE, IgG - scFv, and XmAb incorporate scFv modules.

Summary of the Invention

Means for Solving the Problems

[0004] Some embodiments provide a method for computationally screening a set of single-chain variable fragments (scFvs) based on the predicted thermal stability of the scFvs by a trained machine learning model, the set of scFvs including scFvs having different residue sequences, the method comprising using the trained machine learning model and at least one computer hardware processor to determine the thermal stability metrics of each scFv within the set of scFvs to obtain a plurality of thermal stability metrics, the set of scFvs including a first scFv having a first residue sequence, and determining comprising using information indicative of the three-dimensional (3D) structure of the first scFv to obtain an interaction energy metric for each of a plurality of pairs of residues in the first residue sequence, and generating a first feature set for providing as input to the trained machine learning model, generating comprising including the interaction energy metric in the first feature set, providing the first feature set as input to the trained machine learning model to obtain a corresponding output indicative of the first thermal stability of the first scFv, determining comprising, identifying a subset of the set of scFvs for subsequent manufacture based on the plurality of thermal stability metrics, and manufacturing at least one of the scFvs within the identified subset.

[0005] In some embodiments, the set of scFvs further includes a second scFv different from the first scFv, the second scFv having a second residue sequence. In some embodiments, determining the thermal stability metrics of each scFv within the set of scFvs further comprises obtaining a second interaction energy metric for each of a second plurality of pairs of second residues, the second residues being within the second residue sequence, and generating a second feature set for providing as input to the trained machine learning model, generating comprising including the second interaction energy metric in the second feature set, and providing the second feature set as input to the trained machine learning model to obtain a corresponding output indicative of the second thermal stability of the second scFv.

[0006] In some embodiments, the output indicating the first thermal stability of the first scFv indicates a first temperature at which the first scFv is thermally stable.

[0007] In some embodiments, the first temperature is an estimated value of the temperature corresponding to half of the maximum binding of the first scFv.

[0008] In some embodiments, the output indicating the first thermal stability of the first scFv indicates a first temperature range including at least one temperature at which the first scFv is thermally stable.

[0009] In some embodiments, the first temperature range is an estimated value of a temperature range including the temperature corresponding to half of the maximum binding of the first scFv.

[0010]

[0011] In some embodiments, obtaining the interaction energy metric includes determining information indicating the 3D structure of the first scFv by using protein structure prediction software to generate information indicating the 3D structure from the first residue sequence.

[0012]

[0013] ​​In some embodiments, generating the first feature set comprises generating, for each particular energy metric of the interaction energy metric, a respective two-dimensional (2D) matrix of values of the particular energy metric, wherein the rows and columns of the 2D matrix correspond to each residue within the first residue sequence, and the entry in the i-th row and j-th column of the 2D matrix corresponds to the value of the particular energy metric for the i-th residue and the j-th residue within the first residue sequence; and including the generated 2D matrix in the first feature set.

[0014] In some embodiments, the generated 2D matrix includes rows for at least 75% of the residues within the first residue sequence. In some embodiments, the generated 2D matrix includes rows for at least 90% of the residues within the first residue sequence. In some embodiments, the generated 2D matrix includes rows for at least 95% of the residues within the first residue sequence. In some embodiments, the generated 2D matrix includes rows for at least 99% of the residues within the first residue sequence.

[0015] In some embodiments, the generated 2D matrix includes rows for each residue within the first residue sequence.

[0016] In some embodiments, generating the first feature set further comprises encoding the first residue sequence to obtain an encoded sequence and including the encoded sequence in the first feature set.

[0017] In some embodiments, encoding the first residue sequence comprises one-hot encoding the first residue sequence to obtain an encoded sequence, the encoded sequence including the one-hot encoded version of the first residue sequence.

[0018] In some embodiments, the trained machine learning model includes a trained neural network model.

[0019] In some embodiments, the trained neural network model includes a trained convolutional neural network (CNN) model, and the trained CNN model has a plurality of 2D convolutional layers.

[0020] In some embodiments, the trained CNN model further includes a fully connected layer.

[0021] In some embodiments, the trained CNN model is configured to output a plurality of probabilities that the scFv is thermally stable at each of a plurality of temperature ranges.

[0022] In some embodiments, providing a first set of features as an input to the trained machine learning model to obtain a corresponding output indicative of the first thermal stability of the first scFv includes providing the first set of features to the trained CNN model to obtain a first plurality of probabilities that the first scFv is thermally stable at each of a plurality of temperature ranges, and determining the first thermal stability as either (i) a temperature range within a plurality of temperature ranges associated with the highest probability among the first plurality of probabilities, or (ii) a temperature determined as a weighted linear combination of the average values of the plurality of temperature ranges weighted by the probabilities in the first set of probabilities.

[0023] In some embodiments, identifying a subset of a set of scFvs for subsequent manufacture based on a plurality of determined thermal stability metrics includes determining whether the first thermal stability of the first scFv meets at least one criterion, and after determining that the first thermal stability meets at least one criterion, identifying the first scFv for subsequent manufacture.

[0024] Some embodiments further include testing the thermal stability of at least one scFv in an in vitro assay.

[0025] Some embodiments provide a method for predicting the thermal stability of a single-chain variable fragment (scFv) using a trained machine learning model, the method comprising using a trained machine learning model and at least one computer hardware processor to determine a first thermal stability indicator of a first scFv having a first residue sequence; using information indicative of the three-dimensional (3D) structure of the first scFv to obtain an interaction energy metric for each of a plurality of pairs of residues in the first residue sequence; generating a first feature set for providing as an input to the trained machine learning model, the generating comprising including the interaction energy metric in the first feature set; providing the first feature set as an input to the trained machine learning model to obtain a corresponding output indicative of the first thermal stability of the first scFv.

[0026] Some embodiments provide a system that includes at least one computer hardware processor and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform a method for predicting the thermal stability of a single-chain variable fragment (scFv) using a trained machine learning model. The method includes determining, using the trained machine learning model and the at least one computer hardware processor, a first thermal stability indicator of a first scFv having a first residue sequence; obtaining, for each of a plurality of pairs of residues in the first residue sequence, an interaction energy metric using information indicative of the three-dimensional (3D) structure of the first scFv; generating a first feature set to be provided as an input to the trained machine learning model, the generating including including the interaction energy metric in the first feature set; and providing the first feature set as an input to the trained machine learning model to obtain a corresponding output indicative of the first thermal stability of the first scFv.

[0027] Some embodiments provide at least one non - transitory computer - readable storage medium storing processor - executable instructions that, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to perform a method for predicting the thermal stability of a single - chain variable fragment (scFv) using a trained machine - learning model. The method includes using the trained machine - learning model and at least one computer hardware processor to determine a first thermal - stability indicator for a first scFv having a first residue sequence; using information indicating the three - dimensional (3D) structure of the first scFv to obtain an interaction - energy metric for each of a plurality of pairs of residues in the first residue sequence; generating a first feature set for providing as an input to the trained machine - learning model, the generating including including the interaction - energy metric in the first feature set; and providing the first feature set as an input to the trained machine - learning model to obtain a corresponding output indicating the first thermal stability of the first scFv.

[0028] Some embodiments further include using the trained machine - learning model and at least one computer hardware processor to determine a thermal - stability indicator for each scFv in a set of scFvs to obtain a plurality of thermal - stability indicators, where the set of scFvs includes the first scFv.

[0029] Some embodiments further include identifying a subset of the set of scFvs for subsequent manufacture based on the plurality of thermal - stability indicators.

[0030] Some embodiments provide a method for computationally screening a set of monoclonal antibodies (mAbs) based on the predicted thermal stability of the mAbs by a trained machine learning model, the set of mAbs including mAbs having different residue sequences, the method using a trained machine learning model and at least one computer hardware processor to determine a thermal stability metric for each mAb in the set of mAbs to obtain a plurality of thermal stability metrics, the set of mAbs including a first mAb having a first residue sequence, determining including using information indicative of the three-dimensional (3D) structure of the first mAb to obtain an interaction energy metric for each of a plurality of pairs of residues in the first residue sequence and generating a first feature set to be provided as an input to the trained machine learning model, generating including including the interaction energy metric in the first feature set, providing the first feature set as an input to the trained machine learning model to obtain a corresponding output indicative of the first thermal stability of the first mAb, determining including, based on the plurality of thermal stability metrics, identifying a subset of the set of mAbs for subsequent manufacture and manufacturing at least one of the mAbs in the identified subset.

[0031] Some embodiments further include testing the thermal stability of at least one of the mAbs in an in vitro assay.

[0032] Some embodiments provide a method of computationally screening a set of antibodies based on the predicted thermal stability of the antibodies by a trained machine learning model, the set of antibodies including antibodies having different residue sequences, the method using the trained machine learning model and at least one computer hardware processor to determine the thermal stability metrics of each antibody in the set of antibodies to obtain a plurality of thermal stability metrics, the set of antibodies including a first antibody having a first residue sequence, and determining including using information indicative of the three-dimensional (3D) structure of the first antibody to obtain an interaction energy metric for each of a plurality of pairs of residues in the first residue sequence and generating a first feature set to provide as an input to the trained machine learning model, the generating including including the interaction energy metric in the first feature set, providing the first feature set as an input to the trained machine learning model to obtain a corresponding output indicative of the first thermal stability of the first antibody, identifying a subset of the set of antibodies for subsequent manufacture based on the plurality of thermal stability metrics, and manufacturing at least one of the antibodies in the identified subset.

[0033] Some embodiments further include testing the at least one thermal stability of the antibody in an in vitro assay.

[0034] The accompanying drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component that is illustrated in various figures is represented by a like numeral. For clarity, it is not possible to label every component in every drawing. The drawings are as follows.

Brief Description of the Drawings

[0035]

Figure 1A

Figure 1B

Figure 2A

Figure 2B

Figure 3A

Figure 3B

Figure 3C

Figure 4A

Figure 4B

Figure 5A

Figure 5B

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12A

Figure 12B

Figure 12C

Figure 13A

Figure 13B

Figure 14

Figure 15A

Figure 15B

Figure 15C

Figure 15D

Figure 15E

Figure 15F

Figure 16

Figure 17A

Figure 17B

Figure 18A

Figure 18B

Figure 19

Figure 20

Figure 21A

Figure 21B

Figure 22-1

Figure 22-2

Figure 23

Figure 24

[0036] The inventors have developed machine learning techniques for predicting the thermal stability of single-chain variable fragments (scFvs) or other multispecific constructs. In some embodiments, predicting the thermal stability of an scFv involves processing a feature set generated for the scFv using a trained machine learning model. In some embodiments, the feature set is generated using the residue sequence of the scFv. For example, in some embodiments, the feature set includes interaction energy metrics for residue pairs within the residue sequence. Additionally, or alternatively, the feature set includes the residue sequence of the scFv. In some embodiments, the feature set is provided as an input to a trained machine learning model to obtain an output (e.g., a thermal stability metric) indicating the temperature at which the scFv is thermally stable.

[0037] In some embodiments, the techniques described herein are used to screen a set of scFvs. For example, the set of scFvs may include one or more scFvs that are candidates for subsequent production. In some embodiments, techniques for screening a set of scFvs include determining the thermal stability metric for each scFv included in the set of scFvs and identifying a subset of the set of scFvs based on the determined thermal stability metrics. For example, the identified subset of scFvs may include thermally stable scFvs. Thus, one or more of the scFvs included in the identified subset of scFvs may be subsequently produced (e.g., manufactured).

[0038] As described above, multispecific biologics are becoming increasingly important for accessing novel therapeutically important pathways and mechanisms of action. One of the properties used to evaluate the potential development of multispecific biologics, scFv modules, or multispecifics containing scFv is thermal stability. Thermal stability is a property that can indicate stability under specific environmental conditions. Various factors, such as the specific amino acid sequence of the scFv and / or the resulting structure, can affect the thermal stability of the scFv. For example, interactions between amino acid residues, such as hydrophobic interactions and electrostatic interactions, can affect the thermal stability of the scFv. Additionally, or alternatively, the presence of specific bonds, such as disulfide bonds, can affect the thermal stability of the scFv.

[0039] The reason thermal stability is relevant is that thermostable scFvs can withstand specific environmental conditions and have appropriate adaptations to maintain scFv function under those conditions. For example, such thermostable scFvs can withstand exposure to temperatures within the temperature range that can actually occur. In contrast, scFvs with low thermal stability characteristics can unfold and denature within the temperature range that can actually occur and may lose enzyme activity.

[0040] Conventional approaches for optimizing the thermal stability of scFvs involve experimentally screening thermostable scFv candidates. However, experimental screening requires producing and performing experiments on each scFv within a large set of candidate scFvs to be screened to determine which scFvs have desirable thermal stability characteristics, consuming a large amount of resources, time, and cost.

[0041] The inventors understood that it is particularly useful to computationally screen scFvs for thermal stability. However, the inventors recognized that conventional computational methods for predicting the thermal stability of scFvs are unreliable and inaccurate. For example, some conventional computer techniques involve processing amino acid sequences using a machine learning model trained based on other amino acid sequences and their corresponding thermal stabilities. The machine learning model is trained to predict the thermal stability of an amino acid sequence by identifying similar amino acid sequences used in the training of the model. However, even when two amino acid sequences are highly similar, it does not necessarily mean that their thermal stabilities are the same. For example, two amino acids may be identical except for a single mutation. However, this mutation may cause a dramatic change in thermal stability. Therefore, predictions generated by such machine learning models can be inaccurate.

[0042] To predict thermal stability based on the total energy of a protein or fragment, other conventional computational techniques have been used. However, in such techniques, the structure of the scFv is not considered, which greatly affects thermal stability and results in inaccurate prediction results. In conventional techniques, the structure of the scFv is not considered, so information that could significantly change the resulting prediction of thermal stability is ignored. Therefore, these techniques are also unreliable and inaccurate.

[0043] The inventors recognized the need for an accurate computational method to predict the thermal stability of scFv candidates from their primary amino acid sequences and understood that such a method would guide efforts in thermal stability engineering and be very useful for the development of multispecific drugs. In particular, the inventors recognized that considering the structure of the scFv (e.g., 2D and / or 3D structure) enables more accurate prediction of thermal stability compared to the prior art. Specifically, the inventors developed techniques that consider residue-by-residue interactions when predicting thermal stability. These interactions provide information about the scFv structure and enable its consideration in prediction techniques (e.g., machine learning techniques).

[0044] Accordingly, the inventors developed machine learning techniques for predicting the thermal stability of single-chain variable fragments (scFvs). To determine the thermal stability of a particular scFv, these techniques derive features from the primary amino acid sequence of the particular scFv and provide those features as inputs to a trained machine learning model to generate a corresponding output (a "thermal stability metric") indicative of the thermal stability of the particular scFv. The output may be a measure of thermal stability such as the temperature at which the scFv is thermally stable (e.g., the temperature corresponding to half of the maximum binding of the scFv) or a temperature range that includes such a temperature. In some embodiments, the features provided as inputs to the trained machine learning model include only energy features, e.g., interaction energy metrics between pairs of residues within the residue sequence of a particular scFv. Additionally or alternatively, in some embodiments, the features may include an encoding of the residue sequence.

[0045] One application example of the machine learning techniques for predicting the thermal stability of scFvs is to computationally screen scFvs prior to manufacture to identify scFvs with favorable thermal stability characteristics. Accordingly, the machine learning techniques can be used to determine the thermal stability metric for each scFv included in a set of scFvs to be computationally screened and to identify a subset of the set of scFvs based on the determined thermal stability metrics. At least a portion of the scFvs thus identified can then be manufactured.

[0046] Accordingly, some embodiments provide a method for computationally screening a set of single-chain variable fragments (scFvs) based on the predicted thermal stability of the scFvs by a trained machine learning model, where the set of scFvs includes scFvs having different residue sequences, and the method includes: (A) using the trained machine learning model and at least one computer hardware processor to determine, for each scFv in the set of scFvs, a thermal stability metric for obtaining a plurality of thermal stability metrics, where the set of scFvs includes a first scFv having a first residue sequence, and determining includes: (i) using information indicating the three-dimensional (3D) structure of the first scFv to obtain an interaction energy metric for each of a plurality of pairs of residues in the first residue sequence; (ii) generating a first feature set to be provided as an input to the trained machine learning model, where generating includes including the interaction energy metric in the first feature set; and (iii) providing the first feature set as an input to the trained machine learning model to obtain a corresponding output indicating the first thermal stability of the first scFv. (B) identifying a subset of the set of scFvs for subsequent manufacture based on the plurality of thermal stability metrics; and (C) manufacturing at least one of the scFvs in the identified subset.

[0047] In some embodiments, the output indicating the first thermal stability of the first scFv indicates a first temperature at which the first scFv is thermally stable. The first temperature may be an estimated value of the temperature corresponding to half of the maximum binding of the first scFv (which may be referred to as the "TS50" temperature). In this way, the machine learning model can be configured to operate as a regression model.

[0048] In some embodiments, the output indicating the first thermal stability of the first scFv indicates a first temperature range that includes at least one temperature at which the first scFv is thermally stable. The first temperature range may be an estimated value of the temperature range that includes the temperature corresponding to half of the maximum binding of the first scFv. In some embodiments, providing the first feature set as an input to a machine learning model trained to obtain an output indicating the first thermal stability of the first scFv involves using the trained machine learning model to classify the first scFv into one of a plurality of classes using the first feature set, each of the plurality of classes corresponding to a respective temperature range, including classifying. In this way, the machine learning model can be configured to operate as a classification model.

[0049] In some embodiments, the residue pair interaction energy metric can be obtained in a two-step process in which (1) the residue sequence of the scFv is used to determine information indicating the 3D structure of the first scFv (e.g., using protein structure prediction software as exemplified herein), and (2) the information indicating the 3D structure of the first scFv is used to determine the interaction energy metric. (For example, by using molecular modeling software, an example of which is shown herein).

[0050] The inventors recognized that the way in which the interaction energy metric is provided as an input to a trained neural network model can affect the performance of that model in predicting thermal stability. In particular, the inventors recognized that arranging the interaction energy metric into a two-dimensional array or matrix generates a matrix with a local spatial structure suitable for analysis by a convolutional neural network model, and that in some embodiments, arranging the interaction energy metric in this way may improve performance (e.g., as opposed to providing a linear sequence interaction energy metric).

[0051] Thus, in some embodiments, generating the first feature set includes generating, for each particular energy metric, a respective two-dimensional matrix of values of the particular energy metric, where the rows and columns of the 2D matrix correspond to each residue within the first residue sequence. Accordingly, the entry in the i-th row and j-th column of the 2D matrix corresponds to the value of the particular energy metric between the i-th residue and the j-th residue of the first residue sequence.

[0052] In some embodiments, the interaction energy metrics between all pairs of residues of the scFv are provided as input to a trained machine learning model, while in other embodiments, only the interaction energy metrics between some pairs of residues are provided. For example, the interaction energy metrics between at least 50%, at least 60%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% of the pairs of residues may be provided as input to the trained machine learning model. In some such embodiments, the 2D matrix of interaction energy metrics includes rows for at least 50%, 60%, 75%, 85%, 90%, 95%, and 99% of the residues within the residue sequence of the first scFv.

[0053] In some embodiments, the sequence data may be provided as additional input (in addition to the interaction energy metrics) to the trained machine learning model. Thus, in some such embodiments, generating the first feature set further includes encoding (e.g., one-hot encoding) the first residue sequence to obtain an encoded sequence and including the encoded sequence in the first feature set.

[0054] In some embodiments, the trained machine learning model includes a trained neural network model, such as, for example, a convolutional neural network (CNN) model having a plurality of 2D convolutional layers. The CNN model can have a fully connected layer. Other types of architectures are possible (e.g., one or more dropout layers, one or more pooling layers, etc.).

[0055] In some embodiments, the trained CNN model is configured to output a plurality of probabilities that the scFv is thermally stable at each of a plurality of temperature ranges. These probabilities can be used to identify the single most likely temperature range (e.g., the temperature range associated with the maximum probability value), or, for example, to determine an estimated value of the identified temperature value as a weighted linear combination of the mean values of the temperature ranges weighted by the respective probabilities of the temperature ranges.

[0056] As used herein, the term "single-chain Fv" or "scFv" refers to a single polypeptide chain antibody fragment that includes the variable regions from both the heavy and light chains but lacks the constant regions. Typically, a single-chain Fv further includes a peptide linker between the VH and VL domains that enables the formation of a desired structure capable of binding to an antigen. Such constructs are discussed in detail in Pluckthun in The Pharmacology of Monoclonal Antibodies, vol. 113, Rosenberg and Moore eds., Springer-Verlag, New York, pp. 269-315 (1994), and U.S. Patent No. 7,112,324, entitled "CD 19xCD3 SPECIFIC POLYPEPTIDES AND USES THEREOF," each of which is incorporated herein by reference in its entirety. Structures that may be formed by an scFv include the bispecific T cell engager (BiTE) format structure (a fusion protein consisting of two single-chain variable fragments (scFvs) linked by a peptide linker) described in U.S. Patent No. 7,112,324, entitled "CD 19xCD3 SPECIFIC POLYPEPTIDES AND USES THEREOF," which is incorporated herein by reference in its entirety and begins at the N-terminus with a variable light chain (VL) domain following the variable heavy chain (VH) domain. The term "peptide linker" refers to an amino acid sequence that interconnects the amino acid sequences of one (variable and / or binding) domain of the scFv and the other (variable and / or binding) domain.Suitable peptide linkers include those described in U.S. Patent No. 4,751,180 entitled "EXPRESSION USING FUSED GENES PROVIDING FOR PROTEIN PRODUCT", U.S. Patent No. 4,935,233 entitled "COVALENTLY LINKED POLYPEPTIDE CELL MODULATORS", or International Publication Pamphlet No. WO 88 / 09344 entitled "TARGETED MULTIFUNCTIONAL PROTEINS", the entire contents of each of which are incorporated herein by reference.

[0057] It should be understood that the techniques described herein are applicable not only to any antibody construct having variable domains (VH and VL domains) that include or do not include a construct containing a peptide linker, but also to predicting the thermal stability of scFv. Further, the techniques described herein can be applied to multispecific constructs by utilizing a training data set of multispecific constructs having the same (or very similar, e.g., at least 80%, 85%, 90%, 95%, 99% similarity) sequence and / or structure. Further, the techniques described herein can be applied to any type of protein by utilizing a training data set of proteins having the same (or very similar, e.g., at least 80%, 85%, 90%, 95%, 99% similarity) sequence and / or structure.

[0058] The term "multispecific construct" refers to a molecule whose structure and / or function is based on the structure and / or function of an antibody, such as a full-length or complete immunoglobulin molecule, and / or is derived from the variable heavy (VH) and / or variable light (VL) chain domains of an antibody or a fragment thereof. The definition of the term "multispecific construct" includes monovalent, bivalent and multivalent / polyvalent constructs, thus bispecific constructs that specifically bind to only two antigenic structures, as well as multispecific / polyspecific constructs that specifically bind to more than two antigenic structures, such as three, four or more antigenic structures, via different binding domains. Furthermore, the definition of the term "multispecific construct" includes molecules consisting of only one polypeptide chain, such as a single scFv, as well as molecules consisting of two or more polypeptide chains, which may be identical (homo-dimers, homo-trimers or homo-oligomers) or different (hetero-dimers, hetero-trimers or hetero-oligomers). Examples regarding the molecules and variants or derivatives identified above are described, inter alia, in Harlow and Lane, Antibodies a laboratory manual, CSHL Press (1988) and Using Antibodies: a laboratory manual, CSHL Press (1999), Kontermann and Duebel, Antibody Engineering, Springer, 2nd ed. 2010 and Little, Recombinant Antibodies for Immunotherapy, Cambridge University Press 2009, each of which is incorporated herein by reference in its entirety.As further examples of forms of multi-specific constructs, there are included Fab fragments, which are monovalent fragments having VL, VH, CL, and CH1 domains; F(ab’)2 fragments, which are bivalent fragments having two Fab fragments linked by disulfide bridges in the hinge region; Fd fragment fragments having two VH and CH1 domains; Fv fragments having the VL domain and the VH domain of a single arm of an antibody; dAb fragments having a VH domain (Ward et al., (1989) Nature 341:544-546, which is incorporated herein by reference in its entirety), which have a VH domain, isolated complementarity determining regions (CDRs), and single-chain Fv (scFv).Examples of multispecific constructs are described, for example, in WO 00 / 006605 pamphlet entitled "HETEROMINIBODIES", WO 2005 / 040220 pamphlet entitled "MULTISPECIFIC DEIMMUNIZED CD3-BINDERS", WO 2008 / 119567 pamphlet entitled "CROSS-SPECIES-SPECIFIC CD3-EPSILON BINDING DOMAIN", WO 2010 / 037838 pamphlet entitled "CROSS-SPECIES-SPECIFIC SINGLE DOMAIN BISPECIFIC SINGLE CHAIN ANTIBODY", WO 2013 / 026837 pamphlet entitled "BISPECIFIC T CELL ACTIVATING ANTIGEN BINDING MOLECULES", WO 2013 / 026833 pamphlet entitled "BISPECIFIC T CELL ACTIVATING ANTIGEN BINDING MOLECULES", US Patent Application Publication No. 2014 / 0308285 entitled "HETERODIMERIC BISPECIFIC ANTIBODIES", US Patent Application Publication No. 2014 / 0302037 entitled "BISPECIFIC-FC MOLECULES", WO 2014 / 144722 pamphlet entitled "BISPECIFIC FC MOLECULES", WO 2014 / 151910 pamphlet entitled "HETERODIMERIC BISPECIFIC ANTIBODIES", and WO 2015 / 048272 pamphlet entitled "V-C-FC-V-C ANTIBODY", each of which is incorporated herein by reference in its entirety. Further, the multispecific construct can be a fragment of a full-length antibody such as VH, VHH, VL, (s)dAb, Fv, Fd, Fab, Fab’, F(ab’)2 or "r IgG" ("half antibody").

[0059] Furthermore, the technology described herein can be applied to constructs comprising modified fragments of antibodies, such as single-chain variable fragments (scFv), diabodies (di-scFv) or bispecific single-chain variable fragments (bi(s)-scFv), scFv-Fc, scFv-zipper, scFab, Fab2, Fab3, bispecific antibodies, single-chain bispecific antibodies, tandem bispecific antibodies (Tandab), tandem di-scFv, tandem tri-scFv, trispecific antibodies or quadraspecific antibodies, i.e., "multispecific antibodies", single-domain antibodies such as nanobodies, or single variable domain antibodies comprising only one variable domain that may specifically bind to an antigen or epitope independently of other V regions or domains, such as VHH, VH or VL, and human heavy chain antibodies UniAb® and UniDab®, described in WO 2020 / 206330 A1, the entire content of which is incorporated herein by reference.

[0060] In the context of the present invention, the term "binding domain" refers to a domain that (specifically) binds to / interacts with / recognizes a given target epitope or a given target site on a target molecule (antigen), such as CD33 and CD3, respectively. The structure and function of the first binding domain (recognizing, e.g., CD33), preferably the structure and / or function of the second binding domain (recognizing CD3), are also based on the structure and / or function of an antibody, e.g., a full-length or complete immunoglobulin molecule, and / or are derived from the variable heavy (VH) and / or variable light (VL) chain domains of an antibody or a fragment thereof. Preferably, the binding domain is characterized by the presence of three heavy-chain CDRs (i.e., CDR1, CDR2 and CDR3 of the VH region) following three light-chain CDRs (i.e., CDR1, CDR2 and CDR3 of the VL region). The binding domains used in the multispecific construct and in the training set need to be applied consistently.

[0061] Therefore, as can be understood from the foregoing, the technology described herein is not limited to application to single-chain variable fragments and can also be applied to other constructs described herein. For example, the technology described herein can be applied to monoclonal antibodies (mAbs). For example, the mAb sequence can be input as the VH sequence, followed directly by the VL sequence (without an intervening linker sequence). As another example, the mAb sequence can be input as the VL sequence, followed directly by the VH sequence (without an intervening linker sequence). The input sequence is used to predict the structure, which can be used to calculate an energy metric. Next, the energy metric (and optionally the encoding of the input sequence) can be provided as input to a trained machine learning model to obtain a thermal stability metric for the mAb. As another example, the technology described herein can be applied to one or more types of multispecific constructs of the types exemplified herein.

[0062] In another example, the technology described herein can be applied to any type of antibody. For example, the sequence of the antibody can be used to predict the structure, and the structure can be used to calculate an energy metric. Next, the energy metric (and optionally the encoding of the input sequence) can be provided as input to a trained machine learning model to obtain a thermal stability metric for the antibody. In yet another example, the technology described herein can be applied to any type of protein. For example, the protein sequence can be used to predict the structure, and the structure can be used to calculate an energy metric. Next, the energy metric (and optionally the encoding of the input sequence) can be provided as input to a trained machine learning model to obtain a thermal stability metric for the protein.

[0063] The technology described in this specification is not limited to a specific implementation method and can be implemented in various ways. Detailed examples of embodiments are provided for illustrative purposes only. Furthermore, the technology disclosed in this specification can be used individually or in any suitable combination because the aspects of the technology described herein are not limited to the use of a specific technology or combination.

[0064] FIG. 1A is a diagram of an exemplary technique 100 for determining a thermal stability index 110 of a single-chain variable fragment (scFv) by providing a feature 106 generated from a residue sequence 102 of the scFv as an input to a machine learning model 108.

[0065] In some embodiments, the scFv sequence 102 specifies the residue sequence of the scFv (e.g., the primary amino acid sequence). The scFv sequence can specify the amino acid sequences of the heavy and light chains of the scFv, as well as the linker peptide. The scFv sequence 102 can be of any suitable length. For example, the scFv sequence 102 can have 200 - 300 residues, 225 - 275 residues, or 236 - 254 residues, or any other suitable range within these ranges. In some embodiments, the linker peptide can be composed of 10 - 25 amino acids, and the scFv sequence 102 can include a sub-sequence of that length representing the amino acids within the linker of the scFv.

[0066] In some embodiments, the scFv sequence 102 may be specified by a user. For example, a user can interact with a user interface of a computing device (e.g., the computing device 120 shown in FIG. 1B) to specify the sequence of amino acid residues of the scFv sequence 102. In some embodiments, the scFv sequence 102 can be automatically specified by a computing device (e.g., the computing device 120). For example, the computing device can be programmed to generate the amino acid sequence of the scFv sequence 102 (e.g., by using machine learning to repeatedly change or randomly change amino acids in a predetermined order). This is useful, for example, when the computer is programmed to automatically generate a list of scFvs for computational screening, such as by substituting one or more amino acids in the starting scFv sequence to automatically generate mutations and obtaining candidate scFv sequences that may be evaluated during screening (e.g., based on predicted thermal stability). In some embodiments, the scFv sequence 102 may be generated at least partially automatically (e.g., by programmatically changing one or more amino acids in the sequence) and at least partially manually (e.g., based on user input specifying one or more amino acids in the sequence). The scFv sequence 102 can be specified in any suitable format (e.g., FASTA) since the aspects of the technology described herein are not limited in this regard.

[0067] Next, as part of the technology 100, scFv data 104 is generated from the scFv sequence 102. The scFv data 104 can include any type of data generated from the scFv sequence 102. In the exemplary embodiment shown in FIG. 1A, the scFv data 104 includes sequence data 104a, structural data 104b, and energy data 104c, which are described below. However, this example is illustrative, and it should be understood that in other embodiments, the scFv data 104 may include other suitable data generated from the scFv sequence 102 in addition to or instead of the types of data shown in FIG. 1A.

[0068] In some embodiments, the array data 104a includes the scFv sequence 102 itself or a sub-sequence thereof. Further, the array data 104a may include information regarding the sequence (e.g., amino acid statistics, length, sequence identifier, and / or other information related to the sequence). In some embodiments, the array data 104a is stored in a text-based file such as a FASTA file, and / or in other suitable formats, since the aspects of the techniques described herein are not limited in this regard.

[0069] In some embodiments, the structural data 104b includes information indicating the three-dimensional structure of the scFv sequence 102. The information indicating the 3D structure may include a description and / or annotation of the protein structure including atomic coordinates, secondary structure assignment, and / or atomic connectivity data. The structural data 104b can be in any suitable format for describing the 3D structure of the scFv (e.g., Protein Data Bank (PDB) file format, Crystallographic Information File (CIF) format, macromolecular Crystallographic Information File (mmCIF) format).

[0070] In some embodiments, the structural data 104b is obtained using the array data 104a. For example, obtaining the structural data 104b can include processing the array data 104a using protein structure prediction software (e.g., the protein structure prediction module 160 shown in FIG. 1B) to predict the 3D structure of the scFv. Techniques for obtaining information indicating the 3D structure of the scFv are described herein with respect to at least operation 252 of process 250 shown in FIG. 2B.

[0071] In some embodiments, the energy data 104c includes information indicating the energy levels in the 3D structure of the scFv residue sequence 102. For example, the energy data may include interaction energy metrics for one or more residue pairs within the scFv residue sequence 102. The interaction energy metrics may include the energies of various types of interactions between residue pairs. For example, the interaction energy metric can account for the energy of interactions between non-bonded atom pairs and the statistical potential used to describe the backbone and side-chain torsion preferences of the scFv. In some embodiments, the energy data 104c is included in a delimited text file, such as a comma-separated values (CSV) file, for example.

[0072] In some embodiments, the energy data 104c is obtained using the structural data 104b. For example, obtaining the energy data 104c can include processing the structural data 104b using molecular modeling software (e.g., using the molecular modeling module 162 shown in FIG. 1B) to predict the interaction energy metric for each of a plurality (e.g., some or all) of the residue pairs of the scFv sequence 102. Techniques for obtaining the interaction energy metric are described herein with respect to at least operation 254 of process 250 shown in FIG. 2B.

[0073] After obtaining the scFv data 104, the scFv data 104 is used to generate a feature set 106 of the scFv. In the embodiment of FIG. 1A, the feature set 106 includes an interaction energy metric 106a (e.g., included in the energy data 104c). In some embodiments, the interaction energy metric 106a includes an interaction energy metric for each pair of a plurality of residues of the scFv sequence 102. The interaction energy metric 106a may include interaction energy metrics for at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or all of the pairs of residues of the scFv sequence 102. In other embodiments, the interaction energy metric 106a may include interaction energy metrics for up to 99%, up to 98%, up to 95%, up to 90%, up to 85%, up to 80%, up to 75% or less of the pairs of residues of the scFv sequence 102. Thus, it should be understood that in some embodiments, the interaction energy metric 106a may include interaction energy metrics for a subset (i.e., not all) of the pairs of residues of the scFv sequence 102.

[0074] As shown in FIG. 1A, the feature set 106 optionally includes the encoded array 106b of the scFv. The encoded array 106b can be the encoding of the scFv array 102 (at least a sub-array or the whole). As an example, the encoded array 106b may be obtained by one-hot encoding of the scFv array 102. One-hot encoding is a technique for converting categorical data (each residue in the array is one of 20 possible amino acids) into numerical data (e.g., binary data, integer-valued data, real-valued data). The scFv array 102 may be one-hot encoded by being converted into a series of 20-dimensional vectors, one for each amino acid in the array. Each coordinate of the vector can correspond to one of the 20 amino acids. Then each amino acid can be encoded into a single 20-dimensional vector having 1 at the coordinate of that amino acid and 0 at the other coordinates. However, it should be understood that the aspects of the technology described herein are not limited to using one-hot encoding to encode the scFv array, and other methods for encoding categorical data can also be used.

[0075] In another embodiment, the feature set 106 may include only the encoded array data and may not include the interaction energy metric. In fact, the machine learning techniques described herein can be used to predict the thermal stability of the scFv using only energy features (e.g., the interaction energy metric), using only array features (e.g., the use of one-hot encoding of the scFv array), or using a combination of energy features and array features. As described herein, different machine learning models can be used to perform the prediction depending on the features utilized (e.g., a 2D convolutional neural network if the input includes the interaction energy metric, and a language model if it does not).

[0076] In addition to the array and energy features, the feature set 106 may include one or more additional or alternative features, it should be understood that this is because the aspects of the technology described herein are not limited in this regard. For example, one or more additional or alternative features may be obtained from the scFv data 104 and included in the feature set 106.

[0077] After the features 106 are generated, they are provided as input to the trained machine learning model 108, and the corresponding thermal stability indicators 110 of the scFv are obtained. The machine learning model can be of any suitable type. For example, the machine learning model can be a neural network model such as a convolutional neural network (CNN) model. The CNN model can have one or more convolutional layers (e.g., one or more two-dimensional convolutional layers). The CNN model can have a fully connected layer. As described herein, in some embodiments, the convolutional neural network model can be configured to receive only energy features (e.g., as a 1D or 2D matrix interaction energy metric) or a combination of energy features and array features (e.g., one-hot encoding of the scFv sequence) as input.

[0078] As another example, the machine learning model can be a pre-trained language model adapted to the protein setting. For example, a bidirectional encoder representation from transformers (BERT) model such as the ESM-1b or ESM-1v language model can be used. These types of models are described in A Rives, et al., Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc. Natl. Acad. Sci. 118(2021), and J Meier, et al., Language models enable zero-shot prediction of the effects of mutations on protein function. Adv. Neural Information Processing Systems 35(2021), each of which is hereby incorporated by reference in its entirety. As another example, the UniRep language model can be used, which is described in Ma, Eric J., and Arkadij Kummer. “Reimplementing Unirep in JAX.” bioRxiv(2020), which is hereby incorporated by reference in its entirety.

[0079] In some embodiments, the machine learning model may be formed as an ensemble of multiple machine learning models. For example, the machine learning model may include an ensemble of neural networks. In some embodiments, implementing an ensemble of machine learning models involves training each of a plurality of (e.g., two or more, three or more, etc.) machine learning models on different training datasets, predicting the thermal stability using each of the trained machine learning models, and averaging those predictions. In some embodiments, boosting can be used to combine multiple machine learning models.

[0080] Regardless of the specific type of the machine learning model 108 used as part of the technology 100, the machine learning model outputs an index of the thermal stability index 110. In some embodiments, the index may be the temperature or temperature range at which the scFv is thermally stable. For example, the index may be the temperature corresponding to half of the maximum binding of the scFv (sometimes referred to herein as the "TS50" temperature), or a temperature range including the TS50 temperature. As another example, the index may be the thermal melting temperature (sometimes referred to herein as the "T m " temperature) or a temperature range including the T m temperature. The T m temperature may correspond to the temperature at which the concentration of the folded scFv equals the concentration of the unfolded scFv.

[0081] In some embodiments, the machine learning model 108 is configured to output the probability that the scFv is thermally stable for each of a plurality of temperature ranges (e.g., less than 50 °C, 50 - 60 °C, 60 - 70 °C, greater than 70 °C). The ranges may be closed (e.g., 50 - 60 °C) or open (e.g., less than 50 °C or greater than 70 °C). As a non-limiting example, the machine learning model 108 can output a first probability that the scFv is thermally stable in a first temperature range and a second probability that the scFv is thermally stable in a second temperature range.

[0082] The technique shown in Figure 1A can be used to computationally screen a set of scFvs to identify a subset of scFvs to manufacture. For example, the set of scFvs may be screened based on the thermal stability index generated by the technology 100 for the scFvs within the set. As an example, when the thermal stability index is an index of temperature (e.g., TS50 temperature, Tm temperature), scFvs having a predicted temperature exceeding a specified threshold (e.g., exceeding 50 degrees Celsius (°C), exceeding 55 °C, exceeding 60 °C) can be selected for subsequent manufacture.

[0083] In some embodiments, the thermal stability metric 110 is used to identify scFvs for subsequent manufacture. For example, scFvs having a predicted thermal stability metric that meets one or more criteria can be identified for subsequent manufacture. Techniques for screening and manufacturing scFvs are described herein, at least with respect to FIG. 2A.

[0084] FIG. 1B is a block diagram of an exemplary system 150 for predicting the thermal stability of an scFv and computationally screening scFvs based on such predictions, according to some embodiments of the techniques described herein. System 150 includes a computing device 120 configured to execute software 130 to perform various functions related to predicting the thermal stability of an scFv and computationally screening scFvs based on such predictions.

[0085] In the embodiment shown in FIG. 1B, computing device 120 includes software 130 configured to perform various functions with respect to scFv data (e.g., scFv data 104).

[0086] Computing device 120 can be one or more computing devices of any suitable type. For example, computing device 120 can be a portable computing device (such as a laptop, smartphone, etc.) or a fixed computing device (such as a desktop computer, server, etc.). If computing device 120 includes multiple computing devices, the devices may be physically located in the same location (e.g., within a single room) or distributed across multiple physical locations. In some embodiments, computing device 120 may be part of a cloud computing infrastructure.

[0087] In some embodiments, computing device 120 may be operated by one or more users 172, such as one or more researchers and / or other individuals. For example, user 172 may provide scFv sequence 102 and / or scFv data 104 as input to computing device 120 (e.g., by uploading one or more files), and / or may provide user input specifying a process or other method to be performed on the scFv data.

[0088] As shown in FIG. 1B, software 130 includes a plurality of modules. Each module may include processor-executable instructions that, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to perform the functions of that module. Such modules may also be referred to herein as "software modules." The modules shown in FIG. 1B include processor-executable instructions that, when executed by a computing device, cause the computing device to perform one or more processes, such as the processes described herein with respect to at least FIGS. 2A-2B, 4A, and 5A-5B. The modules shown in FIG. 1B are exemplary, and it should be understood that in other embodiments, software 130 may be implemented using one or more other software modules in addition to, or instead of, the modules shown in FIG. 1B. In other words, software 130 may have an internal configuration different from that shown in FIG. 1B.

[0089] As shown in FIG. 1B, software 130 includes a plurality of software modules for processing scFv data, such as protein structure prediction module 160, molecular modeling module 162, feature generation module 164, thermal stability prediction module 166, and scFv screening module 168. In the embodiment of FIG. 1B, software 130 further includes a user interface module 170 for obtaining user input.

[0090] In some embodiments, the protein structure prediction module 160 obtains scFv data (e.g., scFv data 104) from the scFv data store 146 and / or the user 172 (e.g., by the user uploading scFv data). In some embodiments, the obtained scFv data includes the sequence data of the scFv (e.g., sequence data 104a).

[0091] In some embodiments, the protein structure prediction module 160 may be configured to predict the 3D structure of the scFv using the sequence data. For example, the protein structure prediction module 160 may be configured to generate the structural data of the scFv (e.g., structural data 104b) from the sequence of the scFv. For this purpose, the protein structure prediction module 160 can use protein structure prediction software such as, for example, DeepAb software, SAbPred software, or AlphaFold software. Techniques for obtaining information indicating the 3D structure of the scFv using the structure prediction software are described herein with respect to at least operation 252 of process 250 shown in FIG. 2B.

[0092] In some embodiments, the molecular modeling module 162 obtains scFv data (e.g., scFv data 104) from the scFv data store 146, the user 172 (e.g., by the user uploading scFv data), and / or the protein structure prediction module 160. In some embodiments, the obtained scFv data includes the sequence data (e.g., sequence data 104a) and / or the structural data (e.g., structural data 104b) of the scFv.

[0093] In some embodiments, the molecular modeling module 162 is configured to determine residue interaction energy metrics in the 3D structure of the scFv sequence (e.g., scFv sequence 102). For example, the molecular modeling module 162 can be configured to generate interaction energy metrics for the scFv (which may be part of, for example, energy data 104c). For this purpose, the molecular modeling module 162 can use molecular modeling software such as, for example, Rosetta software, Schrodinger BioLuminate®, or Chemical Computing Group's Molecular Operating Environment (MOE) software. Techniques for determining interaction energy metrics using molecular modeling software are described herein with respect to at least operation 254 of process 250 shown in FIG. 2B.

[0094] In some embodiments, the feature generation module 164 obtains scFv data (e.g., scFv data 104) from the scFv data store 146, the user 172 (e.g., by the user who uploads the scFv data), the molecular modeling module 162, and / or the protein structure prediction module 160, and uses the obtained scFv data to generate a feature set for each scFv. For example, the feature generation module 164 can generate a feature set for an scFv having the scFv sequence 102.

[0095] In some embodiments, the feature generation module 164 generates a feature set by including at least a portion of the acquired data (e.g., scFv data 104) in the feature set. For example, the feature generation module 164 can generate a feature set to include the interaction energy metric of the scFv. For example, the feature generation module 164 can generate a feature set to include a two-dimensional (2D) matrix that stores the value of a specific energy metric between the i-th and j-th residues of the scFv at the (i,j) position for each specific energy metric of the interaction energy metric. The 2D matrix generated in this way can be provided as an input to the trained neural network model. For example, as shown in the example of FIG. 8, if there are 20 interaction energy metrics, 20 such matrices may be generated and provided as inputs to the neural network model via 20 input channels. Further, or alternatively, the feature generation module 164 can generate a feature set that includes the encoded sequence of the scFv. For example, the sequence may be one-hot encoded. However, it should be understood that the feature generation module 164 may include additional features or alternative features in the feature set, because the aspects of the technology described herein are not limited in this regard. The technique for generating the feature set of the scFv is described herein with respect to at least operation 256 of process 250 shown in FIG. 2B.

[0096] In some embodiments, the thermal stability prediction module 166 obtains one or more feature sets from the feature generation module 164, obtains a trained machine learning model from the machine learning model data store 152 (which can be any suitable type of data store), and processes the obtained feature sets using the obtained machine learning model to obtain a thermal stability metric for one or more scFvs. For example, the thermal stability prediction module 166 can use the trained machine learning model 108 to process the feature set generated for the scFv having the sequence 102 and obtain the thermal stability metric 110 of the scFv. Techniques for predicting the thermal stability of scFvs using machine learning are described herein at least with respect to FIG. 2B.

[0097] In some embodiments, the scFv screening module 168 is used to computationally screen a set of scFvs to identify a subset for subsequent manufacture. For this purpose, the scFv screening module 168 obtains the thermal stability metric determined by the thermal stability prediction module 166 (e.g., by calling module 166 to determine the thermal stability metric of the scFvs in the set) and uses the thermal stability metric to identify scFvs for subsequent manufacture. For example, in some embodiments, the scFv screening module 168 compares the thermal stability metric to one or more criteria to determine whether the predicted thermal stability meets one or more criteria. If the thermal stability metric meets the criteria, the scFv screening module 168 can identify the scFv for which the thermal stability has been determined for subsequent manufacture. For example, the thermal stability metric may indicate, for each scFv, the respective temperature or temperature range at which the scFv is thermally stable. Its output is compared to a threshold temperature, and scFvs for which the indicated temperature is higher than the threshold pass the screening step and can be selected for subsequent manufacture.

[0098] In some embodiments, the scFv screening module 168 can perform computational screening based on user input, such as user input provided by user 172 via the user interface module 170. In the user input, one or more criteria (e.g., a threshold temperature) for passing an scFv through the screening can be specified. Further, or alternatively, the user can provide an input to manually select one or more scFvs for subsequent manufacture (e.g., based on determined thermal stability or other factors).

[0099] In some embodiments, the protein structure prediction module 160, the molecular modeling module 162, and / or the feature generation module 164 obtain scFv data via the user interface 170 and / or one or more other interface modules (not shown). The data may be provided by a communication network (not shown), such as the Internet or other suitable network, as the aspects of the technology described herein are not limited in this regard.

[0100] As shown in FIG. 1B, the system 150 also includes an scFv data store 146 and a machine learning model data store 152. In some embodiments, the software 130 obtains data from the scFv data store 146, the machine learning model data store 152, and / or the user 172 (e.g., by uploading the data). In some embodiments, the software 130 further includes a machine learning model training module 154 for training one or more machine learning models (e.g., stored in the machine learning model data store 152).

[0101] In some embodiments, the scFv data is obtained from the scFv data store 146. The scFv data store 146 may be of any suitable type (e.g., a database system, multi-file, flat file, etc.), may store the scFv data in any suitable manner and in any suitable format, as the aspects of the technology described herein are not limited in this regard. The scFv data store 146 may be part of the computing device 120 or external to it.

[0102] In some embodiments, the scFv data store 146 stores scFv data obtained for the scFv, at least as described herein with respect to FIG. 1A. In some embodiments, the stored scFv data may have been previously uploaded by a user (e.g., user 172), or uploaded from one or more public data stores and / or research. In some embodiments, a portion of the scFv data may be processed by the protein structure prediction module 160 to generate information indicative of the structure of the scFv. In some embodiments, a portion of the scFv data may be processed by the molecular modeling module 162 to determine the interaction energy of pairs of residues of the scFv. In some embodiments, a portion of the scFv data may be processed by the feature generation module 164 to generate a set of features of the scFv provided as input to a machine learning model. In some embodiments, a portion of the scFv data may be used to train one or more machine learning models (e.g., using the machine learning model training module 154).

[0103] In some embodiments, the thermal stability prediction module 166 obtains (either retrieves or is provided with) a trained machine learning model from the machine learning model data store 152. The machine learning model may be provided via a communication network (not shown), such as the Internet or other suitable network, as the aspects of the technology described herein are not limited to a particular communication network.

[0104] In some embodiments, the machine learning model data store 152 includes any suitable data store, such as a flat file, a data store, a multi-file, or any suitable type of data storage, because the aspects of the techniques described herein are not limited to a particular type of data store. The machine learning model data store 152 may be part of the software 130 (not shown) or may be excluded from the software 130 as shown in FIG. 1B.

[0105] In some embodiments, the machine learning model data store 152 stores one or more machine learning models used to predict the thermal stability of the scFv. The data store 152 may be of any suitable type (e.g., a database system, a multi-file, a flat file, etc.) and may store machine learning models trained in any suitable manner and in any suitable format, because the aspects of the techniques described herein are not limited in this regard. The data store 152 may be part of the computing device 120 or may be external to it.

[0106] In some embodiments, the machine learning model training module 154 (referred to herein as the training module 154) may be configured to train one or more machine learning models to predict the thermal stability of the scFv. In some embodiments, the training module 154 trains a machine learning model using a training set of scFv data. For example, the training module 154 can obtain training data from the scFv data store 146. In some embodiments, the training module 154 can provide the trained machine learning model to the machine learning model data store 152. Techniques for training machine learning models are described herein at least with respect to FIG. 4A.

[0107] In some embodiments, the predicted thermal stability can be output by the thermal stability prediction module 166. For example, the predicted thermal stability may be output to the user 172 via the user interface 170. Further, or alternatively, the predicted thermal stability may be stored in memory and / or transmitted to one or more other computing devices.

[0108] In some embodiments, the scFv identified for subsequent manufacture can be output by the scFv screening module 168. For example, the identified scFv may be output to the user 172 via the user interface 170. Further, or alternatively, the identified scFv may be stored in memory and / or transmitted to one or more other computing devices.

[0109] The user interface 170 is a graphical user interface (GUI), a text-based user interface, and / or any other suitable type of interface through which a user can provide input and view information generated by the software 130. For example, in some embodiments, the user interface can be a web page or web application accessible through an Internet browser. In some embodiments, the user interface can be the graphical user interface (GUI) of an app running on the user's mobile device. In some embodiments, the user interface may include a number of selectable elements with which the user can interact. For example, the user interface may include a drop-down list, check boxes, text fields, or other suitable elements.

[0110] Figures 2A-2B are flowcharts illustrating an exemplary process for computationally screening a set of scFvs to identify a subset of the set of scFvs to manufacture, according to some embodiments of the techniques described herein.

[0111] Figure 2A is a flowchart of an exemplary process 200 for computationally screening a set of scFvs according to some embodiments of the technology described herein. One or more operations of process 200 may be automatically performed by any suitable computing device. For example, the operations may be performed by a laptop computer, a desktop computer, one or more servers in a cloud computing environment, the computer system 2400 described herein with respect to FIG. 24, and / or any other suitable means. For example, in some embodiments, operation 202 may be automatically performed by any suitable computing device. As another example, operation 204 may be automatically performed by any suitable computing device.

[0112] Process 200 begins at operation 202, where a trained machine learning model is used to determine a thermal stability metric for each scFv within a set of scFvs. As described above, the thermal stability metric can refer to the temperature at which the scFv is stable or a temperature range that includes at least one temperature.

[0113] In some embodiments, determining the thermal stability metric of an scFv includes generating a feature set of the scFv and processing the feature set using a trained machine learning model. The output of the machine learning model may be something that indicates the thermal stability of the scFv (e.g., a thermal stability metric). For example, the output can indicate the temperature at which the scFv is thermally stable. That temperature can be a TS50 temperature, a Tm temperature, or any other type of temperature that indicates the scFv is thermally stable. Further, or alternatively, the output may indicate a temperature range that includes one or more temperatures at which the scFv is thermally stable. Techniques for determining the thermal stability metric of a particular scFv using a trained neural network model are described herein at least with respect to process 250 shown in FIG. 2B.

[0114] The set of scFvs can include any suitable number of scFvs. For example, the set of scFvs can include at least 25 scFvs, at least 50 scFvs, at least 75 scFvs, at least 100 scFvs, at least 200 scFvs, at least 300 scFvs, at least 400 scFvs, at least 500 scFvs, at least 600 scFvs, at least 700 scFvs, at least 800 scFvs, at least 900 scFvs, at least 1,000 scFvs, at least 5,000 scFvs, at least 10,000 scFvs, 100 - 1000 scFvs, 100 - 10,000 scFvs, or any other suitable range within these. Thus, determining the thermal stability index in operation 202 can include determining the thermal stability index at least 25, at least 50, at least 75, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1,000, at least 5,000, at least 10,000, between 100 - 1000, or between 100 - 10,000.

[0115] Next, process 200 proceeds to operation 204, where a subset of the set of scFvs is identified for subsequent manufacture based on the thermal stability index determined in operation 202. In some embodiments, identifying scFvs within the set of scFvs for subsequent manufacture includes identifying scFvs that are thermally stable and manufacturable.

[0116] In some embodiments, identifying a subset of scFvs includes determining whether the thermal stability metric meets one or more criteria. For example, scFvs can be identified whose thermal stability metric includes a temperature exceeding a particular threshold. As another example, scFvs can be included whose thermal stability metric includes a particular temperature range. If the thermal stability metric determined for a particular scFv meets one or more criteria, that scFv can be included in the subset of scFvs for subsequent manufacture. If the thermal stability metric does not meet one or more criteria, the scFv can be excluded from the subset of scFvs for subsequent manufacture.

[0117] As described above, in some embodiments, operation 204 can be performed using a computing device (e.g., computing device 120 shown in FIG. 1B). Additionally or alternatively, operation 204 can be performed by a user. For example, a user can manually select scFvs for subsequent manufacture based on the thermal stability metrics output by a trained machine learning model.

[0118] In some embodiments, the identified subset contains none, some, or all of the scFvs included in the original set of scFvs. For example, the identified subset can contain 0%, less than 10%, less than 25%, less than 50%, less than 75%, less than 90%, or all of the scFvs included in the original set of scFvs.

[0119] Next, process 200 proceeds to operation 206, where at least a portion of the scFvs included in the identified subset of scFvs are manufactured. In some embodiments, the scFvs are manufactured using techniques known in the art.

[0120] In some embodiments, manufacturing at least one scFv within a subset of scFvs includes manufacturing one, some, or all of the scFvs within the subset. For example, in some embodiments, at least 10%, at least 25%, at least 50%, at least 75%, at least 90%, or all of the scFvs included in the identified subset are manufactured in operation 206.

[0121] In some embodiments, implementing process 200 may include additional or alternative steps not shown in FIG. 2A. In some embodiments, process 200 may include only a subset of the operations included in the exemplary flowchart (e.g., only operation 202, only operations 202 and 204).

[0122] FIG. 2B is a flowchart of an exemplary process 250 for determining the thermal stability of a first scFv, according to some embodiments of the techniques described herein. In some embodiments, operation 202 of process 200 may be implemented using process 250. Process 250 can be executed by any suitable computing device (e.g., computing device 120 shown in FIG. 1B).

[0123] Process 250 begins at operation 252, where information indicative of the 3D structure of the first scFv is obtained. In some embodiments, this information has been previously obtained for the first scFv. Thus, in some embodiments, obtaining information indicative of the 3D structure of the first scFv may include accessing the information (e.g., from memory, via a network, via a file provided through a suitable interface, etc.).

[0124] In other embodiments, obtaining information indicative of the 3D structure of the first scFv includes generating this information. Thus, in some embodiments, obtaining information indicative of the 3D structure of the first scFv includes generating the information by processing the residue sequence of the first scFv using protein structure prediction software. The protein structure prediction software can be configured to output information indicative of the 3D structure of the first scFv. Any suitable protein structure prediction software can be used. For example, DeepAb software can be used, aspects of which are described in Ruffolo, Jeffrey A., Jeremias Sulam, and Jeffrey J. Gray. “Antibody structure prediction using interpretable deep learning.” Patterns 3.2 (2022): 100406, the entire content of which is incorporated herein by reference. As another example, SAbPred software can also be used, aspects of which are described in Dunbar, James, et al. “SAbPred: a structure-based antibody prediction server.” Nucleic acids research 44.W1 (2016): W474-W478, the full text of which is incorporated herein by reference. As yet another example, AlphaFold software can also be used, aspects of which are described in Jumper, Johns, et al. “Highly accurate protein structure prediction with AlphaFold.” Nature 596, 583-589 (2021), the full text of which is incorporated herein by reference.

[0125] Next, process 250 proceeds to operation 254, where an interaction energy metric is obtained for each of a plurality of pairs of amino acid residues of the first scFv. In some embodiments, obtaining the interaction energy metric includes generating the energy metric by processing information representing the 3D structure of the first scFv using molecular modeling software. The molecular modeling software can be configured to output the interaction energy metric of the first scFv. Any molecular modeling software capable of estimating residue interaction energy metrics can be used. For example, Rosetta molecular modeling software can be used. The Rosetta software and the techniques used to estimate residue interaction energy metrics are described in Alford, et al. (The Rosetta All-Atom Energy Function for Macromolecular Modeling and Design. 440 J. Chem. Theory Comput. 13, 3031-3048 (2017)), Chaudhury, et al. (PyRosetta: a script-based interface for implementing molecular modeling algorithms using Rosetta, Bioinformatics, 26(5), 689-691 (2010)), and Leaver-Fay, et al. (Rosetta3: An object-oriented software suite for the simulation and design of macromolecules, In Methods in Enzymology, 545-574), the entire texts of which are incorporated herein by reference. As another example, Schrodinger's BioLuminate® and / or Prime software can also be used (Schrodinger Suite 2020-2 release, Schrodinger Inc, New York, NY).As yet another example, the Molecular Operating Environment (MOE) software from Chemical Computing Group can also be used (Molecular Operating Environment (MOE), 2020.09 Chemical Computing Group ULC, 1010 Sherbrooke St. West, Suite #910, Montreal, QC, Canada, H3A 2R7, 2022).

[0126] In some embodiments, interaction energy metrics are obtained for some or all of the residue pairs of the residue sequence of the first scFv. For example, the interaction energy metric can be obtained for at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or all of the residue pairs of the amino acid sequence of the first scFv.

[0127] For pairs of residues, various types of interaction energy metrics can be obtained. For example, the interaction energy metric for a pair of residues can include one or more of the Lennard-Jones attractive and / or repulsive energy, and the van der Waals energy (e.g., "fa_atr": the attractive energy between two atoms on different residues separated by a distance d, "fa_rep": the repulsive energy between two atoms on different residues separated by a distance d, and / or "fa_intra-rep": the repulsive energy between two atoms on the same residue separated by a distance d). As another example, the interaction energy measurement criteria for a pair of residues can include one or more solvation energies (e.g., "fa_sol": the Gaussian exclusion implicit solvation energy between protein atoms of different residues, "fa_intra_sol": the Gaussian exclusion implicit solvation energy between protein atoms of the same residue, and / or "lk_ball_wtd": the orientation-dependent solvation of polar atoms assuming an ideal water shape). As another example, the interaction energy metric for a pair of residues, the electrostatic energy (e.g., "fa_elec": the interaction energy between two non-bonded charged atoms separated by a distance d). As another example, the interaction energy metric for a pair of residues can include one or more hydrogen bonds and / or disulfide bridge energies (e.g., "hbond_lr_bb": the energy of a long-range hydrogen bond, "hbond_sr_bb": the energy of a short-range hydrogen bond, "hbond_bb_sc": the energy of a backbone-side chain hydrogen bond, "hbond_sc": the energy of a side chain-side chain hydrogen bond, "dslf_fa13": the disulfide bridge energy).As another example, the interaction energy metric for pairs of residues can include one or more backbone statistics (e.g., "rama_prepro" - the probability of backbone angles φ and ψ given the amino acid type, "omega" - a backbone-dependent penalty for cis ω dihedral angles deviating from 0° and trans ω dihedral angles deviating from 180°, "p_aa_pp" - the probability of amino acid identity giving backbone φ and ψ angles, "pro_close" - a penalty for open proline rings and proline ω bond energy, "yhh_planarity" - a sine wave penalty for non-planar tyrosine χ3 dihedral angles). As another example, the interaction energy metric for pairs of residues can include knowledge-based rotamer energies (e.g., "fa_dun" - the probability that the selected rotamer is native-like given the backbone values). Thus, the interaction energy metric can include any (some or all) of the aforementioned examples of energy metrics. These aforementioned examples are further described in Alford, et al. (The Rosetta All-Atom Energy Function for Macromolecular Modeling and Design. 440 J. Chem. Theory Comput. 13, 3031 - 3048 (2017)), the entire text of which is incorporated herein by reference. In addition to, or instead of, one or more of the aforementioned energy metric examples, one or more other energy metrics can also be used.

[0128] Next, process 250 proceeds to operation 256, where a first feature set is generated for the first scFv and provided as input to a machine learning model trained with the first feature set, and a corresponding output indicating the first thermal stability of the first scFv is obtained. In some embodiments, generating the first feature set involves including at least a portion of the data obtained in operations 252 and / or 254 of process 250 in the first feature set.

[0129] For example, in operation 256a, the interaction energy metric is included in the first feature set. In some embodiments, the interaction energy metric includes some or all of the interaction energy metrics obtained in operation 254, and examples thereof are provided herein.

[0130] In some embodiments, including the interaction energy metric in the first feature set in operation 256a includes generating one or more matrices of the interaction energy metric. For example, for each specific energy metric, a respective 2D matrix can be generated where the entry in the i-th row and j-th column has the value of the specific metric for the i-th and j-th residues. If the interaction energy metric includes multiple metrics, a 2D matrix is generated for each metric and can thus be included as part of the first feature set generated in operation 256a of process 250. Thus, the 2D matrix can be provided as an input to a trained machine learning model (e.g., via different channels as shown in FIG. 8). Alternatively, a 3D matrix can be generated instead of multiple 2D matrices, since the aspects of the technology described herein are not limited in this regard.

[0131] Furthermore, or alternatively, in operation 256b, the encoded first residue sequence is included in the first feature set. In some embodiments, the encoding of the first residue sequence to obtain the encoded first residue sequence can be performed using any suitable encoding technique such as one-hot encoding. In some embodiments, the first residue sequence is encoded to form a specific input dimension. For example, the first residue sequence can be encoded to form an input dimension of (V H +V L +3)×21, where V H and V L correspond to the heavy chain sequence and the light chain sequence of the scFv residue sequence, respectively.

[0132] Although not shown in FIG. 2B, the first feature set may include one or more additional or alternative features, as it will be understood that the aspects of the technology described herein are not limited in this regard.

[0133] Next, process 250 proceeds to operation 258, where the first feature set is provided as input to the trained machine learning model, and an output indicating the thermal stability of the first scFv is obtained.

[0134] In some embodiments, the machine learning model can be of any suitable type. For example, the machine learning model can be a neural network such as a convolutional neural network (CNN). The CNN can include one or more two-dimensional convolutional layers, one or more three-dimensional convolutional layers, and / or fully connected layers.

[0135] As a specific example, the CNN can have an architecture as shown in FIG. 8, which includes one or more max pooling layers (Adaptive MaxPool 1D layer and Adaptive Max Pool 2D layer in FIG. 8), one or more two-dimensional convolutional layers, one or more non-linear layers (e.g., the rectified linear unit layer "ReLu" in FIG. 8), one or more batch normalization layers, and a dense connection layer. The architecture shown in FIG. 8 shows that both the energy metric and the one-hot encoded array are provided as inputs. However, as described herein, in some embodiments, only the energy metric or only the one-hot encoded array may be provided as input. Therefore, in some embodiments, either the upper branch or the lower branch of the architecture shown in FIG. 8, or both branches, can be used.

[0136] As another example, the machine learning model may be a pre-trained language model such as the Bidirectional Encoder Representations from Transformers (BERT) model (e.g., ESM-1b language model or ESM-1v language model) or UniRep language model described herein (used to perform zero-shot prediction or adjusted by adding a supervised head using transfer learning with a small amount of training data). Such models can be used when the input contains only array features and no energy features (e.g., as shown in FIG. 14).

[0137] In some embodiments, the machine learning model includes a plurality of machine learning models (e.g., an ensemble of machine learning models). For example, the machine learning model may include an ensemble of neural networks. In some embodiments, implementing an ensemble of machine learning models includes training each of a plurality of (e.g., two or more, three or more, etc.) machine learning models on different training data sets, predicting thermal stability using each of the trained machine learning models, and averaging (e.g., weighted averaging) those predictions. The outputs of the plurality of machine learning models can also be combined in any other way, since the aspects of the techniques described herein are not limited in this regard. An ensemble of machine learning models can be obtained in part by using any suitable bagging or boosting technique.

[0138] In some embodiments, the machine learning model is trained to process a first feature set to obtain the probability that the thermal stability of a first scFv belongs to each of a plurality of thermal stability classes. For example, this may include determining the probability that the first scFv is thermally stable at each of a plurality of temperature ranges. As a non-limiting example, the thermal stability classes may include the following four classes: less than 50°C, 50°C to 60°C, 60°C to 70°C, greater than 70°C.

[0139] In some embodiments, based on the determined probabilities, the machine learning model is configured to predict one of a plurality of classes associated with the highest probability. For example, this includes identifying the temperature range associated with the highest probability among a plurality of temperature ranges. The identified temperature range includes one or more temperatures at which the first scFv is thermally stable.

[0140] Further, or alternatively, based on the determined probabilities, the machine learning model may be configured to determine the temperature at which the first scFv is thermally stable. For example, the temperature can be the TS50 temperature. As another example, the temperature may correspond to the thermal melting temperature (Tm). In some embodiments, determining the temperature at which the first scFv is thermally stable includes determining, as the temperature, a weighted linear combination of temperature range averages weighted by the probabilities determined using the machine learning model.

[0141] In operation 260, process 250 includes determining whether another scFv exists within the set of scFvs for which thermal stability can be determined. If it is determined in operation 260 that another scFv for which thermal stability is to be determined exists, operations 252 - 258 are repeated for the other scFv. For example, in the case of the second scFv, this includes determining a second feature set and providing the second feature set as an input to the trained machine learning model to determine a thermal stability metric for the second scFv.

[0142] FIG. 3A is a diagram of an exemplary technique for computationally screening a set of scFvs using a machine learning model trained to generate a thermal stability metric for an scFv from an input including interaction energy metrics for pairs of residues within the scFv, according to some embodiments of the techniques described herein.

[0143] As shown in the embodiment of FIG. 3A, an exemplary technique 300 begins with a set 302 of scFvs. In some embodiments, the set 302 of scFvs is a candidate for manufacture. For example, the scFvs included in the set 302 may have one or more desirable properties (e.g., affinity, specificity, etc.). The set of scFvs can have any suitable size M, examples of which are provided herein, including the reference to FIG. 2A. In the example of FIG. 3A, the set 302 of scFvs includes a first scFv 302-1, a second scFv 302-2, and an Mth scFv 302-M.

[0144] Each scFv within the set 302 has its respective residue sequence. For example, the first scFv 302-1 has a first residue sequence, the second scFv 302-2 has a second residue sequence, and the Mth scFv 302-M has an Mth residue sequence. In some embodiments, the first residue sequence, the second residue sequence, and the Mth residue sequence are different from each other. For example, the sequences may differ from each other by one or more residues.

[0145] As shown in FIG. 3A, the exemplary technique 300 includes generating a feature set for each scFv included in the set 302 of scFvs. This includes, for example, generating a first feature set 304-1 for the first scFv 302-1, generating a second feature set 304-2 for the second scFv 302-2, and generating an Mth feature set 304-M for the Mth scFv 302-M. Techniques for generating the feature sets are described herein with respect to at least operation 256 of process 250 shown in FIG. 2B.

[0146] In some embodiments, the feature set generated for the scFv includes interaction energy metrics for each of a plurality of residue pairs within the scFv. For example, as shown in FIG. 3A, the first feature set 304-1 includes interaction energy metrics 322-1 for each of a plurality of residue pairs of the first residue sequence, the second feature set 304-2 includes interaction energy metrics 322-2 for each of a plurality of residue pairs of the second residue sequence, and the Mth feature set 304-M includes interaction energy metrics 322-M for each of a plurality of residue pairs of the Mth residue sequence. Techniques for obtaining interaction energy metrics for residue pairs are described herein with respect to at least FIG. 2B.

[0147] In the embodiment shown in FIG. 3A, after generating the feature set for each scFv within the set 302 of scFvs, the generated feature sets 304-1, 304-2, … 304-M are provided as inputs to the trained machine learning model 306.

[0148] In some embodiments, the machine learning model 306 is trained to predict a thermal stability metric for the scFv based on the feature set provided as input to the machine learning model 306. For example, the first feature set 304-1 is provided to the machine learning model 306, and an output 308-1 indicating the thermal stability of the first scFv 302-1 can be obtained. The second feature set 304-2 is provided as input to the machine learning model 306, and an output 308-2 indicating the thermal stability of the second scFv 302-2 can be obtained. The Mth feature set 304-M is provided as input to the machine learning model 306, and an output 308-M indicating the thermal stability of the Mth scFv 302-M can be obtained. Techniques for using a machine learning model to predict the thermal stability of an scFv are described herein with respect to at least FIG. 2B.

[0149] In some embodiments, exemplary technique 300 includes identifying a subset 310 of a set 302 of scFvs based on determined thermal stability metrics (e.g., first thermal stability 308-1, second thermal stability 308-2, and Mth thermal stability 308-M). In some embodiments, identifying the scFvs included in subset 310 includes identifying scFvs having a thermal stability that meets one or more criteria. This may include, for example, comparing the thermal stability to a threshold temperature and identifying scFvs having a thermal stability that exceeds the threshold. The identified scFvs may have a thermal stability suitable for manufacturing. Techniques for identifying scFvs for subsequent manufacturing are described herein with respect to at least operation 204 of process 200 shown in FIG. 2A.

[0150] In some embodiments, the subset of scFvs includes one or more of the scFvs included in the original set 302 of scFvs. As shown in the embodiment of FIG. 3A, the identified subset 310 includes N scFvs. For example, subset 310 includes first scFv 302-1, second scFv 302-2, and Nth scFv 302-N. In some embodiments, N is less than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of M, or less than 100%. In some embodiments, N is equal to M.

[0151] In some embodiments, at least one of the scFvs included in the subset of scFvs is subsequently manufactured. For example, one, some, or all of the scFvs included in subset 310 may be subsequently manufactured.

[0152] FIG. 3B is a diagram of an exemplary technique for computationally screening a set of scFvs using a machine learning model trained to generate a thermal stability metric for an scFv from an input that includes the interaction energy metric of pairs of residues within the scFv and features representing the sequence of the scFv, according to some embodiments of the techniques described herein.

[0153] As shown in the embodiment of FIG. 3B, exemplary technique 340 begins with a set 302 of scFvs. An exemplary set of scFvs is described herein at least with respect to FIG. 3A.

[0154] In some embodiments, exemplary technique 340 includes generating a set of features for each scFv included in set 302 of scFvs. This includes, for example, generating a first set of features 344-1 for a first scFv 302-1, generating a second set of features 344-2 for a second scFv 302-2, and generating an Mth set of features 344-M for an Mth scFv 302-M. Techniques for generating sets of features are described herein at least with respect to operation 256 of process 250 shown in FIG. 2B.

[0155] In some embodiments, a set of features includes interaction energy metrics for each of a plurality of residue pairs of each respective scFv. For example, as shown in FIG. 3B, a first set of features 344-1 includes interaction energy metrics 352-1 for each of a plurality of residue pairs of a first residue sequence, a second set of features 344-2 includes interaction energy metrics 352-2 for each of a plurality of residue pairs of a second residue sequence, and an Mth set of features 344-M includes interaction energy metrics 352-M for each of a plurality of residue pairs of an Mth residue sequence. Techniques for obtaining interaction energy metrics for pairs of residues are described herein at least with respect to FIG. 2B.

[0156] Furthermore, in the embodiment shown in FIG. 3B, the feature set includes the encoded residue sequences of the respective scFvs. For example, the first feature set 344-1 includes the encoded sequence 354-1 of the first scFv 302-1, the second feature set 344-2 includes the encoded sequence 354-2 of the second scFv 302-2, and the M-th feature set 344-M includes the encoded sequence 354-M of the M-th scFv 302-M. Techniques for obtaining the encoded residue sequences are described herein at least with respect to FIG. 2B.

[0157] In the embodiment shown in FIG. 3B, after generating the feature sets for each scFv within the set 302 of scFvs, the exemplary technique 340 includes providing the generated feature sets as an input to a trained machine learning model 346.

[0158] In some embodiments, the machine learning model 346 is trained to predict a thermal stability metric of the scFv based on the feature sets provided as an input to the machine learning model 346. For example, the first feature set 344-1 is provided to the machine learning model 346, and an output 348-1 indicating the thermal stability of the first scFv 302-1 can be obtained. The second feature set 344-2 is provided as an input to the machine learning model 346, and an output 348-2 indicating the thermal stability of the second scFv 302-2 can be obtained. The M-th feature set 344-M is provided as an input to the machine learning model 346, and an output 348-M indicating the thermal stability of the M-th scFv 302-M can be obtained. Techniques for using a machine learning model to predict the thermal stability of an scFv are described herein at least with respect to FIG. 2B.

[0159] In some embodiments, exemplary technique 340 includes identifying a subset 350 of the set 302 of scFvs based on determined thermal stability metrics (e.g., first thermal stability 348-1, second thermal stability 348-2, and Mth thermal stability 348-M). Exemplary techniques for identifying a subset of scFvs are described herein with respect to at least FIG. 3A.

[0160] FIG. 3C is a diagram of an exemplary technique for computationally screening a set of scFvs using a machine learning model trained to generate thermal stability metrics of the scFvs from an input that includes features representing the sequences of the scFvs, according to some embodiments of the techniques described herein.

[0161] As shown in the embodiment of FIG. 3C, exemplary technique 360 begins with a set 302 of scFvs. An exemplary set of scFvs is described herein with respect to at least FIG. 3A.

[0162] In some embodiments, exemplary technique 360 includes generating a set of features for each scFv included in the set 302 of scFvs. This includes, for example, generating a first set of features 364-1 for a first scFv 302-1, generating a second set of features 364-2 for a second scFv 302-2, and generating an Mth set of features 364-M for an Mth scFv 302-M. Techniques for generating the set of features are described herein with respect to at least operation 256 of process 250 shown in FIG. 2B.

[0163] In some embodiments, the feature set includes the encoded residue sequences of the respective scFvs. For example, the first feature set 364-1 includes the encoded sequence 372-1 of the first scFv302-1, the second feature set 364-2 includes the encoded sequence 372-2 of the second scFv302-2, and the M-th feature set 364-M includes the encoded sequence 372-M of the M-th scFv302-M. Techniques for obtaining the encoded residue sequences are described herein with respect to at least FIG. 2B.

[0164] In the embodiment shown in FIG. 3C, after generating the feature sets for each scFv within the set 302 of scFvs, the exemplary technique 360 includes providing the generated feature sets as an input to a trained machine learning model 366.

[0165] In some embodiments, the machine learning model 366 is trained to predict the thermal stability index of the scFv based on the feature sets provided as input to the machine learning model 366. For example, the first feature set 364-1 is provided to the machine learning model 366, and an output 368-1 indicating the thermal stability of the first scFv302-1 can be obtained. The second feature set 364-2 is provided as an input to the machine learning model 366, and an output 368-2 indicating the thermal stability of the second scFv302-2 can be obtained. The M-th feature set 364-M is provided as an input to the machine learning model 366, and an output 368-M indicating the thermal stability of the M-th scFv302-M can be obtained. Techniques for using a machine learning model to predict the thermal stability of an scFv are described herein with respect to at least FIG. 2B.

[0166] In some embodiments, exemplary technique 360 includes identifying a subset 370 of the set 302 of scFvs based on determined thermal stability metrics (e.g., a first thermal stability 368-1, a second thermal stability 368-2, and an Mth thermal stability 368-M). Exemplary techniques for identifying a subset of scFvs are described herein, at least with respect to FIG. 3A.

[0167] FIG. 4A is a flowchart of an exemplary process 400 for training a machine learning model to generate thermal stability metrics for scFvs, according to some embodiments of the techniques described herein. Process 400 can be executed by any suitable computing device. For example, the process may be executed on a laptop computer, a desktop computer, one or more servers within a cloud computing environment, the computer system 2400 described herein with respect to FIG. 24, or other suitable means. In some embodiments, software modules such as the machine learning model training module 154 described herein with respect to FIG. 1B include processor-executable instructions that, when executed by a computing device, cause the computing device to execute process 400.

[0168] Process 400 begins at operation 402, where thermal stability metrics for the scFvs are determined experimentally. In some embodiments, determining the thermal stability metrics experimentally includes manufacturing the scFvs and analyzing the manufactured scFvs to determine the experimental thermal stability metrics. This may include, for example, experimentally determining, as thermal stability metrics, the temperature corresponding to half of the maximum binding of the scFv (TS50) and / or the thermal melting temperature (Tm). Techniques for manufacturing the scFvs and experimentally determining the thermal stability metrics are described herein, at least with respect to FIG. 4B.

[0169] Next, process 400 proceeds to operation 404 to obtain information indicating the 3D structure of the scFv using the residue sequence of the scFv. Techniques for obtaining information indicating the 3D structure of the scFv are described herein with respect to at least operation 252 of process 250 shown in FIG. 2B.

[0170] Next, process 400 proceeds to operation 406 to obtain an interaction energy metric for pairs of residues in the residue sequence of the scFv using the information on the 3D structure of the scFv. Techniques for obtaining the interaction energy metric are described herein with respect to at least operation 254 of process 250 shown in FIG. 2B.

[0171] When the machine learning model incorporates a particular type of feature, process 400 includes generating that type of feature and using it for training. The type of feature can include, for example, only energy features, only sequence features, or a combination of energy features and sequence features. For example, in operation 408, process 400 includes generating a feature set for the scFv. Generating the feature set can include, in operation 408a, including the interaction energy metric for pairs of residues in the scFv residue sequence in the feature set. Further, or alternatively, generating the feature set can include, in operation 408b, including the encoded residue sequence of the scFv in the feature set. Techniques for generating a feature set for the scFv are described herein with respect to at least operation 256 of process 250 shown in FIG. 2B.

[0172] In some embodiments, when the machine learning model incorporates sequence features, training the machine learning model can include providing the sequence features to the machine learning model in a particular order. For example, as described herein, the sequence features may include an encoded residue sequence. In the case of an scFv, the encoded scFv sequence may be provided to the machine learning model as the encoded VH sequence followed by the encoded VL sequence ( "VH-VL"), or as the encoded VL sequence followed by the encoded VH sequence ( "VL-VH").

[0173] In some embodiments, the machine learning model can be trained using training data that all contain arrays in the same order. For example, each array in the training data can be composed of an encoded VH array followed by an encoded VL array. As another example, each array in the training data may be composed of an encoded VL array followed by an encoded VH array.

[0174] During training, if all array data are provided to the machine learning model in the same order (e.g., all VH-VL or all VL-VH), new data can also be provided in the same order. If the new array data are provided in a different order, since the machine learning model has not been trained with data provided in that order, the performance may degrade in processing the new data. For example, if the machine learning model is trained with scFV arrays specified only in the VH-VL order, the performance may degrade in processing new arrays provided in the VL-VH order. In fact, for scFv, the VH-VL molecule can be physically different from the VL-VH molecule. By training the machine learning model with scFv arrays specified only in one order, the machine learning model may lose the opportunity to learn about the fundamental physical differences of scFv molecules specified in other orders. As a result, the machine learning model may not perform sufficiently well in processing scFv arrays provided in other orders. In contrast, in the case of an antibody, since VH and VL can be separate chains, whether the model regards it as VH-VL or VL-VH is purely a matter of notation, and the underlying physical entity can be the same. Therefore, the performance difference may be due only to the input order. As a result, when processing an antibody, the performance difference due to the input order may not be as large as when processing scFV.

[0175] In some embodiments, training a machine learning model includes providing array data to the machine learning model in both orders. For example, the array data may be provided in both the order in which an encoded VH array follows an encoded VL array and the order in which an encoded VL array follows an encoded VH array. During training, if the array data is provided to the machine learning model in both orders (e.g., both VH-VL and VL-VH), new data can be provided to the machine learning model in both orders. Since the machine learning model is trained using the array data provided in both orders, its performance can be consistent regardless of the order in which new array data is provided.

[0176] Next, process 400 proceeds to operation 410, where a machine learning model is trained using the thermal stability metric determined by the experiment determined in operation 402 and the feature set generated in operation 408.

[0177] In some embodiments, training the machine learning model in operation 410 includes estimating the parameters of the machine learning model from the training data. In some embodiments, the estimation may be performed iteratively (e.g., using iterative gradient descent techniques). In some embodiments, the estimation can be performed using optimization software to adjust the parameters. For example, in some embodiments, an ADAM optimizer may be used, and the details thereof are described in Kingma, Diederik P., and Jimmy Ba. “Adam: A Method for Stochastic Optimization.” Proceedings of the 3rd International Conference on Learning Representations, ICLR (2015), which is hereby incorporated by reference in its entirety. In some embodiments, estimating the parameters includes estimating the weights of the connections.

[0178] In some embodiments, any suitable hyperparameters may be used during the training in operation 410. The hyperparameters can be set manually or by any other suitable method. Non-limiting examples of hyperparameters include the number of layers, batch size, number of filters, kernel size, epochs, pooling size, learning rate, and the like. However, it should be understood that other suitable hyperparameters may be used during the training in operation 410.

[0179] Figure 4B is a diagram of an exemplary technique 450 for experimentally generating data used to train a model for predicting the thermal stability of scFv452, according to some embodiments of the techniques described herein. As shown, the scFv is produced, for example, at 454 in an E. coli culture. The cells are then lysed (e.g., using freeze / thaw cycles), and the scFv is extracted and cultured at 456. After culturing, the lysate can be cultured with target transfected cells (e.g., CHO cells), and the bound scFv can be detected and analyzed by flow cytometry or other suitable techniques. In some embodiments, the results of the analysis are used to estimate the thermal stability at 458. For example, this can include estimating the TS50 value and / or the T m value of the scFv. Exemplary techniques for producing an scFv and experimentally estimating the thermal stability of the scFv are described herein with respect to at least the sections of "Generation of scFv", "TS50 Screening Assay", and "nanoDSF Method".

[0180] The techniques described herein are not limited to application to single-chain variable fragments and can also be applied to other constructs described herein. For example, as described above, the techniques described herein can be applied to any type of antibody. The antibody sequence is used to predict the structure, and the structure can be used to calculate an energy metric. The energy metric (and optionally the encoding of the input sequence) can then be provided as an input to a trained machine learning model to obtain a thermal stability indicator for the antibody.

[0181] Figure 5A is a flowchart of an exemplary process 500 for computationally screening a set of antibodies according to some embodiments of the techniques described herein. One or more operations of process 500 may be automatically performed by any suitable computing device. For example, the operations may be performed by a laptop computer, a desktop computer, one or more servers in a cloud computing environment, the computer system 2400 described herein with respect to FIG. 24, and / or other suitable means. For example, in some embodiments, operation 502 may be automatically performed by any suitable computing device. As another example, operation 504 may be automatically performed by any suitable computing device.

[0182] Process 500 begins at operation 502, where a trained machine learning model is used to determine a thermal stability metric for each antibody in a set of antibodies. As described above, the thermal stability metric can refer to the temperature at which the antibody is stable or a temperature range that includes at least one temperature.

[0183] In some embodiments, determining the thermal stability metric of an antibody includes generating a feature set of the antibody and processing the feature set using a trained machine learning model. The output of the machine learning model may be something that indicates the thermal stability of the antibody (e.g., a thermal stability metric). For example, the output can indicate the temperature at which the antibody is thermally stable. That temperature can be a TS50 temperature, a Tm temperature, or any other type of temperature that indicates the antibody is thermally stable. Further, or alternatively, the output may indicate a temperature range that includes one or more temperatures at which the antibody is thermally stable. Techniques for determining the thermal stability metric of a particular antibody using a trained neural network model are described herein at least with respect to process 550 shown in FIG. 5B.

[0184] The set of antibodies can include any suitable number of antibodies. For example, the set of antibodies can include at least 25 antibodies, at least 50 antibodies, at least 75 antibodies, at least 100 antibodies, at least 200 antibodies, at least 300 antibodies, at least 400 antibodies, at least 500 antibodies, at least 600 antibodies, at least 700 antibodies, at least 800 antibodies, at least 900 antibodies, at least 1,000 antibodies, at least 5,000 antibodies, at least 10,000 antibodies, 100 - 1,000 antibodies, 100 - 10,000 antibodies, or any other suitable range within these ranges. Thus, determining the thermal stability index in operation 502 can include determining the thermal stability index for at least 25, at least 50, at least 75, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1,000, at least 5,000, at least 10,000, between 100 and 1,000, or between 100 and 10,000.

[0185] Next, process 500 proceeds to operation 504, where a subset of the set of antibodies is identified for subsequent manufacture based on the thermal stability index determined in operation 502. In some embodiments, identifying the antibodies within the set of antibodies for subsequent manufacture includes identifying antibodies that are thermally stable and manufacturable.

[0186] In some embodiments, identifying the subset of antibodies includes determining whether the thermal stability index meets one or more criteria. For example, antibodies can be identified whose thermal stability index includes a temperature above a specific threshold. As another example, antibodies can be included whose thermal stability index includes a specific temperature range. If the thermal stability index determined for a particular antibody meets one or more criteria, that antibody can be included in the subset of antibodies for subsequent manufacture. If the thermal stability index does not meet one or more criteria, the antibody can be excluded from the subset of antibodies for subsequent manufacture.

[0187] As described above, in some embodiments, operation 504 may be performed using a computing device (e.g., computing device 120 shown in FIG. 1B). Additionally, or alternatively, operation 504 may be performed by a user. For example, a user may manually select subsequent antibodies for production based on the thermal stability metrics output by a trained machine learning model.

[0188] In some embodiments, the identified subset may contain none, some, or all of the antibodies included in the original set of antibodies. For example, the identified subset may contain 0%, less than 10%, less than 25%, less than 50%, less than 75%, less than 90%, or all of the antibodies included in the original set of antibodies.

[0189] Next, process 500 proceeds to operation 506, where at least a portion of the antibodies included in the identified subset of antibodies are produced. In some embodiments, the antibodies are produced using techniques known in the art.

[0190] In some embodiments, producing at least one of the antibodies within the subset of antibodies includes producing one, some, or all of the antibodies within the subset. For example, in some embodiments, at least 10%, at least 25%, at least 50%, at least 75%, at least 90%, or all of the antibodies included in the identified subset are produced in operation 506.

[0191] In some embodiments, implementing process 500 may include additional or alternative steps not shown in FIG. 5A. In some embodiments, process 500 may include only a subset of the operations included in the exemplary flowchart (e.g., only operation 502, only operations 502 and 504).

[0192] Figure 5B is a flowchart of an exemplary process 550 for determining the thermal stability of a first antibody according to some embodiments of the technology described herein. In some embodiments, operation 502 of process 500 may be implemented using process 550. Process 550 can be executed by any suitable computing device (e.g., computing device 120 shown in FIG. 1B).

[0193] Process 550 begins at operation 552, where information indicating the 3D structure of the first antibody is obtained. In some embodiments, this information has been previously obtained for the first antibody. Thus, in some embodiments, obtaining information indicating the 3D structure of the first antibody may include accessing the information (e.g., from memory, via a network, via a file provided through a suitable interface, etc.).

[0194] In other embodiments, obtaining information indicating the 3D structure of the first antibody includes generating this information. Thus, in some embodiments, obtaining information indicating the 3D structure of the first antibody includes generating the information by processing the residue sequence of the first antibody using protein structure prediction software. The protein structure prediction software may be configured to output information indicating the 3D structure of the first antibody. Any suitable protein structure prediction software can be used. Examples of protein structure prediction software are described herein with respect to at least operation 252 of FIG. 2B.

[0195] Next, process 550 proceeds to operation 554, where an interaction energy metric is obtained for each of a plurality of pairs of amino acid residues of the first antibody. In some embodiments, obtaining the interaction energy metric includes generating the energy metric by processing information representing the 3D structure of the first antibody using molecular modeling software. The molecular modeling software can be configured to output the interaction energy metric of the first antibody. Any molecular modeling software that can estimate the residue interaction energy metric can be used. An example of molecular modeling software is described herein with respect to at least operation 254 of FIG. 2B.

[0196] In some embodiments, the interaction energy metric is obtained for some or all of the pairs of residues of the residue sequence of the first antibody. For example, the interaction energy metric can be obtained for at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or all of the pairs of residues of the amino acid sequence of the first antibody.

[0197] For pairs of residues, various types of interaction energy metrics can be obtained. Examples of interaction energy metrics are described herein with respect to at least operation 254 of FIG. 2B.

[0198] Next, process 550 proceeds to operation 556, where a first feature set is generated for the first antibody and provided as an input to a machine learning model trained in operation 558, and a corresponding output indicative of the first thermal stability of the first antibody is obtained. In some embodiments, generating the first feature set involves including at least a portion of the data obtained in operations 552 and / or 554 of process 550 in the first feature set.

[0199] For example, in operation 556a, the interaction energy metric is included in the first feature set. In some embodiments, the interaction energy metric includes some or all of the interaction energy metrics obtained in operation 554, examples of which are provided herein.

[0200] In some embodiments, including the interaction energy metric in the first feature set in operation 556a includes generating one or more matrices of the interaction energy metric. For example, for a particular energy metric, a respective 2D matrix can be generated where the entry in the i-th row and j-th column has the value of the particular metric for the i-th and j-th residues. If the interaction energy metric includes multiple metrics, a 2D matrix is generated for each metric and can thus be included as part of the first feature set generated in operation 556a of process 550. Thus, the 2D matrix can be provided as an input to a trained machine learning model (e.g., via different channels as shown in FIG. 8). Alternatively, a 3D matrix can be generated instead of multiple 2D matrices, since the aspects of the techniques described herein are not limited in this regard.

[0201] Furthermore, or alternatively, in operation 556b, the encoded first residue sequence is included in the first feature set. In some embodiments, encoding of the first residue sequence to obtain the encoded first residue sequence can be performed using any suitable encoding technique such as one-hot encoding. In some embodiments, the first residue sequence is encoded to form a particular input dimension. For example, the first residue sequence can be encoded to form an input dimension of (V H +V L +3)×21, where V H and V L correspond to the heavy chain sequence and the light chain sequence of the antibody residue sequence, respectively.

[0202] Although not shown in FIG. 5B, the first feature set may include one or more additional or alternative features, as it will be understood that the aspects of the technology described herein are not limited in this regard.

[0203] Process 550 proceeds to operation 558, where the first feature set is provided as input to the trained machine learning model, and an output indicating the thermal stability of the first antibody is obtained.

[0204] In some embodiments, the machine learning model can be of any suitable type. For example, the machine learning model can be a neural network such as a convolutional neural network (CNN). The CNN can include one or more two-dimensional convolutional layers, one or more three-dimensional convolutional layers, and / or fully connected layers.

[0205] As a specific example, the CNN can have an architecture as shown in FIG. 8, which includes one or more max pooling layers (Adaptive MaxPool 1D layer and Adaptive MaxPool 2D layer in FIG. 8), one or more two-dimensional convolutional layers, one or more non-linear layers (e.g., the rectified linear unit layer "ReLu" in FIG. 8), one or more batch normalization layers, and a densely connected layer. The architecture shown in FIG. 8 shows that both the energy metric and the one-hot encoded sequence are provided as input. However, as described herein, in some embodiments, only the energy metric or only the one-hot encoded sequence may be provided as input. Thus, in some embodiments, either the upper branch or the lower branch of the architecture shown in FIG. 8, or both branches, can be used.

[0206] As another example, the machine learning model may be a pre-trained language model such as the Bidirectional Encoder Representations from Transformers (BERT) model (e.g., the ESM-1b language model or the ESM-1v language model) or the UniRep language model described herein (used to perform zero-shot prediction or adjusted by adding a supervised head using transfer learning with a small amount of training data). Such models can be used when the input contains only array features and no energy features (e.g., as shown in FIG. 14).

[0207] In some embodiments, the machine learning model includes a plurality of machine learning models (e.g., an ensemble of machine learning models). For example, the machine learning model may include an ensemble of neural networks. In some embodiments, implementing an ensemble of machine learning models includes training each of a plurality of (e.g., two or more, three or more, etc.) machine learning models on different training data sets, predicting thermal stability using each of the trained machine learning models, and averaging (e.g., weighted averaging) those predictions. The outputs of the plurality of machine learning models can also be combined in any other way, since the aspects of the techniques described herein are not limited in this regard. An ensemble of machine learning models can be obtained in part by using any suitable bagging or boosting technique.

[0208] In some embodiments, the machine learning model is trained to process a first set of features to obtain the probability that the thermal stability of a first antibody belongs to each of a plurality of classes of thermal stability. For example, this may include determining the probability that the first antibody is thermally stable at each of a plurality of temperature ranges. As a non-limiting example, the classes of thermal stability may include the following four classes: less than 50°C, 50°C to 60°C, 60°C to 70°C, greater than 70°C.

[0209] In some embodiments, based on the determined probabilities, the machine learning model is configured to predict one of a plurality of classes associated with the highest probability. For example, this includes identifying the temperature range associated with the highest probability among a plurality of temperature ranges. The identified temperature range includes one or more temperatures at which the first antibody is thermostable.

[0210] Further, or alternatively, based on the determined probabilities, the machine learning model may be configured to determine the temperature at which the first antibody is thermostable. For example, the temperature can be the TS50 temperature. As another example, the temperature may correspond to the thermal melting temperature (Tm). In some embodiments, determining the temperature at which the first antibody is thermostable includes determining a weighted linear combination of temperature range averages weighted by the probabilities determined using the machine learning model as the temperature.

[0211] In operation 560, process 550 includes determining whether there is another antibody within the set of antibodies that can be determined for thermostability. If it is determined in operation 560 that there is another antibody for which thermostability is to be determined, operations 552 - 558 are repeated for the other antibody. For example, in the case of a second antibody, this includes determining a second set of features and providing that second set of features as an input to the trained machine learning model to determine a thermostability indicator for the second antibody.

Example

[0212] Machine learning techniques for predicting the thermostability of scFv have been developed. In some embodiments, the developed machine learning techniques utilize various features generated for scFv to predict its thermostability. FIG. 6 is a diagram showing an example of an scFv602, a possible set of features 604 generated for the scFv602, and a machine learning model 606 trained to predict the thermostability of the scFv602.

[0213] FIG. 7 is a diagram of an exemplary technique for training and using various machine learning models to predict the thermal stability of scFvs on a dataset (e.g., a labeled TS50 dataset) 702. Branch 704 shows transfer learning using a teacherless network such as a pre-trained language model (PTLM), which can be used to predict thermal stability by zero-shot prediction and fine-tuned prediction. Branch 706 shows a teacher model such as a neural network architecture (e.g., a supervised convolutional neural network (CNN)) trained to predict the thermal stability of scFvs using features derived from the scFvs. Both types of models can be used to predict thermal stability 708, computationally validate experimental designs 710, or for other suitable purposes, since the aspects of the techniques described herein are not limited in this regard.

[0214] A. Supervised Neural Network FIG. 8 is a diagram showing an exemplary architecture of a supervised convolutional neural network (CNN) trained to predict the thermal stability of scFvs using sequence and / or energy features generated for the scFvs. The parameters of the exemplary model were estimated with an ADAM optimizer with categorical cross-entropy (CCE) loss and a learning rate of 10 -3 The model was trained using the datasets described herein, including at least the "Datasets" section.

[0215] As shown in FIG. 8, the input scFv array 802 is processed using protein structure prediction software 804 (e.g., DeepAb, AlphaFold, etc.), and information indicating the 3D structure of the scFv is generated. The structural information is used to evaluate the thermodynamic characteristics 806 (total energy divided into monomeric i-I and dimeric i-j, residue energies) of each scFv using the Rosetta ref2015 energy function. The contribution for each j-th residue of the i-th residue (j ∈ N, N = total number of residues) can be tabulated in an i-j matrix that constitutes the energy feature.

[0216] Furthermore, or alternatively, as shown in FIG. 8, the input scFv array 802 is encoded (e.g., one-hot encoded) to obtain array features 810.

[0217] The energy feature 806 and the array feature 810 are respectively fixed-length embeddings of size LxL and L, where L represents the maximum array length in the dataset, V H and V L are respectively the maximum lengths of the heavy and light chains. Those of the input scFv array 802 that are less than L are zero-padded.

[0218] The energy feature 806 and the array feature 810 are provided as inputs to two parallel branches of the model. Specifically, the energy feature 806 is provided as an input to the 2D convolutional layer 808, and the array feature 810 is provided as an input to the 1D convolutional layer 812. For example, the array and energy features can pass through their respective convolutional layers using batch normalization and ReLU activation.

[0219] As shown in FIG. 8, the sequence input is transformed and concatenated with the energy input. The concatenated matrix passes through another 2D convolutional layer 814, is flattened, and supplied to the dense layer 816, where the logits for each class are output. The class probabilities 818 can be obtained by performing a softmax function on the logits.

[0220] In some embodiments, the class probability 818 represents the probability that the thermal stability of the scFv (e.g., the TS50 measurement or the T m measurement) corresponds to a particular temperature range.

[0221] To estimate the predicted thermal stability (e.g., the TS50 value or the Tm value), the probability is weighted using the average thermal stability (e.g., the average TS50 value or the average Tm value) of each class.

Number

[0222] Further, or alternatively, the technique may include classifying the scFv into the class corresponding to the highest probability output by the machine learning model. The identified class can correspond to a temperature range that includes at least one temperature at which the scFv is thermally stable.

[0223] In some embodiments, the architecture shown in FIG. 8 can be used to predict thermal stability using only one of the energy feature 806 and the sequence feature 810 as an input. In this case, the architecture of the model does not have to be changed. Rather, a tensor of zeros can be provided as an input instead of one of the features. For example, to use only the energy feature 806 to predict thermal stability, a tensor of zeros can be passed to the upper branch of the architecture shown in FIG. 8 as opposed to the sequence feature 810. Similarly, to use only the sequence feature 810 to predict thermal stability, a tensor of zeros can be passed to the lower branch of the architecture shown in FIG. 8 as opposed to the energy feature 806.

[0224] Experiments were conducted to evaluate the performance of various machine learning models (e.g., those with the architecture shown in Figure 8) in predicting the thermal stability in the datasets described herein, including those related to the "Dataset" section. The "Energy theory only" model was configured to predict thermal stability using only energy features (e.g., energy feature 806) as input. The "Array only" model was configured to predict thermal stability using only array features (e.g., array feature 810) as input. The "Energy theory + array" model was configured to predict thermal stability using both array features and energy features (e.g., energy feature 806 and array feature 810).

[0225] The output of the model was used to evaluate whether the experimental sets from which the scFv was derived affected the prediction accuracy. By projecting the embeddings from the dense connection layer of each array into two dimensions via t-distributed stochastic neighbor embedding (t-SNE), the representations learned by the array-only model and the energy-theory-only model can be analyzed. Figure 9 shows the t-SNE generated for each model. The embeddings of the array-only model were clustered by experimental set, as revealed by the aggregation of the shaded points in Figure 9. The embeddings of the energy-theory-only model were independent of clustering based on the experimental set, as demonstrated by the noisy embeddings of the energy theory. Therefore, despite continuously diverse datasets, fine-tuned supervised models trained only with array features can infer the underlying experimental origin of the array and distort the prediction of thermal stability, making generalization to new blind datasets difficult. In contrast, the energy-theory-only model is more generalizable to new blind datasets.

[0226] The performance of the energy-only model was evaluated by constructing receiver operating characteristic (ROC) curves derived from predictions in the over-70 °C class. Figure 10 shows the constructed ROC curves. The ROC was evaluated for four test data sets. Two holdout data sets (Set P and Set Q) representing the test antibody (test Ab) and the separated scFv (separated scFv), and two blind data sets. The area under the ROC exceeds 0.7, indicating high classification accuracy.

[0227] Figure 11 shows a graph comparing the performance of various machine learning models for predicting the thermal stability of scFv. In particular, Figure 11 shows the correlation coefficients for all four test data sets for the energy-only, sequence-only, and energy + sequence models, respectively. In the holdout data sets, the coefficient of the energy-only model exceeds 0.5, and the energy + sequence model shows equally improved performance. However, in the blind data sets, the performance of energy + sequence and sequence-only (coefficient less than 0.1) deteriorates. The energy-only model still shows a relatively high correlation for the blind data sets (0.2 and 0.4, respectively).

[0228] As a control, the weights were randomly initialized in the S-CNN for the classification task. The results showed that the sequences could not be distinguished based on thermal stability. Furthermore, in the test set, weighted random predictions were performed. That is, the class labels were predicted by weighted random selection with the sample size of each class as the weight. In both of these tests, the energy-only S-CNN was able to elucidate some relationship between the energy of the scFv and its thermal stability. The randomly initialized model could not show a distinguishable relationship indicating the importance of the representation learned from the supervised data.

[0229] As described herein, Figure 8 shows one exemplary teacher - aware CNN architecture that can be used to predict the thermal stability of sequences and / or energy features. In the embodiment shown in Figure 8, a 2D - CNN classification model is used for energy features. However, there may be one or more alternative ways to input energy features, such as 1D flattened inputs or 2D inputs with energy values in absolute residue units. These architectures were tested to inform the architecture selection and their performance was compared. The results are shown in Figure 12A. In the case of energy theory only, the 2D - CNN with classified inputs had improved performance compared to the other two architectures.

[0230] Furthermore, or alternatively, to reduce the variance in the performance of the model, the performance of multiple models can be ensembled to average their predictions and generate an ensemble of CNNs. Three teacher - aware CNNs with the same architecture were trained on different dataset splits. The experimental sets were shuffled and three different pairs of training data and hold - out data were obtained. Figures 12B and 12C show that the performance when predicting thermal stability using an ensemble of CNNs is improved compared to the performance of non - ensembled CNNs.

[0231] To test the control, a randomly initialized CNN was employed and the embeddings generated by this randomized model were compared to the ensemble of CNNs. As shown, the model developed by the inventors can better separate sequences based on thermal stability characteristics even when the sequence size and training diversity are limited.

[0232] The performance of the CNN ensemble was also tested on a blind dataset, i.e., scFv arrays separated from the test scFv, with respect to a weighted random prediction model. In this case, the randomization was biased by the weights of the sample sizes of each class observed from the training dataset. This sample size was used as the probability by a random number generator to predict the class of the array. The data points are limited, non-uniform, and highly skewed, but FIGS. 13A - 13B show that the thermal stability predictions obtained using the machine learning techniques developed by the inventors can be used to observe the trends in the thermal stability of scFv, in contrast to the weighted random predictions. In particular, FIGS. 13A - 13B show a confusion matrix with the probabilities of predictions for each class highlighted. The predictions for the top class are biased towards higher temperature regions in the machine learning predictions, as opposed to the weighted random predictions. This means that when predicting the blind arrays, considering the top class, i.e., the class above 70, there is a higher probability of actually selecting arrays that are thermally stable, i.e., arrays located in the top two classes, 60 - 70 or above 70. This is useful for removing redundant and potentially low thermal stability arrays and is important because using a properly curated training set may help a simple supervised network to make robust design estimates.

[0233] B. Pre-trained Language Model A pre-trained language model (PTLM) was evaluated to assess its ability to predict the thermal stability of scFv arrays. FIG. 14 is a diagram showing an exemplary technique 1400 for predicting the thermal stability of scFv using a PTLM, according to some embodiments of the techniques described herein. The pre-trained language model 1402 was evaluated to assess its ability to predict thermal stability using zero-shot prediction 1404 and fine-tuned prediction 1406.

[0234] Three pre-trained language models were evaluated. The first model, UniRep, is an mLSTM with 1900 hidden units pre-trained on the Pfam database. Multiple sequence alignments (MSAs) were collected for each sequence in the TS50 set, which is the "evotuning" approach described by A Bateman et al. in "The pfam protein families database." Nucleic acids research (2004), which is hereby incorporated by reference in its entirety. The sequences were combined into a single dataset, and the model was pre-trained on this set of evolutionarily related sequences. The implementation used is described in Ma, Eric J., and Arkadij Kummer. "Reimplementing UniRep in JAX." bioRxiv (2020), which is hereby incorporated by reference in its entirety.

[0235] Furthermore, transformer models for both ESM-1b and ESM-1v were considered. Both are 33-layer, 650M parameter transformer models, pre-trained using masked language modeling on the Uniref database. ESM-1b was trained on a dataset filtered at 50% sequence identity (Unired50), and ESM-1v was trained on a dataset filtered at 90% sequence identity (Unifer90). The ESM-1b transformer model is described in A Rives, et al., “Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences.” Proc. Natl. Acad. Sci. (2021), the entirety of which is incorporated herein by reference. The ESM-1v transformer model is described in J Meier, et al., “Language models enable zero-shot prediction of the effects of mutations on protein function.” Adv. Neural information processing systems (2021), the entirety of which is incorporated herein by reference.

[0236] i. Zero-shot evaluation FIG. 15A is a diagram showing an exemplary pre-trained language model configured to perform zero-shot thermal stability prediction using sequence data, according to some embodiments of the techniques described herein.

[0237] One approach to predicting thermal stability using a pre-trained language model is to directly use the likelihood or pseudo-likelihood of the model. Sequences that are more likely based on the model are predicted to be more thermally stable. UniRep models the probability of each residue in a residue sequence given all preceding residues. As a result, the likelihood of a sequence can be efficiently evaluated as follows.

Number

[0238] ESM-1v models the probability of masked residues given the unmasked residues. The pseudo-likelihood of the sequence can be obtained as follows.

Number

[0239] In practice, the log-likelihood and pseudo-log-likelihood are evaluated for numerical stability. Additionally, ESM-1v is provided as an ensemble of five models trained with different random seeds. The predictions from all five models are averaged to obtain the final pseudo-log-likelihood.

[0240] ii. Fine-tuned evaluation FIG. 15B shows an exemplary fine-tuned pre-trained language model configured to perform a prediction of thermal stability using sequence data, according to some embodiments of the techniques described herein.

[0241] For the UniRep model, the final hidden state is obtained as a fixed-length vector representation of 3900 hidden units, along with the average of the previous hidden states. For the ESM-1b model, the representation for each residue is projected down to 4 dimensions and then concatenated. As a result, a fixed-length embedding of size 4L is generated. Here, L is the maximum sequence length of the TS50 dataset. If the length of the sequence is less than L, zeros are padded.

[0242] These embeddings pass through a linear layer with 512 hidden dimensions, followed by a tanh activation and are passed to a final layer that predicts class logits. The parameters of the UniRep and ESM-1b models are fixed during training. The parameters of the head model (including the initial down-projection of ESM-1b) are trained with the Adam optimizer and a learning rate of 10 -3It is trained.

[0243] For TS50 data, the model is trained on all targets except one target and evaluated on the holdout target. For data other than TS50, predictions are made using an ensemble of TS50 models (one per holdout target).

[0244] iii. Results Figures 15C, 15D, 15E, and 15F are graphs showing that for some embodiments described herein, the fine-tuned predictions achieve improved correlation with thermal stability compared to zero-shot predictions.

[0245] In particular, Figures 15C and 15E show that zero-shot predictions generally do not correlate well with thermal stability, whether in the TS50 dataset or the blind test set described in the "Dataset" section. In contrast, Figure 15D shows that fine-tuned predictions from both ESM-1b and UniRep achieved moderate to high average Spearman correlations (0.63 and 0.45) on the holdout targets when trained on the TS50 dataset. However, as shown in Figure 15E, these predictions did not generalize well to the blind test set. This suggests that there is some underlying structure in the sequences within the TS50 dataset that the model can utilize to make predictions, but it has not generalized to the new dataset.

[0246] C. Comparison of Machine Learning Model Performance Experiments were conducted to evaluate the ability of the supervised models and pre-trained learning models described herein to distinguish between thermostability mutations and heat denaturation mutations.

[0247] The thermal aggregation experiments regarding point mutations of the anti-VEGF antibody (PDB ID: 2FJG / 2FJF) are described in detail in the studies, P Koenig et al., “Mutational landscape of antibody variable domains reveal a switch modulating the interdomain confirmational dynamics and antigen binding,” Proc. Natl. Acad. Sci. United States Am. (2017), and S Warszawski et al., “Optimizing antibody affinity and stability by the automated design of the variable light-heavy chain interfaces,” PLoS Comput. Biol. (2019), each of which is hereby incorporated by reference in its entirety. In both studies, deep mutational scanning (DMS) experiments were performed on the antibody, and for point mutations that improved binding enrichment over the wild type, the fragment antigen binding (Fab) melting temperature (T m ) was analyzed. These point mutants (20 mutations collected from both studies) serve as test cases to evaluate whether a network trained on TS50 temperature measurements can gain insights into related temperature-dependent properties such as thermal aggregation, and whether it can distinguish between mutations that enhance heat and those that impede heat.

[0248] To determine whether the prediction model described here has potential in protein design, a computational DMS was created for the anti-VEGF antibody (PDBID: 2FJG). Each residue position within the sequence was mutated to 19 other amino acids to obtain mutant sequences. Each sequence was one-hot encoded to obtain sequence data, and an energy theory dataset was generated (e.g., by the techniques described herein, at least with respect to FIG. 2B). This sequence and energy input were input into the model, and point mutants classified into over 70 temperature classes by machine learning prediction were cross-validated with experimental results.

[0249] FIG. 16 is a diagram showing that predictions of thermal stability determined using machine learning techniques according to some embodiments of the techniques described herein are consistent with experimentally determined thermal stability. The spheres shown in FIG. 16 represent mutations verified by experiments that improved the thermal stability of the scFv (e.g., T m ). The starred spheres indicate that the machine learning technique accurately predicted the mutations and residue positions that produced the most thermostable scFv. The lightly shaded spheres indicate that the machine learning technique accurately predicted the residue positions but did not predict the mutations, resulting in the production of the most thermostable scFv. The ring-shaped spheres represent mutations not observed using the machine learning technique.

[0250] As shown, the CNN was able to correctly identify 5 out of 20 mutations. Furthermore, for 18 out of 20 mutations, although different amino acid mutations were predicted to be the most thermostable, the CNN was able to correctly identify the residue positions. Of the 4,540 point mutations analyzed (N res =227 residues, 20 amino acids per residue), experimental data was available for only 20 point mutations. Since the melting temperature was experimentally evaluated for only 0.44% of all possible mutations of the anti-VEGF antibody, the validation dataset for thermal stability is insufficient. Furthermore, despite being an attribute specific to temperature, TS50 and T mare measurement values from different experiments and are not exactly correlated. Therefore, it is notable that the CNN was able to predict the positions of thermostable residues in 90% of the cases, and 25% of them were successfully predicted (correct residue positions as well as amino acid residues). When extrapolating the network trained with TS50 measurements to an alternative heat-intensive experiment (in this case, T m ), it has been demonstrated that the intrinsic heat characteristics can be captured by such a model.

[0251] Furthermore, as shown in FIGS. 17A - 17B, when comparing the positions of residues that violate the germline consensus sequence of anti - VEGFAb, different amino acid mutations are observed, emphasizing that the machine learning techniques described herein have the ability to provide mutations that are orthogonal to conventional germline approaches. With more diverse and large - scale training datasets, more robust models can be developed. The results suggest that these networks may function as useful tools for screening or filtering antibody sequences in a temperature - specific antibody design pipeline.

[0252] FIGS. 18A - 18B are graphs showing that, according to some embodiments of the techniques described herein, the improvement in the correlation with thermostability has been achieved for the thermostability predictions output by the supervised convolutional neural network compared to the thermostability predictions by the unsupervised pre - trained language model.

[0253] FIG. 19 shows a graph comparing the performance of various machine learning models described herein for predicting the thermostability of scFv. As shown, training and prediction using the supervised CNN model are generally more accurate than training and prediction using the pre - trained language model.

[0254] Figure 20 includes a graph showing that training and prediction based on residue pair interaction energy metrics according to some embodiments of the techniques described herein are more accurate than training and prediction based on encoded sequences, and training and prediction based on both encoded sequences and residue pair interaction energy metrics.

[0255] Figures 21A and 21B are graphs showing that training and prediction based on residue pair interaction energy metrics according to some embodiments of the techniques described herein are more accurate than training and prediction based on residue pair interaction energy metrics and encoded sequences.

[0256] D. Dataset To learn temperature-specific context patterns within sequence data, a machine learning model for predicting thermal stability using scFv sequences was developed and trained. Temperature data was collected from various antibody engineering studies for the development of thermostable scFv antibodies. The sequence data included scFv sequences assembled by introducing mutations into the heavy and light chains of multiple germline sequences. The sequence data was constructed by collating 2,700 scFv sequences from 17 germline sequences (hereinafter also referred to as the experimental set). Additionally, an scFv dataset separated from sequences from another scFv study (currently under test) forms a blind test set.

[0257] For each sequence, thermal stability was evaluated using a TS50 measurement that represents the temperature at half the maximum of target binding, and this measurement functions as a temperature annotation. The TS50 data can also be divided into four classes. For example, the TS50 data can be split into less than 50°C, 50°C - 60°C, 60°C - 70°C, and greater than 70°C.

[0258] Since the experimental dataset was not uniform, it may be biased towards the high-temperature classes (i.e., the classes of 60 °C to 70 °C and above 70 °C). The distributions of the training, validation, and test datasets are shown in Figure 22. In the representation of array data, the longer the bar, the higher the consensus is indicated. V H (heavy chain) and V L (light chain), the Gly4 / Ser linker region between them is clear. For the training of the machine learning model, the heavy chain sequence and the light chain sequence were separated from the linker. Figure 23 highlights the temperature distribution of the TS50 measurements to show the skewed nature of the experimental dataset.

[0259] Generation of i.scFv (G4S)3 linker-containing scFv was cloned as a single construct into a pTT vector with a puromycin selection marker. The construct was transfected into the mammalian CHO-K1 cell line and stably expressed at a 4 mL scale. On the 21st day after transfection, the VCD and viability were measured, and the expression level of the secreted protein in the conditioned medium was analyzed by non-reducing SDS PAGE gel. The cells were cultured overnight with magnetic beads conjugated with either proA (for scFv with lambda variable domain) or proL (for scFv with kappa variable domain). The beads were separated from the cell culture medium and then washed 3 times with PBS and 2 times with water. The scFv was eluted from the magnetic beads using a low pH buffer (100 mM glycine, pH 2.7) and neutralized with 3M Tris (pH 11). To determine the melting point of the purified substance, differential scanning fluorimetry (DSF) was performed. Briefly, the molecule was heated at 1.0 °C / min in a nanoDSF instrument. The change in tryptophan fluorescence was monitored to evaluate the unfolding and aggregation of the protein. T m is reported as the midpoint between the start of unfolding and the fully unfolded state.

[0260] ii. TS50 screening assay The thermal stability of the scFv was screened by determining the loss of target binding after high-temperature stress. For this purpose, soluble scFv (VH-(G4S)3-VL) containing a C-terminal FLAG tag (DYKDDDDK) and a 6xHis tag was produced in Escherichia coli TG1 (Agilent, Santa Clara, USA) in 10 mL of LB culture medium. The protein product was induced with 1 mM IPTG. Next, the bacteria were centrifuged and the cell pellet was resuspended in 1 mL of Gibco (registered trademark) DPBS. The cells were lysed by 4 freeze / thaw cycles and residual cells and cell debris were removed by 2 centrifugation steps. 100 μL of these crude extracts were transferred to 0.2 mL tubes and exposed to different temperatures (4 °C, 50 °C, 60 °C, 70 °C) for 5 minutes in a water bath. After incubation, the tubes were transferred directly onto ice and CHO cells transfected with the human target were cultured with 50 μL of the lysate. The bound scFv was detected and analyzed by flow cytometry. The median fluorescence intensity was determined and plotted. The temperature corresponding to half of the maximum binding of each scFv was calculated (TS50). The scFv sequences were further classified into sets (not randomly) based on the ID of the antigen they bind. Since the sets were not curated for the machine learning task, they were not evenly distributed among the sets.

[0261] iii. nanoDSFT m Method The thermal melting temperature (T m ) was determined by performing a Trp shift study on a Prometheus NT.48. A thermal gradient of 1.0 °C / min was applied with a starting temperature of 25 °C and a stop temperature of 95 °C. The unfolding was measured at a fluorescence ratio of 350 nm / 330 nm. The data analysis and the determination of T m were performed using PRThermControl v2.0.4. The samples were normalized to 1.0 mg / mL in the formulation buffer before Tm analysis.

[0262] iv. Curating the Dataset for the Supervised Model To generate the alignment input, a dataset of TS50 measurements of scFvs from all experimental sets was aggregated to form a single dataset. The scFv sequences were composed of heavy and light chains linked by a glycine-serine (G4S) x linker. The dataset was created using the scFv sequences, split into their respective heavy and light chain sequences, and instead of classifying the sequences based on thermal stability, the TS50 measurements were incorporated. The distribution of scFv sequences across the experimental sets and the test dataset is shown in Figure 23. Sets P and Q were removed along with the sequences of scFvs separated from the test antibodies to form the holdout set. The amino acid sequences were one-hot encoded to form an input of dimension (V H +V L +3)×21. Here, V H and V L correspond to the heavy and light chain sequences respectively. The additional tokens to the one-hot encoding of the amino acids correspond to the delimiter characters for the start and end positions of the scFv sequence and the delimiter between the heavy and light chains indicating a chain break.

[0263] To obtain the energy-profile input, the sequences were first passed through a structural module, namely the DeepAb protocol for protein structure prediction. For each predicted structure, Rosetta Relax and an improved protocol (XML script in the supplementary materials) for side-chain repacking were executed. The energy estimation in Rosetta starts with an energy relaxation step (Rosetta Relax) that places constraints on the starting coordinates to reduce steric clashes and ensures that the accuracy of the backbone structure (predicted by DeepAb) does not degrade. The all-atom model was further refined with 4 cycles of side-chain packing to obtain a robust structure, and the lowest energy structure was selected for further calculations. For each refined model, the residue interaction energy metric was estimated using a residue energy decomposition application. The one-body energy and two-body energy were converted into a two-dimensional i-j matrix that serves as energy-profile information for training the supervised CNN model.

[0264] The energy values of the ij matrix were further grouped together into 20 classes between the lower and upper energy limits of |-25, 10| REU each. Additional classes were added for the start token, end token, and chain break token respectively. The dimension of the paired energy data is L×L×21.

[0265] An exemplary implementation of a computer system 2400 that can be used in connection with any embodiment of the techniques described herein (such as the methods of FIGS. 2A - 2B, FIGS. 4A - 4B, and FIG. 6) is shown in FIG. 24. The computer system 2400 includes one or more processors 2410 and one or more articles of manufacture that include a non - transitory computer - readable storage medium (such as memory 2420 and one or more non - volatile storage media 2430). Since the aspects of the techniques described herein are not limited to a particular technique for writing or reading data, the processor 2410 can control the writing of data to and the reading of data from the memory 2420 and the non - volatile storage media 2430 in any suitable manner. To execute any of the functions described herein, the processor 2410 can execute one or more processor - executable instructions stored in one or more non - transitory computer - readable storage media (such as memory 2420) that function as a non - transitory computer - readable storage medium storing the processor - executable instructions executed by the processor 2410.

[0266] The computer system 2400 may also include a network input / output (I / O) interface 2440 through which a computing device can communicate (e.g., via a network) with other computing devices, and may also include one or more user I / O interfaces 2450 through which the computing device can provide output to a user and receive input from the user. The user I / O interface may include devices such as a keyboard, a mouse, a microphone, a display device (e.g., a monitor or a touch screen), a speaker, a camera, and / or other various types of I / O devices.

[0267] The above-described embodiments can be implemented in various ways. For example, the embodiments may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code may be executed by any suitable processor (e.g., a microprocessor) or a set of processors, whether provided on a single computing device or distributed among multiple computing devices. It should be understood that any component or set of components that performs the above functions can generally be considered as one or more controllers that control the above functions. The one or more controllers can be implemented in various ways, such as dedicated hardware or general-purpose hardware (e.g., one or more processors) programmed to perform the above functions using microcode or software.

[0268] In this regard, one implementation of the embodiments described herein, when executed on one or more processors, is a computer program (i.e., a plurality of executable instructions) encoded on at least one computer-readable storage medium (e.g., RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage device, magnetic cassette, magnetic tape, magnetic disk storage device or other magnetic storage device, or other tangible non-transitory computer-readable storage medium) that executes the above functions of one or more embodiments. The computer-readable medium can be transferable so that the program stored thereon can be loaded onto any computing device for implementing the aspects of the technology described herein. Further, it should be understood that the reference to a computer program that executes any of the above functions at runtime is not limited to an application program executed on a host computer. Rather, as used herein, the terms computer program and software are used in a general sense to refer to any type of computer code (e.g., application software, firmware, microcode, or any other form of computer instruction) that can be used to program one or more processors for implementing the aspects of the technology described herein.

[0269] The foregoing description of the implementation is provided for purposes of illustration and description and is not intended to be exhaustive or to limit the implementation to the exact form disclosed. Modifications and variations are possible in light of the above teachings and may also be obtained from practice of the implementation. In other embodiments, the methods shown in these figures may include fewer operations, different operations, operations in a different order, and / or additional operations. Further, independent blocks may be executed in parallel.

[0270] As will be appreciated, the exemplary aspects described above can be implemented in various forms of software, firmware, and hardware in the implementations shown in the figures. Further, certain portions of the implementation may be implemented as "modules" that perform one or more functions. Such modules may include hardware such as a processor, an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA), or a combination of hardware and software.

[0271] Although some aspects and embodiments of the technology described in this disclosure have been described, it will be understood that those skilled in the art will readily conceive of various changes, modifications, and improvements. Such changes, modifications, and improvements are intended to be made within the spirit and scope of the technology described herein. For example, those skilled in the art can readily envision various other means and / or structures for performing the functions described herein and / or obtaining the results and / or one or more advantages, and each such variation and / or modification is considered to be within the scope of the embodiments described herein. Those skilled in the art can recognize or confirm many equivalents to the specific embodiments described herein using only routine experimentation. Thus, the foregoing embodiments are presented for purposes of illustration only, and it will be understood that within the scope of the appended claims and their equivalents, the embodiments of the present invention can be implemented in a manner different from that specifically described. Further, any combination of two or more of the features, systems, articles, materials, kits, and / or methods described herein is included within the scope of this disclosure so long as such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.

[0272] The above-described embodiments can be implemented in various ways. One or more aspects and embodiments of the present disclosure involving a process or method utilize program instructions executable by a device (e.g., a computer, a processor, or other device) to perform or control the execution of the process or method. In this regard, various inventive concepts can be embodied as a computer-readable storage medium (or plural computer-readable storage media) (e.g., a computer memory, one or more floppy disks, compact disks, optical disks, magnetic tapes, flash memories, field programmable gate arrays or other semiconductor device circuit configurations or other tangible computer storage media) encoded by one or more programs that, when executed on one or more computers or other processors, perform a method of implementing one or more of the various embodiments described above. The computer-readable medium can be transferable and the programs stored thereon can be loaded onto one or more different computers or other processors to implement the various aspects described above. In some embodiments, the computer-readable medium can be a non-transitory medium.

[0273] As used herein, the terms “program” or “software” are used in a general sense and refer to any type of computer code or set of computer-executable instructions that can be used to program a computer or other processor to implement the various aspects as described above. Further, according to one aspect, it will be understood that one or more computer programs that execute the methods of the present disclosure at runtime need not be present on a single computer or processor and may be distributed in a modular fashion among a plurality of different computers or processors to implement the various aspects of the present disclosure.

[0274] Computer-executable instructions can take many forms and can be executed by one or more computers or other devices, such as program modules. In general, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. Usually, the functions of program modules can be combined or distributed as needed in various embodiments.

[0275] Also, data structures can be stored in a computer-readable medium in any suitable format. For simplicity of explanation, a data structure can be shown as having fields associated by their positions within the data structure. Such relationships can be similarly realized by allocating storage to fields at locations within the computer-readable medium that convey the relationships between the fields. However, any suitable mechanism can be used, including the use of pointers, tags, or other mechanisms for establishing relationships between data elements, to establish the relationships between the information within the fields of the data structure.

[0276] When implemented in software, the software code can be executed by any suitable processor or set of processors, whether provided on a single computer or distributed among multiple computers.

[0277] Also, a computer can have one or more input devices and output devices. These devices can be used, among other things, to display a user interface. Examples of output devices that can be used to provide a user interface include printers and display screens for visually displaying output, speakers and other sound-generating devices for auditorily displaying output, etc. Examples of input devices that can be used for a user interface include keyboards and pointing devices such as mice, touch pads, digital tablets, etc. As another example, a computer can receive input information in voice recognition or other voice formats.

[0278] Such computers may be interconnected by one or more networks in any suitable form, such as a local area network, or a wide area network such as an enterprise network, and an intelligent network (IN), or the Internet. Such networks may be based on any suitable technology, may operate according to any suitable protocol, and may include wireless networks, wired networks, or fiber optic networks.

[0279] Also, as described, some aspects may be embodied in one or more ways. The operations performed as part of a method can be ordered in any suitable way. Accordingly, embodiments may be constructed in which the operations are performed in an order different from that shown, which may include performing some operations simultaneously, even if shown as sequential operations in the exemplary embodiments.

[0280] All definitions defined and used in this specification should be understood to take precedence over dictionary definitions, definitions of incorporated by reference documents, and / or the ordinary meaning of defined terms.

[0281] The indefinite articles "a" and "an" as used in this specification and the claims should be understood to mean "at least one" unless explicitly indicated otherwise.

[0282] As used herein and in the claims, the phrase "and / or" is to be understood to mean "either or both" of the elements so joined, i.e., elements that may be present conjunctively in some cases and disjunctively in other cases. Multiple elements listed with "and / or" are to be construed in the same manner, i.e., as "one or more" of the elements so joined. Other elements may optionally be present whether or not they are related to those specifically identified in the "and / or" clause, without regard to whether they are specifically identified or not. Thus, by way of non-limiting example, reference to "A and / or B" when used in combination with open-ended language such as "comprising" may refer in one embodiment to only A (optionally including elements other than B), in another embodiment to only B (optionally including elements other than A), and in yet another embodiment to both A and B (optionally including other elements), and so on.

[0283] As used in this specification and the claims, the phrase "at least one" as used with respect to a list of one or more elements is to be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each element specifically recited in the list of elements, and not excluding any combinations of elements in the list of elements. In this definition, elements other than those specifically identified in the list of elements referred to by the phrase "at least one" are optionally present, whether or not they are related to the specifically identified elements. Thus, by way of non-limiting example, "at least one of A and B" (or equivalently, "at least one of A or B", or equivalently, "at least one of A and / or B") can, in one embodiment, refer to at least one (optionally plural) A where B is absent (and optionally includes elements other than B), in another embodiment can refer to at least one (optionally plural) B where A is absent (and optionally includes elements other than A), and in yet another embodiment can refer to at least one (optionally plural) A and at least one (optionally plural) B (and optionally includes other elements), etc.

[0284] In the claims, and in the specification above, all transitional phrases such as "comprising", "including", "carrying", "having", "containing", "involving", "holding", "consisting of", etc. are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases "consisting of" and "consisting essentially of" are closed or semi-closed transitional phrases, respectively.

[0285] The terms "substantially", "essentially", and "about" may be used to mean within ±20% of the target value in some embodiments, within ±10% of the target value in some embodiments, within ±5% of the target value in some embodiments, and within ±2% of the target value in some embodiments. The terms "substantially", "essentially", and "about" may include the target value.

Claims

1. A method for predicting the thermal stability of a single-strand variable fragment (scFv) using a trained machine learning model, wherein the method is: Using the trained machine learning model and at least one computer hardware processor, a first thermal stability index of a first scFv having a first residue sequence is determined. Using information representing the three-dimensional (3D) structure of the first scFv, the interaction energy metric is obtained for each of the multiple pairs of residues in the first residue sequence, Generating a first set of features to be provided as input to the trained machine learning model, wherein the generation includes including the interaction energy metric in the first set of features. The first feature set is provided as input to the trained machine learning model to obtain a corresponding output that shows the first thermal stability of the first scFv, Methods that include...

2. Using the trained machine learning model and the at least one computer hardware processor, determine the thermal stability index of each scFv in the set of scFvs to obtain multiple thermal stability indices. The further includes identifying a subset of the set of scFv for subsequent manufacturing based on the plurality of thermal stability indices, The method according to claim 1, wherein the set of scFv includes the first scFv.

3. Based on the determined thermal stability index, identifying the subset of the set of scFv for subsequent manufacturing is: To determine whether the first thermal stability of the first scFv satisfies at least one criterion, After determining that the first thermal stability satisfies at least one of the criteria, the first scFv for subsequent manufacturing is identified, The method according to claim 2, including the method described in claim 2.

4. The method further includes determining a second thermal stability index of a second scFv having a second residue sequence using the trained machine learning model and the at least one computer hardware processor, wherein the determination is To obtain a second interaction energy metric for each of the second multiple pairs of second residues in the second residue sequence, Generating a second set of features to be provided as input to the trained machine learning model, wherein the generation includes including the second interaction energy metric in the second set of features. To obtain a corresponding output that shows the second thermal stability of the second scFv, the second feature set is provided as input to the trained machine learning model, The method according to claim 1, including the method described in claim 1.

5. The method according to claim 1, wherein the output indicating the first thermal stability of the first scFv indicates a first temperature at which the first scFv is thermally stable, or a first temperature range including at least one temperature at which the first scFv is thermally stable.

6. The method according to claim 5, wherein the first temperature is an estimate of the temperature corresponding to half of the maximum coupling of the first scFv, and the first temperature range is an estimate of the temperature range that includes the temperature corresponding to half of the maximum coupling of the first scFv.

7. Providing the first feature set as input to the trained machine learning model in order to obtain the output indicating the first thermal stability of the first scFv is, Using the trained machine learning model, classify the first scFv into one of a plurality of classes using the first feature set, wherein each of the plurality of classes corresponds to a temperature range. The method according to claim 1, including the method described in claim 1.

8. Obtaining the aforementioned interaction energy metric means Determining the information representing the 3D structure of the first scFv by generating the information representing the 3D structure from the first residue sequence using protein structure prediction software, and / or Determining the interaction energy metric using molecular modeling software, and generating the interaction energy metric using the information showing the 3D structure of the first scFv, The method according to claim 1, including the method described in claim 1.

9. Generating the first set of features is For each specific energy metric of the interaction energy metric, To generate a two-dimensional (2D) matrix for each of the values ​​of the specific energy metric, wherein the rows and columns of the 2D matrix correspond to each residue in the first residue sequence, and the entries in row i and column j of the 2D matrix correspond to the values ​​of the specific energy metric for residue i and residue j in the first residue sequence. The generated 2D matrix is ​​included in the first feature set, The method according to claim 1, including the method described in claim 1.

10. Generating the first set of features further involves, Encoding the aforementioned first residue sequence to obtain the encoded sequence, The encoded sequence is included in the first feature set, The method according to claim 1, including the method described in claim 1.

11. The method according to claim 1, wherein the trained machine learning model includes a trained neural network model.

12. Providing the first feature set as input to the trained machine learning model in order to obtain the corresponding output that shows the first thermal stability of the first scFv is, To obtain a first set of probabilities that the first scFv is thermally stable in each of a plurality of temperature ranges, the first feature set is provided to the trained neural network model. The first thermal stability described above is (i) A temperature range within the plurality of temperature ranges associated with the highest probability in the first plurality of probabilities, (ii) The temperature determined as a weighted linear combination of the mean values ​​of the plurality of temperature ranges weighted by the probabilities in the first set of probabilities. To decide on one of the following, The method according to claim 11, including the method described in claim 11.

13. To manufacture at least one of the scFv within the identified subset, and / or To test the thermal stability of at least one of the scFvs by an in vitro assay. The method according to claim 1, further comprising:

14. It is a system, At least one computer hardware processor, A non-temporary computer-readable storage medium that stores a processor-executable instruction that causes the at least one computer hardware processor to perform the method according to any one of claims 1 to 12 when executed by the at least one computer hardware processor, A system that includes this.

15. A non-temporary computer-readable storage medium that stores a processor-executable instruction, when executed by the at least one computer hardware processor, causing the at least one computer hardware processor to perform the method according to any one of claims 1 to 12.