Estimation of differences between protein sequences implemented by machine learning

By training language models and neural network attribute calculation models, sequence differences between protein molecule pairs are identified, solving the problems of low efficiency and model overfitting in the improvement of macromolecular drug properties in traditional methods, and achieving more efficient molecular design.

CN122029604APending Publication Date: 2026-05-12GENENTECH INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GENENTECH INC
Filing Date
2024-10-04
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently identify and improve the amino acid sequences of macromolecular drugs to enhance their properties, such as affinity, specificity, and bioactivity. Furthermore, traditional computational models overfit under low-data conditions and exhibit poor predictive performance.

Method used

By training a language model and a neural network attribute calculation model, relative embeddings are generated to identify sequence differences between protein molecule pairs, and attribute differences are determined based on these differences. Machine learning methods are then used to identify and improve mutations.

Benefits of technology

It improves the accuracy and efficiency of predicting differences in protein molecular properties under low data conditions, identifies amino acid mutations that improve properties, reduces the need for wet laboratory analysis, and improves the time and resource efficiency of molecular design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122029604A_ABST
    Figure CN122029604A_ABST
Patent Text Reader

Abstract

The present disclosure provides a method that may include generating a training dataset including a pair of sample molecules. The sample molecule pair includes a first molecule and a second molecule. An attribute calculation model is trained to determine a sequence difference that includes a difference between the amino acid sequence of the first molecule and the amino acid sequence of the second molecule. The attribute calculation model is further trained to generate a relative embedding representing the sequence differences and associating the sequence differences with the pairs of sample molecules. The attribute calculation model is further trained to determine an attribute difference corresponding to a difference between a value of the attribute exhibited by the first molecule and a value of the attribute exhibited by the second molecule based on the relative embedding. A trained attribute calculation model may be applied to determine an attribute difference of a pair of input molecules.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This application claims priority to U.S. Provisional Application No. 63 / 588,671, filed October 6, 2023, entitled “MACHINE LEARNING ENABLED ESTIMATION OF DIFFERENCES BETWEEN PROTEIN SEQUENCES”, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0002] The topics described in this article generally involve molecular design, and more specifically, computational models for determining differences in properties exhibited by two different protein molecules and techniques for using computational models to enhance molecular properties. Background Technology

[0003] A molecule is a group of two or more atoms linked together by chemical bonds. Molecules form the smallest identifiable unit, and a pure substance can be broken down into such identifiable units while still retaining the substance's composition and chemical properties. A molecule's various properties, including its ability to function as a therapeutic agent, can depend on its composition and construction (or three-dimensional structure). In contrast, macromolecules (also known as biopharmaceuticals, biologics, or biological agents) are molecules with molecular weights ranging from approximately 3,000 Daltons to 150,000 Daltons. Macromolecular drugs are often derivatives of natural human proteins that regulate many important cellular functions, such as enzymatic reactions, molecular transport, regulation and execution of many biological pathways, cell growth, proliferation, nutrient uptake, morphology, motility, and intercellular communication. Examples of therapeutic proteins include antibodies, chimeric antigen receptors (CARs), enzymes, hormones, and cytokines. A single macromolecule can typically have more than 1,300 amino acid residues linked by peptide bonds to form one or more polypeptides. Due to their size and complexity, macromolecular drugs are produced through engineered cellular recombination rather than chemical synthesis, as is the case with most small molecule drugs. Furthermore, since oral administration is ineffective, macromolecular therapeutics are typically delivered by injection or infusion. The development of macromolecular drugs may require designing one or more sequences of amino acid residues that can bind to targets (e.g., proteins, nucleic acids, etc.), said amino acid residues having sufficient specificity and free from undesirable properties such as immunogenicity, self-association, instability, etc. Summary of the Invention

[0004] Systems, methods, and articles (including computer program products) for estimating differences between protein molecules for machine learning implementations are provided. In one aspect, a system comprising at least one processor and at least one data processor is provided. The at least one memory can store instructions that, when executed by the at least one processor, cause operations. Operations may include: generating a training dataset comprising a pair of sample molecules, the sample molecule pair comprising a first molecule and a second molecule; training an attribute computation model at least based on the training dataset, wherein the attribute computation model is trained to at least determine sequence differences between the first molecule and the second molecule, the sequence differences comprising differences between the amino acid sequences of the first molecule and the amino acid sequences of the second molecule; generating relative embeddings representing the sequence differences and associating the sequence differences with the sample molecule pair; and determining attribute differences for the sample molecule pair at least based on the relative embeddings of the sample molecule pair, wherein the attribute differences for the molecule pair correspond to differences between the values ​​of attributes exhibited by the first molecule and the values ​​of attributes exhibited by the second molecule; receiving a pair of input molecules; and applying the attribute computation model to determine the attribute differences between the pair of input molecules.

[0005] On the other hand, a computer-implemented method for estimating differences between protein molecules for machine learning is provided. The method may include: generating a training dataset comprising a pair of sample molecules, the sample molecule pair comprising a first molecule and a second molecule; training an attribute computation model based at least on the training dataset, wherein the attribute computation model is trained to at least determine sequence differences between the first molecule and the second molecule, the sequence differences including differences between the amino acid sequences of the first molecule and the amino acid sequences of the second molecule; generating relative embeddings representing the sequence differences and associating the sequence differences with the sample molecule pair; and determining attribute differences for the sample molecule pair based at least on the relative embeddings of the sample molecule pair, wherein the attribute differences for the molecule pair correspond to differences between the values ​​of attributes exhibited by the first molecule and the values ​​of attributes exhibited by the second molecule; receiving a pair of input molecules; and applying the attribute computation model to determine the attribute differences between the pair of input molecules.

[0006] In another aspect, a computer program product is provided, comprising a non-transitory computer-readable medium storing instructions. These instructions can cause operations executable by at least one data processor. The operations may include: generating a training dataset comprising a pair of sample molecules, the sample molecule pair comprising a first molecule and a second molecule; training an attribute computation model, at least based on the training dataset, wherein the attribute computation model is trained to at least determine sequence differences between the first molecule and the second molecule, the sequence differences including differences between the amino acid sequences of the first molecule and the amino acid sequences of the second molecule; generating relative embeddings representing the sequence differences and associating the sequence differences with the sample molecule pair; and determining attribute differences for the sample molecule pair, at least based on the relative embeddings of the sample molecule pair, wherein the attribute differences for the molecule pair correspond to differences between the values ​​of attributes exhibited by the first molecule and the values ​​of attributes exhibited by the second molecule; receiving a pair of input molecules; and applying the attribute computation model to determine the attribute differences between the pair of input molecules.

[0007] In some variations, one or more of the features disclosed herein, including the following features, may optionally be included in any feasible combination.

[0008] In some variants, properties include expression, affinity, specificity, exploitability, or bioactivity.

[0009] In some variations, sample molecule pairs are generated by pairing at least a set of sample molecules from a group of sample molecules with known values ​​for the relevant properties.

[0010] In some variations, generating the training dataset involves pairing a first sample molecule and a second sample molecule from a set of sample molecules to generate a sample molecule pair, and pairing the first sample molecule with a third sample molecule from a set of sample molecules to generate another sample molecule pair.

[0011] In some variations, the attribute computation model includes a language model, and training the attribute computation model includes training the language model to generate relative embeddings of sample molecule pairs.

[0012] In some variations, the language model is trained to generate a first embedding for a first sample molecule and a second embedding for a second sample molecule, and a relative embedding for each sample molecule pair is generated by at least determining the difference between the first and second embeddings.

[0013] In some variations, the difference between the first and second embeddings corresponds to a sequence difference that includes the difference between the amino acid sequences of the first molecule and the amino acid sequences of the second molecule.

[0014] In some variations, the attribute computation model further includes a neural network coupled to a language model, and training the attribute computation model includes training the neural network to determine attribute differences for each sample molecule pair based at least on the relative embeddings of each sample molecule pair.

[0015] In some variants, one or more mutations associated with improvements in properties are identified, at least based on differences in the properties of paired input molecules.

[0016] In some variants, one or more mutations include point mutations, where the paired input molecules differ at a single position in each corresponding amino acid sequence.

[0017] In some variants, one or more mutations include a combination of multiple point mutations in which a pair of input molecules differ at multiple positions in each corresponding amino acid sequence.

[0018] In some variants, each sample molecule includes an antibody or a portion thereof.

[0019] In some variants, each sample molecule includes a variable region of antibody, an antigen-binding region, a heavy chain, and / or a light chain.

[0020] In another aspect, a system for identifying mutations is provided. The system may include at least one memory and at least one data processor. The at least one memory may store instructions that, when executed by the at least one processor, cause operations. Operations may include: applying a property calculation model to determine the difference between the values ​​of a property exhibited by a first input molecule and the values ​​of a property exhibited by a second input molecule; determining the amino acid sequences of the first input molecule and the second input molecule; identifying a first mutation by comparing the amino acid sequences of the first input molecule and the second input molecule; and identifying the value of the first mutation-improved property based at least on the difference in the values ​​of the properties exhibited by the first and second input molecules.

[0021] In one aspect, a computer-implemented method for identifying mutations is provided. The method may include: applying a property computation model to determine the difference between the values ​​of a property exhibited by a first input molecule and the values ​​of a property exhibited by a second input molecule; determining the amino acid sequences of the first and second input molecules; identifying a first mutation by comparing the amino acid sequences of the first and second input molecules; and identifying the value of the first mutation-improved property based at least on the difference between the values ​​of the properties exhibited by the first and second input molecules.

[0022] In another aspect, a computer program product is provided, comprising a non-transitory computer-readable medium storing instructions. These instructions may cause operations to be executed by at least one data processor. The operations may include: applying a property calculation model to determine the difference between the values ​​of a property exhibited by a first input molecule and the values ​​of a property exhibited by a second input molecule; determining the amino acid sequences of the first and second input molecules; identifying a first mutation by comparing the amino acid sequences of the first and second input molecules; and identifying the value of the first mutation-improved property based at least on the difference in the values ​​of the properties exhibited by the first and second input molecules.

[0023] In some variations, one or more of the features disclosed herein, including the following features, may optionally be included in any feasible combination.

[0024] In some variants, properties include expression, affinity, specificity, exploitability, or bioactivity.

[0025] In some variations, a property computation model is applied to determine the difference in property values ​​exhibited by the first and third input molecules. The amino acid sequence of the third molecule is determined. A second mutation is identified by comparing at least the amino acid sequences of the first and third input molecules. The second mutation is identified as improving the property value based at least on the difference in property values ​​exhibited by the first and third input molecules.

[0026] In some variations, the first mutation includes an amino acid residue of one type occupying a first position in the amino acid sequence of the first input molecule, and the second mutation includes an amino acid residue of the same type occupying a second position in the amino acid sequence of the first input molecule.

[0027] In some variations, the first mutation includes a position in the amino acid sequence of the first input molecule occupied by a first type of amino acid residue, and the second mutation includes the same position in the amino acid sequence of the first input molecule occupied by a second type of amino acid residue.

[0028] In some variants, at least one output molecule exhibiting the first mutation and / or the second mutation is generated.

[0029] In some variants, each input molecule includes an antibody or a portion thereof.

[0030] In some variants, each sample molecule includes a variable region of antibody, an antigen-binding region, a heavy chain, and / or a light chain.

[0031] In another aspect, a system for identifying mutation combinations is provided. The system may include at least one memory and at least one data processor. The at least one memory may store instructions that, when executed by the at least one processor, cause operations. Operations may include: determining a plurality of point mutations, wherein each of the plurality of point mutations is associated with an improvement in the value of a property; generating a plurality of mutation combinations, wherein each of the plurality of mutation combinations includes point mutations selected from the plurality of point mutations; generating a plurality of input molecules, wherein each of the plurality of input molecules exhibits a mutation combination from the plurality of mutation combinations; applying a trained property computation model to determine the difference between the values ​​of the property exhibited by two input molecules from the plurality of input molecules; and identifying at least one mutation combination as associated with an improvement in the value of the property based at least on the difference between the values ​​of the property exhibited by two input molecules from the plurality of input molecules.

[0032] On the other hand, a computer-implemented method for identifying mutation combinations is provided. The method may include: identifying a plurality of point mutations, wherein each of the plurality of point mutations is associated with an improvement in the value of a property; generating a plurality of mutation combinations, wherein each of the plurality of mutation combinations includes point mutations selected from the plurality of point mutations; generating a plurality of input molecules, wherein each of the plurality of input molecules exhibits a mutation combination from the plurality of mutation combinations; applying a trained property computation model to determine the difference between the values ​​of the property exhibited by two input molecules from the plurality of input molecules; and identifying at least one mutation combination as associated with an improvement in the value of the property based at least on the difference between the values ​​of the property exhibited by two input molecules from the plurality of input molecules.

[0033] In another aspect, a computer program product is provided, comprising a non-transitory computer-readable medium storing instructions. The instructions can cause operations executable by at least one data processor. The operations may include: determining a plurality of point mutations, wherein each of the plurality of point mutations is associated with an improvement in the value of an attribute; generating a plurality of mutation combinations, wherein each of the plurality of mutation combinations includes point mutations selected from the plurality of point mutations; generating a plurality of input molecules, wherein each of the plurality of input molecules exhibits a mutation combination from the plurality of mutation combinations; applying a trained attribute computation model to determine a difference between the values ​​of the attribute exhibited by two input molecules from the plurality of input molecules; and identifying at least one mutation combination as associated with an improvement in the value of the attribute, at least based on the difference between the values ​​of the attribute exhibited by two input molecules from the plurality of input molecules.

[0034] In some variations, one or more of the features disclosed herein, including the following features, may optionally be included in any feasible combination.

[0035] In some variants, each input molecule is generated by applying a combination of mutations to the amino acid sequence of a selected molecule.

[0036] In some variants, properties include expression, affinity, specificity, exploitability, or bioactivity.

[0037] In another aspect, a system for enhancing molecular properties through machine learning is provided. The system may include at least one memory and at least one data processor. The at least one memory may store instructions that, when executed by the at least one processor, cause operations. The operations may include: identifying a first pair of input molecules comprising a first input molecule and a second input molecule, wherein the first input molecule exhibits a mutation relative to the second input molecule; applying a trained property computation model to determine, at least based on a first amino acid sequence of the first input molecule and a second amino acid sequence of the second input molecule, a difference between a first value of a property exhibited by the first input molecule and a second value of a property exhibited by the second input molecule; identifying a mutation as an improvement to the property, at least based on the difference between the first and second values ​​of the property; and generating candidate molecules by modifying at least selected molecules to include the mutation.

[0038] In another aspect, a computer-implemented method for enhancing molecular properties using machine learning is provided. The method may include: identifying a first pair of input molecules comprising a first input molecule and a second input molecule, wherein the first input molecule exhibits a mutation relative to the second input molecule; applying a trained property computation model to determine, at least based on a first amino acid sequence of the first input molecule and a second amino acid sequence of the second input molecule, a difference between a first value of a property exhibited by the first input molecule and a second value of a property exhibited by the second input molecule; identifying a mutation as an improvement to the property based at least on the difference between the first and second values ​​of the property; and generating candidate molecules by modifying at least selected molecules to include the mutation.

[0039] In another aspect, a computer program product is provided, comprising a non-transitory computer-readable medium storing instructions. These instructions can cause operations executable by at least one data processor. The operations may include: identifying a first pair of input molecules comprising a first input molecule and a second input molecule, wherein the first input molecule exhibits a mutation relative to the second input molecule; applying a trained property computation model to determine, at least based on a first amino acid sequence of the first input molecule and a second amino acid sequence of the second input molecule, a difference between a first value of a property exhibited by the first input molecule and a second value of a property exhibited by the second input molecule; identifying a mutation as an improvement to the property, at least based on the difference between the first and second values ​​of the property; and generating candidate molecules by at least modifying selected molecules to include the mutation.

[0040] In some variations, one or more of the features disclosed herein, including the following features, may optionally be included in any feasible combination.

[0041] In some variations, a second pair of input molecules is identified, comprising a first input molecule and a third input molecule. The first input molecule exhibits another mutation relative to the third input molecule. An attribute calculation model is applied to determine the difference between a first value of an attribute exhibited by the first input molecule and a third value of an attribute exhibited by the third input molecule, based at least on the first amino acid sequence of the first input molecule and the third amino acid sequence of the third input molecule. The difference between the first and third values ​​of the attribute, at least based on the first and third values, indicates that the first input molecule exhibits a better value for the attribute than the third input molecule, thus identifying the first pair of input molecules as including the first input molecule rather than the third input molecule.

[0042] In some variants, mutations include multiple point mutations.

[0043] In some variants, multiple point mutations are identified by recognizing at least each individual point mutation as the value that improves the property.

[0044] In some variations, identifying each individual point mutation as improving the value of the attribute includes: generating multiple pairs of input molecules, wherein each pair of input molecules includes two input molecules whose amino acid sequences differ due to a single point mutation; applying an attribute computation model to determine, for each pair of input molecules, that one input molecule exhibits a better value for the attribute than the other input molecule; and for each pair of input molecules, identifying the point mutation exhibited by the input molecule that exhibits the better value for the attribute as one of a plurality of point mutations.

[0045] In some variants, a single point mutation includes an amino acid sequence containing a type 1 amino acid residue at a specific position and another amino acid sequence containing a type 2 amino acid residue at the same position.

[0046] In some variants, a single point mutation involves an amino acid sequence containing a specific type of amino acid residue at a first position and another amino acid sequence containing the same type of amino acid residue at a second position.

[0047] In some variations, the attribute computation model includes a language model with relative embeddings that has been trained to generate a representation of the difference between the first amino acid sequence of the first input molecule and the second amino acid sequence of the second input molecule.

[0048] In some variations, the attribute computation model further includes a neural network coupled to a language model, and the neural network has been trained to determine the difference between a first value and a second value of an attribute, at least based on relative embeddings.

[0049] In some variants, each of the first and second input molecules includes an antibody.

[0050] In some variants, each of the first and second input molecules includes a portion of the antibody.

[0051] In some variations, each of the first and second input molecules includes a variable region of the antibody, an antigen-binding region, a heavy chain, and / or a light chain.

[0052] In some variants, properties include expression, affinity, specificity, exploitability, or bioactivity.

[0053] On the other hand, a system for training an attribute computation model is provided. The system may include at least one memory and at least one data processor. The at least one memory may store instructions that, when executed by the at least one processor, cause operations. Operations may include: identifying a set of sample molecules, wherein each sample molecule in the set is associated with a corresponding value for an attribute; generating a plurality of sample molecule pairs to include in a training dataset, wherein generating the plurality of sample molecule pairs includes pairing a first sample molecule with a second sample molecule selected from the set of sample molecules to generate a first sample molecule pair, and pairing the first sample molecule with a third sample molecule selected from the set of sample molecules to generate a second sample molecule pair; and training the attribute computation model at least based on the training dataset to determine attribute differences between two input molecules based at least on sequence differences between two input molecules, wherein the sequence differences between the two input molecules correspond to differences between amino acid sequences of each input molecule, and wherein the attribute differences between the two input molecules correspond to differences between values ​​of an attribute present in each input molecule.

[0054] On the other hand, a computer-implemented method for training an attribute computation model is provided. The method may include: identifying a set of sample molecules, wherein each sample molecule in the set is associated with a corresponding value for an attribute; generating multiple sample molecule pairs to include in a training dataset, wherein generating multiple sample molecule pairs includes pairing a first sample molecule with a second sample molecule selected from the set of sample molecules to generate a first sample molecule pair, and pairing the first sample molecule with a third sample molecule selected from the set of sample molecules to generate a second sample molecule pair; and training the attribute computation model at least based on the training dataset to determine the attribute difference between two input molecules based at least on the sequence difference between the two input molecules, wherein the sequence difference between the two input molecules corresponds to the difference between the amino acid sequences of each input molecule, and wherein the attribute difference between the two input molecules corresponds to the difference between the values ​​of the attribute present in each input molecule.

[0055] In another aspect, a computer program product is provided, comprising a non-transitory computer-readable medium storing instructions. These instructions can cause operations that can be executed by at least one data processor. The operations may include: identifying a set of sample molecules, wherein each sample molecule in the set is associated with a corresponding value for an attribute; generating a plurality of sample molecule pairs to include in a training dataset, wherein generating the plurality of sample molecule pairs includes pairing a first sample molecule with a second sample molecule selected from the set of sample molecules to generate a first sample molecule pair, and pairing the first sample molecule with a third sample molecule selected from the set of sample molecules to generate a second sample molecule pair; and training an attribute computation model at least based on the training dataset to determine attribute differences between two input molecules based at least on sequence differences between the two input molecules, wherein the sequence differences between the two input molecules correspond to differences between amino acid sequences of each input molecule, and wherein the attribute differences between the two input molecules correspond to differences between values ​​of attributes present in each input molecule.

[0056] In some variations, one or more of the features disclosed herein, including the following features, may optionally be included in any feasible combination.

[0057] In some variations, generating multiple sample molecule pairs further includes pairing a second sample molecule with a third sample molecule to generate a third sample molecule pair.

[0058] In some variations, the attribute computation model includes a language model that has been trained to determine relative embeddings representing sequence differences between two input molecules, based at least on the amino acid sequence of each input molecule.

[0059] In some variations, the attribute computation model includes a neural network coupled to a language model, and the neural network has been trained to determine attribute differences between two input molecules based at least on relative embeddings.

[0060] In some variants, properties include expression, affinity, specificity, exploitability, or bioactivity.

[0061] In some variants, each sample molecule includes an antibody.

[0062] In some variants, each sample molecule includes a portion of an antibody.

[0063] In some variants, each sample molecule includes a variable region of antibody, an antigen-binding region, a heavy chain, and / or a light chain.

[0064] Specific implementations of the present subject matter may include, but are not limited to, methods consistent with the descriptions provided herein, and articles of art comprising a tangible, machine-readable medium operable to cause one or more machines (e.g., computers, etc.) to perform operations implementing one or more of the described features. Similarly, computer systems comprising one or more processors and one or more memories coupled to the one or more processors are also described. Memory that may include a non-transitory computer-readable or machine-readable storage medium may include, encode, store, etc., one or more programs that cause one or more processors to perform one or more of the operations described herein. Computer implementation methods consistent with one or more implementations of the present subject matter may be implemented by one or more data processors existing in a single computing system or multiple computing systems. Such multiple computing systems may be interconnected and may exchange data and / or commands or other instructions via one or more connections, including, for example, direct connections between one or more of the multiple computing systems via a network (e.g., the Internet, wireless wide area network, local area network, wide area network, wired network, etc.).

[0065] Details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Further features and advantages of the subject matter described herein will become apparent from the description, drawings, and claims. While certain features of the currently disclosed subject matter have been described for illustrative purposes in relation to the design of larger molecules, such as protein molecules, it should be readily understood that such features are not intended to be limiting. The claims following this disclosure are intended to define the scope of the protected subject matter. Attached Figure Description

[0066] The accompanying drawings, incorporated in and forming part of this specification, illustrate certain aspects of the subject matter disclosed herein and, together with the specification, help explain some principles associated with the disclosed embodiments. In the drawings, Figure 1 depicts a system diagram illustrating an example of a molecular design system according to some exemplary embodiments; Figure 2A depicts a flowchart illustrating an example of a process for enhancing molecular properties based on machine learning, according to some example embodiments; Figure 2B depicts a flowchart illustrating an example of a process for generating a training dataset for training an attribute calculation model, according to some example embodiments. Figure 3 depicts a flowchart illustrating an example of a process for enhancing molecular properties based on machine learning, according to some example embodiments; Figure 4A depicts a flowchart illustrating an example of a process for identifying mutations that improve properties, according to some example embodiments; Figure 4B depicts a flowchart illustrating an example of a process for identifying mutations that improve properties, according to some example embodiments; Figure 5A depicts a schematic diagram illustrating an example of a workflow for enhancing molecular properties using machine learning, according to some example embodiments. Figure 5B depicts a sampling schematic illustrating point mutations across mutation sites and amino acid residue types according to some example embodiments; Figure 6 depicts a schematic diagram illustrating a comparison of the relative embeddings of amino acid sequences different from point mutations and higher edit distances according to some example embodiments. Figure 7A depicts a graph illustrating the improvement in the prediction and measurement of binding affinity for variants of the lead molecule according to some example embodiments; Figure 7B depicts a graph showing the improvement in the predicted and measured binding affinity of a variant of another lead molecule according to some example embodiments; Figure 7C depicts a graph showing the improvement in the predicted and measured binding affinity of a variant of another lead molecule according to some example embodiments; Figure 7D depicts a graph showing the improvement in the prediction and measurement of binding affinity for variants of the lead molecule according to some example embodiments; Figure 7E depicts a graph showing the improvement in the prediction and measurement of binding affinity for a variant of another lead molecule according to some example embodiments; Figure 7F depicts a graph showing the improvement in the prediction and measurement of binding affinity for a variant of another lead molecule according to some example embodiments; Figure 8A depicts graphs illustrating the affinity improvements according to some example embodiments and the expression and binding rates of molecules designed using mutations identified by utilizing a property computational model; Figure 8B depicts graphs illustrating the affinity improvements according to some example embodiments, as well as the expression and binding rates of molecules designed using mutations identified by utilizing a property computational model; Figure 8C depicts graphs illustrating the affinity improvements according to some example embodiments, as well as the expression and binding rates of molecules designed using mutations identified by utilizing a property computational model; Figure 8D depicts the sequence analysis of molecules with the highest affinity generated by utilizing mutations in the lead molecule identified using a property computation model, according to some example embodiments. Figure 8E depicts a structural analysis of molecules with the highest affinity generated by utilizing a mutation-modified lead molecule identified using a property computation model, according to some example embodiments. Figure 8F depicts a sequence analysis showing the highest affinity molecule generated by modifying another lead molecule with a mutation identified using a property computation model, according to some example embodiments. Figure 8G depicts a structural analysis of a molecule with the highest affinity generated by modifying another lead molecule with a mutation identified using a property computation model, according to some example embodiments. Figure 9 illustrates performance estimates of protein computational models operating with different protein language models (pLM) across three different lead molecules according to some example embodiments; Figure 10 illustrates the impact of removing sample molecules with spurious annotations from the training dataset on the performance of the property computation model according to some example embodiments; Figure 11 depicts a graph illustrating how molecules with mutations identified using a property calculation model, according to some example embodiments, retain the desired properties even at higher edit distances; Figure 12A depicts a surface plasmon resonance (SPR) sensor diagram illustrating the binding interactions of the highest binding variants of a mutant-modified lead molecule identified using a property computation model, according to some example embodiments. Figure 12B depicts a surface plasmon resonance (SPR) sensor plot illustrating the binding interactions of the highest binding variant of another lead molecule with mutations identified using a property calculation model, according to some example embodiments. Figure 12C depicts a surface plasmon resonance (SPR) sensor plot illustrating the binding interactions of the highest binding variant of another lead molecule with mutational modification identified using a property computation model, according to some example embodiments; and Figure 13 depicts a block diagram illustrating an example of a computing system according to some exemplary embodiments.

[0067] In practical applications, similar reference numerals indicate similar structures, features, or elements. Detailed Implementation

[0068] Molecules can be designed to exhibit a variety of desired properties, including, in the case of therapeutics, properties similar to those of drugs, such as affinity, specificity, bioactivity, and developability. Lead optimization is a variation of molecular design in which a lead molecule (or another candidate molecule) is modified to enhance desired properties while minimizing undesirable properties. For example, a lead molecule identified as having clinically useful pharmacological or biological activity during early development may have significant defects, such as chemical or physical instability, poor pharmacokinetics, immunogenicity, etc. In the case of a protein molecule, lead optimization can include modifying the basic sequence of amino acid residues, for example, by changing the type of one or more of the constituent amino acid residues, so that the resulting candidate molecule exhibits better properties than the lead molecule. However, modifications that improve lead molecules, especially those that enhance desired properties without introducing defects, can be difficult to identify, at least due to the vast combinatorial space of possibilities. For example, even for point mutations affecting a single amino acid residue, having… A leader molecule containing only a few amino acid residues can also produce The number of possible modifications increases exponentially with the degree of modification in the lead molecule, for example, from a single amino acid residue to multiple amino acid residues. Due to time and laboratory resource constraints, conventional lead optimization schemes that rely on wet laboratory analyses (e.g., in vitro measurements, in vivo characterization, etc.) to examine the effect of each modification cannot adequately explore the possible modifications that can be made to the lead molecule.

[0069] While traditional wet-lab-based molecular design methods are too resource-intensive to support the extensive exploration of diverse molecular designs aimed at identifying novel molecules with properties superior to known molecules, recent efforts utilizing artificial intelligence and machine learning are challenged by the scarcity of labeled (or annotated) data. Indeed, the cost of acquiring data for developing protein-based therapies is prohibitively high due to time and laboratory resource constraints. Therefore, computational models trained to minimize the need for wet-lab analysis, such as those capable of predicting protein molecule properties, are often limited to low-data models, where the amount of training data is limited and the distribution of the training data fails to capture all possible inputs the computational model might encounter during deployment. Computational models trained in low-data models tend to overfit the training data and thus perform poorly. For example, if the training dataset contains too few protein molecules with known values ​​for a particular property, a computational model trained on that dataset may not accurately predict the value of that property for protein molecules outside the distribution of the training dataset. Consequently, poorly performing computational models are insufficient for wet-lab analysis of modified lead molecules during lead optimization. Computer simulation Alternatives are not available, so the same bottlenecks still exist in the drug discovery process.

[0070] Various embodiments of this disclosure improve the performance of computational models trained in low-data scenarios, particularly those used to predict the effects of modifications to the amino acid residue sequences that form protein molecules. For example, in some example embodiments, the attribute computational model can be trained to determine the differences in attribute values ​​exhibited by different pairs of protein molecules, rather than the absolute values ​​of attributes exhibited by individual protein molecules. In some cases, the attribute computational model may include a language model trained to generate relative embeddings representing the sequence differences of amino acid residues that form individual protein molecules in each pair. Furthermore, in some cases, the attribute computational model may also include a neural network trained to determine the differences in attribute values ​​exhibited by paired protein molecules, at least based on relative embeddings. As described in more detail below, when training the attribute computational model to determine the differences in the values ​​of attributes exhibited by paired protein molecules, small training datasets (e.g., those with known values ​​for the attribute) can be added in combination. (200-1000 protein molecules). For example, combining and increasing the training dataset can increase its size and improve its distribution. Training on a combined increase in the training dataset can prevent overfitting and improve the distribution performance of the property computation model, making it a faster and less expensive alternative to wet lab analysis.

[0071] In some example embodiments, an attribute computation model can be trained based on a training dataset, in which each training sample includes a pair of sample protein molecules and labels indicating the differences between known values ​​of the attribute exhibited by each sample protein molecule. In some cases, the training dataset can be generated by combining a set of sample protein molecules, each having a known value for the attribute, by pairing the same sample protein molecule with multiple other sample protein molecules. For example, in some cases, a training dataset can be generated to include a first training sample pairing a first sample protein molecule with a second sample protein molecule and a second training sample pairing a first sample protein molecule with a third sample protein molecule. In a set of sample protein molecules including In the case of a sample of protein molecules, the above combination can generate Sample protein molecules are used as training samples to be included in the training dataset. It should be understood that these samples have known values ​​for their respective attributes. A single sample of protein molecules may not be sufficient to train a property calculation model to accurately predict the absolute value of the property, especially for Protein molecules outside the distribution of individual sample proteins. In contrast, A sufficiently large and well-distributed training dataset can be constructed from sample protein molecules (with known differences between the attribute values ​​exhibited by each constituent protein molecule) to train an attribute computation model to accurately predict the differences in attribute values ​​exhibited by pairs of protein molecules. As described in more detail below, the attribute computation model trained to predict the differences in attribute values ​​exhibited by pairs of protein molecules can be applied to explore possible modifications to selected molecules (such as lead molecules) in a principled manner, thereby improving the time and resource efficiency of molecular design.

[0072] In some example embodiments, a property computation model can be applied to determine how different mutations in the leader molecule sequence affect the properties exhibited by the leader molecule, thereby identifying mutations that improve a particular property. In this context, improvement can refer to an increase in the desired property value or a decrease in the undesired property value. In some cases, the property computation model can be used to sample different mutations, including point mutations affecting a single amino acid residue in the leader molecule sequence and combinations of multiple point mutations. For example, in some cases, a property computation model can be applied to determine the difference in property values ​​exhibited by one or more pairs of input molecules, each of which includes a first input molecule having a first amino acid sequence and a second input molecule having a second amino acid sequence. In some cases, at least one site in the first amino acid sequence can be occupied by an amino acid residue of a different type than the same site in the second input sequence. In the context of leader optimization, the first amino acid sequence can correspond to the sequence of a leader molecule with a first mutation, and the second amino acid sequence can correspond to the sequence of a leader molecule with a second mutation. A first mutation present in the first amino acid sequence can be identified as improving the property if the output of the property computation model indicates that the first input molecule exhibits a better value for the property than the second input molecule. Therefore, as described in more detail below, candidate molecules can be generated (e.g., for further estimation, such as wet laboratory analysis) to include at least one mutation identified as improving properties.

[0073] In some example embodiments, attribute computation models can be applied to sample different point mutations in order to identify one or more point mutations that cause attribute improvements. In this case, a point mutation between two protein molecules may refer to a difference in a single amino acid residue in the corresponding amino acid sequence. Therefore, in some cases, attribute computation models can be applied to determine the differences in attribute values ​​exhibited by multiple pairs of input molecules, each of which includes a first protein molecule having a first amino acid sequence and a second protein molecule having a second amino acid sequence. In some cases, the first amino acid sequence may remain fixed during sampling, while the second amino acid sequence may be altered to exhibit a variety of different possible point mutations. For example, in some cases, the second amino acid sequence may be altered to exhibit amino acid residues of a different type than those occupying the corresponding sites in the first amino acid sequence for one or more possible sites within the second amino acid sequence. In some cases, the second amino acid sequence may be altered across all possible types of amino acid residues, such as twenty typical types of amino acid residues. Therefore, when a first type of amino acid residue occupies a specific site in the first amino acid sequence, a second type of amino acid residue occupying the same site in the second amino acid sequence may be altered to sample mutations across all other possible types of amino acid residues. In some cases, variations across all possible types of amino acid residues may be repeated at every site within the second amino acid sequence. Doing so could constitute an exhaustive sampling of possible mutation sites and the different types of amino acid residues that could occupy each mutation site.

[0074] In some example embodiments, a property computation model can be applied to sample different combinations of point mutations to identify one or more combinations of point mutations that cause property improvement. For example, in some cases, one or more combinations of point mutations can be generated from a set of point mutations identified as causing property improvement. Each combination of point mutations may include, for example, two or more point mutations randomly selected from the set of point mutations identified as causing property improvement. The property computation model can then be applied to determine the differences in property values ​​exhibited by one or more pairs of protein molecules, each of which exhibits a different combination of point mutations. In some cases, the property computation model can be applied iteratively, for example, during one iteration, a first protein molecule is identified as having a better value for the property than a second protein molecule, and in the next iteration, it is paired with a third protein molecule that exhibits a different combination of point mutations than the second protein molecule. If the first protein molecule is again determined to have a better value for the property than the third protein molecule, the first protein molecule can be iterated again, in which the first protein molecule is paired with a fourth protein molecule that exhibits a different combination of point mutations than the third protein molecule. In some cases, combinations of point mutations present in a protein molecule may exhibit better values ​​for a property than one or more other protein molecules, and can be identified as causing an improvement in the property.

[0075] Figure 1 depicts a system diagram illustrating an example of a molecular design system 100 according to some exemplary embodiments. Referring to Figure 1, the molecular design system 110 may include an attribute calculation engine 110, a training engine 120, a molecular design engine 130, and a client device 140. As shown in Figure 1, the attribute calculation engine 110, the training engine 120, the molecular design engine 130, and the client device 140 may be communicatively coupled via a network 150. The client device 140 may be a processor-based device, including, for example, a workstation, desktop computer, laptop computer, smartphone, tablet computer, wearable device, etc. The network 150 may be a wired network and / or a wireless network, including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, etc.

[0076] In some example embodiments, the attribute calculation engine 110 may include an attribute calculation model 115. In some cases, the attribute calculation model 115 may have been trained, for example, by the attribute calculation engine 120, to determine the attribute calculation result of a pair of molecules (such as...) Figure 1The difference 103 shown is in the attribute values ​​exhibited by the first input molecule 105a and the second input molecule 105b. In some cases, the first input molecule 105a and the second input molecule 105b can be protein molecules used in protein-based therapeutics, including, for example, antibodies, chimeric antigen receptors (CARs), cytokines, hormones, enzymes, etc. Therefore, in some cases, the attribute of interest can be similar to drug-like properties, such as expression, affinity, specificity, immunogenicity, etc. in vivo Stability, self-association, etc. As described in more detail below, in some cases, the property calculation model 115 can be applied to identify one or more mutations, including point mutations affecting a single amino acid residue and combinations of multiple point mutations that cause property improvement. For example, in some cases, the molecular design engine 130 can apply the property calculation model 115 to identify one or more mutations that can be applied to the selected molecule 133 in multiple iterations to generate a candidate molecule 135 with better values ​​for the properties than the selected molecule 133. In some cases, the selected molecule 133 can be a lead molecule of a molecule that has been identified as having clinically useful pharmacological or biological activity in at least some molecular design applications. The selected molecule 133 can undergo iterative property improvement to enhance the desired properties while minimizing undesirable properties, thereby reducing the likelihood that the resulting candidate molecule 135 will exhibit defects that could cause the candidate molecule 135 to fail as a viable therapeutic agent (e.g., chemical or physical instability, poor pharmacokinetics, immunogenicity, etc.). Discovering such a defect late in the drug development cycle can be particularly disastrous, which is why improving the properties of selectable molecule 133 in the manner disclosed in this paper can reduce the long duration, high cost, and unique risks associated with drug development.

[0077] Referring again to Figure 1, the attribute calculation model 115 can be trained based on a training dataset 125 generated by the training engine 120. In some cases, the training dataset 125 may include multiple pairs of sample molecules. Each pair of sample molecules may include a first sample molecule and a second sample molecule. Therefore, in some cases, the attribute calculation model 115 can be trained by at least applying the attribute calculation model 115 to determine, for each pair of sample molecules in the training dataset 125, the attribute difference corresponding to the difference between a first value of the attribute exhibited by the first sample molecule and a second value of the attribute exhibited by the second sample molecule. Once trained, the attribute calculation model 115 can be applied to determine, for each input molecule pair including the first input molecule and the second input molecule, the difference between the first value of the attribute exhibited by the first input molecule and the second value of the attribute exhibited by the second input molecule. As described in more detail below, the trained attribute calculation model 115 can be applied to identify one or a combination of point mutations affecting a single amino acid residue and point mutations that cause attribute improvement. Select molecules, such as lead molecules, can be modified to include these mutations, resulting in candidate molecules (e.g., for wet laboratory analysis) that exhibit better properties than lead molecules.

[0078] In some example embodiments, training engine 120 can generate training dataset 125 by at least collectively increasing a set of sample molecules, each of which has a known value for an attribute, by pairing the same sample molecule with multiple other sample molecules. For example, in Figure 1 In the example shown, a training dataset 125 can be generated to include a first molecule pair 170a and a second molecule pair 170b. The first molecule pair 170a can be generated by pairing the first sample molecule 175a with the second sample molecule 175b. Furthermore, the first sample molecule 175a can also be paired with a third sample molecule 175b to form the second molecule pair 170b. Therefore, with known values ​​for the attribute... When individual sample protein molecules are available, the combinatorial expansion described herein can generate Each pair of molecules is included in the training dataset 175. Although One sample molecule may not be enough to train a property calculation model to accurately predict the absolute value of the property exhibited by the input molecule, but A few molecular pairs can form a sufficiently large and well-distributed training dataset to train the attribute calculation model 115 to accurately predict the differences in attribute values ​​exhibited by one or more pairs of input molecules.

[0079] Figure 2A depicts a flowchart illustrating an example of a process 200 for machine learning-based molecular property enhancement according to some example embodiments. Referring to Figures 1 and 2A, in some example embodiments, process 200 may be performed to train a property computation model 115 based on a training dataset 125.

[0080] At position 202, multiple pairs of sample molecules are generated to be included in the training dataset. In some example embodiments, a training dataset 125 may be generated to include multiple pairs of sample molecules. In some cases, each pair of sample molecules may include a first sample molecule having a first amino acid sequence and a second sample molecule having a second amino acid sequence. In some cases, the first amino acid sequence may exhibit at least one mutation relative to the second amino acid sequence. For example, in some cases, the at least one mutation may include two different types of amino acid residues that occupy the same site position in the first and second amino acid sequences. As described in more detail below, the training dataset can be generated by combining and increasing a set of sample molecules, each of which is associated with a known value for an attribute. Combining and increasing increases the amount of training data and improves its distribution. For example, in some cases, Individual protein molecules can undergo combinatorial augmentation to generate models for training property computation. For the sample protein molecules. Although A single sample of protein molecules may not be sufficient to train a property computation model to accurately predict the absolute value of the property exhibited by the input molecule, but A sufficiently large and uniformly distributed training dataset can be constructed from sample protein molecules to train the attribute computation model to accurately predict the differences in attribute values ​​exhibited by one or more pairs of input molecules. Further details on generating the training dataset for training the attribute computation model are provided at the exemplary Figure 2B.

[0081] At 204, the attribute computation model is trained at least based on a training dataset. In some example embodiments, the attribute computation model 115 may be trained based on training dataset 125 to determine, for each pair of sample molecules, an attribute difference corresponding to the difference between a first value of an attribute exhibited by a first sample molecule and a second value of an attribute exhibited by a second sample molecule. For example, in some cases, the attribute computation model may be trained to generate a relative embedding for each pair of sample molecules, the relative embedding representing a sequence difference corresponding to the difference between a first amino acid sequence of the first sample molecule and a second amino acid sequence of the second sample molecule. Furthermore, in some cases, the attribute computation model may be trained to determine, at least based on the relative embedding of each pair of sample molecules, an attribute difference corresponding to the difference between a first value of an attribute exhibited by a first sample molecule and a second value of an attribute exhibited by a second sample molecule. In some cases, training the attribute computation model may include adjusting one or more parameters of the attribute computation model (e.g., weights, biases, etc.) to reduce (or minimize) errors in the output of the attribute computation model. For example, in some cases, one or more parameters of the attribute calculation model (e.g., weights, biases, etc.) can be adjusted to reduce (or minimize) the difference between the predicted difference of the attribute values ​​determined by the attribute calculation model and the known difference of the attribute values ​​exhibited by each pair of sample molecules (e.g., mean squared error (MSE)).

[0082] In some example embodiments, the attribute computation model may include a language model (e.g., a protein language model) coupled to a neural network (e.g., a convolutional neural network, etc.). For example, in some cases, the language model may be trained to generate a first embedding of a first amino acid sequence and a second embedding of a second amino acid sequence before determining the relative embedding corresponding to the difference between a first embedding and a second embedding. This is done for each pair of amino acid sequences. Generate relative embeddings The paradigm is shown in equation (1) below, where the language model is represented as follows: .

[0083] (1) In equation (1), the language model is trained to generate each individual amino acid sequence. and The embeddings. Alternatively, a language model can be trained to directly generate relative embeddings, meaning that a language model can be trained to generate an output that approximates the difference between a first embedding of the first amino acid sequence and a second embedding of the second amino acid sequence. This alternative paradigm is shown in equation (2).

[0084] (2) Therefore, in some cases, training the attribute computation model may include training a language model to generate relative embeddings for each pair of sample molecules, and training a neural network to determine the differences in attribute values ​​exhibited by each pair of sample molecules, at least based on the relative embeddings. In some cases, training the attribute computation model may include adjusting one or more parameters of the language model (e.g., weights, biases, etc.) so that the language model generates relative embeddings for each pair of sample molecules, thereby enabling the neural network to more accurately predict the differences between the constituent sample molecules. Furthermore, in some cases, training the attribute computation model may include adjusting one or more parameters of the neural network (e.g., weights, biases, etc.) to improve the accuracy of the neural network's predictions based on the relative embeddings generated by the language model.

[0085] At 206, a property computation model is applied to determine the difference in property values ​​exhibited by one or more pairs of input molecules. In some example embodiments, once trained, the property computation model 115 can be applied to determine the difference in property values ​​exhibited by one or more pairs of input molecules. For example, in some cases, a trained property computation model can be applied to determine the difference in property values ​​exhibited by a first input molecule and a second input molecule, wherein the first amino acid sequence of the first input molecule exhibits one or more mutations relative to the second amino acid sequence of the second input molecule. As described in more detail below, in this way, the property computation model can be used to identify one or more mutations that cause property improvement. For example, in some cases, the property computation model can be applied to sample different point mutations affecting a single amino acid residue in order to identify those mutations that cause property improvement. Alternatively and / or additionally, the property computation model can also be applied to sample different combinations of point mutations. One or more candidate molecules can be generated by modifying at least selected molecules (such as lead molecules) to include one or more mutations identified as causing property improvement.

[0086] Figure 2B depicts a flowchart illustrating an example of a process 250 for generating a training dataset for training an attribute computation model, according to some example embodiments. Referring to Figures 1 and 2A-2B, process 250 can generate training dataset 125, for example, by training engine 120. In some cases, process 250 can implement operation 202 of process 200 shown in Figure 2A. As described in more detail below, in some cases, training dataset 125 can be generated by combining a set of sample molecules with known values ​​for the attribute. Combining can include pairing each individual sample molecule with multiple other sample molecules in the set of sample molecules. In doing so, samples with known values ​​for the attribute can be generated. The generation of protein molecules from group samples includes those with known differences in attribute values. Training dataset for sample molecules.

[0087] At position 252, a first sample molecule pair, comprising a first sample molecule and a second sample molecule, is generated to be included in the training dataset. In some example embodiments, the training dataset used to train the attribute computation model to determine the differences in attribute values ​​exhibited by one or more pairs of input molecules can be generated by combining a set of sample molecules having known values ​​for the attribute. In some cases, combining the addition of a set of sample molecules can include pairing each sample molecule with multiple other sample molecules. For example, in some cases, a first pair of sample molecules can be generated to be included in the training dataset by pairing at least the first sample molecule with the second sample molecule. As described in more detail below, the first sample molecule and / or the second sample molecule can be paired with other sample molecules to combine the addition of a set of sample molecules.

[0088] At position 254, a second pair of sample molecules, comprising a first sample molecule and a third sample molecule, is generated to be included in the training dataset. In some example embodiments, a sample molecule that has already been paired with a sample molecule to generate a pair of sample molecules in the training dataset can be paired with other sample molecules to generate additional pairs of sample molecules to be included in the training dataset. For example, a first sample molecule that has already been paired with a second sample molecule to generate a first pair of sample molecules can be further paired with a third sample molecule to generate a second pair of sample molecules to be included in the training dataset.

[0089] At position 256, a third pair of sample molecules, comprising the second and third sample molecules, is generated to be included in the training dataset. In some example embodiments, the combination of a set of sample molecules can be increased for each available sample molecule. For example, in some cases, a second sample molecule paired with a first sample molecule to form a first and a second pair of sample molecules can also be paired with a third sample molecule to form a third pair of sample molecules to be included in the training dataset. Thus, with known values ​​for the attribute... If the sample protein molecules are available, the combined amplification described in this article can generate Each pair of molecules is included in the training dataset. As mentioned above, in In cases where a single sample of protein molecules may not be sufficient to train a property computation model to accurately predict the absolute value of the property exhibited by the input molecule, A sufficiently large and uniformly distributed training dataset of sample protein molecules can be used to train an attribute calculation model to accurately predict the differences in attribute values ​​exhibited by one or more pairs of input molecules.

[0090] Figure 3 depicts a flowchart illustrating an example of a process 300 for machine learning-based molecular property enhancement according to some exemplary embodiments. Referring to Figures 1 and 3, in some exemplary embodiments, a property calculation model 115 may be applied, for example, by a molecular design engine 130 in a workflow in which a lead molecule 133 is modified to generate a candidate molecule 135. For example, in some cases, the property calculation model 115 may be applied to determine which of two input molecules exhibits a better value for drug-like properties, such as expression, affinity, specificity, immunogenicity, etc. in vivo Stability, self-association, etc. As described in more detail below, in some cases, the property calculation model 115 can be applied to identify one or more mutations, including point mutations affecting a single amino acid residue and / or combinations of multiple point mutations that cause property improvement. Therefore, in some cases, candidate molecule 135 can be generated by at least modifying the lead molecule 133 to include one or more mutations identified by the property calculation model 115 as causing property improvement.

[0091] At position 302, a first pair of input molecules is identified, comprising a first input molecule having a first amino acid sequence and a second input molecule having a second amino acid sequence. In some example embodiments, a pair of input molecules comprising a first input molecule and a second input molecule can be generated such that the first amino acid sequence of the first input molecule exhibits one or more mutations relative to the second amino acid sequence of the second input molecule. In some cases, each mutation can be a point mutation, in which two different types of amino acid residues occupy the same site in each of the first and second amino acid sequences. As described in more detail below, a property computation model can be applied to estimate the effect of different point mutations on the property values ​​exhibited by the first and second input molecules. Alternatively and / or additionally, a property computation model can also be applied to estimate the effect of different combinations of point mutations on the property values ​​exhibited by the first and second input molecules. Applying a property computation model in this way allows for the identification of individual point mutations and, in some cases, the identification of combinations of point mutations that cause property improvements.

[0092] At 304, an attribute computation model has been trained to determine the difference in attribute values ​​exhibited by each of the first and second input molecules, based at least on the first and second amino acid sequences. In some example embodiments, the attribute computation model may be applied to determine the difference in attribute values ​​exhibited by each of the first and second input molecules, based at least on the relative embeddings representing the difference between the first and second amino acid sequences. For example, in some cases, the attribute computation model may include a language model (e.g., a protein language model (pLM)) coupled to a neural network (e.g., a convolutional neural network (CNN)). In some cases, the language model may have been trained to generate a first embedding of the first amino acid sequence and a second embedding of the second amino acid sequence. Furthermore, the language model may have been further trained to determine the difference between the first and second embeddings as a relative embedding. In some cases, the neural network may have been trained to determine the difference in attribute values ​​present in each of the first and second input molecules, based at least on the relative embeddings representing the difference between the first and second amino acid sequences.

[0093] At position 306, one or more mutations present in the first amino acid sequence relative to the second amino acid sequence are identified as causing an improvement in the property, at least based on the difference in property value. In some example embodiments, one or more mutations (whether single point mutations or combinations of point mutations) present in one of the first and second input molecules that have a better value for the property can be identified as causing an improvement in the property. For example, if a difference in the property value present in each of the first and second input molecules indicates that the first input molecule exhibits a better value for the property, then one or more mutations present in the first amino acid sequence relative to the second amino acid sequence can be identified as causing an improvement in the property. In some cases, one or more mutations may correspond to the type of amino acid residue occupying one or more sites in the first amino acid sequence that is different from the type of amino acid residue occupying the same sites in the second amino acid sequence.

[0094] At 308, a candidate molecule is generated by modifying at least the selective molecule to include one or more mutations identified as causing property improvement. In some example embodiments, candidate molecules with one or more mutations identified as causing property improvement can be generated for wet laboratory analysis (e.g., in vitro measurements, in vivo characterization, etc.). In some cases, generating a candidate molecule by modifying the selective molecule (such as a lead molecule) to include one or more mutations identified as causing property improvement can enhance one or more desired properties, such as expression, affinity, specificity, etc. in vivoStability, etc. Alternatively and / or additionally, generating candidate molecules by modifying the lead molecule to include one or more mutations recognized as causing property improvements can also minimize one or more undesirable properties, including defects such as chemical or physical instability, poor pharmacokinetics, immunogenicity, etc.

[0095] Figure 4A depicts a flowchart illustrating an example of a process 400 for identifying mutations that improve properties, according to some exemplary embodiments. Referring to Figures 1 and 4A, process 400 can be executed, for example, by a molecular design engine 130, to identify one or more point mutations that cause property improvements, such as expression, affinity, specificity, immunogenicity, etc. in vivo Stability, self-association, etc. As mentioned above, point mutations affect individual amino acid residues, meaning that if two different types of amino acid residues occupy the same site in each of their corresponding amino acid sequences, two input molecules may exhibit point mutations. As described in more detail below, one or more point mutations that cause property improvements can be identified by sampling a variety of different point mutations. For example, in some cases, exhaustive sampling can be performed across all possible mutation sites in the input molecule and all types of amino acid residues (e.g., typical amino acid residues) that can occupy each mutation site.

[0096] At position 402, a first pair of input molecules is generated to include a first input molecule having a first amino acid sequence and a second input molecule having a second amino acid sequence exhibiting a first point mutation, wherein the first site in the second input molecule is occupied by an amino acid residue of a different type than the same site in the first amino acid sequence. In some example embodiments, attribute computation models can be applied to sample different point mutations, including by estimating the effects of having various different types of amino acid residues at individual sites within the protein molecule. For example, in some cases, a first pair of input molecules can be generated to include a first input molecule having a first amino acid sequence that remains fixed during sampling. The first pair of input molecules may further include a second input molecule having a second amino acid sequence, wherein a site in the second input molecule is occupied by an amino acid residue of a different type than the same site in the first amino acid sequence. For example, if the first site in the first amino acid sequence is occupied by leucine (L), the same first site in the second amino acid sequence may be occupied by alanine (A). As described in more detail below, the same first site may be further modified to be occupied by other types of amino acid residues, which in some cases include all possible other types of amino acid residues (e.g., typical amino acid residues).

[0097] At 404, a property calculation model is applied to determine a first difference in the property values ​​exhibited by the first pair of input molecules. In some example embodiments, the effect of a first point mutation present in the second amino acid sequence can be estimated by applying the property calculation model to determine the difference in the property values ​​exhibited by the first pair of input molecules. In some cases, the difference in the property values ​​exhibited by the first pair of input molecules can indicate whether the first or second input molecule exhibits a better value for the property. Furthermore, in some cases, if the second input molecule is determined to exhibit a better value for the property than the first input molecule, the first mutation present in the second amino acid sequence can be identified as causing an improvement in the property. In some cases, operations 402 and 404 can be repeated for each possible type of amino acid residue that can occupy the first point. For example, in some cases, when the first point in the first amino acid sequence is occupied by leucine (L), the property calculation model can be further applied to estimate whether having tryptophan (W) occupy the same first point causes an improvement in the property, or in some cases, whether it causes a greater improvement than alanine (A). As described in more detail below, in addition to sampling across different possible types of amino acid residues at each individual site in the first amino acid sequence of the first input molecule, sampling may further continue to one or more additional sites in the input molecule or, in some cases, to each possible site.

[0098] At position 406, a second pair of input molecules is generated to include the first input molecule and a third input molecule having a third amino acid sequence exhibiting the second point mutation, wherein the second site in the third input molecule is occupied by an amino acid residue of a different type than the same site in the first amino acid sequence. In some example embodiments, to further sample the effect of having different types of amino acid residues at additional sites, a second pair of input molecules may be generated to include the first input molecule and the third input molecule, wherein different types of amino acid residues occupy the second site in each of the amino acid sequences. For example, if the second site in the first amino acid sequence is occupied by glutamine (Q), the second mutation present in the third amino acid sequence could be the same second site occupied by tyrosine (Y).

[0099] At 408, a property calculation model is applied to determine a second difference in the property values ​​exhibited by the second pair of input molecules. In some example embodiments, the property calculation model may be applied to determine whether a second mutation present in the third amino acid sequence causes an improvement in the property. For example, if the second site in the first amino acid sequence is occupied by glutamine (Q), and the same second site in the first amino acid sequence is occupied by tyrosine (Y), then the difference in the property values ​​exhibited by each of the first and third input molecules may indicate whether the presence of glutamine (Q) or tyrosine (Y) at the second site is associated with a better value for the property. In some cases, if the difference in the property values ​​exhibited by the second pair of input molecules indicates that the third input molecule exhibits a better value for the property, then the second mutation may be identified as causing an improvement in the property. Furthermore, in some cases, operations 406 and 408 may be repeated for one or more additional sites in the input molecules, or in some cases for each possible site.

[0100] At 410, the first mutation and / or the second mutation are identified as causing an improvement in the property, based at least on the first and second differences. In some example embodiments, a point mutation present in an amino acid sequence can be identified as causing an improvement in the property if the output of the property calculation model indicates that the amino acid sequence exhibits a better value for the property than another amino acid sequence exhibiting, for example, a different point mutation occupying the same site with different types of amino acid residues. For example, when sampling different types of amino acid residues that can occupy individual sites, a point mutation causing a threshold improvement in the property, or a threshold amount of a point mutation causing the maximum improvement in the property, can be identified. Furthermore, in some cases, such point mutations can be identified for multiple sites in the amino acid sequence or, in some cases, for each possible site. As described in more detail below, a set of point mutations causing an improvement in the property can be identified by combining two or more point mutations, such that combinations of point mutations can be generated for further estimation, in which the property calculation model is applied to identify combinations of one or more point mutations causing an improvement in the property.

[0101] Figure 4B depicts a flowchart illustrating an example of a process for identifying mutations that improve properties, according to some exemplary embodiments. Referring to Figures 1 and 4B, process 450 can be performed, for example, by a molecular design engine 130, to identify combinations of one or more point mutations that cause property improvements, such as expression, affinity, specificity, immunogenicity, etc. in vivoStability, self-association, etc. In some cases, a single combination of point mutations may include two or more point mutations affecting two or more amino acid residues. As described in more detail below, combinations of one or more point mutations can be generated by combining, for example, two or more point mutations identified by performing process 400 as causing property improvement. A pair or more pairs of input molecules, each of which includes a first input molecule and a second input molecule exhibiting different combinations of point mutations, can be estimated by applying a property calculation model. One or more combinations of point mutations can be identified as causing property improvement through an iterative process in which an input molecule having a combination of point mutations identified by the property calculation model as exhibiting a better value for the property during one iteration proceeds to the next iteration, in which the property calculation model is applied to determine the difference in property values ​​exhibited by that input molecule and another input molecule exhibiting a different combination of point mutations.

[0102] At position 452, multiple point mutations causing property improvement are identified. In some example embodiments, multiple point mutations causing property improvement can be identified by sampling different point mutations, for example, by at least changing the type of amino acid residues occupying one or more sites in the amino acid sequence. For example, in some cases, process 400 shown in FIG4A can be iterated once or multiple times to estimate the effect of having different types of amino acid residues at one or more sites in the amino acid sequence. In some cases, multiple point mutations (such as the presence of certain types of amino acid residues at one or more sites in the amino acid sequence) can be identified as causing improvement to the property.

[0103] At 454, multiple combinations of point mutations are generated by combining two or more point mutations selected from a plurality of point mutations. In some example embodiments, multiple combinations of point mutations can be generated, each of which is identified as two or more point mutations that cause an improvement in the property. It should be understood that while a point mutation may individually cause an improvement in the property, combining two or more such point mutations does not necessarily produce an improvement in the property. Therefore, as described in more detail below, combinations of two or more point mutations, each identified as causing an improvement in the property, can be estimated iteratively, for example, by applying a property computation model that has been trained to determine the differences in property values ​​exhibited by a pair of input molecules that exhibit different combinations of point mutations.

[0104] At position 456, a first pair of input molecules is generated to include a first input molecule exhibiting a first combination of point mutations and a second input molecule exhibiting a second combination of point mutations. In some example embodiments, the first and second input molecules may exhibit two different combinations of point mutations, where each point mutation has been identified as causing an improvement to the property, for example, by sampling different point mutations using a property computation model. As described in more detail, a property computation model can also be applied to identify combinations of point mutations that cause property improvements. For example, in some cases, the property computation model can be applied iteratively, in which an input molecule identified as exhibiting a better value for the property during one iteration can proceed to the next iteration, in which the property computation model is applied again to determine whether the same input molecule exhibits a better value for the property than another input molecule exhibiting a different combination of point mutations.

[0105] At position 458, an attribute computation model is applied to determine a first difference in the attribute values ​​exhibited by the first pair of input molecules. In some example embodiments, the attribute computation model may have been trained to determine the differences in attribute values ​​exhibited by two input molecules exhibiting two different combinations of point mutations. For example, in some cases, the attribute computation model may include a language model (e.g., a protein language model (pLM) or the like) that has been trained to generate a representation of the differences between the amino acid sequences of each input molecule in the first pair of input molecules. Furthermore, in some cases, the attribute computation model may include a neural network (e.g., a convolutional neural network (CNN) or the like) that has been trained to determine the differences in attribute values ​​exhibited by the first pair of input molecules based at least on the relative embedding.

[0106] At 460, the first input molecule is identified as exhibiting a better value for the attribute, at least based on a first difference in attribute values. In some example embodiments, the difference in attribute values ​​output by the attribute computation model can identify which of the two input molecules exhibits a better value for the attribute. For example, if the first input molecule exhibits a higher attribute value than the second input molecule, the first input molecule may exhibit a better desired attribute value, such as expression or affinity. Conversely, if the first input molecule exhibits a lower, undesirable attribute value, such as immunogenicity, than the second input molecule, the first input molecule may exhibit a better value for that attribute. As described in more detail below, an input molecule exhibiting a better value for the attribute can proceed to the next iteration, in which another molecule exhibiting a different combination of point mutations is estimated.

[0107] At position 462, a second pair of input molecules is generated to include the first input molecule and a third input molecule exhibiting a third combination of point mutations. In some example embodiments, if the first input molecule in the first pair is identified as exhibiting a better value for the attribute, that first input molecule can be proceeded to the next iteration, in which the first input molecule is paired with a third input molecule exhibiting a different combination of point mutations than the second input molecule. In some cases, the second combination of point mutations may also include individual point mutations that have been identified as causing attribute improvement. As described in more detail below, an attribute computation model can be applied to estimate the second pair of input molecules.

[0108] At 464, a property calculation model is applied to determine a second difference in the property values ​​exhibited by the second pair of input molecules. In some example embodiments, the property calculation model may be applied to determine the difference in the property values ​​exhibited by the second pair of input molecules based at least on the relative embeddings representing the differences in the corresponding amino acid sequences of the first and third input molecules. In some cases, the difference in the property values ​​exhibited by the second pair of input molecules may identify one of the first and third input molecules as exhibiting a better value for the property. In some cases, iterative sampling of different combinations of point mutations may continue, wherein the first or third input molecule enters another iteration, in which it is paired with another input molecule exhibiting another different combination of point mutations. Furthermore, in some cases, iterative sampling may also include applying a property calculation model to identify which of the second combination of second input molecules exhibiting point mutations and the third combination of third input molecules exhibiting point mutations has a better value for the property.

[0109] At 466, at least based on a second difference in attribute values, a first combination or a third combination of point mutations is identified as causing an improvement in the attribute. In some example embodiments, one or more combinations of point mutations may be identified as causing an improvement in the attribute. For example, in some cases, if a first input molecule exhibits a better value for the attribute than a third input molecule exhibiting a third combination of point mutations, a first combination of point mutations present in the first input molecule may be identified as causing an improvement in the attribute. It should be understood that iterative sampling of different combinations of point mutations may continue until one or more criteria are met. For example, in some cases, one or more criteria may include performing sampling iterations of threshold amounts, each of which estimates two different combinations of point mutations. Alternatively and / or additionally, one or more criteria may include combinations of mutations that have been identified as causing a threshold magnitude of attribute improvement.

[0110] At 468, a candidate molecule is generated by modifying the selected molecule to include at least one of a first combination and a third combination of point mutations recognized as causing property improvement. In some example embodiments, a candidate molecule can be generated by modifying the lead molecule to include a combination of point mutations recognized as causing property improvement. In some cases, candidate molecules including combinations of point mutations can be synthesized for wet laboratory analysis, such as in vitro measurements, in vivo characterization, etc. In some cases, generating a candidate molecule by modifying the lead molecule to include a combination of point mutations recognized as causing property improvement can enhance one or more desired properties, such as expression, affinity, specificity, etc. in vivo Stability, etc. Alternatively and / or additionally, generating candidate molecules by modifying the lead molecule to include one or more mutations recognized as causing property improvements can also minimize one or more undesirable properties, including defects such as chemical or physical instability, poor pharmacokinetics, immunogenicity, etc. In some cases, modifying the lead molecule to include different combinations of two or more point mutations can further increase the diversity of the resulting candidate molecules, thereby enabling the selection of a larger number of developable therapeutic candidates.

[0111] Figure 5A depicts a schematic diagram illustrating an example of a workflow 500 for enhancing molecular properties using machine learning, according to some example embodiments. Figure 5A In the example of workflow 500 shown, attribute calculation model 115 can be applied to determine the differences in attribute values ​​exhibited by each of the first amino acid sequence 501 and the second amino acid sequence 503. In some cases, each of the first amino acid sequence 501 and the second amino acid sequence 503 can be an antibody, or alternatively a part of an antibody, such as a parasite, variable region (Fv), antigen-binding fragment (Fab), variable region, complementarity-determining region (CDR), etc. Figure 5A depicts an example of each of the first amino acid sequence 501 and the second amino acid sequence 503 being formed by linking variable domains on the heavy and light chains of an antibody.

[0112] Referring again to Figure 5A, the first amino acid sequence 501 differs from the second amino acid sequence 503. In the example shown in Figure 5A, the second amino acid sequence 503 exhibits a mutation 505 in which alanine (A) occupies the site in the second amino acid sequence 503 that was occupied by leucine (L) in the first amino acid sequence 501. Attribute computation model 115 can be applied to determine the difference in attribute values ​​exhibited by each of the first amino acid sequence 501 and the second amino acid sequence 503, thereby determining whether the presence of alanine (A) instead of leucine (L) causes an improvement in the attribute. For example, Figure 5A shows that attribute computation model 115 includes a language model 510 (e.g., a protein language model (pLM), etc.) that has been trained to generate a first embedding 511 of the first amino acid sequence 501 and a second embedding 513 of the second amino acid sequence 503. In some cases, a relative embedding 515 representing the difference between the first amino acid sequence 501 and the second amino acid sequence 503 can be generated. For example, in some cases, a relative embedding 515 can be generated to correspond to the difference between the first embedding 511 and the second embedding 513. As further illustrated in Figure 5A, in some cases, a convolutional neural network 517 (or other types of neural networks) can be applied to determine the attribute value (in Figure 5A) at least based on the relative embedding 515. Differences in attribute prediction.

[0113] In some example embodiments, the attribute calculation model 115 can be applied to iteratively improve attribute values. For example, the difference in attribute values ​​exhibited by the first amino acid sequence 501 and the second amino acid sequence 503 can indicate whether a mutation 505 present in the second amino acid sequence 503 (e.g., alanine (A) instead of leucine (L)) causes an improvement in the attribute. If the second amino acid sequence 503 is determined to exhibit a better value for the attribute than the first amino acid sequence 501, the mutation 505 can be identified as causing an improvement in the attribute. In some cases, the effect of mutation 505 can be further compared with one or more additional point mutations to identify multiple point mutations causing attribute improvement. Furthermore, in some cases, two or more point mutations causing attribute improvement can be combined, and the attribute calculation model 115 can be applied to estimate the effect of different combinations of point mutations. For example, in some cases, an iterative process can be used to identify one or more combinations of point mutations that cause attribute improvement, in which an input molecule that has been identified as exhibiting a combination of point mutations that shows better values ​​for the attribute during one iteration enters the next iteration, in which a attribute computation model is applied to determine the difference in attribute values ​​exhibited by the input molecule and another input molecule with a different combination of point mutations.

[0114] Figure 5B depicts a schematic diagram illustrating sampling of point mutations across mutation sites and amino acid residue types according to some example embodiments. Referring to Figures 1, 4A, and 5B, in some example embodiments, attribute calculation model 115 can be applied to sample different point mutations, such as in process 400 shown in Figure 4A. Figure 5A further illustrates sampling of different point mutations across amino acid residues and mutation site types. In the example shown in Figure 5A, attribute calculation model 115 can be applied to determine the difference between the attribute values ​​exhibited by amino acid sequence 550 and variants of amino acid sequence 550, where amino acid residue types occupy one or more individual sites. For example, in some cases, attribute calculation model 115 can be applied to determine the difference in attribute values ​​exhibited by a first variant of amino acid sequence 551 where amino acid sequence 550 and alanine (A) instead of leucine (L) occupy a specific site. Furthermore, the attribute calculation model 115 can be applied to determine the difference in attribute values ​​exhibited by the second variant of amino acid sequence 552, which occupies the site of amino acid sequence 550 and aspartic acid (D) instead of leucine (L). The example in Figure 5B further illustrates sampling for point mutations at different mutation sites across amino acid sequence 550. In some cases, sampling can be exhaustive, meaning that the attribute calculation model 115 is applied to determine the difference in attribute values ​​exhibited by amino acid sequence 550 and the variants of amino acid sequence 550 at each mutation site and for each possible type of amino acid residue (e.g., 20 typical amino acid residues). Where a variant of amino acid sequence 550 is determined to exhibit a better value for the attribute than the original amino acid sequence 550, the point mutation present in the variant of amino acid sequence 550 (e.g., alanine (A) or aspartic acid (D) at a specific site instead of leucine (L)) can be identified as causing the improvement in attribute. As described above, in some cases, two or more point mutations can be identified as causing attribute improvement, and attribute computation model 115 can be applied to estimate the effect of the combination of point mutations on attribute values.

[0115] Experimental Example In some example embodiments, labeled (or annotated) data from a Cosmos-Optimized Comprehensive Substitution (COSMO) experiment is used to train the attribute computation model 115. This experiment involves a mutational scan of residues in the antibody complementarity-determining region (CDR) with all natural amino acids (except cysteine) to generate a dataset containing binding affinity markers for ~500-1000 point variants surrounding the lead molecule. These ~500-1000 point variants and their corresponding binding affinity values ​​(…) The data were combined and added to generate a training dataset containing nearly 125,000 unique sample molecule pairs. A property computation model 115 was trained on these ~125,000 sample molecule pairs, each of which had a label (or annotation) corresponding to the differences in binding affinity exhibited by the constituent sample molecules.

[0116] As described above, in some example embodiments, the attribute computation model 115 can operate on paired input molecules such as the amino acid sequences of protein molecules. The amino acid sequences can first be fed through a pre-trained protein language model (pLM), which has been trained to extract the embedding of each individual amino acid sequence, and then a relative embedding representing the difference between them is determined. Figure 6 illustrates the effect of different mutations on the relative embedding representing the difference between two amino acid sequences. In the example shown in Figure 6, each amino acid sequence is a variable region of the heavy chain of an antibody (…). ) and the variable region of the light chain ( The connection of ) is shown in Figure (a). Figure (a) shows the relative embeddings of two amino acid sequences that differ by a single point mutation, while Figure (b) shows the relative embeddings of two amino acid sequences that have a higher edit distance. Although different types of protein language models can be used to implement the attribute computation model 115, it should be understood that in this case, a protein language model (pLM) refers to a language model that has been specifically or primarily trained on protein sequences (e.g., antibody sequences). Such a protein language model (pLM) is coupled with a convolutional neural network (CNN) trained on these relative embeddings and attribute labels to output a prediction of the attribute difference between the two sequences. An example of this setup is shown in Figure 5, where a language model 510 is coupled with a convolutional neural network 517. The attribute computation model 115 can be deployed as a ranking model to score combinations of known mutations, or as shown in Figure 5A, as part of a generative workflow to generate molecules with incrementally better attribute values.

[0117] The performance of the property computation model 115 in determining differences (e.g., improvements) in property values ​​such as binding affinity was estimated on three antibody-antigen binding affinity datasets of varying sizes and complexities, including lead molecule A, anti-EGFR, and anti-IL6. The results are shown in Figures 7A through 7C, which depict graphs illustrating the measured improvements in binding affinity and the improvements predicted by the property computation model 115. In this experiment, binding affinity was defined as the logarithmic transformation of the equilibrium dissociation constant. ,in and represents the measured dissociation and association rate constants, respectively. Each binding affinity dataset includes labels for variants surrounding the leader molecule, in this case, a known conjugate of a given antigen. The difference in binding affinity relative to the leader molecule is then calculated and expressed as . .

[0118] Data preparation for training and testing the attribute computation model 115 included restricting each binding affinity dataset to those antibodies that bind only to the target antigen. As used herein, “adhesive” is defined as having a measurable affinity between 4 and 11. After the first round (R0) design of target A, noisy labels and those with unsuitable surface plasmon resonance (SPR) curves were removed (evaluated by visual inspection). The datasets for lead molecule A and anti-IL6 were randomly divided into 85% / 15% training / test groups. The anti-EGFR training set contained half of all alanine scanpoint mutants and half of randomly selected high edit distance variants. The remaining half was reserved for testing.

[0119] First, the antibody variable domain sequences were aligned with and concatenated using the ANARCI AHo numbering scheme to form variable domain sequences of consistent length (heavy chain followed by light chain). These sequences were then fed into a dataset from one of three pre-trained protein language models, where pLM A is a transformer-based masked language model trained on over 500 million natural antibody sequences from the Observed Antibody Space (OAS), pLM B is a causal protein language model with 11 million parameters trained on antibodies from the OAS and protein complexes from the UniRef50 general protein and protein database (PDB33), and pLM C is a general protein language model trained on ~65M protein sequences from the UniRef5032 dataset.

[0120] The embeddings generated by the protein language model are frozen, resized, normalized to the minimum tensor value, and rescaled according to the maximum tensor value. This is done for each pair of amino acid sequences. Relative embedding It is obtained by subtracting the normalized embedding from equation (3) below.

[0121] (3) in This refers to a protein language model (pLM). That is, in some cases, a protein language model (pLM) can be trained to generate a single embedding, while relative embeddings are the differences between individual embeddings. Alternatively, as mentioned above, it is also possible... Relative embeddings are generated by directly approximating the differences between individual embeddings using a protein language model (PLM). In this case, the protein language model (PLM) does not generate embeddings for each individual amino acid sequence, but rather generates outputs that approximate the differences between embeddings of individual amino acid sequences. For downstream tasks, these relative embeddings... It is used as input to a convolutional neural network (CNN) that has been trained to predict amino acid sequences. and The differences in attributes between them. The ResNet18 architecture implemented in PyTorch and the learning rate were used. The Adam optimizer was used. The model was trained using the mean squared error loss function and a batch size of 32 for 100 epochs. For inference, the model was trained on a given protein language. Model calculation and lead molecule The relative embedding is used to predict property differences relative to the lead molecule.

[0122] Given the above framework, the attribute calculation model 115 is set in the retention test configuration of the first dataset associated with the lead molecule A. It performs well on individual variants and achieves... Pearson correlation coefficient and The Spearman rank correlation coefficient (see Figure 7A). Two other binding affinity datasets for anti-EGFR and anti-IL6 leader molecules are also included, each with only ~100 affinity markers. These binding affinity datasets also include more distant leader molecule variants with an edit distance of 20. Such variants are typically processed through multiple rounds. Computer simulation and / or in vitro Affinity maturation was achieved. For the attribute computation model 115, trained on point variable data from alanine scan mutagenesis and a few higher edit distance variables, the attribute computation model 115 was able to predict anti-EGFR leader molecules with... Pearson correlation coefficient and Differences in binding affinity of variants of Spearman rank correlation coefficient (See Figure 7B). Attribute calculation model 115 performs well even without any point variable data, as shown in Figure 7C for the anti-IL6 lead molecule, with a Pearson correlation coefficient of [value missing]. The Spearman rank correlation coefficient is .

[0123] The performance of the property computation model 115 in generating novel variants of lead molecules with better binding affinity was also estimated. For a set of lead molecules A, all COSMO mutations that individually improved affinity were selected and used to generate combinations with edit distances of 3 to 4. The property computation model 115 then determined the differences in binding affinity, or the variants exhibiting these COSMO mutation combinations relative to the lead molecules. Iteratively generating novel variants of the lead molecule involves using a genetic algorithm (GA) to select and mutate amino acid sequences, sampling a broad design space, and iteratively improving the predicted variants. Forty-six top-level designs generated from the output of attribute computation model 115 were selected for experimental testing, and the results are shown in Figure 8A. As shown in Figure 8A, 72% of the designs in the first round (R0) were successfully expressed in mammalian cells, and 54% bound to the target antigen, comparable to the binding rate of the COSMO point mutant (59%). The performance of attribute computation model 115 on the regression task of the design set is shown in Figures 7D-7F.

[0124] A similar procedure was followed to design variants of the anti-EGFR lead molecule. First, all point mutations in the training set that individually improved affinity (in the alanine scan) or within variants with higher edit distances were selected. Instead of using a genetic algorithm, all combinations of these mutations with edit distances between 3 and 11 were exhaustively sampled and scored using the property computation model 115. Any mutation with a higher affinity than the lead molecule (e.g., The designs with better binding affinity were further narrowed down to 44 final amino acid sequences for experimental testing. In the first round (“R1”), 91% and 89% of the designs expressed and bound to EGFR, respectively (see Figure 8B, R1), with 79% of these binders exhibiting a stronger measured affinity than the lead molecule (3.0 nM). Eleven designs showed at least a tenfold improvement in binding affinity. While the strongest binder had similar affinity (~100 pM) to high edit distance variants in the training set, the binder was generated at a higher rate using property calculation model 115, comparable to the binder of the alanine point mutant. A second round of designs (“R2”) was then developed, incorporating these data into the training set. A single sequence from the second round design (“R2”) was selected for experimental testing (Figure 8B, R2), and this design expressed, bound to EGFR, and exhibited an improved affinity of 63 pM.

[0125] The performance of property calculation model 115 was also tested when given a dataset of ~100 lead molecule variants, such as in the case of an anti-IL6 lead molecule. Similarly, all mutations from affinity-improving variants with edit distances up to 9 were selected and thoroughly combined. The resulting amino acid sequences were selected and scored using property calculation model 115. The output of property calculation model 115 indicates a binding affinity greater than that of the lead molecule (e.g., Eight subsets of the designs were used for experimental testing (see Figure 8C, R2). All selected variants were successfully expressed, bound to IL6, and exhibited better binding affinity than the lead molecule (1.4 nM). The binding affinity of four designs was more than 3-fold higher than that of the lead molecule, exceeding that of all variants in the training set.

[0126] Although the property calculation model 115 is a sequence-based model, analysis of the top designs reveals the assumed structural mechanism behind the observed improved bid affinity. The variable domain (Fab) structures of the anti-EGFR leader molecule and the top first-round variants designed using property calculation model 115 were determined by X-ray crystallography at resolutions of 2.4 Å (PDB entries) and 2.1 Å (PDB entries), respectively. The variable domain overlay is shown in Figure 8F. The output of property calculation model 115 was identified in aEGFR-R1-1 across CDR-H1, CDR-H3, and the framework region. Mutating the Kabat position VH 97 from glycine to aspartic acid – consistent across all top designs (Figure 8E) – altered the construction of CDR-H3. The additional proline at VH 98 further stabilized this binding-compatible construction in aEGFR-R1-1 relative to the leader molecule. All three top designs also mutate VH 34 of the CDR-H1 base from valine to a longer aliphatic amino acid. The highest affinity design (aEGFR-R2) includes nine head editors, six of which are included in the frame. These distal mutations can negatively impact expression because the yield of aEGFR-R2 is less than half that of the lead molecule (Figure 8E).

[0127] Instead of the experimental structures, the structures generating the highest binding affinity anti-IL6 variants using ABodyBuilder219 and property calculation model 115 were predicted computationally (see Figure 8H). While the same VH S98P mutation was observed in the first two designs, the predicted CDR-H3 construction remained unchanged. Conversely, the mutation in CDR-H2 (VH N52aH) drove the loop to extend into solution (see Figure 8G).

[0128] The pairwise framework of attribute computation model 115 can be constructed from various language model embeddings, including domain-specific protein language models (e.g., antibody-specific pLM) and more general protein language models. Figure 9 compares the performance of three different variants of attribute computation model 115, each of which includes a different type of protein language model. In this example, pLM A is an antibody-specific protein language model trained on 500 million natural antibody sequences, pLM B is a hybrid trained on the same antibody sequences plus a general protein universe, and pLM C is a general protein language model trained on a general protein universe. The performance of attribute language model 115 with different protein language models is compared across all three binding affinity datasets (lead molecule A, anti-EGFR, and anti-IL6).

[0129] Figure 9 shows the downstream Performance on prediction tasks varies depending on the dataset. In some cases, using different protein language models outperforms the design. For example, the anti-IL6 R2 design, generated using embeddings from pLM B, outperformed pLM A in three out of four estimation metrics. However, compared to general protein language models (e.g., pLM C), protein language models trained primarily or entirely on antibody libraries (e.g., pLM A and B) produced the most predictive changes to the attribute computation model 115.

[0130] Figure 10 shows the rate of adhesive generation while utilizing the output of the property calculation model 115 to predict the difference in binding affinity. Data cleaning can further improve the performance. For example, Figures 10(a) and (b) show the attribute calculation model after data cleaning between the first design round R0 (Figure 10(a)) and the second design round R1 (Figure 10(b)). The predicted binding affinity differences are comparable. Removing five mislabeled adhesives from the training dataset improves the representation of designs generated using the property calculation model 115 (Fig. 10(c)) and binding rates (Fig. 10(d)). An example surface plasmon resonance (SPR) sensor plot illustrates the incorrect (Fig. 10(d)) and correctly labeled (Fig. 10(d)) adhesives. Figure 10 (e) data points.

[0131] Figure 11 shows that designs generated using attribute computation model 115 maintain affinity and expression at higher edit distances. For the anti-EGFR lead molecule (Figure 11(a)) and the anti-IL6 lead molecule (Figure 11(b)), the expression levels of variants in the training set versus edit distance and the designs generated using attribute computation model 115 are shown.

[0132] Figure 12 depicts the original surface plasmon resonance (SPR) curves of the top design generated using property calculation model 115 with lead molecule A (Figure 12(a)), anti-EGFR lead molecule (Figure 12(b)), and anti-IL6 lead molecule (Figure 12(c)). The measurement mode changes from single-cycle to multi-cycle dynamics between detecting lead molecule A and the anti-EGFR lead molecule.

[0133] Figure 13 depicts a block diagram illustrating an example of a computing system 1300 according to some exemplary embodiments. Referring to Figures 1 through 13, the computing system 1300 may be used to implement an attribute calculation engine 110, a training engine 120, a molecular design engine 110, a client device 140, and / or any of its components.

[0134] As shown in Figure 13, the computing system 1300 may include a processor 1310, a memory 1320, a storage device 1330, and an input / output device 1340. The processor 1310, memory 1320, storage device 1330, and input / output device 1340 may be interconnected via a system bus 1350. The processor 1310 is capable of processing instructions for execution within the computing system 1300. Such executed instructions may implement one or more components, such as an attribute calculation engine 110, a training engine 120, a molecular design engine 110, a client device 140, etc. In some example embodiments, the processor 1310 may be a single-threaded processor. Alternatively, the processor 1310 may be a multi-threaded processor. The processor 1310 is capable of processing instructions stored on the memory 1320 and / or storage device 1330 to display graphical information for a user interface provided via the input / output device 1340.

[0135] Memory 1320 is a computer-readable medium, such as a volatile or non-volatile computer-readable medium, that stores information within computing system 1300. For example, memory 1320 may store a data structure representing a configuration object database. Storage device 1330 provides persistent storage for computing system 1300. Storage device 1330 may be a floppy disk device, hard disk device, optical disk device, magnetic tape device, or other suitable persistent storage device. Input / output device 1340 provides input / output operations for computing system 1300. In some exemplary embodiments, input / output device 1340 includes a keyboard and / or a pointing device. In various specific embodiments, input / output device 1340 includes a display unit for displaying a graphical user interface.

[0136] According to some exemplary embodiments, input / output device 1340 may provide input / output operations for network devices. For example, input / output device 1340 may include an Ethernet port or other networking port to communicate with one or more wired and / or wireless networks (e.g., local area network (LAN), wide area network (WAN), Internet).

[0137] In some exemplary embodiments, the computing system 1300 can be used to execute various interactive computer software applications that can be used to organize, analyze, and / or store data in various formats. Alternatively, the computing system 1300 can be used to execute any type of software application. These applications can be used to perform various functions, such as planning functions (e.g., generating, managing, and editing spreadsheet documents, word processing documents, and / or any other objects), computing functions, communication functions, etc. Applications may include various additional functionalities or may be standalone computing products and / or functions. Once activated within the application, the functionality can be used to generate a user interface provided via the input / output device 1340. The user interface can be generated by the computing system 1300 and presented to the user (e.g., on a computer screen monitor, etc.).

[0138] One or more aspects or features of the subject matter described herein can be implemented as digital electronic circuits, integrated circuits, specially designed ASICs, field-programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof. These aspects or features may be implemented in one or more computer programs that are executable and / or interpretable on a programmable system, which includes at least one programmable processor (which may be dedicated or general-purpose, coupled to receive and send data and instructions to it), a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. Typically, clients and servers are remotely configured to each other and generally interact via a communication network. Relationships between clients and servers arise from computer programs running on their respective computers and the client-server relationships between them.

[0139] These computer programs may also be referred to as programs, software, software applications, applications, components, or code, including machine instructions for a programmable processor, and may be implemented in high-level procedural and / or object-oriented programming languages ​​and / or in assembly / machine language. As used herein, the term "machine-readable medium" refers to any computer product, apparatus, and / or device (such as, for example, a disk, optical disk, memory, and programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor. Machine-readable media may (such as, for example, non-transitory solid-state memory or magnetic hard disk drive or any equivalent storage medium) store such machine instructions non-transitory. Machine-readable media may (such as, for example, a processor cache or other random access memory associated with one or more physical processor cores) optionally or additionally store such machine instructions transiently.

[0140] To provide interaction with a user, one or more aspects or features of the subject matter described herein can be implemented on a computer having a display device (such as, for example, a cathode ray tube (CRT) or liquid crystal display (LCD) or a light-emitting diode (LED) monitor for displaying information to the user) and a keyboard and pointing device (such as, for example, a mouse or trackball, through which the user can provide input to the computer). Other kinds of devices can also be used to provide interaction with the user. For example, feedback provided to the user can be any form of sensory feedback, such as, for example, visual feedback, auditory feedback, or tactile feedback; input from the user can be received in any form, including sound, speech, or tactile input. Other possible input devices include touchscreens or other touch-sensitive devices, such as single-point or multi-point resistive or capacitive trackpads, speech recognition hardware and software, optical scanners, optical indicators, digital image capture devices, and associated interpretation software, etc.

[0141] In the foregoing description and claims, phrases such as “at least one” or “one or more” may appear, followed by a list of combinations of elements or features. The term “and / or” may also appear in a list of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which it is used, the phrase is intended to mean any element or feature listed alone, or any other recounted element or feature in combination with any other recounted element or feature. For example, the phrases “at least one of A and B”; “one or more of A and B”; and “A and / or B” are each intended to mean “A alone, B alone, or A and B together”. A similar interpretation applies to lists comprising three or more items. For example, the phrases “at least one of A, B, and C”; “one or more of A, B, and C”; and “A, B, and / or C” are each intended to mean “A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together”. The use of the term "based on" in the above and claims is intended to mean "at least partially based on," so that undescribed features or elements are also permissible.

[0142] Depending on the desired construction, the subject matter described herein can be embodied in systems, apparatuses, methods, and / or articles of manufacture. The embodiments set forth in the foregoing description do not represent all embodiments consistent with the subject matter described herein. Rather, they are merely some examples consistent with aspects related to the described subject matter. Although some variations have been described in detail above, other modifications or additions are possible. In particular, other features and / or variations may be provided in addition to those features and / or variations set forth herein. For example, the above embodiments may be provided for various combinations and sub-combinations of the disclosed features and / or for combinations and sub-combinations of several further features disclosed above. Furthermore, the logical flows depicted in the drawings and / or described herein do not necessarily require the specific order or sequential order shown to achieve the desired results. Other embodiments may be within the scope of the following claims.

Claims

1. A computer-implemented method, comprising: Generate a training dataset comprising a pair of sample molecules, the sample molecule pair comprising a first molecule and a second molecule; The attribute computation model is trained at least on the training dataset, wherein the attribute computation model is trained to at least The sequence differences between the first molecule and the second molecule are determined, including differences in the amino acid sequences of the first molecule and the second molecule. Generate relative embeddings that represent the sequence differences and associate the sequence differences with the sample molecule pairs, and The property differences for the sample molecule pairs are determined at least based on the relative embedding of the sample molecule pairs, wherein the property differences for the molecule pairs correspond to the difference between the value of the property exhibited by the first molecule and the value of the property exhibited by the second molecule; Receives a pair of input molecules; as well as The attribute calculation model is applied to determine the attribute differences between the pair of input molecules.

2. The method of claim 1, wherein the property includes expression, affinity, specificity, exploitability, or bioactivity.

3. The method according to any one of claims 1 to 2, wherein the sample molecule pair is generated by pairing at least a set of sample molecules having known values ​​for the said property.

4. The method according to any one of claims 1 to 3, wherein generating the training dataset comprises: The sample molecule pairs are generated by pairing a first sample molecule and a second sample molecule from a set of sample molecules; and The first sample molecule is paired with a third sample molecule from the group of sample molecules to generate another sample molecule pair.

5. The method according to any one of claims 1 to 4, wherein the attribute calculation model comprises a language model, and wherein training the attribute calculation model comprises training the language model to generate the relative embeddings of the sample molecule pairs.

6. The method of claim 5, wherein the language model is trained to at least Generate a first embedding for the first sample molecule and a second embedding for the second sample molecule for the sample molecule pair; and The relative embedding of each sample molecule pair is generated by at least determining the difference between the first embedding and the second embedding.

7. The method of claim 6, wherein the difference between the first embedding and the second embedding corresponds to the sequence difference comprising the difference between the amino acid sequence of the first molecule and the amino acid sequence of the second molecule.

8. The method of any one of claims 5 to 7, wherein the attribute calculation model further comprises a neural network coupled to the language model, and wherein training the attribute calculation model comprises training the neural network to determine the attribute differences for each sample molecule pair based at least on the relative embeddings of each sample molecule pair.

9. The method according to any one of claims 1 to 8, further comprising: One or more mutations associated with improvements to the property are identified, at least based on the property differences between the pair of input molecules.

10. The method of claim 9, wherein the one or more mutations comprise point mutations, wherein the pair of input molecules differ at a single position in each corresponding amino acid sequence.

11. The method according to any one of claims 9 to 10, wherein the one or more mutations comprise a combination of multiple point mutations, wherein the pair of input molecules differ at multiple positions in each corresponding amino acid sequence.

12. The method according to any one of claims 1 to 11, wherein each sample molecule comprises an antibody or a portion thereof.

13. The method according to any one of claims 1 to 12, wherein each sample molecule comprises a variable region of an antibody, an antigen-binding region, a heavy chain and / or a light chain.

14. A system comprising: At least one data processor; as well as At least one memory storing instructions that, when executed by the at least one data processor, cause operation including the method according to any one of claims 1 to 13.

15. A non-transitory computer-readable medium storing instructions that, when executed by the at least one data processor, cause operation comprising the method according to any one of claims 1 to 13.

16. A method for identifying mutations, comprising: An attribute calculation model is applied to determine the difference between the value of the attribute exhibited by the first input molecule and the value of the attribute exhibited by the second input molecule; Determine the amino acid sequence of the first input molecule and the amino acid sequence of the second input molecule; The first mutation is identified by comparing the amino acid sequence of the first input molecule with the amino acid sequence of the second input molecule; as well as The value by which the first mutation improves the attribute is identified at least based on the difference in the value of the attribute exhibited by the first input molecule and the second input molecule.

17. The method of claim 16, wherein the property includes expression, affinity, specificity, exploitability, or bioactivity.

18. The method according to any one of claims 16 to 17, further comprising: The attribute calculation model is applied to determine the difference in the value of the attribute exhibited by the first input molecule and the third input molecule; Determine the amino acid sequence of the third molecule; The second mutation is identified by comparing at least the amino acid sequence of the first input molecule with the amino acid sequence of the third input molecule; as well as The second mutation improves the value of the attribute based at least on the difference in the value of the attribute exhibited by the first input molecule and the third input molecule.

19. The method of claim 18, wherein the first mutation comprises an amino acid residue of one type occupying a first position in the amino acid sequence of the first input molecule, and wherein the second mutation comprises an amino acid residue of the same type occupying a second position in the amino acid sequence of the first input molecule.

20. The method of any one of claims 18 to 19, wherein the first mutation comprises a position in the amino acid sequence of the first input molecule occupied by a first type of amino acid residue, and wherein the second mutation comprises the same position in the amino acid sequence of the first input molecule occupied by a second type of amino acid residue.

21. The method according to any one of claims 18 to 20, further comprising: Generate at least one output molecule exhibiting the first mutation and / or the second mutation.

22. The method according to any one of claims 16 to 21, wherein each input molecule comprises an antibody or a portion thereof.

23. The method according to any one of claims 16 to 22, wherein each sample molecule comprises a variable region of an antibody, an antigen-binding region, a heavy chain, and / or a light chain.

24. A system comprising: At least one data processor; as well as At least one memory storing instructions that, when executed by the at least one data processor, cause operation including the method according to any one of claims 16 to 23.

25. A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operation comprising the method according to any one of claims 16 to 23.

26. A computer-implemented method, comprising: Identify multiple point mutations, each of which is associated with an improvement in the value of a property; Generate multiple mutation combinations, wherein each of the multiple mutation combinations includes a point mutation selected from the multiple point mutations; Generate multiple input molecules, wherein each of the multiple input molecules exhibits a mutation combination from the multiple mutation combinations; A trained attribute computation model is applied to determine the difference between the values ​​of the attribute exhibited by two input molecules from the plurality of input molecules; as well as At least one mutation combination is identified as being associated with the improvement of the value of the property, based at least on the difference between the values ​​of the property exhibited by two of the plurality of input molecules.

27. The method of claim 26, wherein each input molecule is generated by applying the one combination of mutations to the amino acid sequence of the selected molecule.

28. The method according to any one of claims 26 to 27, wherein the property includes expression, affinity, specificity, exploitability, or bioactivity.

29. A system comprising: At least one data processor; as well as At least one memory storing instructions that, when executed by the at least one data processor, cause operation including the method according to any one of claims 26 to 28.

30. A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operation comprising the method according to any one of claims 26 to 28.

31. A computer-implemented method, comprising: Identify a first pair of input molecules, including a first input molecule and a second input molecule, wherein the first input molecule exhibits a mutation relative to the second input molecule; A trained attribute calculation model is applied to determine the difference between a first value of an attribute exhibited by the first input molecule and a second value of the attribute exhibited by the second input molecule, based at least on a first amino acid sequence of the first input molecule and a second amino acid sequence of the second input molecule. The mutation is identified as an improvement to the attribute based at least on the difference between the first and second values ​​of the attribute; as well as Candidate molecules are generated by modifying at least selected molecules to include the mutations.

32. The method of claim 31, further comprising: Identify a second pair of input molecules including the first input molecule and the third input molecule, wherein the first input molecule exhibits another mutation relative to the third input molecule; The attribute calculation model is applied to determine, at least based on the first amino acid sequence of the first input molecule and the third amino acid sequence of the third input molecule, the difference between the first value of the attribute exhibited by the first input molecule and the third value of the attribute exhibited by the third input molecule; and The difference between the first and third values ​​of the attribute indicates that the first input molecule exhibits a better value for the attribute than the third input molecule, thus identifying the first pair of input molecules as including the first input molecule rather than the third input molecule.

33. The method according to any one of claims 31 to 32, wherein the mutation comprises a plurality of point mutations.

34. The method of claim 33, further comprising: The plurality of point mutations are identified by recognizing at least each individual point mutation as a value that improves the property.

35. The method of claim 34, wherein identifying each individual point mutation as a value that improves the property includes Generate multiple pairs of input molecules, where each pair of input molecules includes two input molecules whose amino acid sequences differ due to a single point mutation; The attribute calculation model is applied to determine, for each pair of input molecules, an input molecule that exhibits a better value for the attribute than the other input molecule; and For each pair of input molecules, the point mutation exhibited by the one input molecule that shows the better value for the property is identified as one of the plurality of point mutations.

36. The method of claim 35, wherein the single point mutation comprises an amino acid sequence containing a first type of amino acid residue at a specific position and another amino acid sequence containing a second type of amino acid residue at the same position.

37. The method according to any one of claims 35 to 36, wherein the single point mutation comprises an amino acid sequence containing a specific type of amino acid residue at a first position and another amino acid sequence containing the same type of amino acid residue at a second position.

38. The method of any one of claims 31 to 37, wherein the attribute calculation model comprises a language model with relative embeddings trained to generate a representation of the difference between a first amino acid sequence of the first input molecule and a second amino acid sequence of the second input molecule.

39. The method of claim 38, wherein the attribute calculation model further comprises a neural network coupled to the language model, and wherein the neural network has been trained to determine the difference between the first value and the second value of the attribute based at least on the relative embedding.

40. The method according to any one of claims 31 to 39, wherein each of the first input molecule and the second input molecule comprises an antibody.

41. The method according to any one of claims 31 to 40, wherein each of the first input molecule and the second input molecule comprises a portion of an antibody.

42. The method according to any one of claims 31 to 41, wherein each of the first input molecule and the second input molecule comprises a variable region of an antibody, an antigen-binding region, a heavy chain and / or a light chain.

43. The method according to any one of claims 31 to 42, wherein the property includes expression, affinity, specificity, exploitability, or bioactivity.

44. A system comprising: At least one data processor; as well as At least one memory storing instructions that, when executed by the at least one data processor, cause operation including the method according to any one of claims 31 to 43.

45. A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operation comprising the method according to any one of claims 31 to 43.

46. ​​A computer-implemented method, comprising: Identify a set of sample molecules, wherein each sample molecule in the set of sample molecules is associated with a corresponding value for an attribute; Generate multiple sample molecule pairs to be included in the training dataset, wherein generating the multiple sample molecule pairs includes The first sample molecule and the second sample molecule selected from the group of sample molecules are paired to generate a first sample molecule pair, and The first sample molecule is paired with a third sample molecule selected from the group of sample molecules to generate a second sample molecule pair; as well as The attribute calculation model is trained based at least on the training dataset to determine the attribute difference between the two input molecules based at least on the sequence difference between the two input molecules, wherein the sequence difference between the two input molecules corresponds to the difference between the amino acid sequences of each input molecule, and wherein the attribute difference between the two input molecules corresponds to the difference between the values ​​of the attribute present in each input molecule.

47. The method of claim 46, wherein generating the plurality of sample molecule pairs further comprises The second sample molecule is paired with the third sample molecule to generate a third sample molecule pair.

48. The method of any one of claims 46 to 47, wherein the attribute calculation model comprises a language model trained to determine a relative embedding representing the sequence difference between the two input molecules based at least on the amino acid sequence of each input molecule.

49. The method of claim 48, wherein the attribute calculation model comprises a neural network coupled to the language model, and wherein the neural network has been trained to determine the attribute difference between the two input molecules based at least on the relative embedding.

50. The method according to any one of claims 46 to 49, wherein the property includes expression, affinity, specificity, exploitability, or bioactivity.

51. The method according to any one of claims 46 to 50, wherein each sample molecule comprises an antibody.

52. The method according to any one of claims 46 to 51, wherein each sample molecule comprises a portion of an antibody.

53. The method according to any one of claims 46 to 52, wherein each sample molecule comprises a variable region of an antibody, an antigen-binding region, a heavy chain, and / or a light chain.

54. A system comprising: At least one data processor; as well as At least one memory storing instructions that, when executed by the at least one data processor, cause operation including the method according to any one of claims 46 to 53.

55. A non-transitory computer-readable medium storing instructions that, when executed by at least one data processor, cause operation comprising the method according to any one of claims 46 to 53.