Machine learning enabled de novo generation of protein molecules

WO2026198926A1PCT designated stage Publication Date: 2026-09-24GENENTECH INC +9
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/020186
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-21
Filing Date
2026-03-20
Publication Date
2026-09-24

Smart Images

  • Figure US2026020186_24092026_PF_FP_ABST
    Figure US2026020186_24092026_PF_FP_ABST
Patent Text Reader

Abstract

A method for de novo protein generation may include applying a protein design computation model to generate a portion of a protein molecule de novo while conditioned on a conditioner molecule. The protein design computation model may be trained to co-generate an amino acid residue sequence and a three-dimensional structure of the portion of the protein molecule. Experimental data on a de novo portion of the protein molecule generated by the protein design computation model may be obtained. That the de novo portion of the protein molecule satisfies one or more criteria may be determined based at least on the experimental data. In response to the de novo portion of the protein molecule satisfying the one or more criteria, the protein design computation model may be applied to generate a different portion of the protein molecule containing the de novo portion previously generated by the protein design computation model.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1MACHINE LEARNING ENABLED DE NOVO GENERATION OF PROTEIN MOLECULES CROSS REFERENCE TO RELATED APPLICATION

[0001] This application claims priority to U. S. Provisional Application No. 63 / 775,661, entitled “MACHINE LEARNING ENABLED GENERATION OF NOVEL COMPLEMENTARITY DETERMINING REGION LOOPS” and filed on March 21, 2025, the disclosure of which is incorporated herein by reference in its entirety.TECHNICAL FIELD

[0002] The subject matter described herein relates generally to generative artificial intelligence and more specifically to machine learning enabled techniques for de novo generation of protein molecules.INTRODUCTION

[0003] A molecule is a group of two more atoms held together by chemical bonds. Molecules form the smallest identifiable unit into which a pure substance can be divided while still retaining the composition and chemical properties of that substance. Large molecules refer to those molecules that range between approximately 3000 Daltons and 150,000 Daltons in molecular weight. Large molecule therapeutics (also known as biopharmaceuticals, biotherapeutics, biologicals, or biologies) are often derivatives of natural human proteins, which modulate many essential cellular functions such as enzymatic reactions, transport of molecules, regulation and execution of a number of biological pathways, cell growth, proliferation, nutrient uptake, morphology, motility, intercellular communication, and / or the like. A single large molecule can include more than 1,300 amino acid residues linked by peptide bonds to form one or more polypeptide. Due to their size and complexity, large molecule therapeutics are recombinantlyAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1produced by engineered cells instead of being chemically synthesized like the majority of small molecule drugs. Moreover, large molecule therapeutics are usually delivered through injection or infusion due to the ineffectiveness of oral administration. The development of a biotherapeutics may entail designing one or more sequences of amino acid residues capable of binding to a another molecule (e.g., another protein molecule, a nucleic acid, and / or the like) with sufficient specificity and absent undesirable traits such as immunogenicity, self-association, instability, and / or the like.SUMMARY

[0006] Systems, methods, and articles of manufacture, including computer program products, are provided for machine learning enabled structure-sequence co-generation of novel protein molecules or portions thereof, such as one or more of a complementarity determining region (CDR), framework region, variable region (or fragment antigen binding (Fab)), constant region (or fragment crystallizable (Fc)), light chain, heavy chain, and / or the like. In some cases, a protein design computation model may co-generate the amino acid residue sequence and three-dimensional structure of at least a portion of the protein molecule de novo. In some cases, the protein design computation model may co-generate the amino acid residue sequence and three-dimensional structure of the protein molecule instead of first generating the three-dimensional structure of the protein backbone followed by sequence generation. Moreover, in some cases, the protein design computation model may achieve high performance de novo generation of protein molecules by at least implementing a stepwise strategy during training and / or inference while leveraging experimental data where available. For example, in some cases, the stepwise strategy may include applying the protein design computation model to perform incrementally more complex generative tasks, such as co-generating the structure and sequence of incrementally larger portions of the protein molecule.Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1

[0007] In one aspect, there is provided a system for machine learning enabled de novo protein generation. The system may include at least one data processor and at least one memory. The at least one memory may store instructions that cause operations when executed by the at least one data processor. The operations may include: applying a protein design computation model to generate a portion of a protein molecule de novo while conditioned on a conditioner molecule, wherein the protein design computation model is trained to co-generate an amino acid residue sequence and a three-dimensional structure of the portion of the protein molecule; obtaining experimental data on a de novo portion of the protein molecule generated by the protein design computation model; determining, based at least on the experimental data, that the de novo portion of the protein molecule satisfies one or more criteria; and in response to the de novo portion of the protein molecule satisfying the one or more criteria, applying the protein design computation model to generate a different portion of the protein molecule containing the de novo portion previously generated by the protein design computation model.

[0008] In another aspect, there is provided a computer-implemented method for machine learning enabled de novo protein generation. The method may include: applying a protein design computation model to generate a portion of a protein molecule de novo while conditioned on a conditioner molecule, wherein the protein design computation model is trained to co-generate an amino acid residue sequence and a three-dimensional structure of the portion of the protein molecule; obtaining experimental data on a de novo portion of the protein molecule generated by the protein design computation model; determining, based at least on the experimental data, that the de novo portion of the protein molecule satisfies one or more criteria; and in response to the de novo portion of the protein molecule satisfying the one or more criteria, applying the protein designAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1computation model to generate a different portion of the protein molecule containing the de novo portion previously generated by the protein design computation model.

[0009] In another aspect, there is provided a computer program product for machine learning enabled de novo protein generation. The computer program product may include a non-transitory computer readable medium storing instructions. The instructions may cause operations when executed by the at least one data processor. The operations may include: applying a protein design computation model to generate a portion of a protein molecule de novo while conditioned on a conditioner molecule, wherein the protein design computation model is trained to co-generate an amino acid residue sequence and a three-dimensional structure of the portion of the protein molecule; obtaining experimental data on a de novo portion of the protein molecule generated by the protein design computation model; determining, based at least on the experimental data, that the de novo portion of the protein molecule satisfies one or more criteria; and in response to the de novo portion of the protein molecule satisfying the one or more criteria, applying the protein design computation model to generate a different portion of the protein molecule containing the de novo portion previously generated by the protein design computation model.

[0010] In some variations, one or more features disclosed herein including the following features can optionally be included in any feasible combination.

[0011] In some variations, the protein molecule comprises an antibody.

[0012] In some variations, the portion of the protein molecule comprises a complementarity determining region (CDR) of the antibody and the different portion of the protein molecule comprises one or more of a different complementarity determining region (CDR), a heavy chain, a light chain, a fragment antigen binding (Fab), and a variable region (Fv).Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1

[0013] In some variations, the portion of the protein molecule comprises a third heavy chain complementarity determining region (CDR-H3) of the antibody.

[0014] In some variations, the conditioner molecule comprises one or more of another protein molecule, a chemical compound, a nucleic acid molecule, a lipid, or an ion.

[0015] In some variations, a backbone torsion (BBT) representation of the three-dimensional structure of the protein molecule and a logit representation of the amino acid residue sequence of the protein molecule are received. The protein design computation model is applied to operate on the backbone (BBT) torsion representation and the logit representation of the protein molecule in order to co-generate the three-dimensional structure and the amino acid residue sequence of the portion of the protein molecule.

[0016] In some variations, a representation of an amino acid residue sequence and / or a three-dimensional structure of the conditioner molecule is received The protein design computation model is applied to operate on the on the backbone (BBT) torsion representation and the logit representation of the protein molecule and the representation of the conditioner molecule in order to generate the portion of the protein molecule.

[0017] In some variations, the generating the portion of the protein molecule is conditioned one or both of the amino acid residue sequence and the three-dimensional structure of the conditioner molecule.

[0018] In some variations, the backbone (BBT) torsion representation and the logit representation of the protein molecule and the representation of the conditioner molecule are concatenated to form a representation array. The representation array is associated with a register array populated by a plurality of values. Each value of the plurality of values is indicative ofAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1whether a corresponding value in the representation array is associated with the protein molecule or the conditioner molecule.

[0019] In some variations, the backbone torsion representation of the protein molecule specifies, for each amino acid residue included in the protein molecule, a geometric state of a plurality of constituent backbone atoms, and wherein the geometric state comprises at least one of a translation and rotation of the plurality of constituent backbone atoms.

[0020] In some variations, the backbone torsion representation of the protein molecule specifies, for each amino acid residue included in the protein molecule, one or more torsion angles formed by a plurality of constituent sidechain atoms.

[0021] In some variations, the logit representation of the protein molecule includes, for each amino acid residue included in the protein molecule, a logit vector. The logit vector enumerates a probability of a type of the amino acid residue being each of the 20 canonical amino acid residues.

[0022] In some variations, the protein design computation model generates the portion of the protein molecule by at least denoising, at least partially in parallel, an amino acid residue sequence and a three-dimensional structure of the portion of the protein molecule.

[0023] In some variations, the protein design computation model denoises the amino acid residue sequence and the three-dimensional structure of the portion of the protein molecule over a plurality of successive timesteps.

[0024] In some variations, the protein design computation model delays the denoising of the amino acid residue sequence of the portion of the protein molecule until the three-dimensional structure of the portion of the protein molecule has undergone a threshold degree of denoising.Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1

[0025] In some variations, the protein design computation model denoises the amino acid residue sequence and the three-dimensional structure of the portion of the protein molecule while keeping a three-dimensional structure of the conditioner molecule fixed.

[0026] In some variations, the protein design computation model denoises the amino acid residue sequence and the three-dimensional structure of the portion of the protein molecule while simultaneously modifying a three-dimensional structure of the conditioner molecule to reflect one or more conformational changes engendered by a binding interaction with the protein molecule.

[0027] In some variations, the protein design computation model is trained, based at least on a training dataset, to generate the portion of the protein molecule de novo while conditioned on the conditioner molecule.

[0028] In some variations, a training sample is generated to include a protein-protein complex of a sample protein molecule bound to a conditioner molecule in which a portion of the sample protein molecule is masked. An additional training sample is generated to include the same protein-protein complex with a different portion of the sample protein molecule masked. The training dataset is generated to include the training sample and the additional training sample.

[0029] In some variations, the protein-protein complex comprises an immune complex in which the sample protein molecule comprises an antibody and the conditioner molecule comprises an antigen.

[0030] In some variations, the portion of the protein molecule comprises a complementarity determining region (CDR) of the antibody, and the different portion of the protein molecule comprises one or more of a different complementarity determining region (CDR), a heavy chain, a light chain, a fragment antigen binding (Fab), and a variable region (Fv).Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1

[0031] In some variations, the portion of the protein molecule comprises a third heavy chain complementarity determining region (CDR-H3) of the antibody.

[0032] In some variations, the de novo portion of the protein molecule is determined to satisfy the one or more criteria when the experimental data indicates that one or more of a binding affinity, binding specificity, functionality, and developability of the protein molecule including the de novo portion satisfy one or more thresholds.

[0033] In some variations, the experimental data is obtained by at least synthesizing the protein molecule having the de novo portion generated by the protein design computation model and analyzing one or more synthesized samples of the protein molecule.

[0034] In some variations, the experimental data is obtained by at least performing one or more computational analysis on the protein molecule having the de novo portion.

[0035] In some variations, the protein design computation model is applied to generate a first three-dimensional structure of the protein molecule and a second three-dimensional structure of the protein molecule. A docked conformation of a first complex including the first three-dimensional structure of the protein molecule bound to the conditioner molecule is determined. A docked conformation of a second complex including the second three-dimensional structure of the protein molecule bound to the conditioner molecule is determined. A stability of the first complex in the docked conformation is determined. A stability of the second complex in the docked conformation is determined. At least one of the first complex and the second complex is identified as a preferred conformation of the protein molecule for binding with the conditioner molecule based at least on the stability of the first complex and / or the stability of the second complex.Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1

[0036] In some variations, a binding affinity between the protein molecule and the conditioner molecule may be determined based at least on the stability of the first complex and / or the second complex.

[0037] In some variations, the binding affinity between the protein molecule and the conditioner molecule is determined to correspond to a respective value of a metric quantifying the stability of the first complex and the stability of the second complex.

[0038] In some variations, a conformation ensemble for the protein molecule is generated to include the first three-dimensional structure of the protein molecule and the second three-dimensional structure of the protein molecule.

[0039] In some variations, the protein design computation model is applied to generate incrementally larger and / or more complex portions of the protein molecule de novo while conditioned on the conditioner molecule.

[0040] Implementations of the current subject matter can include, but are not limited to, methods consistent with the descriptions provided herein as well as articles that comprise a tangibly embodied machine-readable medium operable to cause one or more machines (e.g., computers, etc.) to result in operations implementing one or more of the described features. Similarly, computer systems are also described that may include one or more processors and one or more memories coupled to the one or more processors. A memory, which can include a non-transitory computer-readable or machine-readable storage medium, may include, encode, store, or the like one or more programs that cause one or more processors to perform one or more of the operations described herein. Computer implemented methods consistent with one or more implementations of the current subject matter can be implemented by one or more data processors residing in a single computing system or multiple computing systems. Such multiple computingAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1systems can be connected and can exchange data and / or commands or other instructions or the like via one or more connections, including, for example, to a connection over a network (e.g. the Internet, a wireless wide area network, a local area network, a wide area network, a wired network, or the like), via a direct connection between one or more of the multiple computing systems, etc.

[0041] The details of one or more variations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features and advantages of the subject matter described herein will be apparent from the description and drawings, and from the claims. While certain features of the currently disclosed subject matter are described for illustrative purposes in relation to the generation of novel heavy chain third complementarity determining regions (CDR-H3), it should be readily understood that such features are not intended to be limiting. The claims that follow this disclosure are intended to define the scope of the protected subject matter.DESCRIPTION OF DRAWINGS

[0042] The accompanying drawings, which are incorporated in and constitute a part of this specification, show certain aspects of the subject matter disclosed herein and, together with the description, help explain some of the principles associated with the disclosed implementations. In the drawings,

[0043] FIG. 1 depicts a system diagram illustrating an example of a protein design system, in accordance with some example embodiments;

[0044] FIG. 2A depicts a flowchart illustrating an example of a process for stepwise training of a protein structure computation model, in accordance with some example embodiments;Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1

[0045] FIG. 2B depicts a flowchart illustrating an example of a process for stepwise binder conditioned protein generation by a protein structure computation model, in accordance with some example embodiments;

[0046] FIG. 3 depicts a flowchart illustrating an example of a process for flexible docking, in accordance with some example embodiments;

[0047] FIG. 4A depicts a screenshot illustrating an example of a protein molecule undergoing center of mass (CoM) rotation, in accordance with some example embodiments;

[0048] FIG. 4B depicts a screenshot illustrating an example of a protein molecule undergoing center of mass (CoM) translation, in accordance with some example embodiments;

[0049] FIG. 5A depicts an example of a conformational ensemble, in accordance with some example embodiments;

[0050] FIG. 5B depicts a structural comparison of a conformer with low root mean squared distance (RMSD) and the corresponding ground truth conformer, in accordance with some example embodiments; and

[0051] FIG. 6 depicts a block diagram illustrating an example of a computing system, in accordance with some example embodiments.

[0052] When practical, similar reference numbers denote similar structures, features, or elements.DETAILED DESCRIPTION

[0053] Biopolymers are large molecules (or macromolecules) naturally produced by living organisms. One example of a biopolymer is a protein molecule, which may include one or more polypeptide chains formed by amino acid residues linked by peptide bonds. Biopolymers,Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1such as protein molecules, may modulate a multitude of biological processes including, for example, enzymatic reactions, molecular transport, biological pathway regulation and execution, cell growth and proliferation, nutrient uptake, morphology, motility, intercellular communication, and / or the like. In some cases, a biopolymer may modulate a biological process by interacting with another molecule, such as another biopolymer. For example, the interaction between antibodies and antigens (e.g., viral antigens, tumor antigens, and / or the like) is a critical part of a body’s immune response in which exposure to antigens triggers the production of antibodies to bind to and neutralize the antigens. In the case of therapeutic proteins, such as antibodies, chimeric antigen receptors (CARs), enzymes, hormones, and cytokines, the specificity of the interaction between a therapeutic protein and its therapeutic binding target may be leveraged for the diagnosis and treatment of a wide range of diseases with fewer side effects. Accordingly, one primary objective of large molecule drug discovery (LMDD) is to engineer therapeutic proteins capable of binding to therapeutic binding targets with sufficient affinity and specificity as well as functional (e g., modulates the binding targets) and with satisfactory developability traits such as clearance, solubility, viscosity, aggregation propensity, storability, expression, manufacturability, polyspecificity (or polyreactivity), pharmacokinetics (PK) (e.g., long in vivo half-life for sustained efficacy), stability (e.g., resistant to heat, pH changes, and storage conditions), and / or the like.

[0054] Designing a protein sequence with one or more properties of interest is a critical task in biomedicine and bioengineering. In the context of drug discovery, the protein sequence may be a biotherapeutic for treating, preventing, or curing diseases and medical conditions. Examples of protein therapeutics include antibodies, peptide hormones, growth factors, plasma proteins, enzymes, or hemolytic factors. The one or more properties of interest may include binding affinity, binding specificity, functionality, and developability. The presence (and absence)Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1of the one or more properties of interest may determine the likelihood of a protein sequence being developed into a viable biotherapeutic. As such, the process of designing a protein sequence with the one or more desired properties typically includes two high-level phases: lead discovery (LD) followed by lead optimization (LO). The objective of lead discovery (LD) is to identify novel protein sequences exhibiting at least some level of fitness as a protein therapeutic. Those novel protein sequences then undergo lead optimization (LO), which aims to increase (or maximize) the fitness of the novel protein sequences identified through lead discovery (LD). The success of lead optimization (LO) is heavily dependent on the quality of lead molecules provided by lead discovery (LD) but conventional lead discovery (LD) techniques fail to consistently deliver high quality lead molecules capable of being developing into viable therapeutics.

[0055] To accelerate drug development and reduce reliance on expensive wet lab resources, lead discovery (LD) as well as lead optimization (LO) efforts may leverage computation tools. For example, in some cases, one or more computation models, such as language models and generative models, may be trained to generate protein sequences and, in some cases, also the corresponding three-dimensional structures. Nevertheless, the capabilities of conventional computation models are limited to, as in the case of lead discovery (LD) and lead optimization (LO), modifying existing protein molecules to enhance desirable properties while minimizing undesirable ones. For example, given a lead molecule, such as an antibody identified through an animal immunization campaign as having binding affinity and specificity towards a target molecule (e.g., a viral antigen, a tumor antigen, and / or the like), a conventional computation model may be applied to improve other properties such as human-ness, poor expression, immunogenicity, and in vivo instability. In other words, conventional computation models are unable to design protein molecules de novo which, in the context of protein engineering, means generating theAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1amino acid residue sequences and corresponding three-dimensional structures of protein molecules to achieve bespoke functionality and predefined binding target specificity outside of the naturally observed protein space. That is, de novo protein molecules are generated from scratch, guided by functionality or binding target alone, rather than modifying an existing protein molecule. As such, de novo protein design may increase structural and functional diversity by at least enabling the discovery of novel protein molecules that are dissimilar from those in the existing repertoire. However, conventional computation models have thus far failed to successfully generate protein molecules that are not the product of modifying a known or natural protein molecule.

[0056] Although a paramount goal in drug discovery, computation models capable of designing therapeutic protein molecules, such as antibodies, based on functionality and binding target alone have remained elusive to this date. While the current state of the art supports designing de novo protein scaffolds (e.g., to support functional motifs such as binding sites) and mini -protein binders (e.g., protein molecules with approximately 37-65 amino acid residues), achieving true de novo protein design remains elusive. For example, to the extent de novo protein scaffolds are designed from scratch, the functional motifs (e.g., variable region (or fragment antibody binding (Fab)) grafted thereon are typically not. In short, de novo protein design is fundamentally difficult due to the sheer vastness and sparsity of the solution space (or combinatorial space) of all possible protein sequences and structures. For a protein molecule containing an N quantity of amino acid residues, approximately 20Npossible protein sequences exist if each of the N quantity of amino acid residues is one of the twenty canonical amino acid residues. In fact, the number of possible protein sequence combinations far exceeds the estimated number of atoms in the observable universe. Whereas conventional lead optimization (LO) reduces the complexity of the problem by effectively limiting the exploration of the solution space (or combinatorial space) to those proteinAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1sequence combinations similar to that of the lead molecule in either sequence space or latent space, de novo protein design requires protein molecules with one or more properties of interest to be generated without guidance from any existing protein molecules. Absent the benefit of lead molecules, such as those known to exhibit clinically useful pharmacological or biological properties, conventional computational models do not generate protein molecules with sufficiently high likelihood of exhibiting requisite properties of viable biotherapeutics. One primary objective of successful de novo protein design is therefore the ability to generate truly novel protein sequences that are also likely to exhibit one or more properties of interest, such as binding affinity, binding specificity, functionality, and developability.

[0057] Various embodiments of the present disclosure achieve the objectives of de novo protein design, including sampling protein molecule with requisite properties from previously unexplored regions of the combinatorial space, by at least adopting a stepwise approach towards generating novel protein molecules against pre-specified targets (e.g., epitopes) with high specificity and affinity. In some example embodiments, the stepwise approach may be implemented during training and / or inference such that a protein design computation model is applied to perform incrementally more complex generative tasks. For example, in some cases, the protein design computation model may be applied in a stepwise fashion to generate, de novo, incrementally larger or more complex (e.g., more conserved) portions of a protein molecule. In some cases, one or more portions of the protein molecule, such as one or both the sequence and structure thereof, may be masked, for example, with a mask token or “noise” in the form of random values. In some cases, the protein design computation model may generate at least a portion of the protein molecule de novo by at least reconstructing the one or more masked portions of the protein molecule. As described in more detail below, in some cases, the protein designAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1computation model may operate on a backbone torsion (BBT) representation of the protein molecule in which the three-dimensional structure of the protein molecule is defined by the geometric state (e.g., translation, rotation, and / or the like) of the backbone atoms of each constituent amino acid residue and the torsion angles (or dihedral angles) formed by the corresponding sidechain angles. In some cases, the protein design computation model may be a diffusion model that generates at last the portion of the protein molecule by at least incrementally denoising, over a series of successive timepoints, the sequence and structure of the masked portions of the protein molecule.

[0058] In some example embodiments, the protein design computation model is trained to generate at least a portion of a protein molecule de novo by at least reconstructing the structure and sequence of one or more masked portions of the protein molecule. In some cases, the protein design computation model may undergo stepwise training in which the protein design computation model is trained to reconstruct the structure and sequence of incrementally larger and / or more complex portions of a protein molecule. For example, in some cases, the protein molecule design computation model may be first trained to generate one or more complementarity determining regions (CDRs) in an antibody before being trained to generate one or both of the heavy chain and light chain. Alternatively and / or additionally, in some cases, the protein molecule design computation model may be first trained to generate a more variable region of a protein molecule, such as the heavy chain third complementarity determining region (CDR-H3), before being trained to generate a more conserved region of the protein molecule, such as the heavy chain first and second complementarity regions (CDR-H1 and H2).

[0059] In some example embodiments, once trained, the protein design computation model may be applied to generate incrementally larger or more complex regions of a proteinAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1molecule while leveraging experimental data. In some cases, a smaller or less complex region of the protein molecule generated by the protein molecule computation model may undergo experimental analysis, including one or both of wetlab and computational analysis. In some cases, where the resulting experimental data indicates that the smaller or less complex region of the protein molecule exhibits one or more properties of interest (or satisfies one or more other criteria), the protein design computation model may leverage the smaller or less complex region of the protein molecule when generating a larger or more complex region of the protein molecule. For example, in some cases, the protein design computation model may be applied to generate the heavy chain third complementarity determining region (CDR-H3) of an antibody. In some cases, where experimental data indicates that a heavy chain third complementarity determining region (CDR-H3) generated by the protein design computation model exhibits one or more properties of interest, the protein design computation model may be applied to generate, based at least on heavy chain third complementarity determining region (CDR-H3), the one or more of a heavy chain first and second complementarity regions (CDR-H1 and H2) of the antibody.

[0060] In some example embodiments, the protein design computation may be applied to co-generate the sequence and structure of at least a portion of a protein molecule while conditioned on at least a portion of another protein molecule. For example, in some cases, the protein design computation model may be applied to go-generate the sequence and structure of at least a portion of an antibody while conditioned on at least a portion of an antigen (e.g., epitope). Where the protein molecule is an antibody, for example, the protein design computation model may be used to generate the sequence and structure of at least a portion of the variable region (Fv) while conditioned on the sequence and structure of a target antigen. In some cases, the protein design computation model may adopt a stepwise approach in which the variable region (Fv) of theAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1antibody is generated by first generating one or more constituent portions, such as the heavy chain third complementarity determining region (CDR-H3). It should be appreciated that the most diversity in both length and amino acid residue sequence is observed in the heavy chain third complementarity determining region (CDR-H3). Accordingly, in some cases, the protein design computation model may undertake a series of incrementally more complex design tasks that includes first generating the heavy chain third complementarity determining region (CDR-H3) before leveraging experimental data to further generate the remaining complementarity determining regions in the heavy and light chains of the antibody. In some cases, the protein design computation model may generate the sequence and structure of the antibody simultaneously (or at least partially in parallel), rather than generating first the backbone structure followed by the corresponding amino acid residue sequence.

[0061] In some example embodiments, the protein design computation model may operate on a backbone torsion (BBT) representation of the three-dimensional structure of a protein molecule. In some cases, the backbone torsion (BBT) representation of the protein molecule may specify the geometric state of the backbone of each amino acid residue in a variety of different ways. For example, in some cases, the geometric state of the backbone of an amino acid residue may be specified based on its translation and rotation as well as a second torsion angle of a rotatable bond between an alpha carbon (Ca) atom and a carbonyl group in the backbone of the amino acid residue. Accordingly, in some cases, the backbone torsion (BBT) representation of the protein molecule may include a frame specifying a rotation and a translation of the backbone of the amino acid residue. For instance, in some cases, that frame may include an affine transformation matrix that includes a rotation matrix specifying the rotation of the backbone of the amino acid residue as well as a displacement vector specifying the translation of the backbone of the amino acid residue.Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1In instances where the backbone torsion (BBT) representation of the protein molecule includes the rotation and the translation of the backbone of each amino acid residue, the backbone torsion (BBT) representation of the protein molecule may include additional frames specifying the torsion angles (e.g., of the rotatable bond between the alpha carbon (Ca) atom and the carbonyl group) present in the backbone of each amino acid residue. Alternatively, in some cases, instead of the geometric state of the backbone being specified based on the torsion angles present in the backbone atoms of each amino acid residue, without specifying the corresponding translations and rotations. Accordingly, in some cases, the backbone torsion (BBT) representation of the protein molecule may include a frame specifying a torsion angle of the rotatable bond between the alpha carbon (Ca) atom and the carbonyl group in the backbone of each amino acid residue in at least a portion of the protein molecule. Moreover, in those instances, the backbone torsion (BBT) representation of the protein molecule may further include a frame specifying the torsion angle of the rotatable bond between the alpha carbon (Ca) atom and the nitrogen (N) atom in the backbone of the amino acid residue as well as an additional frame specifying the torsion angle of the rotatable bond between the carbon (C) atom and the nitrogen (N) atom in the backbone of the amino acid residue.

[0062] In some example embodiments, the protein design computation may operate on a logit representation of the amino acid residue sequence of the protein molecule. In some cases, the logit representation of the protein molecule may include a sequence of probability vectors, each of which enumerating the probability distribution of the type (or identity) of the amino acid residue occupying the corresponding position in the amino acid residue sequence of the protein molecule. For example, the probability vector for a position in the protein sequence may include, for each amino acid residue of the 20 canonical amino acid residues, the probability of that position being occupied that amino acid residue. In some cases, each probability vector may be aAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1probability simplex, which requires the constituent probability across the possible types of amino acid residues to sum to 1. For instance, the probability that the amino acid residue occupying a particular position in the protein sequence is each one of the 20 canonical amino acid residues may add up to 1. Otherwise, where the sum of the probabilities across the possible amino acid residues is not 1, the output of the protein design computation model would be meaningless. Accordingly, in some cases, representing the amino acid residue sequence of the protein molecule as a collection of simplex probabilities may limit the degrees-of-freedom for the protein molecule design computation model to modify the sequence of the protein molecule. The probability simplex may be a mathematical space in which each point represents a probability distribution between a finite number of mutually exclusive events, or categories, which in this case correspond to the set of possible amino acid residues (e.g., the 20 canonical amino acid residues). It should be appreciated that the probability simplex Simplex(D) is a (D — 1) dimensional object occupying a (D — 1) dimensional space, with D corresponding to the quantity of possible amino acid residues (e g., 20 canonical amino acid residues) that can occupy each position in a protein sequence.

[0063] In some example embodiments, the protein design computation model may be implemented using a diffusion model that generates at least a portion of a protein molecule de novo by at least incrementally denoising, over a succession of timepoints, at least the portion of the protein molecule. In some cases, the protein design computation model may denoise the protein molecule by at least removing, at each timepoint over the succession of timepoints, a portion of the noise present in the amino acid residue sequence and / or three-dimensional structure of the protein molecule. In some cases, the denoising that is performed at each timepoint may include an incremental update to the sequence and / or structural representation of the protein molecule. Moreover, in some cases, upon removing a portion of noise from the representation ofAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1the protein molecule at a first timepoint, a second quantity of noise may be removed from the representation of the protein molecule at a second timepoint.

[0064] In some example embodiments, when implemented as a diffusion model, the protein design computation may be trained to approximate and reverse a noise schedule defining the distribution of noise levels across a succession of timepoints. In some cases, during a forward process (or transition), the training samples used to train the protein design computation model may be generated by adding noise to at least a portion of each sample protein molecule. In some cases, the noise may be added incrementally in accordance with a noise schedule governing the quantity (or level) of noise that is added to a sample protein molecule at each successive timepoint. In some cases, the distribution of noise levels may correspond to the degree-of-freedom (DoF) available for the protein design computation model to modify the structure and sequence of the protein molecule. For example, in some cases, more noise may be added to degrees-of-freedom (DoF) where more entropy may be present than to those degrees-of-freedom (DoF) where less entropy is present. In some cases, once trained, the protein design computation model may be applied to perform a reverse process (or transition) that includes denoising the sequence and / or structure of at least a portion of a protein molecule (e g., an antibody) while conditioned on the sequence and / or structure of another protein molecule (e.g., antigen). In some cases, the reverse process (or transition) may be performed in accordance with a noise schedule that determines, for example, when at least partially parallel denoising of the structure and sequence of the protein molecule occurs. For instance, in some cases, the denoising of the protein sequence may start after the structure of the protein molecule has undergone sufficient denoising at least because the meaningful determination of the type (or identity) of each amino acid residue may require at least some semblance of structure.Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1

[0065] FIG. 1 depicts a system diagram illustrating an example of a protein design system 100, in accordance with some example embodiments. Referring to FIG. 1, the protein design system 100 may include a protein design engine 110, a data store 120, and a client device 130 including a user interface 135. In the example shown in FIG. 1, the protein design engine 110, the data store 120, and the client device 130 may be communicatively coupled via a network 140. In some cases, the data store 120 may be a database including, for example, a relational database, a NoSQL database, a columnar database, an objected-oriented database, a key -value database, a hierarchical database, a document database, a graph database, and / or the like. In some cases, the client device 130 may be a processor-based device including, for example, a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable apparatus, and / or the like. In some cases, the network 140 may be a wired network and / or a wireless network including, for example, a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, and / or the like.

[0066] In some example embodiments, the protein design engine 110 may include a protein design computation model 115 trained to co-generate the sequence and structure of a protein molecule. For example, as shown in FIG. 1, in some cases, the protein design computation model 115 may receive an amino acid sequence 112 and a three-dimensional structure 113 of an input molecule 111. In some cases, the protein design engine 110 may also receive a conditioner molecule 114. In some cases, the conditioner molecule 114 may be another molecule with which the input protein molecule Ill is engaged in a binding interaction. In some cases, the conditioner molecule 114 may be another protein molecule, a chemical compound, a nucleic acid, one or more ions, lipids, and / or the like. It should be appreciated that the conditioner molecule 114 may be aAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1binding target of the input protein molecule 111 but it is also possible forthe input protein molecule 111 to be the binding target of the conditioner molecule 114 instead.

[0067] In instances where the conditioner molecule 114 is another protein molecule, such as an antigen, the protein design engine 110 may receive the sequence and structure of the conditioner molecule 114. Alternatively, in cases where the conditioner molecule 114 is a chemical compound, the protein design engine 110 may receive a chemical structure of the conditioner molecule 114. It should be appreciated that the conditioner molecule 114 may also be a nucleic acid molecule (e.g., DNA, RNA, and / or the like), one or more ions, lipids, and / or the like.

[0068] In some cases, the protein design computation model 115 may generate, based at least on the amino acid residue sequence 112 and the three-dimensional structure 113 of the input molecule 111, an output molecule 117 having an amino acid residue sequence 118 and a three-dimensional structure 119. In some cases, the protein design computation model 115 may be trained based on a training dataset that includes one or more protein-protein complexes 125. For instance, in some cases, each training sample may include a ground truth protein-protein complex including a sample protein molecule bound to a sample conditioner molecule. In some cases, the protein design computation model 115 may be trained to co-generate the sequence and structure of at least a portion of the sample protein molecule while conditioned on the sequence and structure of the sample conditioner molecule. In some cases, the training of the protein design computation model 115 may include adjusting, for each batch of one or more training samples processed by the protein design computation model 115, one or more parameters (e.g., weights, biases, and / or the like) of the protein design computation model 115 to reduce (or minimize) the difference betweenAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1the sequence and structure generated by the protein design computation model 115 and that of the sample protein molecule.

[0069] As described in more detail below, in some example embodiments, the protein design computation model 115 may achieve de novo generation of the sequence 118 and structure 119 of the output molecule 117 by at least adopting, during training and inference, a stepwise strategy that includes performing incrementally more complex generative tasks. For example, in some cases, the protein design computation model 115 may be applied to generate incrementally larger or more complex portions of the output molecule 117. In some cases, the co-generating of the amino acid residue sequence 118 and the three-dimensional structure 119 of the output molecule 117 may be conditioned on the conditioner molecule 114. Moreover, in some cases, the protein design computation model 115 may co-generate the amino acid residue sequence 118 and the three-dimensional structure 119 of the output molecule 117 by at least denoising, at least partially in parallel, the amino acid sequence 112 and the three-dimensional structure 113 of the input molecule 111.

[0070] FIG. 2A depicts a flowchart illustrating an example of a process 200 for stepwise training of a protein structure computation model, in accordance with some example embodiments. Referring to FIGS. 1 and 2A, in some example embodiments, the process 200 may be performed to train the protein design computation model 115 to co-generate the sequence and structure of a protein molecule. For example, in some cases, the protein design computation model 115 may be trained, in a stepwise fashion, to co-generate the sequence 118 and structure 119 of the output molecule 117 based at least on the sequence 112 and structure 113 of the input molecule 111. In some cases, the protein design computation model 115 may be trained to co-generate the amino acid residue sequence 118 and the three-dimensional structure 119 of the output molecule 117 byAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1at least denoising, at least partially in parallel, the amino acid sequence 112 and the three-dimensional structure 113 of the input molecule 111. In some cases, the protein design computation model 115 may be trained to generate the sequence 118 and structure 119 of the output protein molecule 117 while conditioned on the conditioner molecule 114. In some cases, the protein design computation model 115 may be trained in a stepwise fashion to perform incrementally more complex generative tasks, including the co-generation of incrementally larger and / or more complex portions of the output protein molecule 117.

[0071] At 202, a training sample is generated to include a protein-protein complex of a sample protein molecule bound to a conditioner molecule in which a portion of the sample protein molecule is masked. In some example embodiments, the protein-protein complex may be a crystallized structure determined through protein crystallization methodologies, such as X-ray diffraction / X-ray crystallography, cryogenic electron microscopy (CryoEM) (including electron crystallography and microcrystal electron diffraction (MicroED)), small-angle X-ray scattering, neutron diffraction, and / or the like. Alternatively and / or additionally, one or both of the sample protein molecule and the conditioner molecule in the protein-protein complex may be determined computationally. In some cases, the protein-protein complex may be an immune complex with the sample protein molecule being an antibody and the conditioner molecule being an antigen. In some cases, a portion of the amino acid residue sequence and / or the three-dimensional structure of the sample protein molecule may be masked. Furthermore, in some cases, the unmasked amino acid residue sequence of the sample protein molecule and its three-dimensional structure in the protein-protein complex may serve as ground-truths.

[0072] At 204, an additional training sample is generated to include the same proteinprotein complex with a different portion of the sample protein molecule masked. In some exampleAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1embodiments, one or more additional training samples may be generated to include the same protein-protein complex but with different portions of the sample protein molecule masked. In some cases, the additional training sample may be generated to include the sample protein molecule with the amino acid residue sequence and / or the three-dimensional structure of a larger and / or more complex portion of the sample protein molecule masked. For example, where the sample protein molecule is an antibody, one training sample may be generated to include the antibody with one or more complementarity determining regions (CDR) of the antibody masked while another training sample may be generated to include the antibody with the entirety of its heavy chain, light chain, or variable region (Fv) masked. In some cases, the ground-truth for these additional training samples may include the unmasked amino acid residue sequence of the sample protein molecule and its three-dimensional structure in the protein-protein complex.

[0073] At 206, a training dataset is generated to include the training sample and the additional training sample. In some example embodiments, the training dataset may be generated to include at least some crystallized structures of protein-protein complexes determined through protein crystallization methodologies, such as X-ray diffraction / X-ray crystallography, cryogenic electron microscopy (CryoEM) (including electron crystallography and microcrystal electron diffraction (MicroED)), small-angle X-ray scattering, neutron diffraction, and / or the like. In some cases, the quantity of training samples in the training dataset may be augmented using proteinprotein complexes in which one or both of the sample protein molecule and the conditioner molecule are determined computationally. In some cases, that different portions of the sample protein molecule from a single protein-protein complex may be masked to yield multiple training samples may further augment the number of training samples available for inclusion in the training dataset.Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1

[0074] At 208, a protein design computation model is trained on the training dataset to generate, de novo, at least a portion of a protein molecule while conditioned on a conditioner molecule. In some example embodiments, the protein design computation model may be trained to de novo generate at least the portion of the protein molecule by at least denoising the corresponding portion of an input protein molecule. In some cases, the protein design computation model may co-generate the amino acid residue sequence and the three-dimensional structure of at least a portion of the protein molecule by at least denoising the amino acid residue sequence and / or the three-dimensional structure of the corresponding portion of the input protein molecule. In some cases, the protein design computation model may be trained to operate on a backbone torsion (BBT) representation of the three-dimensional structure of the input protein and a logit representation of the corresponding amino acid residue sequence, In some cases, the denoising of the input protein molecule may be conditioned on the sequence and / or structure of the conditioner molecule. Furthermore, in instances where the protein design computation model is implemented using a diffusion model, the protein design computation model may denoise the input protein molecule incrementally, over a succession of timesteps, with a portion of the noise present in the sequence and / or structure of the input molecule removed at each timestep. In some cases, the denoising of the amino acid residue sequence and the three-dimensional structure of the input protein molecule may be performed at least partially in parallel. For example, in some cases, the protein design computation model may first denoise the three-dimensional structure of the input protein molecule before starting the denoising of the corresponding amino acid residue sequence. In some cases, the denoising of the amino acid residue sequence of the input protein molecule may be delayed until the three-dimensional structure of the input protein molecule has undergone sufficient denoising to achieve some semblance of a protein-like structure. For instance, theAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1denoising of the amino acid residue sequence of the input protein molecule may commence when a threshold quantity of noise has been removed from the three-dimensional structure of the input protein molecule and that the noise scale of what remains satisfies one or more thresholds.

[0075] FIG. 2B depicts a flowchart illustrating an example of a process 250 for stepwise binder conditioned protein generation by a protein structure computation model, in accordance with some example embodiments. Referring to FIGS. 1 and 2B, in some example embodiments, the process 250 may be performed to apply the protein design computation model 115 to perform incrementally more complex generative tasks while leveraging experimental data. For example, as described in more detail below, the process 250 may include applying the protein design computation model 115 to generate a first portion of the sequence 118 and the structure 119 of the output protein molecule 117. In some cases, experimental data on the first portion of the sequence 118 and the structure 119 of the output protein molecule 117 may be obtained. In some cases, where the experimental data indicates that the first portion of the sequence 118 and the structure 119 of the output protein molecule 117 satisfies one or more criteria, such as exhibiting one or more properties of interest, the protein design computation model 115 may be applied to determine a second portion of the sequence 118 and the structure 119 of the output molecule 117 with the first portion of the sequence 118 and the structure 119 of the output molecule 117 incorporated therein.

[0076] At 252, a protein design computation model is applied to generate a portion of a protein molecule de novo while conditioned on a conditioner molecule. In some example embodiments, the de novo portion of the protein molecule may be generated by the protein design computation model co-generating an amino acid residue sequence and a three-dimensional structure of a portion of the protein molecule. In some cases, the protein design computation modelAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1may be a diffusion model that co-generates the sequence and structure of the protein molecule by at least denoising, at least partially in parallel, the sequence and structure of the protein molecule. In some cases, the denoising of the protein molecule may be conditioned on the sequence and / or structure of the conditioner molecule. In some cases, the protein design computation model may operate on a backbone torsion (BBT) representation of the three-dimensional structure of the protein molecule and a logit representation of the amino acid residue sequence of the protein molecule. In some cases, the input of the protein design computation model may further include the sequence and / or structure of the conditioner molecule. In some cases, the sequence and structure of the protein molecule may be concatenated with that of the conditioner molecule to form a representation array (e.g., of numerical values). For example, the backbone torsion (BBT) and logit representations of the protein molecule may be concatenated with the backbone torsion (BBT) and logit representations of the conditioner molecule. In some cases, to enable differentiation between the values in the representation array corresponding to the protein molecule and those corresponding to the conditioner molecule, the representation array may be associated with a separate register array. In some cases, each value in the register array may be associated with a value in the representation array. Furthermore, each value in the register array may indicate whether the corresponding value in the representation array is associated with the protein molecule or the conditioner molecule. It should be appreciated that where there are two molecules in the system (e.g., a single protein molecule and a single conditioner molecule), the register array may be a binary array. However, where there are more than two molecules in the system (e.g., a single protein molecule and two conditioner molecules), the register array may contain additional values to enable differentiation between the molecules.Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1

[0077] As described in more detail below, in some cases, the protein design computation model may be applied to perform incrementally more complex generative tasks. For example, in some cases, the protein design computation model may be applied to co-generate the amino acid residue sequence and the three-dimensional structure of a smaller or less complex portion of the protein molecule before being applied to co-generate the amino acid residue sequence and the three-dimensional structure of a larger or more complex portion of the protein molecule. In instances where the protein molecule is an antibody, the protein design computation model may be applied to co-generate the amino acid residue sequence and the three-dimensional structure of a complementarity determining region (CDR) of the antibody while conditioned on an antigen before being applied to co-generate the amino acid residue sequence and the three-dimensional structure of the entirety of the heavy chain, light chain, fragment antigen binding (Fab), or variable region (Fv) of the antibody.

[0078] At 254, experimental data on a de novo portion of the protein molecule generated by the protein design computation model is obtained. In some example embodiments, the experimental data may indicate the presence (or absence) of one or more properties of interest in the protein molecule or the magnitude thereof. In some cases, the experimental data may be obtained (or derived) from wet lab and / or computational analysis of the protein molecule. For example, in some cases, the protein molecule may be synthesized and subject to various assays using a variety of laboratory equipment. In some cases, the laboratory equipment may include any wetlab and dry lab equipment capable of synthesis, purification, and / or analysis. Examples of such laboratory equipment may include synthesizers, including standard equipment (e.g., fume hoods, glassware, heating and cooling devices, stirrers), automated and specialized synthesis platforms (e.g., automated synthesizers, parallel synthesis workstations, and high-throughputAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1experimentation (HTE) for efficient reaction optimization, and specialized reactors (e.g., microwave, flow, photochemistry). In some cases, the laboratory equipment may also include tools to support various purification techniques such as chromatography (e.g., silica gel chromatography, preparative high performance liquid chromatography (Prep HPLC)), centrifugal, crystallization / recrystallization, and / or the like. In some cases, the laboratory equipment may also include analytical tools such as microscopes, spectroscopes, spectrometers (e.g., nuclear magnetic resonance (NMR) spectrometers, mass spectrometers), balances, pH meters, elemental analyzers, and / or the like.

[0079] At 256, the de novo portion of the protein molecule is determined, based at least on the experimental data, to satisfy one or more criteria. In some example embodiments, the experimental data may indicate that the protein molecule including the portion generated by the protein design computation model exhibits one or more properties of interest, such as binding affinity, binding specificity, functionality, and developability. For example, where the protein molecule is an antibody and the protein design computation model is applied to generate the third heavy chain complementarity determining region (CDR-H3) de novo, the experimental data may indicate whether the third heavy chain complementarity determining region (CDR-H3) exhibits sufficient binding affinity and specificity towards the conditioner molecule. Accordingly, as described in more detail below, the portion of the protein molecule may be leveraged when the protein design computation model is applied again to generate a larger or more complex portion of the protein molecule.

[0080] At 258, in response to the de novo portion of the protein molecule satisfying the one or more criteria, the protein design computation model is applied to generate a different portion of the protein molecule containing the de novo portion previously generated by the protein designAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1computation model. In some example embodiments, where the de novo portion of the protein molecule previously generated by the protein design computation model satisfies one or more criteria, such as exhibiting one or more properties of interest, the de novo portion of the protein molecule may be leveraged to further generate, de novo, a larger or more complex portion of the protein molecule. In some cases, the protein design computation model may operate on the sequence and structure of the protein molecule having the de novo portion previously generated by the protein design computation model. Moreover, in some cases, the protein design computation model may be applied to co-generate the amino acid residue sequence and the three-dimensional structure of the different portion of the protein molecule while conditioned on the sequence and / or structure of the conditioner molecule. In some cases, the de novo portion of the protein molecule, which has been experimentally validated as satisfying one or more criteria, may provide useful context for the de novo generation of a different portion of the protein molecule. In some cases, the protein design computation model may be applied in a stepwise fashion, meaning that the different portion of the protein molecule may be a larger or more complex portion of the protein molecule than the de novo portion previously generated by the protein design computation model. For instance, in the example where the de novo portion is the third heavy chain complementarity determining region (CDR) of an antibody, the different portion of the protein molecule may be another, more conserved complementarity determining region (CDR), the heavy chain, the light chain, the entire variable region (Fv) of the antibody, or the antibody as a whole.

[0081] Referring again to FIG. 1, in some example embodiments, the protein design engine 110 may perform flexible docking by at least applying the protein design computation model 115 to co-generate the sequence 112 and structure 113 of the input protein molecule 111 based at least on the sequence and / or structure of the conditioner molecule 114. For example, inAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1cases where the protein design computation model 115 is applied to perform flexible docking, the protein design computation model 115 may be trained to operate on a backbone torsion (BBT) representation of the three-dimensional structure 113 of the input protein molecule 111 to generate the three-dimensional structure 119 of the output molecule 117 in a docked conformation (or in complex with) the conditioner molecule 114. In some cases, the protein design computation model 115 may generate the structure 119 of the output protein molecule 117 by at least denoising the backbone torsion (BBT) representation of the structure 113 of the input protein molecule 111. In some cases, the denoising of the backbone torsion (BBT) representation of the structure 113 of the input protein molecule 111 may be tantamount to modifying the geometric state of the backbone and sidechains of the input protein molecule 111 while rotating and / or translating a center of mass (CoM) of the input protein molecule 111 in its entirety. In some cases, doing so may alter the orientation of the entire input protein molecule 111 relative to the conditioner molecule 114 while one or more of the sequence 112 and structure 113 of the input protein molecule 111 and, in some cases, the structure of the conditioner molecule 114 are modified to reflect the conformational changes engendered by the binding interaction between the two molecules. This behavior is further illustrated in FIGS. 4A-B. In FIG. 4A, an example of a protein molecule is shown being rotated about its center of mass (CoM) while in FIG. 4B another example of a protein molecule is shown being translated about its center of mass (CoM). However, it should be appreciated that the protein design computation model 115 may also modify the sequence 112 and the structure 113 of the input molecule 111 while keeping the sequence and structure of the conditioner molecule 114 fixed.

[0082] In some example embodiments, the protein design engine 110 may apply the protein design computation model 115 to generate, based on the input protein molecule 111,Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1multiple output molecules 117, each having at least one of a different sequence 118 and structure 119. In some cases, the different output protein molecules 117 may correspond to individual conformers of the same underlying protein molecule. Accordingly, in some cases, the protein design engine 110 may generate multiple output protein molecules 117 in order to generate a conformational ensemble (CE) of a protein molecule. FIG. 5 A depicts an example of conformation ensemble 500 that includes multiple target three-dimensional structure of a same protein molecule. FIG. 5B depicts a structural comparison of a conformer 525 with low root mean squared distance (RMSD) and the corresponding ground truth conformer 550, in accordance with some example embodiments. In some cases, an ensemble of the ground truth conformers 550 may be generated through molecular dynamics simulations performed based on the observed crystal structures of protein molecules.

[0083] In some example embodiments, the protein design engine 110 may perform flexible docking based on multiple the three-dimensional structures 119 (or the conformational ensemble (CE)) of the same output protein molecule 117. It should be appreciated that the molecular properties of the output protein molecule 117 may vary across its conformational ensemble at least because the molecular properties of the output protein molecule 117 may often be dependent on its three-dimensional structure 119. In the case of binding affinity between two molecules, for example, the ability of a protein molecule to bind to another molecule may depend on the ability of the molecule to adopt a three-dimensional structure, or conformational shape, that is complementary to the three-dimensional structure of the other molecule. Accordingly, in some cases, the protein design engine 110 may perform molecular docking to generate, for the structure 119 of each output protein molecule 117, a docked conformation of a complex including the output protein molecule 117 bound to another molecule, such as the conditioner molecule 114. In someAttomey Ref.: 14786-094-228 (103963-228094) / P60064-W0-1cases, the protein design engine 110 may identify, based at least on the stability of each complex, the three dimensional structure 119 of one or more output protein molecules 117 as a preferred conformation of the output protein molecule 117 for binding with the conditioner molecule 114. Alternatively and / or additionally, the protein design engine 120 may determine, based at least on the stability of each complex, the binding affinity between the output protein molecule 117 and the conditioner molecule 114.

[0084] FIG. 3 depicts a flowchart illustrating an example of a process 300 for flexible docking, in accordance with some example embodiments. Referring to FIGS. 1 and 3, the process 300 may be performed by the protein design engine 110 applying, for example, the protein design computation model 115. As described in more detail below, in some cases, the process 300 may be performed to determine the docked conformation of, for example, the output protein molecule 117 in a binding interaction or in complex with the conditioner molecule 114. Alternatively and / or additionally, the process 300 may be performed to determine the binding affinity between the output protein molecule 117 and the conditioner molecule 114.

[0085] At 302, a protein structure computation model is applied to generate a first three-dimensional structure and a second three-dimensional structure of a protein molecule. In some example embodiments, the protein design computation model may be applied to generate multiple three-dimensional structures (or conformers) of the same protein molecule. For example, in some cases, the protein design computation model may be applied to generate a first three-dimensional structure of the protein molecule by at least operating on a backbone torsion (BBT) representation of the three-dimensional structure of the protein molecule. Furthermore, the protein design computation model may also be applied to generate a second three-dimensional structure of the same protein molecule by again modifying the backbone torsion (BBT) representation of the three-Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1dimensional structure of the protein molecule. In some cases, the first three-dimensional structure and the second three-dimensional structure may be different conformations (or molecular conformers) of the same protein molecule, which may exist as an ensemble of conformations due to the flexible nature of the protein molecule.

[0086] At 304, a docked conformation of a first complex including the first three-dimensional structure of the protein molecule bound to a conditioner molecule is determined. In some example embodiments, one or more rounds of molecular docking may be performed to determine the docked conformation of the first three-dimensional structure of the protein molecule bound to the conditioner molecule. For example, in some cases, the one or more rounds of molecular docking may be performed to determine a preferred orientation of the protein molecule having the first three-dimensional structure relative to the conditioner molecule when the two molecules are engaged in a binding interaction or bound to form a complex. In some cases, rigid docking may be performed, meaning that the first three-dimensional structure of the protein molecule is rotated and / or translated as a single rigid body relative to that of the conditioner molecule without any changes to the constituent atoms.

[0087] At 306, a docked conformation of a second complex including the second three-dimensional structure of the protein molecule bound to the conditioner molecule is determined. In some example embodiments, one or more additional rounds of molecular docking may be performed to determine the docked conformation of the second three-dimensional structure of the protein molecule bound to the conditioner molecule. In some cases, the one or more additional rounds of molecular docking may be performed to determine a preferred orientation of the protein molecule having the second three-dimensional structure relative to the conditioner molecule when the two molecules are engaged in a binding interaction or bound to form a complex. Here again,Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1rigid docking may be performed, which includes rotating and / or translating the second three-dimensional structure of the protein molecule as a single rigid body relative to the conditioner molecule.

[0088] At 308, a stability of the first complex in the docked conformation is determined. In some example embodiments, a value of a metric quantifying the stability of the first complex in the docked conformation determined in operation 304 may be determined. For example, in some cases, the metric may quantify the stability of the interface between the first three-dimensional structure of the protein molecule and the conditioner molecule when the first three-dimensional structure of the protein molecule and the conditioner molecule are in a docked conformation. In some cases, the metric may be determined by applying a variety of scoring functions including, for example, a force field scoring function, an empirical scoring function, a knowledge-based scoring function, a machine learning based scoring function, and / or the like.

[0089] At 310, a stability of the second complex in the docked conformation is determined. In some example embodiments, a value of the metric quantifying the stability of the second complex in the docked conformation determined in operation 306 may be determined. In some cases, the same metric may also be determined to quantify the stability of the interface between the second three-dimensional structure of the protein molecule and the conditioner molecule when the second three-dimensional structure of the protein molecule and the conditioner molecule are in a docked conformation. As noted, the metric may be determined by applying a variety of scoring functions including, for example, a force field scoring function, an empirical scoring function, a knowledge-based scoring function, a machine learning based scoring function, and / or the like.Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1

[0090] At 312, one or more downstream tasks based at least on the metric. In some example embodiments, one or both of the first three-dimensional structure or the second three-dimensional structure may be identified, based at least on the respective values of the metric, as a preferred conformation of the protein molecule for binding with the conditioner molecule. For example, in cases where a comparison of the respective values of the metric indicates that the first complex is more stable than the second complex, the first three-dimensional structure of the protein molecule may be identified as the preferred conformation of the protein molecule for binding with the conditioner molecule.

[0091] Alternatively and / or additionally, the binding affinity between the protein molecule and the conditioner molecule may be determined based on the respective values of the metric for the first complex and the second complex. For example, in some cases, a corresponding mean value, median value, mode value, maximum value, minimum value, and / or range may be determined for the values of the metric across the different three-dimensional structures (or conformation ensemble (CE)) of the protein molecule in a binding interaction or in complex with the conditioner molecule. In some cases, the binding affinity between the two molecules may be determined based on the mean, median, mode, maximum, minimum, and / or range of the different values of the metric quantifying the stability of the interface in the complexes formed by the different three-dimensional structures (or conformers) of the protein molecule. For instance, in cases where the value of the metric of a threshold quantity of three-dimensional structures (or conformers) of the protein molecule satisfies a first threshold, the binding affinity between the protein molecule and the conditioner molecule may also be determined to satisfy a second threshold. Contrastingly, in cases where fewer than the threshold quantity of the different three-dimensional structures (or conformer) of the protein molecule have a metric satisfying the firstAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1threshold, the binding affinity between the protein molecule and the conditioner molecule may fail to satisfy the second threshold.

[0092] As noted, various example embodiments of the protein design computation model described herein may operate on a backbone torsion (BBT) representation of the three-dimensional structure of a protein molecule and a logit representation of the corresponding amino acid residue sequence. In some cases, the amino acid identity and geometry of one or more groups of residues in a protein molecule may be parameterized as follows. In some cases, each amino acid residue may be assigned a logit vector gmi∈ ℝ21whose dimension corresponds to the 20 canonical amino acid types and an unknown residue type. In some cases, the logit vectors may be constrained to a subspace with zero-identity component such that there is a smooth bijective mapping (i.e., softmax) defined between this subspace and a probability simplex of the same dimension. At diffusion time t = 0, the logit vector may be initialized to assign probability p0= 0.99 to the class given by data and probability (1 — po) / 2O to the other classes.

[0093] In some cases, the backbone torsion (BBT) representation may include a backbone rigid frame using Ca, N, and C heavy atoms. In some cases, a translation tmi∈ ℝ3and a rotation rmi∈ SO(3) may be assigned to each amino acid residue. Regardless of residue type, each residue is assigned five torsion angles xmiq∈ SO(2) with q G {0,1, 2, 3,4}, one for the backbone oxygen atom torsion angle x0= > and the rest for the four side chain angles (A.i<X..2<..3<X. A) = ( i<x2>x3>x)- Insome cases, the uniform number of torsion angles over all residues may be necessary for sequence design tasks, since the state of residue identity represented by a logit vector is ambiguous. Note that this results in a coupling of logit and torsion angles in the diffusion modeling. In some cases, to initialize the extra torsion angles given a data point at diffusion time t = 0, random angles in [0, 2π) may be assigned. In instances where only structureAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1prediction task are considered, the extra torsion angles may be masked according to the fixed residue identities.

[0094] In some cases, group of residues may also be assigned a group-level translation tm∈ ℝ3and a rotation rm∈ SO(3), which determine the group’s global orientation. In some cases, the parameters {gmi},Rm)and {fin}may collectively specify the state of a protein molecule, including its sequence and structure.

[0095] Various example embodiments of the protein design computation model described herein may implement diffusion generative modeling to the aforementioned parameters, where each set of parameters is modeled independently. For example, in some cases, rotations and translations may be modeled with the translation parameters tmiand tmconstrained to have zero center-of-mass at residue-group and inter-group levels, or= 0 and Σmtm= 0. In some cases, the modeling of torsion angles may handle specially the angles with 7i-periodicity given by the ground truth residue identity in loss calculations.

[0096] In some cases, a variance preserving (VP) process may be applied for translations and logits while a variance exploding (VE) process may be applied for rotations and torsions. At high noise level, the distributions of logits may approach multivariate normal with unit diagonal covariance matrix, distributions of translations to multivariate normal with a more general covariance matrix, and that of torsion angles and rotations to uniform distributions over SO(2) and SO(3), respectively. In some cases, the convention where diffusion time t = 0 corresponds to the data distribution and t > 0 to noised distributions is followed.

[0113] As there are multiple types of parameters, optimally setting their noise schedules for diffusion generative modeling simultaneously may be non-trivial. In some cases, the heuristicAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1strategy described in more detail below may be applied. For example, in some cases, the minimum noise levels for the different parameters may be set according to their expected data noise scales beyond which the model is not expected to learn to denoise accurately given noisy data, or there is no benefit to training a model to denoise beyond those scales.

[0114] In some cases, residue-level and group-level translation noise schedules may be set by fixing the minimum scale, for example, to 300.0 angstroms. In some cases, noise schedules may be set for the different parameters such that their symmetry breaking phase transition scales occur around the same time of diffusion process. In some cases, the approximate transition may scale to be 1.0 angstrom for the translation parameters, 0.2 radians for the rotation and torsion parameters, and 1.0 for logits. Given the minimum and transition noise scales, the maximum noise scales may be determined based on the functional forms of the schedules. In practice, one or both of the minimum and maximum values may be adjusted based on experiments performed at small scale.

[0115] In some cases, while there may be one set of noise schedules for group-level translations and rotations, this formulation may permit separate noise schedules for residue-level parameters per each residue group. Separate noise schedules may be useful in certain settings, such as binder design where a subset of residue groups may be designated as “target” groups and should be keep relatively fixed, or only partially noise and denoise, compared to the “binder” group whose (subset of) sequence and structure should undergo substantive denoising.

[0116] To further illustrate the various operations of various example embodiments of the protein design computation model describe herein, examples of algorithms with pseudocode implementing a protein design computation model trained to perform de novo generation by denoising at least a portion of an input protein molecule are provided below. In some exampleAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1embodiments, various example embodiments of the protein design computation model described herein may, at inference time, operate on a backbone torsion (BBT) representation of a three-dimensional structure of an input protein molecule and a logit representation of the corresponding amino acid residue sequence to generate de novo at least a portion of the input protein molecule while conditioned on a conditioner molecule (Algorithm 1). In some cases, the InputFeatureEmbedder (Algorithm 2) may act as a shallow encoding network to initialize single (ID) and pair (2D) representations for the primary conditioning trunk. In some cases, relative position encodings may be used to break symmetries across identical residues and chains. In some cases, scaffold sequence information may be provided by one-hot encoded amino-acid identities as well as protein language model embeddings, while residue-level binary indicator variables,{rm }>an<3 auxiliary and task features are used as modeling-task specific conditioning information.Algorithm 1 Main inference1: Input: features {fmJ; scaffold sequence and structure parameters Zmput=c input) r, input) f input) r input) r, input) r input)-> i • i-. - • •,t. ryfiyilffmi J’ mi }xmiq J’ Vmi h Emrm ]}; binary conditioning variable indicators Znx=inference start time point t E (0,1], number of time steps lVsteps, number of corrector steps per time step A / corr, hybrid Langevin sampling parameters A()and I / J2: Output: Generated sequence and structure parametersJ- W3:{s™PUt}' InputFeaturesEmbedder({fmi}, {^}, {tx}, {x^q}, {r *}) S”putGSinpme4: {Sm, {Zminj} Inputconditioner ({S^put}, }) SmiE RCs, {ZminJ) G R%5: <- SampleConditional (ZinPut, Zfix, t)6: Z(°) DiffusionModule(ZW, Zfix, {Smi}, {Zminj}, {S^put}, {Z^}- Nsteps, Ncorr, k0,Algorithm 2 Input Feature Embedder _1: Input: (f,), foS), (tS). {xS, J. {rSJ:2: Output: (S* {z*“J. _res_noise_level i _ _i-r i„ f-res_noise_levelAami" ’ -L00KUPlrmt" ’ JAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1input > --■ / Irres-type c res.confidencefres_at_interfacefpLM res_noise_level C fix! ix 1 ix') ix'll A4- mi -COnCaIlLTmi 'rmi 'Tmi 'Imi ’ami ■ [9mi’5: S^ut«— Linear No Bias (S™put)6: Z RelativePositionEncoding(f^-indcx, fg™“P-index,fchain.index)

[0117] In some cases, the Inputconditioner stack (Algorithm 3) may constitute the bulk of the computation in the protein design computation model, which passes the single (ID) and pair (2D) embeddings initialized by the InputFeatureEmbedder (Algorithm 2) through a simple 1D-2D mixing layer followed by multiple Pairformer-style attention blocks. In some cases, in each Pairformer block, full O(iVr3esiclue) triangle attention may be used across rows and columns of thepair embeddings. In some cases, the final ID and 2D embeddings, {Sm, {Zminj, updated throughAcyciePairformer stacks, along with their initial representationsmaY providethe primary conditioning information for the diffusion module (Algorithm 4)Algorithm 3 Input Conditioneri: mput:2: Output: {Smi},{Zmin7}3: Smi= Linear NoBias ({S^”put})j. 7 > yinputLminj ~Lminj5: Zmin7+= LinearNoBias(SJ,”?ut) + Linear No Bias (s“put)6: {Zmin7} < — 0,0E RCs, Zmm;, E RCz7: for b «— 1 to IVcyclesdo8:9: {SmJ <— {SmJ + LinearNoBias(LayerNorm({Smi)})10:minjZminj+ LinearNoBias(LayerNorm(Zmi7l7))11: T < — FuUyConnectedGraphQ12: {SmJ, {Zmin7} Pairformer({S„, {Zmi„7}, F)13: end for>Algorithm 4 Diffusion Module1: Input: Z, Zfi* {SmJ, [ZminJ], {S^put}, t, Nsteps, Ncorr, Ao,2: Output: Z^3: St ^ t / NstepsAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-14: K7VL Kz} - O’0cev. e e 5. for Upr ec| jC(or* 1 to J^steps do 6: {SmJ, {ZminJ} DiffusionConditioner({SmJ, {ZminJ}, {S^put}, {zX7}' {^T}’ {^S}' 0Smi<e^CsZmtnjeI®Cz7: < D, {S^v}, {ZP7J7} DiffusionScoreModule(Z^\ {Smi}, {Zminj}, t) 8: < — HybridLangeviriReverseStep(|t>, Z(f\ Zflx, t, 8t, 70, x / j) 9. for U-correCfOr* 1 to Ncorrdo 10: < — DiffusionScoreModule(Z(-t\{Smj], {Zmjnyj, t)11: < — HybridLangevinCorrectorStep(, Z^, Zflx, t, 8t, Ao, ip) 12: end for13: t <— t — 8t14: end for15: Z(0)<— Z(t)

[0118] In some cases, Sampl eConditional module in Algorithm 1 may sample noised sequence and structure parameters Z^ at time t conditioned on the input parameters Zinput; these are efficiently implemented as they do not require simulation. In some cases, the module may also receive a set of binary conditioning variable indicators Zflx, which instructs the module to treat the corresponding input variables as conditioning variables and keep them unchanged if the indicator values are set to True. In some cases, for de novo structure prediction tasks, typically all conditioning indicators may be set to False except for those corresponding to the sequence parameters, which are set to True. For binder design task, one plausible example strategy may be to set the indicators corresponding to the target residue groups’ parameters to True, to keep the target sequence and structure fixed, while for the binder being designed, they are set to False for all except for those corresponding to the logit parameters of the portions of the binder sequence that are not designed. In some cases, sampleConditional module may also implement different noise schedules given as input additional non-negative integer-valued noise levels.

[0119] In some cases, conditioning ID and 2D embeddings produced by the InputFeatureEmbedder and the Inputconditioner stack may be further updated by theAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1Diffusionconditioner (Algorithm 5) and combined with embeddings of the current diffusion time step t before being passed to the DiffusionScoreModule (Algorithm 6). In addition, in some cases, self-conditioning may be used where the output ID and 2D embeddings of the previous time step are used in diffusion conditioning.Algorithm 5 Diffusion Conditioner1: Input: {Smi}. (Sir '}. {zZ“'}. {^7}. K4 < 2: Output:3: Smi= concat([S?pnv, S^put])4: Smi<— LinearNoBias(LayerNorm(Smi))5: Smj + = LinearNoBias (LayerN orm(S^fev))6: = TimeEmbedding(t)7: Smj+= LinearNoBias(LayerNorm(nmi))8: for b < — 1 to Astepsdo9: S)n;- + = Transition(Smj,n = 2)10: end for11: Zminj= concat([ZPnZ*])12: Zmin / LinearNoBias(LayerNorm(ZjnjTl7))13: Zmmy + = Linear N oBias (LayerN o rm (Zp^.))14: for b «— 1 to Mstepsdo15: Zminj+= Transition(Zminy,n = 2)16: end forAlgorithm 6 Diffusion Score Module (direct score prediction)1: Input: Z®:= {{5«}, {t®}, {r^}, {t^}, {SmJ, {Zmin7}, t. Output: Scores for all sequence and structure parameters45'■= ( mi }- mi ’ miq}’ mi ’ {< Pm m }3: {fm / na1}’ {rminbalj ComputeG]obalFrames(Z(t))4: g <— KNNRandomGraph([tgl°bal], / c = 32, / = 32)5:6-[Zminj J < Pairformei ({Smi}, {Zmin7}, £ / )7:{Xm7n0)}’ jXmTn0} GlobalFramesEmbedder({r„gl'°ba1}, j^’},8: g ^ KNNGraph({tgl°lbal}, / c = 8)Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-19: for U <— 1 to do11: end for1?. fv((=o)l, JYO°)1 JYG=1)1 (Ami J’(Mnt J (MniO J’ [ 'mill J 13:£X^M°J}' * — GroupNodesEmbedder(Ngroups) X^“o)G RCs', X^j'JG RCsX314: 15: Q < — KNNRandomGraph({^\°0balj, / c = 32, f = 32) 16: Jf < — ResiduesToGroupsGraph() 17: J < — GroupsToResiduesGraphQ 18: for u < — 1 to do 19: {x£=0)}, {x£=1)} EquivariantNN({x^=0)}, {x£=1)}, {t£>}, {x£=0)}, {x^}, {^°ba1}, 0, Jf) K?'”} > K=1}} - EquivariantNN({x«=0)}, {x«=n}, {^oba1}, {x^=0)} U {x£=0)}, 20:(x"-1’) u {x"-1’}, {e“) « (4MS 21: end for

[0120] In some cases, Algorithm 6 presents the diffusion score module that computes scores used in denoising of sequence and structure parameters. In some cases, the function KNNRandomGraph({%J, A, ) may return a set of edges per each input node, including: -nearest neighbor edges and edges involving distant nodes sampled with inverse cubic propensity. The function ComputeGlobalFrames may construct global frames.

[0121] In some cases, the architecture for the score may utilize alternating layers of Pairformer-style attention blocks and equivariant neural networks. In some cases, to reduce computational complexity of the full Pairformer block, triangle attention updates within each Pairformer block may be restricted to the edges of the local neighborhood graph G. In some cases, these triangle attention updates may be further biased with additional geometric information derived from the backbone distogram, planar angles, and dihedral angles.Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1

[0122] In some cases, a form of hybrid Langevin dynamics may be implemented to control the tradeoff between sample diversity and sample quality. In some cases, the target data distribution may be rescaled with a factor analogous to an inverse temperature, PAO(^('°')) po(^('°'))(2(-o-))Ao, where a low-temperature regime, Ao> 1, places greater emphasis high-likelihood states. Doing so may induce a time-dependent temperature rescaling factor of the scores of the form At= F(A0, ot), where the functional form of F depends on the choice of noise schedule and the manifold of the given degrees of freedom. In some cases, to compensate for the fact that this simple rescaling of the learned scores does not adequately reweigh relative sample populations across multimodal distributions, the dynamics may be further modified with an effective equilibration factor tr This heuristic choice may be tantamount to introducing a component of slow, annealed Langevin dynamics, where ip > 0 controls the rate at which the system approximately equilibrates over a sufficiently small time interval. A reverse step under this hybrid dynamics is detailed in Algorithm 7.

[0123] In some cases, when guidance potentials are included at generation time, it may be necessary to include additional corrector operations in the reverse dynamics (Algorithm 8). In some cases, the corrector operations follow annealed Langevin dynamics to allow the reverse diffusion to sufficiently mix under the introduction of additional gradients. In some cases, using up to 5 corrector steps per predictor step was sufficient to generate high quality samples.Algorithm 7 Hybrid Langevin Reverse Step _1: Input: 0>, Z Zfix, t, 6t, Ao, ip2: Output: Z(t"5t)3:4: F = -f(Z&, t) + g(t)2([xt+ $ + EaVt / aguide)5: e Z~SampleNoiseLike(Z('t'))6: w = pcom[F]rit +7: i = PcomAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-18: / (tgt) = expz(t)|W]Algorithm 8 Hybrid Langevin Corrector Step _1: Input:2: Output: Z(t)3: 4: F = |f? CO2C«f + StM“'de) 5: e Z~SampleNoiseLike(Zff))>6:w = Pcom[F]6t + ^(t)VdTG Z 7: W = Pcom[W]8: Z(t:)expz(t) [V / ]

[0124] FIG. 6 depicts a block diagram illustrating an example of a computing system 600, in accordance with some example embodiments. Referring to FIGS. 1-6, the computing system 600 may be used to implement the protein design engine 110, the data store 120, the client device 130, and / or any components therein.

[0125] As shown in FIG. 6, the computing system 600 can include a processor 610, a memory 620, a storage device 630, and input / output devices 640. The processor 610, the memory 620, the storage device 630, and the input / output devices 640 can be interconnected via a system bus 650. The processor 610 is capable of processing instructions for execution within the computing system 600. Such executed instructions can implement one or more components of, for example, the protein design engine 110, the data store 120, the client device 130, and / or the like. In some example embodiments, the processor 610 can be a single-threaded processor.Alternatively, the processor 610 can be a multi -threaded processor. The processor 610 is capable of processing instructions stored in the memory 620 and / or on the storage device 630 to display graphical information for a user interface provided via the input / output device 640.

[0126] The memory 620 is a computer readable medium such as volatile or non-volatile that stores information within the computing system 600. The memory 620 can store dataAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1structures representing configuration object databases, for example. The storage device 630 is capable of providing persistent storage for the computing system 600. The storage device 630 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, or other suitable persistent storage means. The input / output device 640 provides input / output operations for the computing system 600. In some example embodiments, the input / output device 640 includes a keyboard and / or pointing device. In various implementations, the input / output device 640 includes a display unit for displaying graphical user interfaces.

[0127] According to some example embodiments, the input / output device 640 can provide input / output operations for a network device. For example, the input / output device 640 can include Ethernet ports or other networking ports to communicate with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).

[0128] In some example embodiments, the computing system 600 can be used to execute various interactive computer software applications that can be used for organization, analysis and / or storage of data in various formats. Alternatively, the computing system 600 can be used to execute any type of software applications. These applications can be used to perform various functionalities, e.g., planning functionalities (e.g., generating, managing, editing of spreadsheet documents, word processing documents, and / or any other objects, etc.), computing functionalities, communications functionalities, etc. The applications can include various add-in functionalities or can be standalone computing products and / or functionalities. Upon activation within the applications, the functionalities can be used to generate the user interface provided via the input / output device 640. The user interface can be generated and presented to a user by the computing system 600 (e.g., on a computer screen monitor, etc.).Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1

[0129] One or more aspects or features of the subject matter described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs, field programmable gate arrays (FPGAs) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship between client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship with each other.

[0130] These computer programs, which can also be referred to as programs, software, software applications, applications, components, or code, include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, apparatus and / or device, such as for example magnetic discs, optical disks, memory, and Programmable Logic Devices (PLDs), used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. The machine-readable medium can store such machine instructions non-transitorily, such as for example as would a non-transient solid-state memory orAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1a magnetic hard drive or any equivalent storage medium. The machine-readable medium can alternatively or additionally store such machine instructions in a transient manner, such as for example, as would a processor cache or other random access memory associated with one or more physical processor cores.

[0131] To provide for interaction with a user, one or more aspects or features of the subject matter described herein can be implemented on a computer having a display device, such as for example a cathode ray tube (CRT), a liquid crystal display (LCD), a light emitting diode (LED) monitor, or an organic light emitting diode (OLED) monitor for displaying information to the user and a keyboard and a pointing device, such as for example a mouse or a trackball, by which the user may provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well. For example, feedback provided to the user can be any form of sensory feedback, such as for example visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including acoustic, speech, or tactile input. Other possible input devices include touch screens or other touch-sensitive devices such as single or multi-point resistive or capacitive track pads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, and the like.

[0132] In the descriptions above and in the claims, phrases such as “at least one of’ or “one or more of’ may occur followed by a conjunctive list of elements or features. The term “and / or” may also occur in a list of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which it is used, such a phrase is intended to mean any of the listed elements or features individually or any of the recited elements or features in combination with any of the other recited elements or features. For example, the phrases “at leastAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1one of A and B;” “one or more of A and B;” and “A and / or B” are each intended to mean “A alone, B alone, or A and B together.” A similar interpretation is also intended for lists including three or more items. For example, the phrases “at least one of A, B, and C;” “one or more of A, B, and C;” and “A, B, and / or C” are each intended to mean “A alone, B alone, C alone, A and B together, A and C together, B and C together, or A and B and C together.” Use of the term “based on,” above and in the claims is intended to mean, “based at least in part on,” such that an unrecited feature or element is also permissible.

[0133] The subject matter described herein can be embodied in systems, apparatus, methods, and / or articles depending on the desired configuration. The implementations set forth in the foregoing description do not represent all implementations consistent with the subject matter described herein. Instead, they are merely some examples consistent with aspects related to the described subject matter. Although a few variations have been described in detail above, other modifications or additions are possible. In particular, further features and / or variations can be provided in addition to those set forth herein. For example, the implementations described above can be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of several further features disclosed above. In addition, the logic flows depicted in the accompanying figures and / or described herein do not necessarily require the particular order shown, or sequential order, to achieve desirable results. Other implementations may be within the scope of the following claims.

Claims

Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1CLAIMSWhat is claimed is:

1. A computer-implemented method, comprising:applying a protein design computation model to generate a portion of a protein molecule de novo while conditioned on a conditioner molecule,wherein the protein design computation model is trained to co-generate an amino acid residue sequence and a three-dimensional structure of the portion of the protein molecule;obtaining experimental data on a de novo portion of the protein molecule generated by the protein design computation model;determining, based at least on the experimental data, that the de novo portion of the protein molecule satisfies one or more criteria; andin response to the de novo portion of the protein molecule satisfying the one or more criteria, applying the protein design computation model to generate a different portion of the protein molecule containing the de novo portion previously generated by the protein design computation model.

2. The method of claim 1, wherein the protein molecule comprises an antibody.

3. The method of claim 2, wherein the portion of the protein molecule comprises a complementarity determining region (CDR) of the antibody, and wherein the different portion of the protein molecule comprises one or more of a different complementarity determining region (CDR), a heavy chain, a light chain, a fragment antigen binding (Fab), and a variable region (Fv).Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-14. The method of any of claims 2 to 3, wherein the portion of the protein molecule comprises a third heavy chain complementarity determining region (CDR-H3) of the antibody.

5. The method of any of claims 1 to 4, wherein the conditioner molecule comprises one or more of another protein molecule, a chemical compound, a nucleic acid molecule, a lipid, or an ion.

6. The method of any of claims 1 to 5, further comprising:receiving a backbone torsion (BBT) representation of the three-dimensional structure of the protein molecule and a logit representation of the amino acid residue sequence of the protein molecule; andapplying the protein design computation model to operate on the backbone (BBT) torsion representation and the logit representation of the protein molecule in order to co-generate the three-dimensional structure and the amino acid residue sequence of the portion of the protein molecule.

7. The method of claim 6, further comprising:receiving a representation of an amino acid residue sequence and / or a three-dimensional structure of the conditioner molecule; andapplying the protein design computation model to operate on the on the backbone (BBT) torsion representation and the logit representation of the protein molecule and the representation of the conditioner molecule in order to generate the portion of the protein molecule.

8. The method of claim 7, wherein the generating the portion of the protein molecule is conditioned one or both of the amino acid residue sequence and the three-dimensional structure of the conditioner molecule.Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-19. The method of any of claims 7 to 8, wherein the backbone (BBT) torsion representation and the logit representation of the protein molecule and the representation of the conditioner molecule are concatenated to form a representation array, wherein the representation array is associated with a register array populated by a plurality of values, and wherein each value of the plurality of values is indicative of whether a corresponding value in the representation array is associated with the protein molecule or the conditioner molecule.

10. The method of any of claims 6 to 9, wherein the backbone torsion representation of the protein molecule specifies, for each amino acid residue included in the protein molecule, a geometric state of a plurality of constituent backbone atoms, and wherein the geometric state comprises at least one of a translation and rotation of the plurality of constituent backbone atoms.

11. The method of any of claims 6 to 10, wherein the backbone torsion representation of the protein molecule specifies, for each amino acid residue included in the protein molecule, one or more torsion angles formed by a plurality of constituent sidechain atoms.

12. The method of any of claims 6 to 11, wherein the logit representation of the protein molecule includes, for each amino acid residue included in the protein molecule, a logit vector, and wherein the logit vector enumerates a probability of a type of the amino acid residue being each of the 20 canonical amino acid residues.

13. The method of any of claims 1 to 12, wherein the protein design computation model generates the portion of the protein molecule by at least denoising, at least partially in parallel, an amino acid residue sequence and a three-dimensional structure of the portion of the protein molecule.Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-114. The method of claim 13, wherein the protein design computation model denoises the amino acid residue sequence and the three-dimensional structure of the portion of the protein molecule over a plurality of successive timesteps.

15. The method of any of claims 13 to 14, wherein the protein design computation model delays the denoising of the amino acid residue sequence of the portion of the protein molecule until the three-dimensional structure of the portion of the protein molecule has undergone a threshold degree of denoising.

16. The system of any of claims 13 to 15, wherein the protein design computation model denoises the amino acid residue sequence and the three-dimensional structure of the portion of the protein molecule while keeping a three-dimensional structure of the conditioner molecule fixed.

17. The system of any of claims 13 to 16, wherein the protein design computation model denoises the amino acid residue sequence and the three-dimensional structure of the portion of the protein molecule while simultaneously modifying a three-dimensional structure of the conditioner molecule to reflect one or more conformational changes engendered by a binding interaction with the protein molecule.

18. The method of any of claims 1 to 17, further comprising:training, based at least on a training dataset, the protein design computation model to generate the portion of the protein molecule de novo while conditioned on the conditioner molecule.

19. The method of claim 17, further comprising:Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1generating a training sample to include a protein-protein complex of a sample protein molecule bound to a conditioner molecule in which a portion of the sample protein molecule is masked;generating an additional training sample to include the same protein-protein complex with a different portion of the sample protein molecule masked; andgenerating the training dataset to include the training sample and the additional training sample.

20. The method of claim 19, wherein the protein-protein complex comprises an immune complex in which the sample protein molecule comprises an antibody and the conditioner molecule comprises an antigen.

21. The method of claim 20, wherein the portion of the protein molecule comprises a complementarity determining region (CDR) of the antibody, and wherein the different portion of the protein molecule comprises one or more of a different complementarity determining region (CDR), a heavy chain, a light chain, a fragment antigen binding (Fab), and a variable region (Fv).

22. The method of any of claims 20 to 22, wherein the portion of the protein molecule comprises a third heavy chain complementarity determining region (CDR-H3) of the antibody.

23. The method of any of claims 1 to 22, wherein the de novo portion of the protein molecule is determined to satisfy the one or more criteria when the experimental data indicates that one or more of a binding affinity, binding specificity, functionality, and developability of the protein molecule including the de novo portion satisfy one or more thresholds.

24. The method of any of claims 1 to 23, wherein the experimental data is obtained by at least synthesizing the protein molecule having the de novo portion generated by the proteinAttorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-1design computation model and analyzing one or more synthesized samples of the protein molecule.

25. The method of any of claims 1 to 24, wherein the experimental data is obtained by at least performing one or more computational analysis on the protein molecule having the de novo portion.

26. The method of any of claims 1 to 25, further comprising:applying the protein design computation model to generate a first three-dimensional structure of the protein molecule and a second three-dimensional structure of the protein molecule;determining a docked conformation of a first complex including the first three-dimensional structure of the protein molecule bound to the conditioner molecule;determining a docked conformation of a second complex including the second three-dimensional structure of the protein molecule bound to the conditioner molecule;determining a stability of the first complex in the docked conformation; determining a stability of the second complex in the docked conformation; and identifying, based at least on the stability of the first complex and / or the second complex, at least one of the first complex and the second complex as a preferred conformation of the protein molecule for binding with the conditioner molecule.

27. The method of claim 26, further comprising:determining, based at least on the stability of the first complex and / or the second complex, a binding affinity between the protein molecule and the conditioner molecule.Attorney Ref.: 14786-094-228 (103963-228094) / P60064-W0-128. The method of claim 27, wherein the binding affinity between the protein molecule and the conditioner molecule is determined to correspond to a respective value of a metric quantifying the stability of the first complex and the stability of the second complex.

29. The method of any of claims 26 to 28, further comprising:generating, for the protein molecule, a conformation ensemble to include the first three-dimensional structure of the protein molecule and the second three-dimensional structure of the protein molecule.

30. The method of any of claims 1 to 29, wherein the protein design computation model is applied to generate incrementally larger and / or more complex portions of the protein molecule de novo while conditioned on the conditioner molecule.

31. A system, comprising:at least one data processor; andat least one memory storing instructions, which when executed by the at least one data processor, result in operations comprising the method of any of claims 1 to 30.

32. A non-transitory medium storing instructions, which when executed by at least one data processor, result in operations comprising the method of any of claims 1 to 30.