A method for antibody design based on geometric features and molecular simulation

The reinforcement learning framework efficiently predicts antibody binding events by iteratively modifying sequences and structures, addressing the limitations of existing methods to achieve precise and cost-effective antibody design.

WO2026158810A1PCT designated stage Publication Date: 2026-07-30NEC ONCOIMMUNITY AS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NEC ONCOIMMUNITY AS
Filing Date
2025-05-16
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Current methods for antibody design face challenges in efficiently navigating the vast configuration space and accurately predicting binding events due to limitations in energy functions and the lack of biological priors, leading to time-intensive and costly laboratory screenings that often overlook functionally critical epitopes.

Method used

A computer-implemented method using a reinforcement learning framework iteratively changes biomolecule sequences and structures to predict binding events by employing an agent and environment system, with a reward function based on simulations and proxy functions like normal mode analysis to accelerate the process.

Benefits of technology

This approach allows for the rapid identification of molecules with high binding affinity, reducing computational time and cost while effectively targeting specific epitopes, enabling precise antibody design and production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025063547_30072026_PF_FP_ABST
    Figure EP2025063547_30072026_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method of identifying one or more molecules that are predicted to be likely to instigate a binding event with a target molecule is disclosed. The method comprises: accessing a biomolecule sequence and / or structure of an initial molecule; inputting the biomolecule sequence and / or structure of the initial molecule into a reinforcement learning framework to predict a modified molecule that is predicted to instigate a binding event with the target molecule, wherein the reinforcement learning framework is configured to iteratively change the biomolecule sequence and / or structure of the initial molecule to generate the modified molecule, and comprises an agent and an environment, wherein the agent is configured to propose a change to the biomolecule sequence and / or structure of the molecule in its current state based at least in part on a calculated reward, and the environment is configured to implement the change proposed by the agent and to simulate the molecule, wherein a reward is calculated based on the simulation, for feedback to the agent. The present invention can be used in a variety of applications including, but not limited to, several anticipated use cases in drug development, optimization of antibodies and medical / healthcare.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] A METHOD FOR ANTIBODY DESIGN BASED ON GEOMETRIC FEATURES AND MOLECULAR SIMULATION

[0002] FIELD OF THE INVENTION

[0003] The present invention relates to computer implemented methods of identifying one or more molecules that are predicted to be likely to instigate a binding event with a target molecule. The invention has particular application for antibody-antigen binding prediction.

[0004] BACKGROUND

[0005] Antibodies (Abs) represent a rapidly expanding class of biological therapeutics in the human pharmaceutical market. Consequently, the need for optimal selection, design, and engineering of antibodies has not only grown significantly but has also emerged as a key competitive factor. Despite substantial interest from the pharmaceutical industry, the discovery of novel Abs remains a complex and costly process, often necessitating extensive laboratory screenings to ensure specific targeting. This screening process is inherently time-intensive and typically prioritizes antibodies with the highest binding affinities, which are usually associated with immunodominant epitopes. However, this approach can inadvertently overlook antibodies that, despite exhibiting lower affinities, target functionally critical sites. Furthermore, viral adaptation mechanisms may result in the elimination of these immunodominant epitopes, thereby potentially compromising antibody efficacy.

[0006] The computational design of Abs provides a promising alternative to mitigate these challenges by significantly reducing the time and costs associated with conventional laboratory-based screenings. This approach facilitates controlled screening of desired biophysical properties and enables precise targeting of specific epitopes of interest. Nonetheless, protein design for specific targets remains a complex endeavour, often yielding low success rates and requiring multiple rounds of affinity maturation experiments [DO - Fischman, S. & Ofran],Challenges in protein design largely arise from the limitations of current energy functions, which are used to calculate free energies but often fail to account for factors such as entropy and molecular flexibility, as they are based on static models that represent only a single conformational state. This is particularly relevant for Ab design, given that Abs typically feature extended, highly flexible loops known as complementarity-determining regions (CDRs). Given the marginal stability of proteins, with unfolding free energies generally ranging from 5-15 kcal / mol, accurate computation of these energies is essential for achieving precision in protein design.

[0007] Recent machine learning approaches have provided partial solutions, focusing either on inverse or forward design strategies. However, these methods face challenges in navigating the vast configuration space required for effective Ab design. Without prioritization mechanisms, these models necessitate prolonged computational times to evaluate potential configurations. Furthermore, current models often lack biological priors — language models, for instance, do not incorporate knowledge of underlying biological processes — resulting in limitations in obtaining and validating biologically feasible structures within this expansive design space.

[0008] The problem that is addressed by our method is to find good candidates with high binding energy. The search space is exponential in the number of amino acids (AA) to modify (about 20 to the power of the number of positions). Existing methods either use complex proxy functions (e.g. using protein folding) or weak proxy functions (e.g. QED).

[0009] SUMMARY OF THE INVENTION

[0010] In accordance with a first aspect of the invention there is provided a computer-implemented method of identifying one or more molecules (e.g. “binder molecules”) that are predicted to be likely to instigate a binding event with a target molecule, the method comprising:

[0011] accessing a biomolecule sequence and / or structure of an initial molecule (e.g. that is predicted to instigate a binding event with a target molecule);inputting the biomolecule sequence and / or structure of the initial molecule into a reinforcement learning framework to predict a modified molecule that is predicted to instigate a binding event with the target molecule, wherein

[0012] the reinforcement learning framework is configured to iteratively change the biomolecule sequence and / or structure of the initial molecule to generate the modified molecule, and comprises an agent and an environment, wherein the agent is configured to propose a change to the biomolecule sequence and / or structure of the molecule in its current state based at least in part on a calculated reward, and

[0013] the environment is configured to implement the change proposed by the agent and to simulate the molecule, wherein a reward is calculated based on the simulation, for feedback to the agent.

[0014] In this way, the method of the present invention utilises a reinforcement learning framework for molecule (e.g. antibody) design by incrementally updating the initial molecule, based on a reward function. The molecule(s) that are identified by and used in the method may be described herein as “binder molecules” (in the sense that they are predicted to bind to the target molecule), or “query molecules”. The identified molecules may be part of larger structures. Preferably the identified molecules are antibodies.

[0015] The initial molecule is preferably a known structure (for example an existing antibody) that has been obtained from a database (e.g. the protein data bank, PDB). Typically, the initial molecule is predicted to instigate a binding event with the target molecule. In some examples, we can start from a neutral or simple amino acid (e.g. glycine) by substituting the existing amino acids.

[0016] The method may be used to generate a plurality of candidate molecules that are predicted to be likely to instigate a binding event with a target molecule, and further filtering and / or ranking steps may be applied to the plurality of candidate molecules (e.g. using longer simulations) to generate a subset of molecules.Preferably, the method may be used to generate an amino acid sequence (e.g. protein) that is potentially part of a larger structure (e.g. protein, in some cases antibodies) that binds to a target (typically another protein).

[0017] The agent preferably comprises a machine learning model. The agent preferably comprises an action / critic architecture comprising an action network (e.g. configured to implement an action policy) and a critic network (e.g. configured to predict a value function).

[0018] As discussed, the environment is configured to implement the change proposed by the agent and to simulate the molecule (e.g. in its updated state after the change has been implemented), wherein a reward is calculated based on the simulation. Thus, the environment is configured to perform a (computational) simulation. The reward function is typically a cost function (e.g. with respect to a target value) that needs to be optimised.

[0019] In some embodiments, the simulation simulates at least one of: a structure of the molecule; a binding event between the molecule and the target molecule.

[0020] Preferably, the simulation comprises computing an equilibrated structure (e.g. of or comprising the molecule). In other words, preferably the simulation comprises an equilibration step. This may be referred to as a first, or “initial”, simulation phase. In such embodiments, the change proposed by the agent of the reinforcement learning framework is implemented by the environment and the updated molecule is simulated to compute an equilibrated structure, for example at the temperature and pressure of interest. Typically, the equilibrated structure is used to calculate the reward for feedback to the agent. In other words, the equilibration step provides a sufficiently stable (e.g. “realistic” or “reasonable”) sample from which the reward may be calculated. For example, following the implemented change, the system may initially be far from equilibrium and the equilibration step provides a sufficiently stable sample for calculation of the reward. The structure is an in-distribution structure at the temperature of interest. In this way, the reward is calculated based on the simulation.Thus, the first aspect of the invention may provide a computer-implemented method of identifying one or more molecules that are predicted to be likely to instigate a binding event with a target molecule, the method comprising:

[0021] accessing a biomolecule sequence and / or structure of an initial molecule; inputting the biomolecule sequence and / or structure of the initial molecule into a reinforcement learning framework to predict a modified molecule that is predicted to instigate a binding event with the target molecule, wherein

[0022] the reinforcement learning framework is configured to iteratively change the biomolecule sequence and / or structure of the initial molecule to generate the modified molecule, and comprises an agent and an environment, wherein the agent is configured to propose a change to the biomolecule sequence and / or structure of the molecule in its current state based at least in part on a calculated reward, and

[0023] the environment is configured to implement the change proposed by the agent and to compute an equilibrated structure (e.g. of or comprising the molecule), wherein a reward is calculated based on the equilibrated structure, for feedback to the agent.

[0024] In some embodiments an energy minimisation step may be performed in addition to (e.g. before) or as part of computing the equilibrated structure. The step of equilibration (and typically energy minimisation, when performed) may generally be referred to as relaxation of the molecule following the implemented change. As an alternative to energy minimisation, conservation of energy may be used.

[0025] Typically, the equilibrated structure comprises the molecule (e.g. antibody) and the target molecule (e.g. antigen) in a bound state.

[0026] Typically, the simulation models the movement and interaction of the atoms of the molecule over a plurality of (discrete) time steps.Preferably, the simulation is not a “full” simulation (e.g. in the sense that the present invention uses relatively short timeframe simulations of each state of the molecule for the iterative reinforcement learning approach). The simulation may be based on molecular dynamics. The simulation may be based on accelerated molecular dynamics. In some examples the simulation may be used to provide a direct estimation of the binding free energy of the molecule.

[0027] The reward may be indicative of the binding affinity of the molecule in its current state to the target molecule.

[0028] The reward may comprise (e.g. an approximation of) the binding free energy of the molecule. In some embodiments, the reward function may be a combination of the binding free energy (or approximation thereof) with additional reward metrics, for example one or more of binding time, and frequency of attachment and detachment. The reward may comprise a calculation or prediction of the binding affinity of the molecule in its current state to the target molecule. Herein, the terms “binding affinity” and “binding free energy” are typically used to mean the same reward property.

[0029] In embodiments, the simulation may comprise substantially only computing an equilibrated structure. In some embodiments, the simulation may simulate a binding event between the molecule and target molecule. For example, following equilibration, the simulation may comprise a further, second, simulation step (e.g. further simulation “phase”) from which the binding free energy between the molecule and the target molecule may be directly calculated or estimated. The further simulation step typically simulates a binding event between the molecule and target molecule and may for example include simulating a binding time and / or and frequency of attachment and detachment of the molecule with the target molecule. These metrics are indicative of the free binding energy of the molecule: for example a higher frequency of detachment may indicate a relatively lower free binding energy. The further simulation step may generate additional configurations after the initial equilibration step.The reward may be (e.g. alternatively or further) based on a proxy function, preferably based on normal mode analysis.

[0030] To advantageously accelerate the method, the method may make use of a proxy function (e.g. instead of or together with a direct estimation of the binding free energy) such as a cost function based on normal mode analysis. In some examples, the reward is calculated based on the simulation jointly with a proxy cost function (preferably based on a prediction model that uses normal mode analysis).

[0031] Typically, a proxy function is used as a less computationally expensive alternative to simulating a binding event (e.g. by performing a further simulation step following equilibration) and directly calculating the reward from the further simulation step. In other words, the term “proxy” is used in the context of calculating the reward without performing a further simulation phase following the initial simulation phase to compute the equilibrated structure. Thus, the use of a proxy function advantageously speeds up the method by predicting the reward property without a further simulation phase and allows for increased exploration of the molecular landscape by the reinforcement learning algorithm.

[0032] In preferred embodiments where the reward is based on a proxy function, the reward (e.g. “proxy reward”) comprises a prediction of the binding affinity of the molecule to the target molecule. Typically, the prediction of the binding affinity (or other proxy reward) is based on the calculated equilibrated structure from the simulation. Preferably, the proxy function is based on a normal mode analysis of the equilibrated structure.

[0033] Preferably, (e.g. where the reward is based on a proxy function) the reward is generated using a prediction model configured to predict a metric indicative of the binding affinity of the molecule in its current state to the target molecule, based on the simulation (e.g. based on the equilibrated structure). In other words, the proxy function may be generated using a prediction model. Preferably, the prediction model is a trained machine learning model.The proxy function is typically generated based on (e.g. one or more parameters of) the equilibrated structure computed in the simulation. For example, following the calculation of the equilibrated structure in the simulation, one or more parameters of the equilibrated molecule (such as the biomolecule sequence and structural data) may be input into a trained machine learning model to calculate a reward function. An example of such a trained model is the ANTIPASTI model

[0014] which is configured to predict a binding affinity based on normal mode analysis (e.g. of the equilibrated structure).

[0034] Alternative or additional proxy functions (e.g. proxy cost functions) may be used to evaluate the reward, for example based on predicted template modelling.

[0035] In some cases, the molecule may be optimised over multiple reward properties. For example, a plurality of different property predictions (“rewards”) may be generated per molecule, or separate predictors per molecule may be used.

[0036] The biomolecule sequence may be an amino acid sequence. The biomolecule sequence may be a nucleotide sequence.

[0037] The change to the biomolecule sequence and / or structure may comprise one or more of:

[0038] addition of an amino acid or nucleotide

[0039] removal of an amino acid or nucleotide

[0040] substitution (e.g. “modification”) of an amino acid or nucleotide.

[0041] The change to the biomolecule sequence and / or structure of the molecule is preferably proposed by an action network of the agent. The molecule in its current state implements the most recent (e.g. iterative) change to the biomolecule sequence and / or structure.Preferably, the change to the biomolecule sequence occurs on a single location of the sequence. In this way, the present invention provides an incremental construction of the modified molecule(s), starting from the initial molecule.

[0042] The simulation may be a continuous simulation. By providing a continuous simulation, advantageously no datasets for the reward are required, and the final molecule may be evaluated directly.

[0043] The use of a continuous simulation is particularly advantageous, as it means that the method does not require a new simulation beginning from first principles following each change proposed by the agent and implemented by the environment. This allows a greater number of iterations to be performed by the learning framework (for example of the order of 106iterations) for a given timeframe compared to if individual discrete simulations (e.g. beginning from first principles) were performed following each proposed change, thereby allowing a greater exploration of the molecular landscape.

[0044] Preferably, in a continuous simulation, the simulation of the molecule in its current state is based on the simulation of the molecule performed for a previous state. In other words, information from previous iterations of the reinforcement learning framework may be re-used as the simulation transitions from one state of the molecule to the next. For example, node coordinates from previous iterations can be re-used following an implemented change. In another example, the simulation following an implemented change may re-start from the closest structure, for example using geometric vicinity. In other words, the changes made in previous iterations may be “merged”, rather than re-starting the simulation from first principles each time a change to the molecule is implemented. The continuous simulation may be considered to be continuous across multiple iterations of the reinforcement learning framework.

[0045] In particularly preferred embodiments, the simulation is a continuous simulation, and the simulation comprises computing an equilibrated structure. In this way, the method advantageously performs continuous equilibration of the molecule as it isiteratively updated in the reinforcement learning framework. Therefore, the method of the invention is speeded up in comparison with performing individual “standalone” simulations following each implemented change, especially when a proxy reward is used based on the equilibrated structure.

[0046] The simulation may be accelerated using at least one of relaxation, metadynamics, population-annealing or weighted-ensemble method. Herein, relaxation is preferably used to merge the change(s) in previous simulations (e.g. iterations) so that the same simulation can continue running.

[0047] The agent may comprise an action network and a critic network, preferably wherein the action network and the critic network are implemented as a transformer architecture.

[0048] The input to the agent may include the biomolecule sequence and preferably geometric features of the molecule in its current state; and the output of the agent includes the location and change type of the biomolecule sequence.

[0049] In some embodiments, the agent (e.g. action network) comprises a trained machine learning model, wherein the machine learning model is trained to reconstruct existing molecule structures. For example, the machine learning model may be trained using a training dataset comprising a plurality of existing molecular structures. The model is trained to reconstruct existing molecules from an input that comprises a modification (perturbation) to an existing structure.

[0050] The molecule(s) may be a first protein having an amino acid sequence, and the target molecule may be a second protein (e.g. having a second amino acid sequence). In general, the protein may be a protein, protein-domain, or a protein sub-unit. The term “protein” may include any protein subsequence that may have a viable 3D or functional structure. The term "protein" may include a complex of multiple proteins.In some embodiments, the molecule is or comprises an antibody (or part thereof) and the target molecule is or comprises an antigen and / or an epitope. In some examples the target antigen may comprise a neoantigen and / or neoepitope. The target antigen may be a tumor associated antigen such as human epidermal growth factor receptor 2 (HER2).

[0051] Further disclosed herein is a method of ranking a plurality of candidate molecules that are predicted to be likely to instigate a binding event with a target molecule, the plurality of candidate molecules having been identified using any of the methods described above.

[0052] In accordance with a second aspect of the invention there is provided method of creating a biologically active molecule, comprising

[0053] identifying one or more molecules that are predicted to be likely to instigate a binding event with a target molecule, using a method of any of the examples described above, and

[0054] manufacturing a biologically active molecule based on the identified molecule(s).

[0055] The biologically active molecule may be referred to as a biomolecule. The method may manufacture a pharmaceutical compound. The biologically active molecule may comprise antibodies. The biologically active molecule may comprise protein(s) nucleic acids, DNA, RNA, and / or mRNA. The biologically active molecule may be used as a medicine. The biologically active molecule may be used in a diagnostic assay. For example, a lateral flow test may be manufactured that includes the biologically active molecule.

[0056] Further disclosed herein is a computer program product comprising instructions which, when executed by a computer, cause the computer to perform any of the methods described above.

[0057] Further described herein is a system for identifying one or more molecules that are predicted to be likely to instigate a binding event with a target moleculecomprising at least one processor in communication with at least one memory device, the at least one memory device having stored thereon instructions for causing the at least one processor to perform any of the methods described above.

[0058] The invention may further comprise designing, by the methods as described herein, and subsequently producing, one or more antibodies that are predicted to bind to an epitope and / or a predicted epitope.

[0059] Antibodies that are predicted by the present invention to bind to an epitope and / or a predicted epitope will find use for a variety of industrial purposes, such as therapeutics and diagnostics, biotechnological applications, or research settings, as would be understood by the skilled person.

[0060] Accordingly, the present invention may find use in one of more of the following use cases:

[0061] • Disease treatment or prevention, including that for infectious diseases or other diseases such as autoimmune or immune-related diseases and cancer, in particular immunotherapy or passive immunisation.

[0062] • Drug delivery in the context of antibody-drug conjugates.

[0063] • To assist in scientific research, through immunoprecipitation studies or immunohistochemistry studies.

[0064] • Diagnostics through disease detection assays, medical imaging, and biosensors.

[0065] In some embodiments, the method may further comprise generating a corresponding protein or peptide sequence for one or more antibodies that are predicted to bind to an epitope and / or a predicted epitope. In some embodiments, the method may further comprise generating a DNA or RNA sequence that encodes one or more antibodies that are predicted to bind to an epitope and / or a predicted epitope.

[0066] The antibodies as disclosed herein may be of therapeutic use, particularly in therapeutics and diagnostics. To administer the antibody to a subject, the proteinor peptide may be delivered to an individual through a variety of routes, including but not limited to intravenous, intramuscular, intradermal, intraaural, intraarterial, intraocular, intravitreal, intranasal, subcutaneous, or inhalation.

[0067] An alternative way of ensuring the production of antibodies in a subject is by using various biotechnological techniques which involves using a nucleic acid sequence (i.e. DNA or RNA) which encodes a desired antibody. One such technique is the use of CRISPR, which can modify a subject’s DNA by inserting a DNA sequence into cells (i.e. B-cells) which encodes for the desired antibody, thereby facilitating a means by which the subject is able to produce the desired antibody from the transcription of this new genetic sequence and the subsequent translation of the RNA transcribed from it. Alternatively, DNA may be delivered as an exogenous molecule, typically as part of a larger vector construct with elements to facilitate its transcription to result in the production of the desired antibody. mRNA may also be delivered in a similar way, with its translation producing the desired antibody.

[0068] Thus, in accordance with a further aspect of the present invention, there is provided a method of creating an antibody, comprising: identifying one or more antibodies that are predicted to be likely to instigate a binding event with a target molecule, such as an epitope, from any of the examples discussed above; and synthesising the antibody, or encoding the antibody, or predicted or simulated variants thereof, into a corresponding protein, peptide, DNA or RNA sequence.

[0069] Such a protein, peptide, DNA or RNA sequence may be delivered in a naked or encapsulated form or incorporated into a genome or cell of a bacterial or viral delivery system. In addition, bacterial vectors can be used to deliver the DNA in to vaccinated host cells.

[0070] In accordance with a further aspect of the present invention, there is provided a method of creating and / or designing a diagnostic assay to determine whether a patient has or has had a cancer or prior infection with a pathogen, wherein the diagnostic assay is carried out on a biological sample obtained from a subject, comprising identifying at least one antibody that is predicted to bind to an epitopeand / or a predicted epitope which characterises a certain cancer or pathogen, using a method according to any of the examples discussed above; wherein the diagnostic assay comprises the utilisation within the biological sample of the antibody that is predicted to bind to an epitope and / or a predicted epitope which characterises a certain cancer or pathogen.

[0071] In this way, the present invention may advantageously be used to create a diagnostic test or assay, by means of a rapid and / or automated antibody discovery.

[0072] The term utilisation as used herein is intended to mean that the at least one antibody that is predicted to bind to an epitope and / or a predicted epitope is used in an assay to identify the presence of a particular epitope and / or predicted epitope in a subject.

[0073] The in vitro diagnostic assay may comprise identification of a biological component (i.e. target molecule) within the biological sample. In a preferred embodiment, the biological component may be an epitope or a predicted epitope. For example, the assay may comprise the identification of epitopes and / or predicted epitopes that may be bound by an antibody that is predicted to bind such epitopes.

[0074] As an example of such a diagnostic use, a sample, preferably a blood sample, isolated from a patient may be analysed for the presence of antibodies that are predicted to bind to an epitope and / or a predicted epitope within the biological sample that recognise and bind to such an epitope(s), the antibodies of which have been identified as part of the present invention and that are contained within the assay.

[0075] Suitable diagnostic assays would be appreciated by the skilled person, but may include enzyme-linked immune absorbent spot (ELISPOT) assays, enzyme-linked immunosorbent assays (ELISA), cytokine capture assays, intracellular staining assays, tetramer staining assays, microfluidic devices, lab-on-a-chip, microarrays, flow cytometry, CyTOF, or limiting dilution culture assays.In a method of creating a diagnostic test, the amino acid sequence of the one or more antibodies that are predicted to bind to an epitope and / or a predicted epitope may be chosen based on the desired epitope and / or predicted epitope to be tested. For example, the one or more source antibodies (i.e. the initial binder molecule) may be one or more source antibodies which have binding affinity towards an epitope of any selected pathogen and / or virus (or fragments thereof). In such a case, the present invention may be used to create a diagnostic test for determining whether a patient has or has had prior infection with the virus, and / or its variants and / or related viral species. However, as will be appreciated by the skilled person, the one or more source antibodies may have binding affinity towards any epitope of any pathogen (e.g. any virus, parasite, bacterium, or cancer cell).

[0076] Further disclosed herein is a diagnostic assay to determine whether a patient has or has had prior infection with a pathogen and / or cancer, wherein the diagnostic assay is carried out on a biological sample obtained from a subject, and wherein the diagnostic assay comprises the utilisation of at least one antibodies that are predicted to bind to an epitope and / or a predicted epitope, wherein the epitope is from a pathogen or tumour, using any of the methods as described herein. The diagnostic assay may comprise identification of a biological component within the biological sample that that serves as an epitope that may be bound by the antibody as described herein.

[0077] The invention may further comprise designing, by the methods as described herein, and subsequently producing one or more nucleic acids that are predicted to bind to a target molecule, such as a protein. Such nucleic acid molecules may comprise aptamers or mRNA.

[0078] Aptamers that are predicted by the present invention to bind to a target protein will find use for a variety of industrial purposes, such as therapeutics and diagnostics, biotechnological applications, or research settings, as would be understood by the skilled person.Accordingly, the present invention may find use in one of more of the following use cases:

[0079] • Disease treatment or prevention, including that for infectious diseases or other diseases such as autoimmune or immune-related diseases and cancer, in particular for anti-cancer therapies and antiviral treatments. • Drug delivery in the context of aptamer-drug conjugates.

[0080] • Drug discovery when used to mimic ligands for the study of receptorligand interactions.

[0081] • Diagnostics through disease detection assays, medical imaging, and biosensors.

[0082] The potential implementation of the aptamers designed by the methods as described herein and subsequently produced will be similar to the foregoing embodiments described for antibodies, and therefore the skilled person would be able to adjust these methods to adapt them for use with aptamers.

[0083] The methods of the present invention may be used to design, and subsequently produce, RNA (including mRNA) and / or RNA binding proteins after predicting specific RNA-protein interactions important for RNA processing, stability, functionality, and transport. Certain sequence and structural motifs comprised within RNA molecules can be recognised by RNA binding proteins, and therefore the design of RNA molecules and / or RNA binding proteins is critical to ensure their correct functioning in biological systems.

[0084] Thus, in accordance with a further aspect of the present invention, there is provided a method of creating a nucleic acid molecule, such as an aptamer, comprising: identifying one or more nucleic acid molecules, or aptamers, that are predicted to be likely to instigate a binding event with a target molecule, such as a protein, from any of the examples discussed above; and synthesising the nucleic acid molecule, or aptamer, or encoding the aptamer, or predicted or simulated variants thereof, into a corresponding DNAor RNA sequence.

[0085] BRIEF DESCRIPTION OF THE DRAWINGSEmbodiments of the invention will be briefly described with reference to the drawings, in which:

[0086] Figure 1 illustrates a schematic diagram of an antibody-antigen binding;

[0087] Figure 2 illustrates an example of an antibody design workflow;

[0088] Figure 3 illustrates a de-novo design workflow;

[0089] Figure 3A illustrates the principal steps of a method according to an embodiment of the invention;

[0090] Figure 4 illustrates an example of Reinforcement Learning (RL) for antibody design;

[0091] Figure 5 illustrates a schematic diagram of an accelerated dynamic using local changes and relaxation to generate in-distribution samples;

[0092] Figure 6 illustrates a schematic diagram of adding a residual;

[0093] Figure 7 illustrates a schematic diagram of removing a residual;

[0094] Figure 8 illustrates an example architecture where the features from the context / environment and the features from the current design are input into the model and then the action is predicted;

[0095] Figure 9 illustrates an example architecture where the features from the context / environment and the features from the current design are input into the model and then the action is predicted;

[0096] Figure 10 illustrates an overview diagram illustrating an embodiment of the invention;

[0097] Figure 11 schematically illustrates a system suitable for implementing embodiments of the invention; and

[0098] Figure 12 schematically illustrates a server suitable for implementing embodiments of the invention.

[0099] DETAILED DESCRIPTION

[0100] Disclosed herein is reinforcement method for antibody design that uses efficient continuous simulation and additive evolution, where the reward approximates the binding free energy using normal mode analysis. The method allows for generation of antibodies in complex environments.Although the following description is directed primarily to antibody design, it will be appreciated that the methods of the present invention may be applied to other molecule binding events and use cases, for example small molecules / RNAs that bind to proteins, aptamers, CAR design and protein / peptide design more generally.

[0101] Figure 1 is a schematic diagram illustrating antibody-antigen binding. The antibody 101 has a binding area, shown generally at 102. The antigen 105 has a binding area shown generally at 106. A binding event will occur when the binding area 102 of the antibody 101 corresponds with the binding area 106 of the antigen 105. We want to design amino-acid sequences of antibodies 101 that produce effective binding affinities (binding free energies) with the target antigen 105. The space of the possible amino-acid sequences is very large (potentially infinite); thus, reinforcement learning may be used to address this problem.

[0102] Figure 2 shows an antibody design workflow example 200. We show the sequence of steps for the antibody design, where first we select the target, then predict the binding area of the antigen; then we add the environment (water and lipid layer) and simulate with Molecular Dynamics (MD) to explore the binding areas. The next step is to generate the antibody for the highest priority binding spots. We can then select the most promising and validate with simulation for each of the candidate.

[0103] In more detail, in step 201 antibody target selection is performed based on a relevant biological metabolism (e.g. the relevant biological pathway being addressed). In step 202 static epitope prediction is performed. In step 203 dynamic epitope prediction is performed based on a Molecular Dynamics simulation of the epitome. Reference numeral 203a represents the Molecular Dynamics (MD) simulation 203a. Step 204 corresponds to de-novo antibody design based on protein design. This step includes antibody generation 204a. Step 205 corresponds to quantum mechanics (QM) candidate ranking. This step includes Molecular Dynamics simulation 205a.Figure 3 shows a design workflow 300 for antibody design according to an embodiment of the invention. The workflow takes a target 301 (e.g. such as a target epitope) and environment 302 and generates (e.g. antibody) candidates 307. The process includes epitope prediction 303, epitope evaluation 304, de-novo design 305, and candidate ranking 306.

[0104] In Figure 3, the de-novo design workflow is composed of:

[0105] 1) from the target, predict the epitope / paratope (steps 303 and 304). This may be performed using conventional techniques known in the art;

[0106] 2) from the epitope, design the new protein / anti-body (step 305); and

[0107] 3) ranking multiple candidates based on a longer simulation (step 306).

[0108] Figure 3A is a flow diagram outlining the principal steps of a method 500 according to an embodiment of the invention.

[0109] At step S501, the amino acid sequence and / or structure of an initial antibody is accessed. The initial antibody is typically a known structure that is predicted to bind a target antigen. The amino acid sequence and / or structure of the known initial antibody may be obtained from a publicly available database such as the protein data bank, PDB.

[0110] At step S503, the amino acid sequence and / or structure obtained in step S501 is input into a reinforcement leaning (RL) framework. As discussed in further detail herein, the RL framework is configured to iteratively change the amino acid sequence and / or structure of the initial antibody to generate a modified antibody that is predicted to bind the target antigen.

[0111] At step S505, the RL framework outputs one or more modified antibodies (i.e. modified with respect to the initial antibody) that are predicted to bind the target antigen. Then, in step S507 the one or more modified antibodies identified in step S505 are simulated to model their binding affinity with the target molecule in order to generate a filtered / ranked list of candidate antibodies. The (moleculardynamics) simulations performed in step S507 are generally additional “standalone” simulations outside of the RL framework. The simulations in step S507 are typically more complete simulations (in the sense that they are run for longer) than the simulations for each molecule iteration in the RL framework.

[0112] In some implementations, a plurality of candidate antibodies are generated, and further filtering and / or ranking steps may be applied to the plurality of candidate antibodies to generate a subset of antibodies.

[0113] The principal steps of the reinforcement learning framework of step S503 are illustrated at S510. As will be described in further detail herein, the RL framework comprises an agent and an environment. In step S512, the agent proposes a change to the amino acid sequence and / or structure of the antibody in its current state (this may be the first action performed on the initial antibody, or a proposed change to the antibody in its current, modified state). The change may typically be the addition, removal and / or change of an amino acid of the antibody sequence. At step S514, the environment of the RL framework implements the change proposed by the agent in step S512 and simulates the updated antibody using a continuous simulation. The simulation computes an equilibrated structure (e.g. an equilibrate combined structure of the antibody in its current state and the target antigen). The equilibration step is performed at the temperature and pressure of interest, and may be performed using techniques known in the art.

[0114] The simulation is a continuous simulation in that the changes made in previous iterations may be “merged”, rather than re-starting the simulation each time a change is implemented. Hence, the environment is continuously equilibrating the structure as changes proposed by the agent are implemented by the environment. The simulation is typically based on molecular dynamics.

[0115] At step S516, a reward is calculated based on the simulation of the antibody in its current state. The reward is fed back to the agent which may then propose a further modification for simulation based on the reward. The RL framework therefore iteratively performs the steps S512, S514 and S516 to update the antibody. The reward calculated in step S516 may comprise a calculation orapproximation of the free binding energy of the antibody in its current state. This may be achieved by performing a further “binding” simulation phase using the equilibrated structure to directly calculate the free binding energy. Alternatively or additionally, the reward may be based on a proxy function, preferably based on normal mode analysis of the equilibrated structure (without performing the further “binding” simulation phase), which may advantageously accelerate the method compared to directly calculating the binding free energy.

[0116] The steps S512 to S516 of the RL framework may be iteratively performed for a user-defined number of iterations, or until a threshold reward value is achieved.

[0117] The steps of the method 500 will now be described in further detail.

[0118] The Reinforcement Learning (RL) approach

[0119] We consider a RL method, where the agent generates a proposal for changing an initial antibody. Figure 4 schematically illustrates an example of RL for antibody design.

[0120] Figure 4 schematically illustrates an example of a reinforcement learning (RL) architecture 400 according to an embodiment of the invention. The RL architecture 400 includes an agent 410 and an environment 430. The agent 410 performs an action (for example a modification to the initial antibody) which is implemented and simulated in the environment 430. The simulation includes an initial equilibration step to generate a stable structure from which a reward indicative of the binding affinity of the antibody in the current state to the target molecule (e.g. target antigen) can be calculated. The reward may be based on a proxy function as is described further herein. Based on the simulation of the modified antibody, the updated state of the antibody and a reward are fed back to the agent 410 which utilises this information for proposing the subsequent action to the antibody. This process, described in further detail below, iterates using a continuous simulation until the final antibody design is achieved.

[0121] The RL framework may include a replay buffer to store data obtained from interactions between the agent and environment for training of the agent. Thefeatures indicated at 450 may optionally be used to increase the speed of the simulation, or provide additional possibilities, as will be discussed herein.

[0122] The agent 410 has a policy (action network) 402 and a critic (prediction network) 404, for example using PPO (Proximal Policy Optimization). At each iteration, it will decide whether to:

[0123] 1) Add an AA (amino acid)

[0124] 2) Remove an AA

[0125] 3) Modify an AA

[0126] 4) Do nothing

[0127] 5) Stop

[0128] To perform this action, the agent 410 first predicts the location of the modification and then the type of modification.

[0129] The main idea is that the model (in particular the environment of the RL framework) is continuously equilibrating the structure, starting from the initial known structure. Since the changes performed by the RL iteration only change part of the molecule, the equilibration is fast. For example, the initial structure target structure is taken from protein database (PDB). If the target is new an initial simulation is run or a forward folding method is used, possibly followed by a simulation.

[0130] Since a full simulation (e.g. for each iteration of the RL framework) is not possible, we consider:

[0131] 1. When we apply a change, we first equilibrate locally

[0132] 2. We then expand the equilibration, up to a certain distance I area (e.g. a minimum number of amino acids, such as 2 neighbouring amino acids) 3. We may first energy minimize the move if the change is too high

[0133] 4. We then use the structure as input

[0134] 5. For acceleration we also consider the use of a proxy function, for example based on normal mode analysisIn step 4 above, the equilibrated, or “relaxed” structure from steps 1-3 may be input into a further simulation step (“production step”) which directly calculates the binding free energy of the binding event between the antibody and antigen. For the purposes of the RL method, this further simulation step is significantly shorter than a “full” simulation. For example, in the further simulation step, a single attachment / de-attachment interaction between the antibody and the antigen may be performed to calculate the free binding energy. This is in contrast to a “full” simulation which might typically model many attachment / de-attachment interactions between the antigen and the antibody, but for which the timeframe of doing so would be unsuitable for the iterative method of the present invention. Such a further simulation step may be performed using molecular dynamics techniques know in the art.

[0135] In order to speed up the method, as shown in step 5 above, instead of (or in some cases in addition to) inputting the equilibrated structure into a further simulation step, a proxy function may be used instead of a direct calculation of the free binding energy as a reward indicative of the binding affinity between the antibody and antigen. One example of a proxy function is an evaluation of binding affinity based on normal mode analysis of the equilibrated structure.

[0136] Normal mode analysis is an alternative method to MD, that can be used to evaluate the binding affinity of two or more structures by analyzing the low frequency modes of the combined structure.

[0137] A preferred way of evaluating the binding affinity using normal mode analysis is to use a predictive model trained to predict a binding affinity based on the normal mode analysis. An example of such a trained predictive model is the ANTIPASTI model

[0014] ,

[0138] Thus, as a reward function for the RL, we consider either the direct estimation of the free binding energy from the samples computed in the short simulation or other proxy function. A preferred reward function is one based on normal mode analysis. Given a structure, the modal analysis produces the Hessian and lowereigenvalues associated and we use this information with the structure of the system to map the structure to a prediction of free binding energy. The prediction function is trained on an existing dataset and updated during training when the sample(s) have high accuracy. Additional optional proxy measure(s) can be added, for example when the uncertainty of the model is high. There include predictive template modelling (pTM) given the structure or a proxy function trained on the pTM of available structures.

[0139] Reward based on molecular dynamics

[0140] Figure 5 schematically illustrates how accelerated dynamic simulation uses local changes and relaxation to generate in-distribution samples. Figure 5 schematically illustrates the processes performed by the environment 430 of the reinforcement learning framework.

[0141] Figure 5 shows how at regular intervals we reconsider evaluating a new action. After equilibration we take the last structure and use it as input for the next prediction. Reward is the target used by the RL algorithm, for us the reward is preferably the free binding energy (reward / cost function). An alternative reward(s) could be a combination of free binding energy, binding time and frequency of detach and reattach.

[0142] Moving from left to right through the figure, firstly at step S550 (“Implement Change”) the change required by a latest action proposed by the agent 410 is implemented. As discussed, this may be adding, removing or modifying an amino acid, or stop / do nothing. Once the change is implemented, the environment performs an energy minimisation step (S552). The energy minimised structure is then input to simulation at step S554 (“simulation”) to compute or approximate the free binding energy of the modified antibody by modelling the binding interaction between the antigen and antibody. This will typically involve equilibrating the structure before modelling the binding interaction. The computed free binding energy from the dynamic simulation S554 is used as a reward function (S580) for feedback to the agent 410. Alternatively, the simulation at S554 may equilibrate the antibody / antigen structure and normal mode analysis is performed on theequilibrated structure to generate a proxy reward indicative of the binding free energy. As discussed above, in such cases the proxy award is typically calculated using a trained predictive model such as the ANTIPASTI model.

[0143] The process iterates as schematically indicated at steps S550a and S552a where the next change proposed by the agent, based on reward S580, is implemented and the structure energy minimised.

[0144] Optionally, in addition to computing or approximating the binding free energy from simulation S554, at step S560 baricenter structures are computed (“compute baricenter structures”). From this a predictive template modelling (pTM) score may be calculated (step S562) and included in the reward for feedback to the agent.

[0145] Predictive template modelling (pTM) is a measure of the shape of a protein (see reference

[0010] ). pTM is the predicted template modelling (pTM) score that is derived from template modelling (TM) score, which measures the accuracy of the global structure of the protein and is relatively insensitive to localized inaccuracies as defined in

[0010] , A pTM score may be calculated and used together with the binding energy to generate the reward (S580). The pTM may be an additional reward score or integrated with the reward based on the computed binding free energy.

[0146] Optionally, at step S570, geometric features (“geo-features”) of the antibody in the current state may be calculated, and may be fed back to the agent.

[0147] In some embodiments, clustering (e.g. k-means clustering) of equilibrium configurations may be performed to cluster structures that are similar in structure. This may be based on the calculated baricenters of the structures.

[0148] Adding, Removing, or modifying the residual and accelerated simulation

[0149] Figures 6 and 7 schematically illustrate how actions proposed by the agent 410 of the reinforcement learning framework may be implemented.When we decide to add (e.g. add an amino acid), we first add the new residual and then minimize the energy or the free energy.

[0150] Figure 6 shows a schematic diagram for adding a residual.

[0151] At step 602, a decision is made (e.g. by the action network 402 of the agent 410) to add an amino acid to the current state of the antibody. Here the action is to add an amino acid in order to modify the molecular structure of the antibody, illustrated at 604 (“Implement change”). Following the implementation of the change at 604, energy minimization (schematically shown at 606) is performed on the modified structure.

[0152] A simulation is performed (step 608) to generate the reward function for the current state of the antibody. In the example shown in Figure 6, the dynamic simulation is performed using population annealing metadynamics. Following the simulation at 608, the reward is returned to the agent, and the next change decided by the agent is implemented at 604a, followed by energy minimisation (606a). In this way, the method iteratively updates the antibody.

[0153] Alternatively, the action may be to remove an amino acid. Figure 7 shows a schematic diagram for removing a residual. The steps are similar to those shown in Figure 6. Firstly, at step 702 a decision is made by the action network 402 of the agent 410 to remove a residual from the current state of the antibody. The change is implemented by removing the amino acid in order to modify the molecular structure of the antibody, illustrated at 704 (“Implement change”). Following the implementation of the change at 704, the updated structure is equilibrated schematically shown at step 706. Following equilibration 706, a further simulation step (708) is performed to calculate the reward function for the current state of the antibody. The method then iterates as schematically shown at 704a and 706a.

[0154] Architecture

[0155] As an example, the action and critic architecture can be implemented as shown in Figure 8, where the core model is a transformer-based architecture, the output isthe location and type of change, plus the expected reward or value function for the critic model, and the input includes the context / environment, the current design in both sequential (amino acids) and position (configuration).

[0156] Figure 8 illustrates a schematic example architecture where the features from the context / environment and the features from the current design are input into the model and then the action is predicted.

[0157] The output includes the location on the current amino acid sequence where the action is to take place, the action to be taken (e.g. add, remove, stop, change, leave), and the amino acid that is to be added, removed or changed. These outputs are vector encoded. The output also includes the value function 805 of the critic.

[0158] The input to the model includes the current sequence 801 (letter is vector encoded) together with features (e.g. of the current state of the antibody). The input includes geometric features 803 of the current structure (e.g. C_\alpha coordinates, Orientation). The antibody sequence is vector encoded. The context input to the model is shown at 807. This includes the target molecule (e.g. antigen) sequence, together with geometric features. The target molecule sequence is typically vector encoded.

[0159] Figure 9 illustrates an example architecture 900 similar to that shown in Figure 8 and where like reference numerals indicate like features. However, in Figure 9 the context input into the model comprises the target molecule (e.g. antigen) sequence only. In some examples, geometric features could be derived from the sequence information using known structural prediction models.

[0160] Entropy minimization

[0161] As an alternative to energy minimization we can require conservation of energy. If we have one sample, we ask that

[0162] p(%)£(%) « p(x' E(x'with x' the previous sample. In general:

[0163]

[0164] which leads to the minimization problem

[0165]

[0166] and the update to be

[0167]

[0168] where (3 = y. kBis the Boltzmann constant and Tis the temperature.

[0169] In the above equations, p(x) is the probability of the configuration x, whose potential energy is £(%). Typically we assume a Boltzmann distribution, p(x) oce- / ?£■(%)N|S the number Of configurations of x’. x~ and x+are the configurations before and after the update respectively. Vx£(x) is the gradient of the gradient in the coordinates of the potential energy. (E) is the mean of the potential energy with respect to the probability distribution, (E) = J p(x)£(x)d(x).

[0170] Figure 10 illustrates an overview diagram illustrating an embodiment of the invention, together with various aspects thereof. The boxes with dotted shading represent the recursive loop of the reinforcement learning framework, with the unshaded boxes representing off-line processing steps. The numbered features represent advantageous aspects of the invention as described below.

[0171] Reference 1 : As discussed, embodiments of the invention may use a proxy reward based on normal mode analysis. For example, normal mode analysis may be performed on the equilibrated structure following simulation, with the normal mode analysis used to generate a proxy reward indicative of the binding affinity of the antibody in its current state with the target antigen. The proxy reward may be generated using a trained predictive model, such as the ANTIPASTI model. Thepredictive model may be trained using an existing dataset (Reference 1b). Using a proxy reward advantageously speeds up embodiments of the method in comparison with performing a further simulation step based on the equilibrated structure.

[0172] Reference 2: Embodiments of the invention utilise a continuous simulation to compute the structure during the iterative process of the reinforcement learning. The continuous simulation allows the environment to continuously equilibrate the structure rather than starting the simulation “from scratch” at each iteration, thereby advantageously speeding up the recursive loop.

[0173] Reference 3: As discussed, embodiments of the invention advantageously use both (e.g. continuous) simulation and a proxy reward to provide additive evolution for antibody design.

[0174] Reference 4: In embodiments, the agent utilises geometric features (e.g. backbone and residual geometric features) together with sequence information. In embodiments the agent model is based on a transformer architecture. (See also Reference 9.)

[0175] Reference 5: In some embodiments the inverse design of the antibody may be used as a feature for the agent model. For example, embodiments of the invention may use the inverse problem on the protein backbone as sequence input: at the end of the simulation, the backbone is classified, and this sequence is used as input.

[0176] Reference 6: In embodiments the action policy (e.g. action network 402) is trained to reconstruct existing protein (e.g. antibody) structures based on a dataset comprising existing structures. This is done using imitation learning, where the policy is trained as a supervised learning problem: the state is the input and the action is the correct amino acid. The input will be changed locally but the label is fixed from the dataset. Typically the existing structures of the training dataset are obtained from the protein data bank (PDB). These may be existing antibodyantigen structures. The action network is trained to reconstruct an existingantibody from an input that includes one or more modifications to the original amino acid sequence. In this way the action network is trained to learn an action (e.g. an amino acid addition, removal or change) between a given amino acid sequence and an existing protein structure.

[0177] Reference 7: The output of the recursive loop comprises generated structures with uncertainty and associated reward and energy. These structures may be further analysed, for example by performing longer “standalone” simulations outside of the recursive loop. The generated structures may be used to fine tune the policy (Reference 8).

[0178] Reference 9: In embodiments, an uncertainty of the current structure may be calculated. This may be used to improve (e.g. fine tune) the reward function.

[0179] In the above discussion we have focussed on the embodiments for antibody design. Further embodiments of the invention may include the following:

[0180] • Small molecules / RNAs that bind to proteins, aptamers, CAR design

[0181] • Embodiments for Protein design: Since antibodies are proteins, the same method can be applied to generic protein design.

[0182] • Embodiments for Peptide design: Peptides are small proteins, and the method could be applied.

[0183] Embodiments of the invention may include a method or methods comprising the following steps:

[0184] 1) (Target definition) We first identify the target system and an initial Ab structure; this can be obtained from the protein database (e.g. PDB) or by first forward folding the sequence and then running molecular dynamic (MD); epitope and paratope prediction can be used to limit the search domain.

[0185] 2) (Optimization) We start the main loop, where the agent receives the current structure and then generates the proposed change, at each new proposal we run an accelerated simulation (Figure 5) and query a proxy rewarddefined based on the normal mode analysis. The end of the simulation (after user defined number of iterations) provides the candidates antibody.

[0186] a. optional: The simulation is further accelerated using meta-dynamics or population-annealing or weighted-ensemble method.

[0187] b. optional: We pre-train the action and possibly the critic, using existing data and introducing random changes.

[0188] 3) (Candidate validation) After candidate antibody candidates have been generated, they will be evaluated with a longer simulation to improve the estimation of the free binding energy. This step is optional.

[0189] Features of embodiments of the invention and their advantages are set out as follows:

[0190] 1) Reward function for antibody design and incremental updates to existing structure:

[0191] a. Use of accelerated simulation to compute an approximate binding free energy or an equilibrate structure

[0192] i. In an alternative, use the entropy minimization to generate realistic sample

[0193] ii. The acceleration is obtained by inserting or removing nodes (i.e. atoms) from the previous configuration in a progressive way (see Figure 6 and Figure 7)

[0194] b. Use of geometric feature and latent transformer to encode the action and critic network

[0195] c. Use of a proxy reward based on normal mode analysis (where the predictive model, which takes as input the structure and the normal model analysis, is trained separately from existing dataset) d. One-sentence: generate in distribution sample and query the normal-mode analysis as a reward for the RL

[0196] e. Accelerated simulation: the use of relaxation (see Figure 6 And Figure 7) and meta-dynamics (see Figure 5), or another acceleration method (see definition below)Further features of embodiments of the invention and their advantages are detailed below:

[0197] a. Use of continuous simulation to evaluate the reward

[0198] a. Advantage:

[0199] i. No need to have a dataset for the reward

[0200] ii. Direct evaluation of the final property

[0201] b. Faster simulation: restrict the simulation dynamically to reduce computation time: variable radius simulation

[0202] a. Can increase over time and epochs (e.g. with improvements to the model)

[0203] b. Include uncertainty in the estimation of the reward (binding energy)

[0204] c. Scheduling the transition of the simulation (at minimum is energy-min + equilibration)

[0205] d. ML for binding-affinity prediction (proxy) e. Restart from the closest structure (either previous step or some other configuration, using for example geometric vicinity)

[0206] f. Only energy minimization or scheduled annealing (high / medium / low / medium), where the scheduling is trained as a bilevel problem

[0207] c. Incremental build of the AA sequence using actions (add / remove / change) on a single location, but starting from a correct sequence (existing antibody)

[0208] d. Use transformer network for the action / critic with geometric features a. Use of the clusters of equilibrium configurations for features (k-means)

[0209] e. Online estimation of binding energy using accelerated Molecular Dynamic, Meta Dynamic, Population Annealing or Weighted Ensamplef. Accelerate simulation: simulation needs to be short: ignore H (hydrogen atoms) and include constraints in Backbones and Sidechains (simplified simulation) or focus on the active area g. Hot restart: if the action brings a configuration already seen, the old output will be used as starting point

[0210] h. Use inverse problem on the backbone as sequence input: at the end of the simulation, the backbone is “classified”, and this sequence is used as input

[0211] i. Off-policy: action policy is first trained to reconstruct existing structures based on PDB (protein DataBase) dataset; pretraining action by random change of existing antibody: we train the action model, by perturbing an existing antibody in one position and ask the model to predict the position and the action, for example we modify one amino acid of the original sequence. We can also apply multiple modification and the model need to be able to predict in sequence

[0212] j. Train proxy function for the reward function, for example using ANTIPASTI model, and available data

[0213] a. Use backbone / residual structure to compute the reward function

[0214] b. Based on residual Normal Mode Analysis k. Loop on computed samples: query costly models I longer simulation: continuous simulation

[0215] l. Additional properties of the structure (e.g. surface accessible surface area, SAS) may be used as part of the reward feedback in the RL framework

[0216] Additional information is provided below:

[0217] Accelerated Simulation: Meta-Dynamic, Parallel Tempering, Population Annealing, Weighted Ensemble, MM / PBSA and MM / GBSA (molecular mechanics energies with Poisson-Boltzmann or generalized Born and surface areacontinuum solvation), Thermodynamic Integration (Tl), Bennett's acceptance ratio, Free-Energy perturbation, Alchemic path integration methods are method used to accelerate simulation in molecular dynamics. The methods either introduce biases based on collective variable, uses a known reference system or perform a thermal expansion to explore larger configuration spaces, thus reducing time to compute the desoldered properties.

[0218] Reward function: reward function is used in Reinforcement learning as a cost function that needs to be optimized. For each state, the reward function returns the value or cost associated.

[0219] An advantage of the embodiments of the invention described herein is the ability to incrementally build an antibody starting from a known structure and evaluate binding free energy and the structure to be used as input in an efficient way.

[0220] Applications

[0221] X. Antibody applications

[0222] The optimal selection, design, and engineering of antibodies tailored to specific purposes (i.e. epitopes) requiring immunisation or natural sources leads to major improvements in therapeutics, diagnostics, drug design, research and development, and cost. They are of particular importance in the fields of infectious disease and cancer. Examples are provided below.

[0223] X.1 Therapeutic antibodies

[0224] Antibodies are the key effector molecules of the humoral immune response, produced by plasma B-cells and serving the purpose of identifying and binding to foreign antigens to assist in eliminating and preventing the spread of pathogens in the body. The effector functions of antibodies are principally neutralization (i.e. arresting a biological effect), opsonization (i.e. marking the target for phagocytosis), and antibody-dependent cellular cytotoxicity (i.e. attracting effector cells to eliminate the threat).Computational in silico and machine learning approaches have greatly advanced our understanding of antibody binding and has improved the pipeline to produce desired antibody candidates which have specificity for a target epitope and exhibit high binding energy.

[0225] The binding event of an antibody to an antigen / epitope is a key event in activating the humoral B Cell immune response against any pathogenic threat, such as viral and bacterial infections, in addition to cancer antigens in malignant tumour cells. The production of antibodies towards target epitopes with good binding properties is essential for continuing therapeutic antibody development. The production of antibodies that are predicted to bind to an epitope and / or a predicted epitope, as done by the methods as described herein not only provides an enhanced functional understanding of the critical residues involved antibody-antigen binding; it also aids in the selection of therapeutically relevant epitopes, particularly those which are comprised within functionally critical sites but are neglected by the natural immune system due to immunodominance.

[0226] The antibodies produced by the methods as described herein may be produced with the intention of being therapeutic towards any one of a plethora of potential diseases. Such indications include, for example, hematologic, oncologic, immunologic, neurologic, musculoskeletal and pulmonary diseases.

[0227] The methods as described herein may also be suitable for the generation of chimeric antigen receptors (CARs). CARs are receptor proteins that are engineered into T cells to give them the ability to target antigens. This is done by fusing an single-chain variable fragment (scFv) domain (which is derived from the variable region of an antibody and has affinity towards an antigen / epitope) to an intracellular signalling module which can induce T cell activation upon antigen / epitope binding. This allows T cells to directly detect antigens and become activated, rather than requiring an initial processing and presentation of antigens / epitopes by MHC molecules for T cells to recognise and bind to. As such, they are particularly useful in detecting cell surface antigens, such as those expressed on cancer cells which represent therapeutic targets, such as CD19 and BCMAfor B-cell cancers.X.2 Diagnostics

[0228] Antibodies designed by the methods as described herein may prove useful in developing diagnostic tests towards target epitopes. For example, antibodies may be generated that exhibit higher affinity towards a target epitope than naturally occurring alternatives, thus increasing assay sensitivity and lessening the about of biomarker required to produce a positive test result. In practice, this may result in the detection of certain indications at an earlier stage when compared to using antibodies with lower affinities (i.e. a less sensitive assay). This could, for example, involve the detection of early cancer markers.

[0229] Since the production of novel antibodies is a complex and costly process, often necessitating extensive laboratory screenings to ensure specific targeting, computational design of antibodies provides a promising alternative to mitigate these challenges by significantly reducing the time and costs associated with conventional laboratory-based screenings. This factor, combined with the ability to develop antibodies with high affinities towards target epitopes, will expedite and improve the development of novel biosensors and diagnostic platforms like lateral flow assays or ELISAs which may be required in the event of disease outbreak or pandemic.

[0230] X.3 Research and development and drug design

[0231] Various research applications may be improved by the antibodies produced by the methods of the present invention. The improved specificity and affinity of such antibodies will improve techniques such as immunoprecipitation for the isolation of proteins are isolated from mixtures, and immunohistochemistry using fluorescently or enzymatically labelled antibodies for observing the localisation of proteins in cells or tissues.

[0232] Additionally, the blocking or neutralizing of target proteins by specifically engineered antibodies can be used to study the functions of select proteins in vitro or in vivo.Antibodies designed by the methods as described herein will also have use in drug discovery and design, due to the potential of such antibodies to be developed towards specific epitopes which result in protein activation or inhibition, thereby confirm specific proteins’ roles in diseases, which will subsequently expedite drug discovery pipelines.

[0233] Y. Other protein applications

[0234] Since antibodies are proteins, the methods as described herein are also applicable to generic protein design. For example, how a protein interacts with other molecules, such as small molecules, ligands, DNA, RNA, or other proteins, may be improved or optimised.

[0235] For example, protein-ligand binding concerns the binding of a selected protein to a ligand, such as a small molecule or biomolecule, with high affinity and selectivity. These studies are relevant in the field of drug discovery. Another application would be for protein-protein interactions, which can result in the development of therapies intended to disrupt pathological protein-protein interactions, for example in cancer or neurodegenerative diseases. A further application would be found in protein-nucleic acid binding studies. As an example, ribosome binding proteins (RBPs) can bind RNA targets through the interaction between their amino acid residues and the nucleotides of the RNA. These binding proteins serve to regulate various activities and properties of RNA transcripts, such as splicing, stability, localization, and translation. RBPs may be designed to enhance the binding properties to these RNA motifs and structures, such as the 5’ cap, 3’ UTR, internal ribosome entry sites (IRES), or stem-loop and secondary / tertiary structures. The optimisation of affinity of the RNA-binding protein for the target mRNAcan be done by adjusting the binding interface, which would then impart its effect onto the target RNA.

[0236] Z. Nucleotide applications

[0237] Aptamers are single-stranded nucleic acid molecules that bind to target molecules with high specificity and affinity and comprise either DNA or RNA nucleotides. Likeantibodies, the effective binding to their target molecule exhibited by aptamers is due to their folding into a 3D structure specific to the target molecule. Their target molecules may include peptides, proteins, organic compounds, small molecules, metal ions. Additionally, they can be made to target biological entities such as mammalian cells, viruses, bacteria, and yeast.

[0238] Due to their non-immunogenic properties, and additionally being easier to synthesise and modify, aptamers exhibit several notable advantages. As such, they have widespread use in industry and represent a promising therapeutic modality to treat various diseases and conditions.

[0239] Due to being a molecule which has binding properties, they share many applications with antibodies, from therapeutics and drug delivery to diagnostics and medical imaging.

[0240] Messenger RNA (mRNA) is a single stranded RNA molecule transcribed from a gene that is read by the ribosome to synthesise a protein. Upon transcription from the gene, mRNA undergoes various processing reactions remove non-translated elements (introns) and protected from degradation. The processed mRNA molecule is then translated into a protein through being read by a ribosome which catalyses the polymerisation of amino acids corresponding to the encoded sequence of the mRNA. mRNAs comprise specific sequence motifs or secondary structures, such as the 5’ cap, 3’ untranslated region, stem-loop structures, and internal ribosome entry sites (IRES). RBPs can bind RNA targets through the interaction between their amino acid residues and the nucleotides of the RNA. These binding proteins serve to regulate various activities and properties of RNA transcripts, such as splicing, stability, localization, and translation. As such, the methods as described herein may be used to design, and subsequently produce, RNAs, including mRNAs, which optimally interact with RBPs.

[0241] Computational System

[0242] Figure 11 schematically illustrates an example of a system suitable for implementing embodiments of the invention. The system comprises at least oneserver 710 which is in in communication with data store 700, such as a database. The server is also in communication with a client device 730 over a communications network 720 such the internet. The client device may be a desktop computer or mobile device for example. In use, a request to perform the methods of any of the embodiments of the invention may be input by a user to the client device 730. The client device may send the request to the server 710 over the communications network 720. The server 710 implements embodiments of the invention as described herein, and may retrieve data, such as amino acid sequences, from the data store 700. The resulting molecular structure(s) predicted by performing the present invention may then be communicated to the client device 730 over network 720.

[0243] An example of a suitable server 700 for implementing embodiments of the invention is shown in Figure 12. In this example, the server includes at least one microprocessor 701 , a memory 702, and an external interface 704, interconnected via a bus 705 as shown. In this example the external interface 704 can be utilised for connecting the server 700 to peripheral devices, such as the communications networks, data stores, other storage devices, or the like. The server may also optionally include an input / output device 703, such as a keyboard and / or display,

[0244] In use, the microprocessor 701 executes instructions in the form of applications software stored in the memory 702 to allow the required processes to be performed, including identifying one or more molecules that are likely to instigate a binding event with a target molecule as described herein. The applications software may include one or more software modules, and may be executed in a suitable execution environment, such as an operating system environment, or the like. Accordingly, it will be appreciated that the server 700 may be formed from any suitable processing system, such as a suitably programmed client device, PC, web server, network server, or the like. Whilst the term server is used, this is for the purpose of example only and is not intended to be limiting.

[0245] Whilst the server 700 is a shown as a single entity, this is not essential and it will be appreciated that the server 700 can be distributed over a number ofgeographically separate locations, for example as part of a cloud-based environment.

[0246] References

[0247] References useful for understanding the invention are listed below, all of which are incorporated by reference herein in their entirety.

[0248] [0] DO - Fischman, S. & Ofran, Y. Computational design of antibodies. Current Opinion in Structural Biology 51, 156-162 (2018).

[0249] [1] D1 - Michalewicz, Kevin, Mauricio Barahona, and Barbara Bravi. “ANTIPASTI: Interpretable Prediction of Antibody Binding Affinity Exploiting Normal Modes and Deep Learning.” bioRxiv, December 23, 2023. https: / / doi.Org / 10.1101 / 2023.12.22.572853.

[0250] [2] D2 - Subramanian, Jithendaraa, Shivakanth Sujit, Niloy Irtisam, Umong Sain, Derek Nowrouzezahrai, Samira Ebrahimi Kahou, and Riashat Islam. “Reinforcement Learning for Sequence Design Leveraging Protein Language Models.” arXiv, July 3, 2024. https: / / doi.orq / 10.48550 / arXiv.2407.03154.

[0251] [3] D3 - Watson, J. L. et al. De novo design of protein structure and function with RFdiffusion. Nature 620, 1089-1100 (2023).

[0252] [4] D4 - Bennett, N. R. et al. Atomically accurate de novo design of single-domain antibodies. 2024.03.14.585103 Preprint at : / / doi.orq / 10 1101 / 2024.03 14.585103 (2024).

[0253] [5] D5 - Olivecrona M, Blaschke T, Engkvist O, Chen H. Molecular de novo design through deep reinforcement learning. J Cheminform 2017;9:48

[0254] [6] D6 - Mason DM, Friedensohn S, Weber CR, Jordi C, Wagner B, Meng SM, et al. Optimization of therapeutic antibodies by predicting antigen specificity from antibody sequence via deep learning. Nat Biomed Eng 2021;5:600-12.[7] D7 - Mason, D. M. et al. Deep learning enables therapeutic antibody optimization in mammalian cells by deciphering high-dimensional protein sequence space. 617860 Preprint at https: / / doi.orq / 10.1101 / 617860 (2019).

[0255] [8] D8 - Ohue, M., Kojima, Y. & Kosugi, T. Generating Potential Protein-Protein Interaction Inhibitor Molecules Based on Physicochemical Properties. Molecules 28, 5652 (2023).

[0256] [9] D9 - Bickerton, G.R.; Paolini, G.V.; Besnard, J.; Muresan, S.; Hopkins, A.L. Quantifying the chemical beauty of drugs. Nat. Chem. 2012, 4, 90-98.

[0257] pTM definition pTM: predicted template modelling (pTM) score; derived from a measure called the template modelling (TM) score. This measures the accuracy of the global structure of the protein and is relatively insensitive to localised inaccuracies (

[0010] Xu and Zhang, 2010).

[0010] D10 - Xu, J. & Zhang, Y. How significant is a protein structure similarity with TM-score = 0.5? Bioinformatics 26, 889-895 (2010).

[0258]

[0011] D11 - Pereira, T. O., Abbasi, M. & Arrais, J. P. Enhancing reinforcement learning for de novo molecular design applying self-attention mechanisms. Briefings in Bioinformatics 24, bbad368 (2023).

[0259]

[0012] D12 - Xu, X. et al. AB-Gen: Antibody Library Design with Generative Pretrained Transformer and Deep Reinforcement Learning. Genomics, Proteomics & Bioinformatics 21, 1043-1053 (2023).

[0260]

[0013] D13 - Kosugi, T. & Ohue, M. Quantitative Estimate Index for Early-Stage Screening of Compounds Targeting Protein-Protein Interactions. International Journal of Molecular Sciences 22, 10925 (2021).

[0261]

[0014] Michalewicz K, Barahona M, Bravi B. ANTIPASTI: Interpretable prediction of antibody binding affinity exploiting normal modes and deep learning. Structure.

[0262] 2024 Dec 5;32(12):2422-2434.e5. doi: 10.1016 / j.str.2024.10.001. Epub 2024 Oct 25. PMID: 39461331

Claims

CLAIMS1. A computer-implemented method of identifying one or more molecules that are predicted to be likely to instigate a binding event with a target molecule, the method comprising:accessing a biomolecule sequence and / or structure of an initial molecule; inputting the biomolecule sequence and / or structure of the initial molecule into a reinforcement learning framework to predict a modified molecule that is predicted to instigate a binding event with the target molecule, whereinthe reinforcement learning framework is configured to iteratively change the biomolecule sequence and / or structure of the initial molecule to generate the modified molecule, and comprises an agent and an environment, wherein the agent is configured to propose a change to the biomolecule sequence and / or structure of the molecule in its current state based at least in part on a calculated reward, andthe environment is configured to implement the change proposed by the agent and to simulate the molecule, wherein a reward is calculated based on the simulation, for feedback to the agent.

2. The computer-implemented method of claim 1, wherein the simulation simulates at least one of:a structure of the molecule;a binding event between the molecule and the target molecule.

3. The computer-implemented method of claim 1 or claim 2, wherein the simulation comprises computing an equilibrated structure.

4. The computer-implemented method of any of the preceding claims, wherein the reward is indicative of the binding affinity of the molecule in its current state to the target molecule.

5. The computer-implemented method of any of the preceding claims, wherein the reward comprises the binding free energy of the molecule.

6. The computer-implemented method of any of the preceding claims, wherein the reward is based on a proxy function, preferably based on normal mode analysis.

7. The computer-implemented method of any of the preceding claims, wherein the reward is generated using a prediction model configured to predict a metric indicative of the binding affinity of the molecule in its current state to the target molecule, based on the simulation.

8. The computer-implemented method of any of the preceding claims, wherein the change to the biomolecule sequence and / or structure comprises one or more of:addition of an amino acid or nucleotideremoval of an amino acid or nucleotidesubstitution of an amino acid or nucleotide..

9. The computer-implemented method of any of the preceding claims, wherein the simulation is a continuous simulation.

10. The computer-implemented method of any of the preceding claims, wherein the agent comprises an action network and a critic network, preferably wherein the action network and the critic network are implemented as a transformer architecture.

11. The computer-implemented method of any of the preceding claims, wherein the input to the agent includes the biomolecule sequence and geometric features of the molecule in its current state; and the output of the agent includes the location and change type of the biomolecule sequence.

12. The computer-implemented method of any of the preceding claims, wherein the agent comprises a trained machine learning model, wherein the machine learning model is trained to reconstruct existing molecule structures.

13. The computer-implemented method of any of the preceding claims, wherein the molecule(s) is a first protein having an amino acid sequence, and the target molecule is a second protein.

14. The computed-implemented method of claim 13, wherein the molecule is or comprises an antibody and the target molecule is or comprises an antigen and / or an epitope.

15. A method of creating a biologically active molecule, comprising identifying one or more molecules that are predicted to be likely to instigate a binding event with a target molecule, using a method of any of the preceding claims, andmanufacturing a biologically active molecule based on the identified molecule(s).