Fine tuning for adaptive multi-scale (ML / mm / CG) force field for large biological systems and antibody design

A multiscale model with varying simulation granularity levels optimizes computational load and accuracy for antibody design, addressing inefficiencies in current methods by using machine learning and classical force fields to simulate molecules accurately and efficiently.

WO2026158809A1PCT designated stage Publication Date: 2026-07-30NEC ONCOIMMUNITY AS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NEC ONCOIMMUNITY AS
Filing Date
2025-05-16
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Current methods for antibody design are time-intensive and costly, often prioritizing high-binding affinities that overlook functionally critical sites, and lack biological priors, leading to inefficient and inaccurate protein design.

Method used

A multiscale model with varying simulation granularity levels and an assignment model to optimize computational load while maintaining accuracy, using machine learning and classical force fields to simulate molecules, including antibodies, antigens, and epitopes.

Benefits of technology

The method reduces computational time and cost while ensuring high accuracy in simulating large biological systems, enabling precise targeting of specific epitopes and efficient antibody design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025063545_30072026_PF_FP_ABST
    Figure EP2025063545_30072026_PF_FP_ABST
Patent Text Reader

Abstract

A computer implemented method of training a molecular model for simulating a molecule, the molecular model comprising: a multiscale model for simulating properties of the molecule, the multiscale model having a plurality of simulation granularity levels, and an assignment model configured to assign each portion of the molecule to one of the simulation granularity levels, the method comprising: training the multiscale model and / or the assignment model based on a loss function representing computational complexity and / or accuracy of the molecular model. The present invention can be used in a variety of applications including, but not limited to, several anticipated use cases in drug development, material synthesis, and medical / healthcare.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] FINE TUNING FOR ADAPTIVE MULTI-SCALE (ML / MM / CG) FORCE FIELD FOR LARGE BIOLOGICAL SYSTEMS AND ANTIBODY DESIGN

[0002] FIELD OF THE INVENTION

[0003] The present invention relates to models for simulating molecules, and more specifically simulating molecules for antibody design.

[0004] BACKGROUND TO THE INVENTION

[0005] Antibodies (Abs) represent a rapidly expanding class of biological therapeutics in the human pharmaceutical market. Consequently, the need for optimal selection, design, and engineering of antibodies has not only grown significantly but has also emerged as a key competitive factor. Despite substantial interest from the pharmaceutical industry, the discovery of novel Abs remains a complex and costly process, often necessitating extensive laboratory screenings to ensure specific targeting. This screening process is inherently time-intensive and typically prioritizes antibodies with the highest binding affinities, which are usually associated with immunodominant epitopes. However, this approach can inadvertently overlook antibodies that, despite exhibiting lower affinities, target functionally critical sites. Furthermore, viral adaptation mechanisms may result in the elimination of these immunodominant epitopes, thereby potentially compromising antibody efficacy.

[0006] The computational design of Abs provides a promising alternative to mitigate these challenges by significantly reducing the time and costs associated with conventional laboratory-based screenings. This approach facilitates controlled screening of desired biophysical properties and enables precise targeting of specific epitopes of interest. Nonetheless, protein design for specific targets remains a complex endeavor, often yielding low success rates and requiring multiple rounds of affinity maturation experiments [DO - Fischman, S. & Ofran],

[0007] Challenges in protein design largely arise from the limitations of current energy functions, which are used to calculate free energies but often fail to account forfactors such as entropy and molecular flexibility, as they are based on static models that represent only a single conformational state. This is particularly relevant for Ab design, given that Abs typically feature extended, highly flexible loops known as complementarity-determining regions (CDRs). Given the marginal stability of proteins, with unfolding free energies generally ranging from 5-15 kcal / mol, accurate computation of these energies is essential for achieving precision in protein design.

[0008] Recent machine learning approaches have provided partial solutions, focusing either on inverse or forward design strategies. However, these methods face challenges in navigating the vast configuration space required for effective Ab design. Without prioritization mechanisms, these models necessitate prolonged computational times to evaluate potential configurations. Furthermore, current models often lack biological priors — language models, for instance, do not incorporate knowledge of underlying biological processes — resulting in limitations in obtaining and validating biologically feasible structures within this expansive design space.

[0009] We want to be able to scale the simulation of large system given a limited time, computation and memory budget but still get high quantum level accuracy where most matters.

[0010] SUMMARY OF THE INVENTION

[0011] According to an aspect of the present invention there is provided a computer implemented method of training a molecular model for simulating a molecule, the molecular model comprising: a multiscale model for simulating properties of the molecule, the multiscale model having a plurality of simulation granularity levels, and an assignment model configured to assign each portion of the molecule to one of the simulation granularity levels, the method comprising: training the multiscale model and / or the assignment model based on a loss function representing computational complexity and / or accuracy of the molecular model.The molecule comprises a plurality of separate portions. Each portion may be represented with a node. Each node may include a single atom, or a group of atoms. Each node may include a single edge or bond (e.g., between atoms or groups of atoms).

[0012] The term “training” may refer to adjustment or fine-tuning of an already-trained model. The multiscale model and / or the assignment model may be at least partially pre-trained. As used herein, the term “simulation granularity level” preferably refers to different ways to compute the energy and forces for each atom (or groups of atoms). The multiscale model may include a plurality of models (which may be referred to herein as simulation models). The simulation models may communicate with each other, such as to share information such as forces and energies, and other information such as constraints. The different simulation granularity levels may be provided by different simulation models. The different simulation granularity levels have different levels of computational complexity. Each level may have a different cost, such as a computation cost and / or an accuracy cost. Therefore, by training an assignment model that assigns different parts of a molecule to different simulation granularity levels, some parts of the molecule can be precisely simulated where required, with other parts being more coarsely represented. This enables the molecular model to represent key parts of a molecule accurately while minimizing computational load. In some examples, the assignment model may be partially or entirely based on heuristics (e.g., templates), rather than being trained. In such examples, the training may only involve the multiscale model and not the assignment model. Where the multiscale model is trained, the loss function may represent accuracy of the molecular model (and does not necessarily represent computation complexity). Where the assignment model is trained, the loss function may represent both computational complexity and accuracy of the molecular model. Where both the assignment model and the multiscale model are trained, a single loss function may be used (e.g., for joint training), or separate loss functions may be used.

[0013] Assigning different portions of the molecule to different levels (e.g., using different levels or simulation models for different parts of the molecule) may affect theoperation of the simulation models (e.g., compared to using only a single model or simulation granularity level for the whole molecule). Therefore, the multiscale model may be trained to maintain accuracy of the molecular model as a whole. In this way, division of the molecule into separate portions is optimized, as well as adjusting operation of the plurality of simulation granularity levels to accurately simulate properties of the molecule.

[0014] The accuracy of the molecular model may be represented using a prediction error. The loss function may be based on other parameters, such as an uncertainty of an output of the molecular model. The molecule may comprise, represent or model a protein. The molecule may comprise an antibody, an antigen, and / or an epitope. Alternatively or additionally, the molecular model may be used to simulate other structures, such as combinations of molecules (which may interact together).

[0015] The multiscale model may be jointly trained with the assignment model. In this way, the molecular model not only reduces the computation time and / or computation load, but also ensures that the simulation accuracy is maintained. Alternatively, the multiscale model may be trained (e.g., adjusted or fine-tuned) separately to adjustment of the assignment model.

[0016] The assignment model may be trained based on constraints (e.g., a maximum or bound on computation cost) and / or heuristics (e.g., considering a local neighbourhood of each atom or group of atoms in the molecule). The heuristics may help the assignment and may avoid large errors, such as by promoting assigning local nodes belonging to the same atomic system. For example, in a simple case, the assignment model may be implemented with heuristics such as distance, and / or number of atoms assigned to each model. The assignment model may have its own parameters. One way to parameterize is to assign a probability of assigning at each of the models based on the neighbours (e.g., graph-based prediction).The plurality of simulation granularity levels may be provided by: a machine learning model trained on a quantum mechanical level dataset or by query with a QM solver; a classical force field model for modelling atoms of a molecule; or a coarse-grained force field model modelling groups of two or more atoms of the molecule.

[0017] The plurality of simulation granularity levels are preferably different from each other. For example, the plurality of simulation models includes at least two models that are different to each other (e.g., having different computational complexity). The machine learning model may have a higher computational complexity than the classical force field model. The classical force field model may have a higher computation complexity than the coarse-grained force field model. The machine learning model may have a higher computational complexity than the coarsegrained force field model. Preferably, the plurality of simulation granularity levels include the machine learning model and the classical force field model.

[0018] The multiscale model may be at least partially pre-trained. As used herein, the term “training” may be used to refer to training from a randomly initialized or untrained model. The term training may be used to refer to further training or adjustment. The term “fine-tuning” preferably refers to adjustment starting from an already pre-trained model, such as on an existing dataset.

[0019] The multiscale model may include the coarse-grained force field model, and the method may include mapping at least one coarse-grained node to a set of individual atoms. The coarse-grained nodes may contain a plurality of atoms grouped together. The mapping (referred to herein as “backmapping”) provides a plurality of individual atoms based on properties of the course-grained node. The backmapping may include performing a first energy minimization around a current solution where other atoms of the molecule are fixed, and performing a second energy minimization or Langevin simulation to equilibrate the system. Where more than one coarse-grained node needs to be backmapped, these steps may be repeated several times until convergence to an equilibrium occurs. Alternatively, a generative model may be trained to perform the backmappingprocess (e.g., by proposing a backmapping atoms’ configuration fora given local configuration).

[0020] The multiscale model may be trained based on a dataset sample, a quantum mechanics (QM) solver, and / or a quantum mechanics I molecular mechanics (QM / MM) solver. The training may include active learning (AL) using the dataset sample, the QM solver and / or the QM / MM solver. For example, the multiscale model may be adjusted so that an output of the molecular model (e.g., a simulated property of the molecule) is consistent with a value determined based on the dataset sample, the QM solver and / or the QM / MM solver. The adjustment may be based on a loss function. This may enable the molecular model to provide an accurate output, with portions of the molecule assigned to different simulation granularity levels (thereby reducing computation load of the molecular model as a whole).

[0021] The method may further comprise outputting a property of the molecule using the molecular model. The molecular model may comprise a property prediction model. The property prediction model may output the property of the molecule. The property prediction model may be an equivariant model. The method may comprise training the property prediction model (e.g., the equivariant model) to predict a property of the molecule. The equivariant model may comprise a graph or transformer network.

[0022] The molecular model may include an acceleration model. The method may include training and / or adjusting the acceleration model to perform accelerated simulation of a molecule.

[0023] The training of the multiscale model and / or the assignment (e.g., adjustment) model may be based on a loss calculated between the outputted property of the molecule and validation data of the molecule. In this way, high level features or properties of the molecule are respected during training of the molecular model. The validation data may include experimental data. The validation data may include data from a QM / MM solver.The molecular model may include a generative model trained to generate a plurality of molecules for the training step. The method may include training the generative model. The generative model may generate a sequence and structure. Advantageously, the generative model may be used to provide further candidate molecules. The candidate molecules may be used for training. The candidate molecules may be assessed to determine their binding affinity to a target molecule, such as to identify antibodies for an antigen. The generative model may be a diffusion model. As used herein, the term “diffusion model” preferably refers to a model where, during training, samples from the dataset are progressively corrupted by gaussian noise and a model is trained to regenerate the original data, and at test time the model is used to remove the noise from a randomly generated sample to obtain a valid sample. Other generative models may be used other than diffusion models.

[0024] The generative model may be trained jointly with the multiscale model and / or the assignment model. In this way, the generative model can be fine-tuned such that it produces valid configurations of maximal uncertainty for the multi-scale system. Once the generative model is trained, it may be used to generate candidate molecules. The generative model may be trained to generate a variety of candidate molecules having predicted properties. For example, the generative model may be trained to produce candidate molecules having a particular binding affinity. Such candidates may be ranked to assess their viability as antibodies.

[0025] The method may further comprise estimating an uncertainty on the outputted property. The uncertainty may be estimated using gradient estimation.

[0026] According to another aspect of the present invention, there is provided a molecular model trained using the method of any preceding claim. Disclosed herein is a molecular model for simulating a molecule, the molecular model comprising a multiscale model for simulating properties of the molecule, the multiscale model having a plurality of simulation granularity levels, and an assignment model configured to assign each portion of the molecule to one of the simulationgranularity levels; wherein the multiscale model and / or the assignment model have been trained to minimize processing complexity while maximizing accuracy. The multiscale may be deployed together with the assignment model. One or more of the models of the molecular model may be used separately. For example, the generative model may be used separately to the assignment model and / or the multiscale model. For example, the generative model may be used together with the property prediction module. Alternatively, the multiscale model may be used separately and / or independently to the assignment model and vice versa.

[0027] According to another aspect of the present invention there is provided a method of selecting a candidate molecule that is predicted to bind to a target molecule, the method comprising: applying the molecular model as described above and herein to a plurality of candidate molecules to determine a property of each candidate molecule; and ranking the plurality of candidate molecules based on the determined property of each candidate molecule.

[0028] The property may be a binding affinity (e.g., binding free energy) between the candidate molecule and the target molecule. The property may include a binding location on the molecule. Since the molecular model is trained to produce both accurately simulate properties of molecules as well as reduce the computational load, the molecular model can be quickly applied to a large number of candidate molecules without substantial computation burden.

[0029] The candidate molecule and / or the target molecule may comprise a protein. In general, the protein may be a protein, protein-domain, or a protein sub-unit. The term “protein” may include any protein subsequence that may have a viable 3D or functional structure.

[0030] The candidate molecule may comprise an antibody (or part thereof), and / or the target molecule may comprise an antigen and / or epitope. In some examples the target antigen may comprise a neoantigen and / or neoepitope.According to another aspect of the present invention there is provided a method of creating a biologically active molecule, comprising: identifying a candidate molecule using the method as described above and herein; manufacturing a biologically active molecule based on the candidate molecule.

[0031] The biologically active molecule may be referred to as a biomolecule. The method may manufacture a pharmaceutical compound. The biologically active molecule may comprise antibodies. The biologically active molecule may comprise protein(s), nucleic acids, DNA, RNA, and / or mRNA. The biologically active molecule may be used as a medicine. The biologically active molecule may be used in a diagnostic assay. For example, a lateral flow test may be manufactured that includes the biologically active molecule.

[0032] Also described herein is a computer program product comprising instructions which, when executed by a computer, cause the computer to perform the method as described above and herein.

[0033] Also described herein is a system for simulating a molecule, comprising a processor in communication with at least one memory device, the at least one memory device have stored thereon instructions for causing the at least one processor to perform a method as described above and herein.

[0034] The invention may further comprise designing, by the methods as described herein, and subsequently producing, one or more antibodies that are predicted to bind to an epitope and / or a predicted epitope.

[0035] Antibodies that are predicted by the present invention to bind to an epitope and / or a predicted epitope will find use for a variety of industrial purposes, such as therapeutics and diagnostics, biotechnological applications, or research settings, as would be understood by the skilled person.

[0036] Accordingly, the present invention may find use in one of more of the following use cases:• Disease treatment or prevention, including that for infectious diseases or other diseases such as autoimmune or immune-related diseases and cancer, in particular immunotherapy or passive immunisation.

[0037] • Drug delivery in the context of antibody-drug conjugates.

[0038] • To assist in scientific research, through immunoprecipitation studies or immunohistochemistry studies.

[0039] • Diagnostics through disease detection assays, medical imaging, and biosensors.

[0040] In some embodiments, the method may further comprise generating a corresponding protein or peptide sequence for one or more antibodies that are predicted to bind to an epitope and / or a predicted epitope. In some embodiments, the method may further comprise generating a DNA or RNA sequence that encodes one or more antibodies that are predicted to bind to an epitope and / or a predicted epitope.

[0041] The antibodies as disclosed herein may be of therapeutic use, particularly in therapeutics and diagnostics. To administer the antibody to a subject, the protein or peptide may be delivered to an individual through a variety of routes, including but not limited to intravenous, intramuscular, intradermal, intraaural, intraarterial, intraocular, intravitreal, intranasal, subcutaneous, or inhalation.

[0042] An alternative way of ensuring the production of antibodies in a subject is by using various biotechnological techniques which involves using a nucleic acid sequence (i.e. DNA or RNA) which encodes a desired antibody. One such technique is the use of CRISPR, which can modify a subject’s DNA by inserting a DNA sequence into cells (i.e. B-cells) which encodes for the desired antibody, thereby facilitating a means by which the subject is able to produce the desired antibody from the transcription of this new genetic sequence and the subsequent translation of the RNA transcribed from it. Alternatively, DNA may be delivered as an exogenous molecule, typically as part of a larger vector construct with elements to facilitate its transcription to result in the production of the desired antibody. mRNA may also be delivered in a similar way, with its translation producing the desired antibody.Thus, in accordance with a further aspect of the present invention, there is provided a method of creating an antibody, comprising: identifying one or more antibodies that are predicted to be likely to instigate a binding event with a target molecule, such as an epitope, from any of the examples discussed above; and synthesising the antibody, or encoding the antibody, or predicted or simulated variants thereof, into a corresponding protein, peptide, DNAor RNA sequence.

[0043] Such a protein, peptide, DNA or RNA sequence may be delivered in a naked or encapsulated form or incorporated into a genome or cell of a bacterial or viral delivery system. In addition, bacterial vectors can be used to deliver the DNA in to vaccinated host cells.

[0044] In accordance with a further aspect of the present invention, there is provided a method of creating and / or designing a diagnostic assay to determine whether a patient has or has had a cancer or prior infection with a pathogen, wherein the diagnostic assay is carried out on a biological sample obtained from a subject, comprising identifying at least one antibody that is predicted to bind to an epitope and / or a predicted epitope which characterises a certain cancer or pathogen, using a method according to any of the examples discussed above; wherein the diagnostic assay comprises the utilisation within the biological sample of the antibody that is predicted to bind to an epitope and / or a predicted epitope which characterises a certain cancer or pathogen.

[0045] In this way, the present invention may advantageously be used to create a diagnostic test or assay, by means of a rapid and / or automated antibody discovery.

[0046] The term utilisation as used herein is intended to mean that the at least one antibody that is predicted to bind to an epitope and / or a predicted epitope is used in an assay to identify the presence of a particular epitope and / or predicted epitope in a subject.

[0047] The in vitro diagnostic assay may comprise identification of a biological component (i.e. target molecule) within the biological sample. In a preferred embodiment, the biological component may be an epitope or a predicted epitope.For example, the assay may comprise the identification of epitopes and / or predicted epitopes that may be bound by an antibody that is predicted to bind such epitopes.

[0048] As an example of such a diagnostic use, a sample, preferably a blood sample, isolated from a patient may be analysed for the presence of antibodies that are predicted to bind to an epitope and / or a predicted epitope within the biological sample that recognise and bind to such an epitope(s), the antibodies of which have been identified as part of the present invention and that are contained within the assay.

[0049] Suitable diagnostic assays would be appreciated by the skilled person, but may include enzyme-linked immune absorbent spot (ELISPOT) assays, enzyme-linked immunosorbent assays (ELISA), cytokine capture assays, intracellular staining assays, tetramer staining assays, microfluidic devices, lab-on-a-chip, microarrays, flow cytometry, CyTOF, or limiting dilution culture assays.

[0050] In a method of creating a diagnostic test, the amino acid sequence of the one or more antibodies that are predicted to bind to an epitope and / or a predicted epitope may be chosen based on the desired epitope and / or predicted epitope to be tested. For example, the one or more source antibodies (i.e. the initial binder molecule) may be one or more source antibodies which have binding affinity towards an epitope of any selected pathogen and / or virus (or fragments thereof). In such a case, the present invention may be used to create a diagnostic test for determining whether a patient has or has had prior infection with the virus, and / or its variants and / or related viral species. However, as will be appreciated by the skilled person, the one or more source antibodies may have binding affinity towards any epitope of any pathogen (e.g. any virus, parasite, bacterium, or cancer cell).

[0051] Further disclosed herein is a diagnostic assay to determine whether a patient has or has had prior infection with a pathogen and / or cancer, wherein the diagnostic assay is carried out on a biological sample obtained from a subject, and wherein the diagnostic assay comprises the utilisation of at least one antibodies that arepredicted to bind to an epitope and / or a predicted epitope, wherein the epitope is from a pathogen or tumour, using any of the methods as described herein. The diagnostic assay may comprise identification of a biological component within the biological sample that that serves as an epitope that may be bound by the antibody as described herein.

[0052] The invention may further comprise designing, by the methods as described herein, and subsequently producing one or more nucleic acids that are predicted to bind to a target molecule, such as a protein. Such nucleic acid molecules may comprise aptamers or mRNA.

[0053] Aptamers that are predicted by the present invention to bind to a target protein will find use for a variety of industrial purposes, such as therapeutics and diagnostics, biotechnological applications, or research settings, as would be understood by the skilled person.

[0054] Accordingly, the present invention may find use in one of more of the following use cases:

[0055] • Disease treatment or prevention, including that for infectious diseases or other diseases such as autoimmune or immune-related diseases and cancer, in particular for anti-cancer therapies and antiviral treatments. • Drug delivery in the context of aptamer-drug conjugates.

[0056] • Drug discovery when used to mimic ligands for the study of receptorligand interactions.

[0057] • Diagnostics through disease detection assays, medical imaging, and biosensors.

[0058] The potential implementation of the aptamers designed by the methods as described herein and subsequently produced will be similar to the foregoing embodiments described for antibodies, and therefore the skilled person would be able to adjust these methods to adapt them for use with aptamers.

[0059] The methods of the present invention may be used to design, and subsequently produce, RNA (including mRNA) and / or RNA binding proteins after predictingspecific RNA-protein interactions important for RNA processing, stability, functionality, and transport. Certain sequence and structural motifs comprised within RNA molecules can be recognised by RNA binding proteins, and therefore the design of RNA molecules and / or RNA binding proteins is critical to ensure their correct functioning in biological systems.

[0060] Thus, in accordance with a further aspect of the present invention, there is provided a method of creating a nucleic acid molecule, such as an aptamer, comprising: identifying one or more nucleic acid molecules, or aptamers, that are predicted to be likely to instigate a binding event with a target molecule, such as a protein, from any of the examples discussed above; and synthesising the nucleic acid molecule, or aptamer, or encoding the aptamer, or predicted or simulated variants thereof, into a corresponding DNAor RNA sequence.

[0061] BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Embodiments of the invention will be briefly described with reference to the drawings, in which:

[0063] Figure 1 illustrates a schematic diagram of an antibody-antigen binding;

[0064] Figure 2 illustrates an example of an antibody design workflow;

[0065] Figure 3 illustrates a de-novo design workflow for generating and ranking candidates;

[0066] Figure 4 illustrates a selection mechanism applied to a neighborhood of a node where each node is simulated at a different level;

[0067] Figure 5 illustrates an architecture of the adaptive multi-scale selection mechanism and the fine-tuning of the ML FF potential;

[0068] Figure 6 illustrates a workflow for extending the fine-tuning and including the (accelerated) simulation;Figure 7 illustrates an extension of the fine-tuning that includes a generative model;

[0069] Figure 8 illustrates a workflow for deployment of the model;

[0070] Figure 9 schematically illustrates a system suitable for implementing embodiments of the invention; and

[0071] Figure 10 schematically illustrates a server suitable for implementing embodiments of the invention.

[0072] DETAILED DESCRIPTION

[0073] Disclosed herein is a multi-scale method for modelling biological system and antibody design that uses an adaptive selection of granularity for the simulation of a molecular level biological system and the fine-tuning of machine learning potential for maximal accuracy. The method allows for prioritisation of antibodies candidates in complex environments with a predefined budget.

[0074] It is advantageous to design amino-acid sequences of antibodies that produce effective binding affinities (binding free energies) with the target antigen. Figure 1 shows a schematic diagram of an antibody-antigen binding. The antibody 101 has a binding area 102. The antigen 105 has a binding area 106 that corresponds to the binding area 102 of the antibody 101.

[0075] The space of the possible amino-acid sequences is very large (potentially infinite); thus, reinforcement learning may be used to address this problem.

[0076] An antibody workflow 200 is shown in Figure 2. Step 201 corresponds to antibody target selection based on a relevant biological mechanism. Step 202 corresponds to ML static epitope prediction. Step 203 corresponds to ML accelerated dynamic epitope prediction and evaluation. This step includes Molecular Dynamics (MD) simulation 203a. Step 204 corresponds to ML de-novo antibody protein design. This step includes antibody generation 204a. Step 205 corresponds to ML based candidate ranking. This step includes MD simulation 205a.We show the sequence of steps for the antibody design, where first we select the target, then predict the binding area of the antigen; then we add the environment (water and lipid layer) and simulate with Molecular Dynamics (MD) to explore the binding areas. The next step is to generate the antibody for the highest priority binding spots. We can then select the most promising and validate with simulation for each of the candidate.

[0077] Figure 3 shows a design workflow 300. The workflow takes a target 301 and environment 302 and generates (antibody) candidates 307. The process includes epitope prediction 303, epitope evaluation 304, de-novo design 305, and candidate ranking 306.

[0078] In Figure 3, the de-novo design workflow is composed of:

[0079] 1) from the target, predict the epitope / paratope (step 303). Subsequently epitope evaluation may occur (step 304);

[0080] 2) from the epitope, design the new protein / anti-body (step 305); and

[0081] 3) ranking multiple candidates based on a longer simulation (step 306).

[0082] Adaptive Multi-scale selection is used to select the level of simulation of a target system (steps 203 and 205 in Figure 2). While a QM solver may be considered as providing the most accurate output for a molecule, such solvers require a large amount of computational power. Therefore, it is only feasible to use (solely) a QM solver for small molecules. Particularly where a molecule contains more than about 1000 atoms, another approach is beneficial to maintain accuracy while reducing computation time. Disclosed herein are different levels of simulation of a molecular system:

[0083] 1) ML: machine learning potential or machine learning interatomic potential are machine learning models that model the energy and forces of a molecular system trained either on a quantum mechanical (QM) level dataset or by querying a QM solver. For example, a QM solver may be based on density functional theory (DFT). Approaches such as a coupledcluster (CC) technique may be used for higher accuracy, or semi-empirical methods may be used for increased computation speed;

[0084] 2) MM: classical force field (FF) are models whose functional formula and parameters are tabulated and are characterized by fast simulations, the FF typically model each atom of the system;

[0085] 3) CG: coarse grained force field (CG) model forces and energy after aggregating atoms together in beads. This potential results in a soother energy function and reduced number of nodes that allows to accelerate considerably the simulation time. Atoms and groups of atoms far from interface can be approximated by aggregating, thus, reducing the effective number of simulated particles.

[0086] By selecting for the best compromise between computation time and accuracy, larger systems can be simulated at a higher accuracy. Introducing an adaptive selection of granularity also affects the final accuracy, indeed the potentials are defined in the hypothesis that all other nodes are at the same level of accuracy. By introducing the adaptive mechanism, then the machine learning (ML) potential and in theory also the FF, need to be retrained or fine-tuned to achieve at least the same original accuracy level. To this end, disclosed herein is an adaptive selection of the granularity and re-training the ML FF to achieve the maximal performance at various budgets (see Figure 4 and Figure 5).

[0087] Figure 4 shows the selection mechanism applied to a neighborhood of a node, each node is simulated at different level; when a neighbor is selected, this model should have the maximal accuracy against the QM data or solver. In Figure 4, each node is represented as a circle or triangle, and bonds are illustrated as straight lines between nodes. Course grain (CG) nodes are depicted as triangles, Machine Learning (ML) nodes are depicted as solid circles, and Molecular Mechanics (MM) nodes are depicted as empty circles. The portion of the molecule in the dotted oval is a fine-tuning fragment. The fine-tuning of the ML model (and / or the MM / CG models) for this fragment may be performed during training against either a dataset sample or using AL using an external QM / MM solver.Since ML techniques are usually used for entire models, it cannot be expected that the ML model will automatically provide valid outputs when used in combination with other types of models (e.g., MM or CG). In other words, each individual model is not aware that other nodes are modelled differently, and thus the multiscale model as a whole should be fine-tuned to maintain accuracy when combining different simulation granularity levels. For example, an existing ML model may be used and fine-tuned during training of the overall system to ensure that its outputs continue to match validation data (e.g., from the QM data or a QM solver) even when the molecular system is divided into nodes. The term “fine tuning” is used herein to refer to training of a model that has already been at least partially trained (e.g., with other data).

[0088] The division of the molecule into nodes, and the assignment of each node to different models is carried out by an assignment model. Each node may be a single atom (such as for the ML and MM models). Alternatively, a node may represent a group of atoms (such as for the CG model). In some instances, only two simulation granularity levels are used, such as ML and MM. In such cases, the CG model is not required. However, in some instances the CG model may be used as well as the ML and MM models. CG is beneficial where there is a need to simulate very large systems that would require too much memory to handle with only ML and MM and / or which would take too long to simulate. However, since CG requires a greater level of approximation (i.e. , by grouping atoms together), it is generally preferable not to CG.

[0089] When including the CG, a mechanism, backmapping, is needed to map the coarse-grained nodes to the base atoms. Backmapping may be required where the specific configuration of atoms is required as an output of the molecular model; in situations where the specific configuration of the atoms is not required (e.g., where only properties of the molecule as a whole are desired), then backmapping is not essential.

[0090] Backmapping is used whenever the assignment of atoms into different nodes is adjusted so that a plurality of atoms previously grouped as a course-grained node are instead represented as separate nodes for each of the plurality of atoms (i.e.,so that the ML or MM model can be applied to each atom). The term “backmapping” can be considered to refer to reconstruction of the position of individual atoms from a single CG node. There are several ways that this may be achieved; for example, the backmapping may use heuristics (e.g., templates) and / or may be based on historical data. Alternatively, the atoms may be randomly (or semi-randomly) assigned to positions within the CG node. These techniques may provide an initial estimate for the position of each of the atoms within the CG node. However, the initial estimate does not typically correspond to an energy minimum, and further adjustment is usually required to provide the “real” positions of the atoms. The simplest way to address this is to do energy minimization around the current solution, where all the other atoms are kept fixed and then another short energy minimization or Langevin simulation to equilibrate the system is performed before proceeding with the normal simulation. This can be considered as running a local simulation. When multiple nodes go from CG to ML or MM, this operation needs to be repeated multiple times, until convergence to equilibrium. In an alternative, we can train a separate generative model that proposes a backmapping of the atoms’ configuration for a given local configuration. In generative implementations of backmapping, the current local configuration (or global) provides conditioning of a generative model (such as a diffusion model) and the model is trained to generate the “correct” atom coordinate.

[0091] Figure 5 shows an architecture 500 of the Adaptive multi-scale selection mechanism and the fine-tuning of the ML FF potential. The architecture 500 may be referred to as a “system”, or “molecular model”. This architecture 500 performs fine-tuning ML / MM and adaptive integration mapping. This allows faster and more accurate evaluation of protein / antibody candidates. The architecture 500 may use one or more pre-trained models 510, such as Force Field (FF) models (e.g., ML, MM, CG). A QM level dataset 511 (or datasets), a QM solver or a QM / MM solver 512 can be interfaced. In Figure 5 the loss is with respect to the QM dataset or the QM / MM solver.Module 501 is an assignment model 501, as discussed above. The assignment model 501 is trained to assign each atom in the molecule to a different simulation granularity level (e.g., one of the pre-trained models 510, such as ML, MM, orCG). In other words, the assignment model 501 is trained to decide how each atom in the molecule is modelled. Module 501 corresponds to adaptive ML / MM / CG integration (back mapping). Module 501 may be at least partially trained to perform this assignment, and / or may be based on heuristics, which means that the module 501 does not need to be trained end-to-end. Training of module 501 may include training the module 501 to perform backmapping; the input for backmapping training may include one or more of: a local environment of an atom (or each atom), coordinates of the atoms (optionally with respect to time), the energy of the atoms, and / or forces between atoms. The module 501 may be trained to provide an assignment including individual positions of atoms based on an earlier assignment including CG nodes.

[0092] Module 502 corresponds to a multiscale model 502 having a plurality of simulation granularity levels (such as ML / MM / CG). Training the multiscale model 502 corresponds to join ML / MM / CG (fine-tuning) training with Active learning (AL). In this step, one or more of the pre-trained models 510 may be fine-tuned. This fine-tuning may use data from the QM datasets 511 and / or a QM / MM solver 512. The fine-tuning ensures that the multiscale model continues to accurately simulate properties of the molecule when it is represented (e.g., simplified) as interconnected nodes. The multiscale model 502 takes a configuration of nodes (as assigned by assignment model 501) and calculates the forces between the nodes and / or the energy based on the assignment.

[0093] We can extend to include some global property or some property that depends on short simulation, see Figure 6. We can extend the fine-tuning and include the (accelerated) simulation such that the overall system performance with respect to high level features or properties is respected, for example coming from a QM / MM solver or from experiments. In other words, the architecture 600 (or “molecular model”) in Figure 6 performs joint local and global property fine-tuning training.Similarly to as discussed above in relation to Figure 5, the architecture 600 may use one or more pre-trained models 610, such as Force Field (FF) models (e.g., ML, MM, CG). A QM / MM solver 611 and / or QM datasets 612 may be used. Furthermore, experimental data 613 may be provided.

[0094] Module 601 is an assignment model 601 and may correspond to adaptive ML / MM / CG integration (back mapping). Module 602 is a multiscale model 602 and may correspond to joint ML / MM / CG (fine-tuning) training with active learning. Any features discussed above in relation to modules 501 and 502 may apply to modules 601 and 602.

[0095] Module 603 is an acceleration model 603, which may perform accelerated simulation. Based on the forces and / or energy determined by the multiscale model 602, the acceleration model 603 may adjust the position of each of the nodes of the molecule. The acceleration model 603 may calculate a property or parameter of the simulated molecule. The property or parameter may be calculated by integrating the forces and / or energy, and may be a local property of the molecule. Once the acceleration model is trained for a specific system, then the acceleration model 603 can be applied to similar systems to calculate corresponding properties with a decreased run-time. For example, the training may include training time steps of a Hamiltonian Monte Carlo algorithm (e.g., ST-HMC), or training Molecular Dynamics (MD) steps or masses of the atoms to be used during simulation. Module 604 is a loss calculation module 604, which is configured to calculate a loss. The loss may be with respect to the property from experimental data and uses the estimated gradient through the accelerated simulation. For example, the loss calculation module 604 may compare the calculated property or parameter to an actual value of the property or parameter in the experimental data 613 to calculate the loss and adjust the assignment model and / or the pre-trained models 610. The calculated property or parameter is typically a global property or parameter of the molecule as a whole. The property may be the binding free energy and / or an association time. This global property may be determined by integrating local properties (e.g., energy and forcesbetween nodes, such as calculated by the multiscale model 602) with respect to time and / or space.

[0096] Multi-system generalization

[0097] We can finally also extend to the case of fine-tuning for a large system class. In this case a generative model is used, and the generative model can be trained jointly to the fine-tuning of the ML FF, as shown in Figure 7.

[0098] Figure 7 shows an extension of the fine-tuning that includes a generative model. This figure is a schematic diagram of an architecture 700 (or “molecular model” 700) indicating the relevant modules, and any inputs and outputs that are relevant to each module. The generative model produces a system where the ML-MM-CG needs to be fine-tuned, but at the same time, the generative model can be finetuned such that it produces a valid configuration of maximal uncertainty for the multi-scale system.

[0099] As shown in Figure 7, the architecture 700 may use one or more pre-trained models 710, such as classical force field models (e.g., ML, MM, CG). The architecture may use one or more QM datasets 711. The architecture may use a QM / MM solver 712. In this workflow, a pre-trained generative model 715 may be used. A global property and pre-trained model 716 may be used. The global property and pre-trained model 716 may provide labels for supervised training of parameters of other models. For example, the model 716 may provide global properties such as binding free energy.

[0100] Module 701 is a generative module 701. The generative module 701 is configured to generate a molecular system such as a molecule. This may include generating both sequence and structure of a molecular system. The generation may be performed (at least in part) using the pre-trained generative model 715. The generative model 715 may be a diffusion model with guided loss. As discussed below, in addition to use of the pre-trained generative model 715, this generation may include guided generation 722, where the output is guided by a property prediction model to provide more promising candidate molecular systems. Theguided generation 722 may use a diffusion model with guided loss. In order to generate more promising novel structures, the structures need to be different to those already known (e.g., within training data), but not so different as to be irrelevant. As well as generating structures, the generative module 701 provides a confidence or uncertainty associated with the generated structure. In other words, the module 701 can indicate the types of structure that it is not good at predicting. Therefore, generation can be guided by maximising this uncertainty, which increases the likelihood of generating novel molecular structures (i.e., which are not present in the training data).

[0101] Module 702 is an assignment model 702 and may correspond to adaptive ML / MM / CG integration (back mapping). This module 702 may correspond to modules 501 and 601 discussed previously. This module 702 may use a ML / MM / CG assignment network 721 with heuristic and loss. This module 702 also uses the pre-trained models 710. The nodes or edges are assigned to a specific model (ML, MM, CG) dynamically based on heuristics (e.g., local neighborhood) and minimization of loss that represents computation complexity. This assignment network, once trained, provides an assignment model 731 which may be considered an output 730 of the architecture 700.

[0102] Module 703 is a multiscale model 703 and may correspond to joint ML / MM / CG (fine-tuning) training with active learning. This module 703 may correspond to modules 502 and 602 discussed above. This module 703 may use the QM datasets 711 and / or the QM / MM solver712. Following training (fine-tuning) of this module, fine-tuned ML / MM / CG models 732 are produced, which can be considered an output 730 of the architecture 700.

[0103] Module 704 is an acceleration model 704, which may perform accelerated simulation. This module 704 may correspond to module 603 described above. The acceleration model 704 adjusts the position of each of the nodes of the molecule based on the forces and / or energy calculated by the multiscale model 703. The acceleration model 704 may calculate a property or parameter of the simulated molecular system. Gradient estimation 723 may be used for a “long” sequence to estimate the uncertainty on the property or parameter beingcalculated. Following the training during this step, an acceleration model 733 may be produced, which may be considered an output 730 of the architecture 700.

[0104] Module 705 is a property prediction model 705. The property prediction model 705 comprises a regression equivariant network and / or a message passing based neural network. This module 705 may use the global property and pre-trained model 716 and the output of the accelerated simulation step 704. This module 705 predicts properties of the static molecular system (without requiring any simulation). However, the module 705 is less effective at generalizing to samples outside of distribution, so where this is required, the acceleration model 704 may be used to calculate properties or parameters using a more in-depth simulation. Therefore, properties of the molecule can be predicted using the property prediction model 705 and / or the acceleration model 704. An output 730 of the architecture 700 may be property prediction 734 (e.g., of the molecular system).

[0105] Module 706 is a loss calculation module 705, which is configured to calculate a loss. The loss is calculated based on uncertainty and expected improvement. This module 705 may correspond to the module 604 described previously. This loss may be used for the guided generation 722, so that more relevant molecular systems can be subsequently generated in step 701. The guided generation may use the generative model 701 and the property prediction model (regression equivariant network) 705. The former generates physical structures, while the latter guides this generation towards a specific property. In this example, these models are combined by performing a sum of gradients from the two models to update the generated sample. As a result, the output of the generated model can be moved closer towards producing a desired property.

[0106] While the description above relates to different architectures 500, 600, 700, it will be appreciated that several training steps and procedures are also used in order to provide the final system. For example, disclosed herein is a computer implemented method of training a molecular model 500, 600, 700, the molecular model comprising: a multiscale model 502, 602, 703 for simulating properties of the molecule, the multiscale model having a plurality of simulation granularity levels, and an assignment model 501 , 601 , 702 configured to assign each portionof the molecule to one of the simulation granularity levels, the method comprising: training the multiscale model and / or the assignment model based on a loss function representing computational complexity and / or accuracy of the molecular model. In some instances, the assignment model need not be trained and can instead be programmed based on heuristics or templates. Where both the assignment model and the multiscale model are trained, they may be jointly trained. Where the multiscale model is trained, the loss function represents accuracy of the molecular model. Where the assignment model is trained, the loss function represents accuracy of the molecular model and computational complexity. The steps of the method may be iterated as required such as during training.

[0107] Deployment

[0108] The description above generally relates to molecular models 500, 600, 700 during training. However, not all the modules discussed above need to be present during deployment. At deployment time, a different output can be used. At the basic level the adaptive assignment module and the connected potentials (ML / MM / CG) are used to perform the simulation. For example, a trained molecular model may include only the assignment model 501, 601, 702 and the multiscale model 502, 602, 703.

[0109] As shown in Figure 8a, at test time the fine-tuned model may be used for ranking the candidates. More specifically, a first example of a deployed system 800a may include an adaptive assignment ML / MM module 801, which may correspond to the assignment model (e.g., output 731 in Figure 7). This module 801 may take any input molecule or molecular structure and assign portions (e.g., nodes) of the molecule to different simulation granularity levels (e.g., ML, MM, or CG). The system 800a may include a fine-tuned ML / MM / CG module 802 (e.g., corresponding to output 732 in Figure 7). This module 802 may be referred to as a multiscale model, since it is able to simulate portions of the molecule with different granularity levels (depending on the assignment by the assignment model). The system 800a includes an acceleration model 803, (e.g., corresponding to the acceleration model 733 in Figure 7). This may enable aproperty 804 to be calculated for the molecule. Where several molecules are input to the system 800a, they may be ranked based on the calculated properties.

[0110] As shown in Figure 8b, we can also use the generative model for the generation of new structures. In this example of a deployed system 800b, a generative model 810 is included (e.g., as described previously in relation to Figure 7). Then, based on a generated structure, a property prediction module 811 (e.g., corresponding to property prediction 734 in Figure 7) may be applied. This may lead to a structure / sequence candidate 812. These candidates may be tested and ranked, such as using system 800a.

[0111] Further embodiments of the invention may include the following:

[0112] Embodiments for Antibody, Protein or mRNA design by Nucleotide and Aminoacid modelling (ML / MM): Since antibodies are proteins, the same method can be applied to generic protein design, but also for mRNA design.

[0113] Embodiments for Antibody design (includes protein-protein design) (ranking of atomic system) ML / MM: The adaptive multi-scale mechanism can be used to rank candidates for Antibody design.

[0114] Complex environment modelling (membrane / multi-system / non protein elements) needs ML / MM / CG (complex interactions): When considering complex scenarios, such as the presence of other tissues or membrane, then CG is required because of the size of the system.

[0115] Some embodiments of the invention include a method comprising the following steps:

[0116] 1) System setup: Specify the class of systems depending on the application;

[0117] add constraints, heuristics and computation budget for the assignment function; select pre-trained models; define QM / MM solver and / or QM dataset, (see Figure 5.)a. (optional:) the architecture of the regression / classification model; the accelerated simulation (Figure 6.) and its parameters; select the pre-trained generative model (Figure 7.) (e.g. diffusion model) 2) Run the fine-tuning of the multi scale model (ML / MM / CG) with the assignment function

[0118] a. (optional:) train the property prediction using an equivariant model;

[0119] train a generative model (as a diffusion model) (Figure 7.); guided by the property prediction model to generate system with maximal uncertainty and expected improvement of the property (to generate uncertainty systems with interesting properties) (Figure 7.); optimally training the acceleration model to accelerate the simulation (Figure 6. and 7.)

[0120] b. (optional:) jointly train the assignment function during fine tuning 3) Use the fine-tuned multi-scale models (output of step 2) for ranking candidates with the adaptive selection multi-scale model (Figure 5. and Figure 8.)

[0121] a. (optional:) using accelerate simulation (Figure 6. and 7.); use generative models guided by property prediction model to generate on specific property (Figure 7.)

[0122] In some embodiments, we distinguish between fine-tuning and training, where training starts from un-trained models (randomly initialized) and fine-tuning starts from an already pre-trained model on existing dataset.

[0123] Assignment function or adaptive assignment is the function that assigns each atom (node) to a specific model (MM, ML, CG).

[0124] Diffusion models are one of the existing generative models, where during training, samples from the dataset are progressively corrupted by gaussian noise and a model is trained to regenerate the original data, and at test time the model is used to remove the noise from a randomly generated sample to obtain a valid sample.

[0125] Further features of embodiments of the invention and their advantages are set out as follows:Adaptive Assignment network for (ML / MM / CG and) fine-tuning ML / MM / CG (Figure 7, point 1.)

[0126] • A neural network (or heuristics) adaptively assigns nodes (or edges / bonds) to a specific level (multi-scale simulation) considering a loss function to minimize properties including computation cost (for example as a bound) and the prediction error or uncertainty. The assignment contains heuristics to help the assignment and to avoid high error, for example by promoting assigning local nodes belonging to the same atomic system (for binding problems). We can also improve the assignment by propagation through the simulation, by estimating the gradient thus changing from exploration to optimization / exploitation.

[0127] • Advantage: optimal assignment for a new system that minimizes the expected error / uncertainty and with bounded computation cost

[0128] Guided generation of uncertainty and relevant (has the property of interest, e.g. binding energy) structures; the generation is guided by a property model (Figure 7, point 2.)

[0129] • The objective is to train a generative model which generates new structures to train the multi-scale model (Generative model sequence + structure)

[0130] • Use loss based on equivariant loss, which includes the property prediction. The property prediction uses an equivariant (graph or transformer) network

[0131] • Advantage: fast exploration and guarantee on the error / uncertainty Compute gradient estimation to evaluate the uncertainty & error on the original ML model (Figure 7, point 3.)

[0132] • Fine tuning ML / MM / CG driven by uncertainty on the property (3a);

[0133] gradient is computed with respect the gradient estimated from the simulation but could also be computed from the property model • Accelerate simulation based on global property: parameters adapt to the specific system (3b)o Run simulation: compute the correlation function on the full simulation and then estimate a gradient of the acceleration function’s parameters that reduces the correlation among samples

[0134] • Advantage: does not require propagating through time, only simulation points

[0135] An advantage of the invention in contrast to the current state of the art is the ability to scale simulation to very large systems and adapt to computational cost and accurately modelling the most important part of the system.

[0136] Applications

[0137] X. Antibody applications

[0138] The optimal selection, design, and engineering of antibodies tailored to specific purposes (i.e. epitopes) requiring immunisation or natural sources leads to major improvements in therapeutics, diagnostics, drug design, research and development, and cost. They are of particular importance in the fields of infectious disease and cancer. Examples are provided below.

[0139] X.1 Therapeutic antibodies

[0140] Antibodies are the key effector molecules of the humoral immune response, produced by plasma B-cells and serving the purpose of identifying and binding to foreign antigens to assist in eliminating and preventing the spread of pathogens in the body. The effector functions of antibodies are principally neutralization (i.e. arresting a biological effect), opsonization (i.e. marking the target for phagocytosis), and antibody-dependent cellular cytotoxicity (i.e. attracting effector cells to eliminate the threat).

[0141] Computational in silico and machine learning approaches have greatly advanced our understanding of antibody binding and has improved the pipeline to produce desired antibody candidates which have specificity for a target epitope and exhibit high binding energy.The binding event of an antibody to an antigen / epitope is a key event in activating the humoral B Cell immune response against any pathogenic threat, such as viral and bacterial infections, in addition to cancer antigens in malignant tumour cells. The production of antibodies towards target epitopes with good binding properties is essential for continuing therapeutic antibody development. The production of antibodies that are predicted to bind to an epitope and / or a predicted epitope, as done by the methods as described herein not only provides an enhanced functional understanding of the critical residues involved antibody-antigen binding; it also aids in the selection of therapeutically relevant epitopes, particularly those which are comprised within functionally critical sites but are neglected by the natural immune system due to immunodominance.

[0142] The antibodies produced by the methods as described herein may be produced with the intention of being therapeutic towards any one of a plethora of potential diseases. Such indications include, for example, hematologic, oncologic, immunologic, neurologic, musculoskeletal and pulmonary diseases.

[0143] The methods as described herein may also be suitable for the generation of chimeric antigen receptors (CARs). CARs are receptor proteins that are engineered into T cells to give them the ability to target antigens. This is done by fusing an single-chain variable fragment (scFv) domain (which is derived from the variable region of an antibody and has affinity towards an antigen / epitope) to an intracellular signalling module which can induce T cell activation upon antigen / epitope binding. This allows T cells to directly detect antigens and become activated, rather than requiring an initial processing and presentation of antigens / epitopes by MHC molecules for T cells to recognise and bind to. As such, they are particularly useful in detecting cell surface antigens, such as those expressed on cancer cells which represent therapeutic targets, such as CD19 and BCMAfor B-cell cancers.

[0144] X.2 Diagnostics

[0145] Antibodies designed by the methods as described herein may prove useful in developing diagnostic tests towards target epitopes. For example, antibodies maybe generated that exhibit higher affinity towards a target epitope than naturally occurring alternatives, thus increasing assay sensitivity and lessening the about of biomarker required to produce a positive test result. In practice, this may result in the detection of certain indications at an earlier stage when compared to using antibodies with lower affinities (i.e. a less sensitive assay). This could, for example, involve the detection of early cancer markers.

[0146] Since the production of novel antibodies is a complex and costly process, often necessitating extensive laboratory screenings to ensure specific targeting, computational design of antibodies provides a promising alternative to mitigate these challenges by significantly reducing the time and costs associated with conventional laboratory-based screenings. This factor, combined with the ability to develop antibodies with high affinities towards target epitopes, will expedite and improve the development of novel biosensors and diagnostic platforms like lateral flow assays or ELISAs which may be required in the event of disease outbreak or pandemic.

[0147] X.3 Research and development and drug design

[0148] Various research applications may be improved by the antibodies produced by the methods of the present invention. The improved specificity and affinity of such antibodies will improve techniques such as immunoprecipitation for the isolation of proteins are isolated from mixtures, and immunohistochemistry using fluorescently or enzymatically labelled antibodies for observing the localisation of proteins in cells or tissues.

[0149] Additionally, the blocking or neutralizing of target proteins by specifically engineered antibodies can be used to study the functions of select proteins in vitro or in vivo.

[0150] Antibodies designed by the methods as described herein will also have use in drug discovery and design, due to the potential of such antibodies to be developed towards specific epitopes which result in protein activation or inhibition, therebyconfirm specific proteins’ roles in diseases, which will subsequently expedite drug discovery pipelines.

[0151] Y. Other protein applications

[0152] Since antibodies are proteins, the methods as described herein are also applicable to generic protein design. For example, how a protein interacts with other molecules, such as small molecules, ligands, DNA, RNA, or other proteins, may be improved or optimised.

[0153] For example, protein-ligand binding concerns the binding of a selected protein to a ligand, such as a small molecule or biomolecule, with high affinity and selectivity. These studies are relevant in the field of drug discovery. Another application would be for protein-protein interactions, which can result in the development of therapies intended to disrupt pathological protein-protein interactions, for example in cancer or neurodegenerative diseases. A further application would be found in protein-nucleic acid binding studies. As an example, ribosome binding proteins (RBPs) can bind RNA targets through the interaction between their amino acid residues and the nucleotides of the RNA. These binding proteins serve to regulate various activities and properties of RNA transcripts, such as splicing, stability, localization, and translation. RBPs may be designed to enhance the binding properties to these RNA motifs and structures, such as the 5’ cap, 3’ UTR, internal ribosome entry sites (IRES), or stem-loop and secondary / tertiary structures. The optimisation of affinity of the RNA-binding protein for the target mRNAcan be done by adjusting the binding interface, which would then impart its effect onto the target RNA.

[0154] Z. Nucleotide applications

[0155] Aptamers are single-stranded nucleic acid molecules that bind to target molecules with high specificity and affinity and comprise either DNA or RNA nucleotides. Like antibodies, the effective binding to their target molecule exhibited by aptamers is due to their folding into a 3D structure specific to the target molecule. Their target molecules may include peptides, proteins, organic compounds, small molecules,metal ions. Additionally, they can be made to target biological entities such as mammalian cells, viruses, bacteria, and yeast.

[0156] Due to their non-immunogenic properties, and additionally being easier to synthesise and modify, aptamers exhibit several notable advantages. As such, they have widespread use in industry and represent a promising therapeutic modality to treat various diseases and conditions.

[0157] Due to being a molecule which has binding properties, they share many applications with antibodies, from therapeutics and drug delivery to diagnostics and medical imaging.

[0158] Messenger RNA (mRNA) is a single stranded RNA molecule transcribed from a gene that is read by the ribosome to synthesise a protein. Upon transcription from the gene, mRNA undergoes various processing reactions remove non-translated elements (introns) and protected from degradation. The processed mRNA molecule is then translated into a protein through being read by a ribosome which catalyses the polymerisation of amino acids corresponding to the encoded sequence of the mRNA. mRNAs comprise specific sequence motifs or secondary structures, such as the 5’ cap, 3’ untranslated region, stem-loop structures, and internal ribosome entry sites (IRES). RBPs can bind RNA targets through the interaction between their amino acid residues and the nucleotides of the RNA. These binding proteins serve to regulate various activities and properties of RNA transcripts, such as splicing, stability, localization, and translation. As such, the methods as described herein may be used to design, and subsequently produce, RNAs, including mRNAs, which optimally interact with RBPs.

[0159] Computational System

[0160] Figure 9 schematically illustrates an example of a system suitable for implementing embodiments of the invention. The system comprises at least one server 910 which is in in communication with data store 900, such as a database. The server is also in communication with a client device 930 over a communications network 920 such the internet. The client device may be adesktop computer or mobile device for example. In use, a request to perform the methods of any of the embodiments of the invention may be input by a user to the client device 930. The client device may send the request to the server 910 over the communications network 920. The server 910 implements embodiments of the invention as described herein, and may retrieve data, such as amino acid sequences, from the data store 900. The resulting (e.g., candidate) molecular structure(s) that are simulated or predicted by performing the present invention may then be communicated to the client device 930 over network 920.

[0161] An example of a suitable server 900 for implementing embodiments of the invention is shown in Figure 10. In this example, the server includes at least one microprocessor 901 , a memory 902, and an external interface 904, interconnected via a bus 905 as shown. In this example the external interface 904 can be utilised for connecting the server 900 to peripheral devices, such as the communications networks, data stores, other storage devices, or the like. The server may also optionally include an input / output device 903, such as a keyboard and / or display,

[0162] In use, the microprocessor 901 executes instructions in the form of applications software stored in the memory 902 to allow the required processes to be performed, including simulation of molecules using the molecular model described herein. The applications software may include one or more software modules, and may be executed in a suitable execution environment, such as an operating system environment, or the like. Accordingly, it will be appreciated that the server 900 may be formed from any suitable processing system, such as a suitably programmed client device, PC, web server, network server, or the like. Whilst the term server is used, this is for the purpose of example only and is not intended to be limiting.

[0163] Whilst the server 900 is a shown as a single entity, this is not essential and it will be appreciated that the server 900 can be distributed over a number of geographically separate locations, for example as part of a cloud-based environment.References

[0164] References useful for understanding the invention are listed below, all of which are incorporated by reference herein in their entirety.

[0165] [0] DO - Fischman, S. & Ofran, Y. Computational design of antibodies. Current Opinion in Structural Biology 51 , 156-162 (2018).

[0166] [1] D1 - Pultar, Felix, Moritz Thuerlemann, Igor Gordiy, Eva Doloszeski, and Sereina Riniker. "Neural Network Potential with Multi-Resolution Approach Enables Accurate Prediction of Reaction Free Energies in Solution." arXiv preprint arXiv:2411.19728 (2024).

[0167] [2] D2 - Wang, Yuanqing, Kenichiro Takaba, Michael S. Chen, Marcus Wieder, Yuzhi Xu, Tong Zhu, John ZH Zhang et al. "On the design space between molecular mechanics and machine learning force fields." arXiv preprint arXiv:2409.01931 (2024).

[0168] [3] D3 - Watson, J. L. et al. De novo design of protein structure and function with RFdiffusion. Nature 620, 1089-1100 (2023).

[0169] [4] D4 - Bennett, N. R. et al. Atomically accurate de novo design of single-domain antibodies. 2024.03.14.585103 Preprint at https: / / doi.Org / 10.1101 / 2024.03.14.585103 (2024).

Claims

36CLAIMS1. A computer implemented method of training a molecular model for simulating a molecule, the molecular model comprising:a multiscale model for simulating properties of the molecule, the multiscale model having a plurality of simulation granularity levels, andan assignment model configured to assign each portion of the molecule to one of the simulation granularity levels, the method comprising:training the multiscale model and / or the assignment model based on a loss function representing computational complexity and / or accuracy of the molecular model.

2. The method of claim 1 , wherein the multiscale model is jointly trained with the assignment model.

3. The method of claim 1 or claim 2, wherein the plurality of simulation granularity levels are provided by:a machine learning model trained on a quantum mechanical level dataset or by query with a QM solver;a classical force field model for modelling atoms of a molecule; or a coarse-grained force field model modelling groups of two or more atoms of the molecule.

4. The method of any preceding claim, wherein the multiscale model includes the coarse-grained force field model, and the method includes mapping at least one coarse-grained node to a set of individual atoms.

5. The method of any preceding claim, wherein the multiscale model is trained based on a dataset sample, a quantum mechanics (QM) solver, and / or a quantum mechanics I molecular mechanics (QM / MM) solver.

6. The method of any preceding claim, further comprising outputting a property of the molecule using the molecular model.

377. The method of claim 6, wherein the training of the multiscale model and / or the assignment model is based on a loss calculated between the outputted property of the molecule and validation data of the molecule.

8. The method of any preceding claim, wherein the molecular model includes a generative model trained to generate a plurality of molecules for the training step.

9. The method of claim 8, wherein the generative model is trained jointly with the multiscale model and / or the assignment model.

10. The method of any of claims 6 to 9, further comprising estimating an uncertainty on the outputted property.

11. A molecular model trained using the method of any preceding claim.

12. A method of selecting a candidate molecule that is predicted to bind to a target molecule, the method comprising:applying the molecular model of claim 11 to a plurality of candidate molecules to determine a property of each candidate molecule; andranking the plurality of candidate molecules based on the determined property of each candidate molecule.

13. The method of claim 12, wherein the candidate molecule and / or the target molecule comprise a protein.

14. The method of claim 12 or 13, wherein the candidate molecule comprises an antibody, and / or the target molecule comprises an antigen and / or epitope.

15. A method of creating a biologically active molecule, comprising:identifying a candidate molecule using the method of any of claims 12manufacturing a biologically active molecule based on the candidate molecule.