Drug design method using active learning over synthon space

The method of active learning over synthon space with surrogate models efficiently identifies high-scoring molecules for drug discovery, addressing inefficiencies in current techniques by reducing synthesis and testing requirements, thus optimizing the drug discovery process.

WO2025229023A1PCT designated stage Publication Date: 2025-11-06RECURSION PHARMACEUTICALS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/061768
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-07
Filing Date
2025-04-29
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

Existing drug discovery methods are inefficient and costly, often requiring the synthesis and testing of thousands of compounds to identify optimal candidates, with current computational techniques struggling to effectively explore and score the vast chemical space needed for identifying molecules with desired properties.

Method used

A method utilizing active learning over synthon space, involving the use of surrogate models to select and score a reduced subset of synthons based on structural features, reducing the number of molecules to be synthesized and tested while optimizing for desired properties, using machine learning models to predict molecule scores and iteratively refine the selection process.

Benefits of technology

This approach significantly reduces the number of compounds needed for synthesis and testing, enhancing the efficiency and effectiveness of drug discovery by identifying high-scoring molecules with desired properties, thereby accelerating the process and minimizing resource expenditure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025061768_06112025_PF_FP_ABST
    Figure EP2025061768_06112025_PF_FP_ABST
Patent Text Reader

Abstract

The invention provides an active learning method for drug design that involves training surrogate models on synthon space – rather than molecule space or product space – to predict synthons that have a high probability of scoring favourably for a given objective function. Importantly, this means that combinatorial space is reduced by orders of magnitude prior to enumeration of molecules. In turn, this means that the total number of molecules that need to be scored via evaluation of an expensive scoring function is decreased drastically. The method achieves these benefits while also being then able to thereby identify optimised molecules – using the optimally identified synthons – in an enumerated molecule space.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DRUG DESIGN METHOD USING ACTIVE LEARNING OVER SYNTHON SPACE

[0002] TECHNICAL FIELD

[0003] The invention relates to a method for drug design and, in particular, to a method for drug design using active learning over synthon space.

[0004] BACKGROUND

[0005] Drug discovery is the process of identifying candidate compounds I molecules for progression to the next stage of drug development, e.g. pre-clinical trials. Such candidate compounds are required to satisfy certain criteria for further development. Modern drug discovery involves the identification and optimisation of initial screening ‘hit’ compounds. In particular, such compounds need to be optimised relative to required criteria, which can include the optimisation of a number of different properties. The properties to be optimised can include, for instance: efficacy / potency against a desired target; selectivity against nondesired targets; low probability of toxicity; and, good drug metabolism and pharmacokinetic properties (ADME). Only compounds satisfying the specified requirements become candidate compounds that can continue to the drug development process.

[0006] The drug discovery process can involve making / synthesising a significant number of compounds during the optimisation from initial screening hits to candidate compounds. In particular, those compounds which are synthesised are measured to determine their properties, such as biological activity. However, the number of compounds that could be made as part of a particular drug discovery project will far outnumber - likely by orders of magnitude - the number of compounds that can be synthesised and tested. The results of the measurements of synthesised compounds are therefore analysed and used to inform a decision on which compounds to synthesise next to maximise the likelihood of obtaining compounds with further improved properties relative to the various criteria required by a candidate compound.

[0007] The synthesis and subsequent measurement of biological activity of one or more compounds at a particular stage is referred to as a design cycle (or iteration) of the drug discovery process. Typically, a set of compounds is synthesised and tested at each design cycle of the process as this is more efficient than synthesising and testing a single compound at a time. However, a level of available resources usually means that there is an upper limit on the number of compounds in a set that can be synthesised at any given design cycle.

[0008] During a drug discovery project, many hundreds or even thousands of compounds typically are synthesised across several design cycles before a candidate compound is found. This is a lengthy, expensive, and inefficient process: synthesis of a single compound can cost thousands of pounds, and it can take three to five years on average to obtain a single candidate compound.

[0009] The identification of initial screening hit compounds may be performed in different ways. For instance, high-throughput screening is a technique for conducting a relatively large number of automated biological or chemical tests in a relatively short period of time to identify relevant compounds, e.g. compounds that are active against a defined biological target. This allows large compound libraries to be screened relatively quickly, but is limited to compounds included in such libraries.

[0010] Another technique for identifying initial screening hit compounds is virtual screening. This is a computational technique to search libraries of compounds to identify those structures which are most likely to exhibit desired properties. This technique may be useful when the virtual compounds are readily available or easy to synthesise. However, only a small fraction of drug-like chemical space may be searched according to such techniques, even with up-to-date computational resources.

[0011] A computational technique for identifying and optimising hit compounds is de novo design. This is a technique for exploring chemical space using search or optimisation procedures to explicitly generate compounds. This can involve the incremental construction of compounds from scratch, for instance by sequentially selecting and adding fragments to be included in the designed compound.

[0012] Fragment-based drug discovery is an established method to discover starting points to create drug-like molecules that can inhibit or activate the function of a protein I target of interest. However, as a result of the inherently small size of such fragments, they make few interactions to the target protein, which results in low binding affinity. They also have low specificity, leading to significant off-target interactions. As a consequence, significant medicinal chemistry effort is required to produce higher affinity and more selective lead compounds.

[0013] One strategy to develop a lead compound from a fragment screen is elaboration, which involves adding functional groups to a fragment vector to increase favourable interactions with the protein I target and to create a more drug-like molecule. Typically, the fragment substructure remains present, and growth occurs from vectors on the fragment.

[0014] How specific vectors are chosen, and which functional groups are added, varies based on the elaboration method. Vector elaborations - sometimes referred to as R-group elaborations - in in silico generative methods generally allow for growth at any position on a fragment. The functional groups that are added to those vectors relies on what chemical transformations a generative method has access to.

[0015] However, synthesizability is a common issue with generative methods given that many do not have an innate awareness of building blocks or chemical reactions that are available. Reaction-based enumeration - a synthetically-aware elaboration method - grows from vector positions by identifying available chemistry that can be performed on the fragment (or a similar building block, usually one that contains the fragment substructure) and retrieves the complementary building blocks from vendor catalogues or in-house compound databases for that chemistry. The building block can be enumerated onto the fragment to create molecules that are the product of the chemistry. Thus, synthesizability is embedded into the elaboration as chemistry and functional group additions (building blocks) are the means for elaboration. Large combinatorial, enumerated libraries like Pfizer’s global virtual library (PGVL), Eli Lily’s Proximal Collection (PLC), Merck’s Accessible Inventory (MASSIV), and Enamine’s “Readily Accessible" (REAL) Database use reaction-enumeration methods for the creation of their de novo chemical spaces. Reaction-based enumeration involves a transformation that is a computational I virtual form of a chemical reaction that is performed on a fragment I intermediate molecule to obtain a product molecule.

[0016] The number of molecules generated from a single vector reaction -based enumeration increases linearly with the number of building blocks that are compatible with that vector position. However, elaboration from multiple vectors is oftentimes desired, particularly in fragment elaboration regimes, since the fragment is relatively small, multiple pockets exist in the target protein, and the target protein has not yet been fully explored.

[0017] To complete an exhaustive multi-vector reaction-based enumeration, every combination of available building blocks at every vector position would need to be considered. This means the number of total enumerated molecules is the Cartesian product of the number of building blocks at each vector, i.e. two vectors with 105compatible building blocks each results in an enumerated space of 1O10molecules. This can result in a ‘combinatorial explosion’, whereby the number of combinations of building blocks increases exponentially given a linear increase in the number of building blocks.

[0018] In an exhaustive in silico drug discovery screen, all enumerated molecules would be scored based on their properties relative to one or more desired properties of a candidate compound I molecule, e.g. docking with a target, shape similarity, etc., and top-scoring molecules would be chosen for further testing. However, scoring functions for scoring molecules in such a manner tend to be computationally expensive, which makes it infeasible to enumerate and score all of the molecules in a typical enumerated molecule space. For instance, the size of a typical chemical space that is considered for a particular drug discovery project may be of the order of 1012molecules, whereas it may typically be feasible to score only a fraction of the molecules in this space, e.g. 106molecules.

[0019] To address this, methods for selecting and scoring only some molecules in a particular molecule space of interest have been developed. However, it is clear that in such approaches, molecules that are optimised against a desired property profile of a specific drug discovery project under consideration may not be selected and scored, meaning that optimal molecules are not identified as candidate compounds, thereby reducing the probability of the drug discovery project successfully developing a drug.

[0020] There is a need for an approach in which optimal molecules in an enumerated molecule space are identified for synthesis and testing as part of a drug discovery process. It is against this background to which the present invention is set.

[0021] SUMMARY OF THE INVENTION According to an aspect of the present invention there is provided a method for drug design. The method comprises computer-implemented steps. The method comprises defining an intermediate molecule having a first enumeration vector and optionally a second enumeration vector, defining a first synthon set comprising a plurality of first synthons each for enumerating the first enumeration vector of the intermediate molecule, and optionally defining a second synthon set comprising a plurality of second synthons each for enumerating the second enumeration vector of the intermediate molecule. The method comprises selecting a first synthon subset comprising a subset of the plurality of first synthons in the first synthon set, and optionally selecting a second synthon subset comprising a subset of the plurality of second synthons in the second synthon set.

[0022] The method comprises determining a molecule set comprising a plurality of product molecules each obtained by enumerating the first, and optionally second, enumeration vectors of the defined intermediate molecule. The molecule set comprises, for each of the first synthons in the first synthon subset, each product molecule obtained by enumerating the first enumeration vector with the respective first synthon and enumerating the optional second enumeration vector with each of a plurality of defined second proxy synthons. The molecule set may comprise, for each of the second synthons in the second synthon subset, each product molecule obtained by enumerating the second enumeration vector with the respective second synthon and enumerating the first enumeration vector with each of a plurality of defined first proxy synthons.

[0023] The method comprises selecting a molecule subset comprising a subset of the product molecules in the molecule set, and determining, by executing a defined scoring function, a molecule score for each of the product molecules in the molecule subset. The molecule score is indicative of a proximity of a property of the respective product molecule to a desired property of a drug molecule to be designed.

[0024] The method comprises training a first surrogate model to output predicted molecule score as a function of structural features of first synthons. The first surrogate model is trained using the first synthon and corresponding determined molecule score for each product molecule in the molecule subset. The method may comprise training a second surrogate model to output predicted molecule score as a function of structural features of second synthons. The second surrogate model may be trained using the second synthon and corresponding determined molecule score for each product molecule in the molecule subset.

[0025] The method comprises, for each of the first synthons in the first synthon set, executing the trained first surrogate model to determine a first predicted molecule score as a function of structural features of the respective first synthon. The method may comprise, for each of the second synthons in the second synthon set, executing the trained second surrogate model to determine a second predicted molecule score as a function of structural features of the respective second synthon.

[0026] The method comprises selecting a final first synthon subset comprising a subset of the plurality of first synthons in the first synthon set. The selection is based on the determined first predicted molecule scores of the respective first synthons in the first synthon set. The method may comprise selecting a final second synthon subset comprising a subset of the plurality of second synthons in the second synthon set. The selection may be based on the determined second predicted molecule scores of the respective second synthons in the second synthon set.

[0027] The method comprises determining one or more high-scoring product molecules, wherein each high-scoring product molecule is determined by enumerating the first enumeration vector of the intermediate molecule with one of the first synthons in the final first synthon subset and optionally enumerating the second enumeration vector of the intermediate molecule with one of the second synthons in the final second synthon subset.

[0028] The first synthon subset may comprise a defined number of the first synthons in the first synthon set. The second synthon subset may comprise a defined number of the second synthons in the second synthon set.

[0029] The defined number of synthons in the first and second synthon subsets may be at least one order of magnitude less than a total number of synthons in the first and second synthon sets. Optionally, the defined number may be at least two orders of magnitude less than the total number. The total number of first synthons in the first synthon set may be at least 105synthons, and optionally at least 106synthon. The total number of second synthons in the second synthon set may be at least 105synthons, and optionally at least 106synthons.

[0030] The plurality of defined first proxy synthons may be the first synthon subset. The plurality of defined second proxy synthons may be the second synthon subset.

[0031] The plurality of defined first proxy synthons may be a randomly sampled subset of first synthons from the first synthon set. The plurality of defined second proxy synthons may be a randomly sampled subset of second synthons from the second synthon set.

[0032] Selecting the first synthon subset may comprise performing a random selection of first synthons in the first synthon set. Selecting the second synthon subset may comprise performing a random selection of second synthons in the second synthon set.

[0033] Selecting the first synthon subset may comprise selecting first synthons to maximise a value of a diversity metric indicative of a diversity of structural features present in the first synthon subset. Selecting the second synthon subset may comprise selecting second synthons to maximise a value of the diversity metric indicative of a diversity of structural features present in the second synthon subset.

[0034] The steps of selecting the first and second synthon subsets, determining the molecule set, and selecting the molecule subset may collectively comprise: determining a master molecule set comprising a plurality of product molecules each obtained by enumerating the first and second enumeration vectors of the defined intermediate molecule, wherein the master molecule set comprises: for each of the first synthons in the first synthon set, each product molecule obtained by enumerating the first enumeration vector with the respective first synthon and enumerating the second enumeration vector with each of a plurality of defined second master proxy synthons; and for each of the second synthons in the second synthon set, each product molecule obtained by enumerating the second enumeration vector with the respective second synthon and enumerating the first enumeration vector with each of a plurality of defined first master proxy synthons; and selecting the molecule subset by performing a hypercube sampling of molecules in the master molecule set, wherein the hypercube sampling comprises sampling molecules based on the first and second synthons present therein.

[0035] The plurality of defined first master proxy synthons may be the first synthon set. The plurality of defined second master proxy synthons may be the second synthon set.

[0036] A number of product molecules in the selected molecule subset may be at least one order of magnitude less, and optionally at least two orders of magnitude less, than a number of product molecules in the molecule set.

[0037] Selecting the molecule subset may comprise performing a random selection of the product molecules in the molecule set.

[0038] For each of the first synthons in the first synthon subset, the molecule subset may be selected to include at least a defined number of product molecules obtained by enumerating the first enumeration vector with the respective first synthon. Optionally, the defined number may be 1 , 2, 3, 4, 5, 10, 15, 20, 25, 50 or any other suitable number.

[0039] For each of the second synthons in the second synthon subset, the molecule subset may be selected to include at least a defined number of product molecules obtained by enumerating the second enumeration vector with the respective second synthon. Optionally, the defined number may be 1 , 2, 3, 4, 5, 10, 15, 20, 25, 50 or any other suitable number.

[0040] The defined scoring function may comprise a docking score function. The molecule score may be indicative of a prediction of a binding affinity of the respective product molecule when it is docked with a defined target molecule.

[0041] The defined scoring function may comprise a shape similarity function. The molecule score may be indicative of a prediction of a shape similarity of the respective product molecule to a defined target molecule.

[0042] The defined scoring function may be a multi-parameter optimisation function. That is, a molecule may be scored based on a combination of a plurality of parameters or properties of the molecule. In one example, the scoring function may include a docking score and a chemical property score to obtain a combined score. The combined score of each product molecule may then be used to train the surrogate model(s).

[0043] The first surrogate model may be a first machine learning model. The second surrogate model may be a second machine learning model.

[0044] The first machine learning model may be a first random forest model. The second machine learning model may be a second random forest model.

[0045] The first machine learning model may be a first feed forward neural network model. The second machine learning model may be a second feed forward neural network model.

[0046] The first machine learning model may be a first message passing neural network model. The second machine learning model may be a second message passing neural network model.

[0047] The method may comprise, prior to the steps of selecting the final first synthon subset and the final second synthon subset, iteratively repeating steps of: selecting a further first synthon subset comprising a subset of the plurality of first synthons in the first synthon set; selecting a further second synthon subset comprising a subset of the plurality of second synthons in the second synthon set; determining a further molecule set comprising a plurality of product molecules each obtained by enumerating the first and second enumeration vectors of the defined intermediate molecule, wherein the further molecule set comprises: for each of the first synthons in the further first synthon subset, each product molecule obtained by enumerating the first enumeration vector with the respective first synthon and enumerating the second enumeration vector with each of a plurality of defined second proxy synthons; and for each of the second synthons in the further second synthon subset, each product molecule obtained by enumerating the second enumeration vector with the respective second synthon and enumerating the first enumeration vector with each of a plurality of defined first proxy synthons; selecting a further molecule subset comprising a subset of the molecules in the further molecule set; determining, using the defined scoring function, the molecule score for each of the molecules in the further molecule subset; retraining the first surrogate model using the first synthon and corresponding determined molecule score for each product molecule in the further molecule subset to obtain an updated first surrogate model; retraining the second surrogate model using the second synthon and corresponding determined molecule score for each product molecule in the further molecule subset to obtain an updated second surrogate model; for each of the first synthons in the first synthon set, executing the updated first surrogate model to determine the first predicted molecule score; and for each of the second synthons in the second synthon set, executing the updated second surrogate model to determine the second predicted molecule score, until a stop condition is satisfied.

[0048] The further first synthon subset may comprise a defined number of the first synthons in the first synthon set. The further second synthon subset may comprise a defined number of the second synthons in the second synthon set. Optionally, the defined number of synthons in the further first and second synthon subsets may be at least one order of magnitude less than a total number of synthons in the first and second synthon sets, respectively. Further optionally, the defined number may be at least two orders of magnitude less than the total number.

[0049] Selecting the further first synthon subset may comprise performing a greedy selection of first synthons in the first synthon set according to determined first predicted molecule score. Selecting the further second synthon subset may comprise performing a greedy selection of second synthons in the second synthon set according to determined second predicted molecule score.

[0050] Selecting the further first synthon subset may comprise performing a random selection of first synthons in the first synthon set. Selecting the further second synthon subset may comprise performing a random selection of second synthons in the second synthon set. Selecting the further first synthon subset may comprise only selecting first synthons not selected in a previous iteration of the method. Selecting the further second synthon subset may comprise only selecting second synthons not selected in a previous iteration of the method.

[0051] The stop condition may be satisfied when a first rate of improvement of first predicted molecule scores of the first synthons in the further first synthon subset of a current iteration of the method relative to a previous iteration of the method is below a defined first threshold rate of improvement. Alternatively, or in addition, the stop condition may be satisfied when a second rate of improvement of second predicted molecule scores of the second synthons in the further second synthon subset of the current iteration of the method relative to the previous iteration of the method is below a defined second threshold rate of improvement. Optionally, the defined first threshold rate of improvement may be equal to the defined second threshold rate of improvement.

[0052] The first rate of improvement may be determined according to a calculation of average first predicted molecule score of the first synthons in the further first synthon subset of the current iteration of the method relative to the previous iteration of the method. The second rate of improvement may be determined according to a calculation of average second predicted molecule score of the second synthons in the further second synthon subset of the current iteration of the method relative to the previous iteration of the method.

[0053] The stop condition may be satisfied when the repeated steps have been performed a defined number of times. Optionally, the defined number of times may be any one of 1 , 2, 3, 4, 5, 10, 20, 50, 100, or any other suitable number.

[0054] The plurality of defined first proxy synthons may be the further first synthon subset. The plurality of defined second proxy synthons may be the further second synthon subset.

[0055] A number of product molecules in the selected further molecule subset may be at least one order of magnitude less, and optionally at least two orders of magnitude less, than a number of product molecules in the molecule set.

[0056] Selecting the further molecule subset may comprise performing a random selection of the product molecules in the further molecule set. For each of the first synthons in the further first synthon subset, the further molecule subset may be selected to include at least a defined number of product molecules obtained by enumerating the first enumeration vector with the respective first synthon. Optionally, the defined number may be 1 , 2, 3, 4, 5, 10, 15, 20, 25, 50, or any other suitable number.

[0057] For each of the second synthons in the further second synthon subset, the further molecule subset may be selected to include at least a defined number of product molecules obtained by enumerating the second enumeration vector with the respective second synthon. Optionally, the defined number may be 1 , 2, 3, 4, 5, 10, 15, 20, 25, 50, or any other suitable number.

[0058] Selecting the final first synthon subset may comprise: ranking the first synthons in the first synthon set according to determined first predicted molecule score; defining a first ranked subset comprising a defined number of the top ranked first synthons; and selecting at least some of the top ranked first synthons in the first ranked subset to be included in the final first synthon subset.

[0059] Selecting the final second synthon subset may comprise: ranking the second synthons in the second synthon set according to determined second predicted molecule score; defining a second ranked subset comprising a defined number of the top ranked second synthons; and selecting at least some of the top ranked second synthons in the second ranked subset to be included in the final second synthon subset.

[0060] The defined number of synthons in the final first and second synthon subsets may be at least one order of magnitude less than a total number of synthons in the first and second synthon sets, respectively. Optionally, the defined number may be at least two orders of magnitude less than the total number.

[0061] The determined high-scoring product molecules may form a high-scoring molecule set. The high-scoring molecule set may comprise, for each of the first synthons in the final first synthon set, each high-scoring product molecule obtained by enumerating the first enumeration vector with the respective first synthon and enumerating the second enumeration vector with each of the plurality of second synthons in the final second synthon set. The method may comprise selecting at least some of the high-scoring product molecules for synthesis.

[0062] Selecting the high-scoring product molecules for synthesis may comprise: determining, by executing the defined scoring function, or retrieving, from a database storing previously-determined molecule scores, the molecule score for each of the high-scoring product molecules; ranking the high-scoring product molecules according to determined molecule score; and, selecting a defined number of the top-ranked high-scoring product molecules for synthesis.

[0063] The method may comprise, prior to selecting the defined number of the top-ranked high- scoring product molecules, performing one or more filters to remove one or more high- scoring product molecules from the ranking.

[0064] The one or more filters may be configured to remove high-scoring product molecules having one or more undesirable properties. Optionally, the undesirable properties may include high toxicity or that the high-scoring molecule is a duplicate of a previously-tested molecule.

[0065] The defined number of the top-ranked high-scoring product molecules may be less than 100, less than 50, less than 25, less than 10, less than 5, and / or less than another other suitable number.

[0066] The method may comprise synthesising the selected high-scoring product molecules, and testing the synthesised high-scoring product molecules to determine one or more biological properties of the high-scoring product molecules.

[0067] The determined one or more biological properties may include the property associated with the molecule score of the first and second surrogate models.

[0068] The method may comprise determining the molecule score of the respective tested high- scoring product molecules. The method may comprise: retraining the trained first surrogate model, or the updated first surrogate model, using the first synthon and corresponding determined molecule score for each tested high- scoring product molecule to obtain a tested first surrogate model; and optionally retraining the trained second surrogate model, or the updated second surrogate model, using the second synthon and corresponding determined molecule score for each tested high-scoring product molecule to obtain a tested second surrogate model.

[0069] The method may comprise: selecting a new first synthon subset comprising a subset of the plurality of first synthons in the first synthon set; selecting a new second synthon subset comprising a subset of the plurality of second synthons in the second synthon set; determining a new molecule set comprising a plurality of product molecules each obtained by enumerating the first and second enumeration vectors of the defined intermediate molecule, wherein the new molecule set comprises: for each of the first synthons in the new first synthon subset, each product molecule obtained by enumerating the first enumeration vector with the respective first synthon and enumerating the second enumeration vector with each of the plurality of defined second proxy synthons; and for each of the second synthons in the new second synthon subset, each product molecule obtained by enumerating the second enumeration vector with the respective second synthon and enumerating the first enumeration vector with each of the plurality of defined first proxy synthons; selecting a new molecule subset comprising a subset of the product molecules in the new molecule set; determining, by executing the defined scoring function, the molecule score for each of the product molecules in the new molecule subset; retraining the tested first surrogate model using the first synthon and corresponding determined molecule score for each product molecule in the new molecule subset to obtain a new first surrogate model; retraining the tested second surrogate model using the second synthon and corresponding determined molecule score for each product molecule in the new molecule subset to obtain a new second surrogate model; for each of the first synthons in the first synthon set, executing the new first surrogate model to determine a first predicted molecule score as a function of structural features of the respective first synthon; for each of the second synthons in the second synthon set, executing the new second surrogate model to determine a second predicted molecule score as a function of structural features of the respective second synthon; selecting a new final first synthon subset comprising a subset of the plurality of first synthons in the new first synthon set, wherein the selection is based on the determined first predicted molecule scores of the respective first synthons in the first synthon set; selecting a new final second synthon subset comprising a subset of the plurality of second synthons in the second synthon set, wherein the selection is based on the determined second predicted molecule scores of the respective second synthons in the second synthon set; and determining one or more new high-scoring product molecules, wherein each high- scoring product molecule is determined by enumerating the first enumeration vector of the intermediate molecule with one of the first synthons in the new final first synthon subset and enumerating the second enumeration vector of the intermediate molecule with one of the second synthons in the new final second synthon subset.

[0070] The method may comprise selecting at least some of the new high-scoring product molecules, and synthesising the selected new high-scoring product molecules. Optionally, the method may comprise testing the synthesised new high-scoring product molecules to determine one or more biological properties of the new high-scoring product molecules.

[0071] The method may comprise defining a machine learning model for approximating one or more biological properties of molecules as a function of the one or more structural features of said molecules. The method may comprise training the machine learning model using the tested high-scoring product molecules.

[0072] The machine learning model may be at least one of: a Bayesian optimisation model; a regression model; a clustering model; a decision tree model; a random forest model; and, a neural network model. The method may comprise executing the machine learning model, after the training step, to predict one or more molecules in a defined population of molecules having one or more desired biological properties.

[0073] The method may further comprise synthesising at least one of the one or more predicted molecules.

[0074] The intermediate molecule may be defined in dependence on a defined target molecule with which it is desired that determined product molecules interact. Optionally, the defined target molecule may be associated with a defined disease of interest.

[0075] One or more of the high-scoring product molecules, or predicted molecules, may be a candidate drug or therapeutic molecule having a desired biological, biochemical, physiological and / or pharmacological activity against a defined target molecule.

[0076] The predetermined target molecule may be an in vitro and / or in vivo therapeutic, diagnostic or experimental assay target.

[0077] The candidate drug or therapeutic molecule may be for use in medicine; for example, in a method for the treatment of an animal, such as a human or non-human animal.

[0078] The structural features of the first and second synthons may correspond to fragments present in the respective first and second synthons.

[0079] The fragments present in each of the first and second synthons may be represented as a respective molecular fingerprint; Optionally, the molecular fingerprint may be an Extended Connectivity Fingerprint, ECFP; further optionally, ECFPO, ECFP2, ECFP4, ECFP6, ECFP8, ECFP10 or ECFP12.

[0080] According to another aspect of the invention there is provided a method for drug design. The method comprises computer-implemented steps. The method comprises: defining a synthetic route having a plurality of chemical reaction steps including a first chemical reaction step and a second chemical reaction step to be performed after the first chemical reaction step; defining a first synthon set comprising a plurality of first synthons each for combining with an initial intermediate molecule at the first chemical reaction step to obtain a further intermediate molecule; defining a second synthon set comprising a plurality of second synthons each for combining with the further intermediate molecule at the second chemical reaction step to obtain a product molecule; selecting a first synthon subset comprising a subset of the plurality of first synthons in the first synthon set; selecting a second synthon subset comprising a subset of the plurality of second synthons in the second synthon set; determining a molecule set comprising a plurality of product molecules each obtained by implementing the defined synthetic route starting at the initial intermediate molecule, wherein the molecule set comprises: for each of the first synthons in the first synthon subset, each product molecule obtained by combining the initial intermediate molecule with the respective first synthon at the first chemical reaction step and combining the further intermediate molecule, obtained from the first chemical reaction step, with each of a plurality of defined second proxy synthons at the second chemical reaction step; and for each of the second synthons in the second synthon subset, each product molecule obtained by combining the initial intermediate molecule with each of a plurality of defined first proxy synthons at the first chemical reaction step and combining the further intermediate molecule, obtained from the first chemical reaction step, with the respective second synthon at the second chemical reaction step; selecting a molecule subset comprising a subset of the product molecules in the molecule set; determining, by executing a defined scoring function, a molecule score for each of the product molecules in the molecule subset, wherein the molecule score is indicative of a proximity of a property of the respective product molecule to a desired property of a drug molecule to be designed; training a first surrogate model to output predicted molecule score as a function of structural features of first synthons, wherein the first surrogate model is trained using the first synthon and corresponding determined molecule score for each product molecule in the molecule subset; training a second surrogate model to output predicted molecule score as a function of structural features of second synthons, wherein the second surrogate model is trained using the second synthon and corresponding determined molecule score for each product molecule in the molecule subset; for each of the first synthons in the first synthon set, executing the trained first surrogate model to determine a first predicted molecule score as a function of structural features of the respective first synthon; for each of the second synthons in the second synthon set, executing the trained second surrogate model to determine a second predicted molecule score as a function of structural features of the respective second synthon; selecting a final first synthon subset comprising a subset of the plurality of first synthons in the first synthon set, wherein the selection is based on the determined first predicted molecule scores of the respective first synthons in the first synthon set; selecting a final second synthon subset comprising a subset of the plurality of second synthons in the second synthon set, wherein the selection is based on the determined second predicted molecule scores of the respective second synthons in the second synthon set; and determining one or more high-scoring product molecules, wherein each high- scoring product molecule is obtained by implementing the defined synthetic route starting at the initial intermediate molecule, comprising combining the initial intermediate molecule with one of the first synthons in the final first synthon subset at the first chemical reaction step and combining the further intermediate molecule, obtained from the first chemical reaction step, with one of the second synthons in the final second synthon subset at the second chemical reaction step.

[0081] The steps of selecting the first and second synthon subsets, determining the molecule set, and selecting the molecule subset may collectively comprise: determining a master molecule set comprising a plurality of product molecules each obtained by implementing the defined synthetic route starting at the initial intermediate molecule, wherein the master molecule set comprises: for each of the first synthons in the first synthon set, each product molecule obtained by combining the initial intermediate molecule with the respective first synthon at the first chemical reaction step and combining the further intermediate molecule, obtained from the first chemical reaction step, with each of a plurality of defined second master proxy synthons at the second chemical reaction step; and for each of the second synthons in the second synthon set, each product molecule obtained by combining the initial intermediate molecule with each of a plurality of defined first master proxy synthons at the first chemical reaction step and combining the further intermediate molecule, obtained from the first chemical reaction step, with the respective second synthon at the second chemical reaction step; and selecting the molecule subset by performing a hypercube sampling of molecules in the master molecule set, wherein the hypercube sampling comprises sampling molecules based on the first and second synthons present therein.

[0082] The plurality of defined first master proxy synthons may be the first synthon set. The plurality of defined second master proxy synthons may be the second synthon set.

[0083] A number of product molecules in the selected molecule subset may be at least one order of magnitude less than a number of product molecules in the molecule set. Optionally, this may be at least two orders of magnitude less.

[0084] Selecting the final first synthon subset may comprise: ranking the first synthons in the first synthon set according to determined first predicted molecule score; defining a first ranked subset comprising a defined number of the top ranked first synthons; and selecting at least some of the top ranked first synthons in the first ranked subset to be included in the final first synthon subset.

[0085] Selecting the final second synthon subset may comprise: ranking the second synthons in the second synthon set according to determined second predicted molecule score; defining a second ranked subset comprising a defined number of the top ranked second synthons; and selecting at least some of the top ranked second synthons in the second ranked subset to be included in the final second synthon subset.

[0086] The method may comprise selecting at least some of the high-scoring product molecules for synthesis.

[0087] Selecting the high-scoring product molecules for synthesis may comprise: determining, by executing the defined scoring function, or retrieving, from a database storing previously-determined molecule scores, the molecule score for each of the high-scoring product molecules; ranking the high-scoring product molecules according to determined molecule score; and, selecting a defined number of the top-ranked high-scoring product molecules for synthesis.

[0088] The method may comprise synthesising the selected high-scoring product molecules, and testing the synthesised high-scoring product molecules to determine one or more biological properties of the high-scoring product molecules.

[0089] The determined one or more biological properties may include the property associated with the molecule score of the first and second surrogate models. The method may comprise determining the molecule score of the respective tested high-scoring product molecules.

[0090] The method may comprise: retraining the trained first surrogate model, or the updated first surrogate model, using the first synthon and corresponding determined molecule score for each tested high- scoring product molecule to obtain a tested first surrogate model; and retraining the trained second surrogate model, or the updated second surrogate model, using the second synthon and corresponding determined molecule score for each tested high-scoring product molecule to obtain a tested second surrogate model.

[0091] The method may comprise: selecting a new first synthon subset comprising a subset of the plurality of first synthons in the first synthon set; selecting a new second synthon subset comprising a subset of the plurality of second synthons in the second synthon set; determining a new molecule set comprising a plurality of product molecules each obtained by implementing the defined synthetic route starting at the initial intermediate molecule, wherein the new molecule set comprises: for each of the first synthons in the new first synthon subset, each product molecule obtained by combining the initial intermediate molecule with the respective first synthon at the first chemical reaction step and combining the further intermediate molecule, obtained from the first chemical reaction step, with each of the plurality of defined second proxy synthons at the second chemical reaction step; and for each of the second synthons in the new second synthon subset, each product molecule obtained by combining the initial intermediate molecule with each of the plurality of defined first proxy synthons at the first chemical reaction step and combining the further intermediate molecule, obtained from the first chemical reaction step, with the respective second synthon at the second chemical reaction step; selecting a new molecule subset comprising a subset of the product molecules in the new molecule set; determining, by executing the defined scoring function, the molecule score for each of the product molecules in the new molecule subset; retraining the tested first surrogate model using the first synthon and corresponding determined molecule score for each product molecule in the new molecule subset to obtain a new first surrogate model; retraining the tested second surrogate model using the second synthon and corresponding determined molecule score for each product molecule in the new molecule subset to obtain a new second surrogate model; for each of the first synthons in the first synthon set, executing the new first surrogate model to determine a first predicted molecule score as a function of structural features of the respective first synthon; for each of the second synthons in the second synthon set, executing the new second surrogate model to determine a second predicted molecule score as a function of structural features of the respective second synthon; selecting a new final first synthon subset comprising a subset of the plurality of first synthons in the new first synthon set, wherein the selection is based on the determined first predicted molecule scores of the respective first synthons in the first synthon set; selecting a new final second synthon subset comprising a subset of the plurality of second synthons in the second synthon set, wherein the selection is based on the determined second predicted molecule scores of the respective second synthons in the second synthon set; and determining one or more new high-scoring product molecules, wherein each high- scoring product molecule is obtained by implementing the defined synthetic route starting at the initial intermediate molecule, comprising combining the initial intermediate molecule with one of the first synthons in the new final first synthon subset at the first chemical reaction step and combining the further intermediate molecule, obtained from the first chemical reaction step, with one of the second synthons in the new final second synthon subset at the second chemical reaction step.

[0092] The method may comprise selecting at least some of the new high-scoring product molecules, and synthesising the selected new high-scoring product molecules; optionally, testing the synthesised new high-scoring product molecules to determine one or more biological properties of the new high-scoring product molecules.

[0093] The method may comprise: defining a machine learning model for approximating one or more biological properties of molecules as a function of the one or more structural features of said molecules; and, training the machine learning model using the tested high-scoring product molecules.

[0094] The method may comprise executing the machine learning model, after the training step, to predict one or more molecules in a defined population of molecules having one or more desired biological properties.

[0095] The method may further comprise synthesising at least one of the one or more predicted molecules.

[0096] The initial intermediate molecule may be an initial synthon from an initial synthon subset comprising a plurality of initial synthons. The initial synthon subset may be selected from an initial synthon set comprising a plurality of initial synthons.

[0097] The molecule set may include product molecules including different initial synthons from the initial synthon subset.

[0098] The method may comprise training an initial surrogate model to output predicted molecule score as a function of structural features of initial synthons. The initial surrogate model may be trained using the first synthon and corresponding determined molecule score for each product molecule in the molecule subset. For each of the initial synthons in the first synthon set, the method may comprise executing the trained initial surrogate model to determine a first predicted molecule score as a function of structural features of the respective initial synthon. Selecting a final initial synthon subset may comprise a subset of the plurality of initial synthons in the initial synthon set. The selection may be based on the determined initial predicted molecule scores of the respective initial synthons in the initial synthon set. Each of the one or more high-scoring product molecules may be determined using one of the initial synthons from the final initial synthon subset.

[0099] Defining the first and second synthon sets may comprise: providing an example product molecule and performing an atom mapping process, based on the example product molecule, to identify atoms involved in each of the first and second reaction steps; and determining molecular substructures, based on the identified atoms, for performing each of the first and second reaction steps.

[0100] The synthons in the first and second synthon sets may be selected based on the determined molecular substructures for the respective first and second reaction steps.

[0101] Defining the first and second synthon sets may comprise retrieving synthons, having the determined molecular substructures for the respective first and second reaction steps, from a defined database.

[0102] Defining the first and second synthon sets may comprise retrieving synthons, having molecular substructures similar to the determined molecular substructures for the respective first and second reaction steps, from the defined database. Molecular similarity may be determined according to a defined molecular similarity metric.

[0103] According to another aspect of the invention there is provided a method for drug design. The method comprises computer-implemented steps of: defining a synthetic route having a plurality of defined chemical reaction steps to be performed in a defined order to obtain a product molecule from an initial intermediate molecule; for each of the defined chemical reaction steps of the defined synthetic route: defining a synthon set comprising a plurality of synthons each for combining, at the respective chemical reaction step, with an intermediate molecule, obtained from a previous chemical reaction step of the defined synthetic route; selecting a synthon subset comprising a subset of the plurality of synthons in the synthon set; determining a molecule set comprising a plurality of product molecules each obtained by implementing the defined synthetic route starting at the initial intermediate molecule and using one of the synthons from the respective synthon subset at each respective chemical reaction step of the synthetic route; selecting a molecule subset comprising a subset of the product molecules in the molecule set; determining, by executing a defined scoring function, a molecule score for each of the product molecules in the molecule subset, wherein the molecule score is indicative of a proximity of a property of the respective product molecule to a desired property of a drug molecule to be designed; for each of the defined chemical reaction steps of the defined synthetic route: training a respective surrogate model to output predicted molecule score as a function of structural features of synthons, wherein the respective surrogate model is trained using the synthon, from the respective synthon subset, and corresponding determined molecule score for each product molecule in the molecule subset; for each of the synthons in the respective synthon set, executing the respective trained surrogate model to determine a predicted molecule score as a function of structural features of the respective synthon; selecting a respective final synthon subset comprising a subset of the plurality of synthons in the respective synthon set, wherein the selection is based on the determined predicted molecule scores of the respective synthons in the respective synthon set; and determining one or more high-scoring product molecules, wherein each high- scoring product molecule is obtained by implementing the defined synthetic route starting at the initial intermediate molecule and using one of the synthons from the respective final synthon subset at each respective chemical reaction step of the synthetic route.

[0104] The plurality of synthons sets may include a synthon set comprising a plurality of synthons for use as the initial intermediate molecule. According to another aspect of the invention there is provided a method for drug design. The method comprises computer-implemented steps of: defining a plurality of synthon sets each comprising a respective plurality of synthons; for each of the plurality of synthon sets, selecting a synthon subset comprising a subset of the plurality of synthons in the respective synthon set; determining a molecule set comprising a plurality of product molecules each obtained by combining together one of the synthons from each of the respective synthon subsets; selecting a molecule subset comprising a subset of the product molecules in the molecule set; determining, by executing a defined scoring function, a molecule score for each of the product molecules in the molecule subset, wherein the molecule score is indicative of a proximity of a property of the respective product molecule to a desired property of a drug molecule to be designed; for each of the plurality of synthon sets: training a surrogate model to output predicted molecule score as a function of structural features of synthons, wherein the surrogate model is trained using the synthon, from the synthon subset of the respective synthon set, and corresponding determined molecule score for each product molecule in the molecule subset; for each of the synthons in the respective synthon set, executing the respective trained surrogate model to determine a predicted molecule score as a function of structural features of the respective synthon; selecting a final synthon subset comprising a subset of the plurality of synthons in the respective synthon set, wherein the selection is based on the determined predicted molecule scores of the synthons in the respective synthon set; and determining one or more high-scoring product molecules, wherein each high-scoring product molecule is obtained by combining together one of the synthons from each of the respective final synthon subsets.

[0105] Combining together one of the synthons from each of the respective synthon subsets may comprise one of: enumerating a plurality of enumeration vectors of an intermediate molecule with a respective one of the synthons from the respective synthon subsets; and implementing a defined synthetic route having a plurality of chemical reaction steps performed in order, wherein each chemical reaction step comprises enumerating a respective intermediate molecule with a respective one of the synthons from the respective synthon subsets.

[0106] According to another aspect of the invention there is provided a non-transitory, computer- readable storage medium storing instructions thereon that when executed by a computer processor causes the computer processor to perform the computer-implemented steps of the method defined above.

[0107] According to another aspect of the present invention there is provided a computing device comprising a processor that is configured to perform the computer-implemented steps of the method defined above.

[0108] BRIEF DESCRIPTION OF THE DRAWINGS

[0109] Examples of the invention will now be described with reference to the accompanying drawings, in which:

[0110] Figure 1 is a schematic illustration of an active learning loop for selecting and scoring synthons used to enumerate molecules;

[0111] Figure 2 shows the steps of a method in accordance with examples of the invention, wherein the method includes the active learning loop of Figure 1 ;

[0112] Figure 3 illustrates an example of an intermediate fragment with two enumeration vectors for enumeration with synthons that are selected and scored according to the active learning loop of Figure 1 ;

[0113] Figure 4 illustrates how the intermediate fragment of Figure 3 receives a linker atom at each of the two enumeration vector positions when building blocks for elaborating the enumeration vectors are transformed into synthons; Figure 5 illustrates examples of building blocks in Figure 4 being transformed into synthons;

[0114] Figures 6(a) and 6(b) show experimental result plots of percentage of top scoring molecules identified, against number of molecules used to train a surrogate model, for different synthon acquisition and replacement strategies, when using the active learning loop of Figure 1 , wherein the molecules are scored using docking and shape similarity metrics, respectively;

[0115] Figures 7(a) and 7(b) show experimental result plots of percentage of top scoring molecules identified, against number of molecules used to train a surrogate model, for different numbers of synthons being selected in each iteration, when using the active learning loop of Figure 1 , wherein the molecules are scored using docking and shape similarity metrics, respectively;

[0116] Figures 8(a) and 8(b) show experimental result plots of percentage of top scoring molecules identified, against number of molecules used to train a surrogate model, for different types of surrogate models, when using the active learning loop of Figure 1 , wherein the molecules are scored using docking and shape similarity metrics, respectively;

[0117] Figures 9(a) and 9(b) show experimental result plots of percentage of top scoring molecules identified, against number of molecules used to train a surrogate model, for different enumeration and selection strategies, when using the active learning loop of Figure 1 , wherein the molecules are scored using shape similarity and docking metrics, respectively;

[0118] Figures 10(a) and 10(b) show experimental result plots of percentage of top scoring molecules identified, against number of molecules used to train a surrogate model, for different enumeration strategies, when using the active learning loop of Figure 1 , wherein the molecules are scored using shape similarity and docking metrics, respectively;

[0119] Figures 11 (a) and 11 (b) show experimental result plots of percentage of top scoring molecules identified, against number of molecules used to train a surrogate model, for different numbers of synthons being acquired at each iteration, when using the active learning loop of Figure 1 , wherein the molecules are scored using shape similarity and docking metrics, respectively;

[0120] Figure 12 illustrates results of a further experiment using the active learning loop of Figure 1 ;

[0121] Figures 13 and 14 illustrate results of further experiments using the active learning loop of Figure 1 when a multiparameter scoring function is used to score molecules;

[0122] Figure 15 illustrates an overview of a method according to the invention for obtaining desirable product molecules based on a known synthetic route; and

[0123] Figure 16 schematically illustrates an example of a product molecule obtained by implementing a synthetic route as in Figure 15.

[0124] DETAILED DESCRIPTION

[0125] The present invention provides an approach for identifying optimal molecules in a nonenumerated molecule space for synthesis and testing as part of a drug discovery process. In particular, the described approach addresses the above-outlined issue of ‘combinatorial explosion’ that arises with multi-vector reaction-based enumeration. The approach of the present invention is advantageous in that molecules are identified that are optimised against a desired property profile of a lead compound or hit compound for the specific drug discovery project under consideration, e.g. affinity against a specified protein target, that would not be identified using previous methods (as in previous methods it is typically not feasible to enumerate and score all of the molecules in a chemical space of interest). The invention is therefore beneficial in that the approach results in the identification of molecules that are better candidates to be drugs, i.e. have therapeutic effect, against a disease under consideration in a specified drug discovery project. Furthermore, these optimised molecules are identified in a manner that requires fewer steps of (wet lab) synthesis and testing of molecules, because identified molecules are easier to make, thereby reducing the time and cost associated with identifying suitable candidate compounds. An approach to reduce the issue of combinatorial explosion from multi-vector reaction enumeration in a docking regime has previously been described in ‘Synthon-based ligand discovery in virtual libraries of over 11 billion compounds’, Sadybekov et al., Nature, vol. 601 , 452-459 (2022). In this approach, synthons are used to first enumerate off a single vector, cap the other vector(s) with a methyl group, dock those molecules, and select the best scoring molecules. As is known in the art, synthons are building blocks that are modified to appear as if they are already in the product molecule. In this previous approach, the best scoring compounds from a previous run would be de-methylated and the process would be repeated until all vectors are enumerated.

[0126] Such an elaboration technique may be particularly suited for docking because the areas for growth are well-defined. The vector on any particular molecule is enumerated only if there is an area in the binding pocket to accommodate that enumeration. However, when trying to assess this for different scoring functions, such an approach may not translate well. For instance, scoring with molecular weight or other QED (Quantitative Estimate of Drug likeness) properties would be difficult to evaluate with this enumeration strategy. Expressed differently, this previous approach enumerates one vector at a time and then filters for molecules that score best, before completing subsequent enumeration on the next vectors. However, this approach does not allow for scoring the complex interplay that synthons may have on each other and how they contribute / balance molecule properties at large.

[0127] Active learning is a technique that has been shown to greatly reduce computational costs associated with virtual screening of molecules. For instance, Bayesian optimisation may be used for this purpose, which is a model-based approach that uses a surrogate function to evaluate molecular descriptors so as to learn molecular spaces more quickly. Such an approach may be used not only for docking, but also other objective functions, including more complex multi-parameter optimisation functions.

[0128] In essence, active learning strategies use a series of acquiring and scoring molecules, followed by training surrogate models using molecular I structural features of molecules, to predict the best molecules and navigate the uncertainty of a given chemical space. However, previous active learning tools for screening combinatorial spaces may require evaluating all product molecules. The present invention advantageously provides an active learning method that trains surrogate models on synthon space - rather than molecule space or product space - to predict synthons that have a high probability of scoring favourably for a given objective function. Importantly, this means that combinatorial space is reduced by orders of magnitude prior to enumeration of molecules. In turn, this means that the total number of molecules that need to be scored via evaluation of an expensive scoring function is decreased drastically. The method achieves these benefits while also being then able to thereby identify optimised molecules - using the optimally identified synthons - in an enumerated molecule space.

[0129] Broadly, the method may start by considering all enumeration vectors from a core or intermediate molecule and retrieve synthons that are compatible with each vector given their available chemistry. These synthons may be regarded as being part of synthon sets, with each vector having a corresponding synthon set. Examples described herein focus on fragment-like molecules that are enumerated with two reaction vectors. However, it will be understood that this described method is also applicable to fragment-like molecules that are enumerated with more than two reaction vectors, e.g. three reaction vectors, as well as to fragment-like molecules that are enumerated with a single reaction vector.

[0130] Once the synthon sets have been determined I defined, the active learning loop may include steps of selecting, enumerating, sampling, scoring, training and predicting. Importantly, all of the steps in this loop are performed on individual synthon sets, with the exception of the enumerating, sampling and scoring steps. Enumeration will combine the synthon sets together, along with the fragment, creating ‘product’ molecules, while sampling will choose I select molecules from that space.

[0131] Scoring is then performed on the sampled product molecules. The scores will be associated to the synthons which created those scores to train independent models for each synthon set. At the end of the active learning loop, a set of synthons that has an increased likelihood of scoring well on the given objective function is output. These synthons can then be enumerated to product molecules, whereby the total enumerated spaces may be decreased by orders of magnitude. Moreover, further optimisation methods may be used in this product space to increase computational efficiency even further. Figure 1 schematically illustrates a computer-implemented active learning loop or method 10 in accordance with examples of the present invention. In the described example, enumeration of a fragment-like molecule (or intermediate molecule or reaction intermediate or, simply, intermediate) 101 having two enumeration vectors is considered. An intermediate is a fragment or molecular entity arising within the sequence of a chemical reaction. The intermediate 101 may include a core scaffold that molecules are to be enumerated from (as described below). Specifically, the intermediate molecule has one or more (typically two or three) vectors that can be leveraged to generate novel molecules.

[0132] The intermediate 101 may be selected from a source of defined, synthesisable intermediates, e.g. from the biomedical literature or a database. In an example, selecting or determining the intermediate 101 may involve identifying a compound that has previously been made for one or more drug targets, and then search for / identify building blocks that are available to act as the intermediate. The intermediate 101 may be selected based on the specific drug discovery project under consideration. It will be understood that the methods described herein are applicable to any intermediate determined either manually (from an existing I known synthesis scheme) or obtained automatically, e.g. from retrosynthesis software.

[0133] A set or pool of synthons for each of the enumeration vectors of the intermediate 101 is defined, namely, first and second synthon sets 102a, 102b. These synthons may be transformed building blocks. The building blocks may be obtained from a catalogue or database, e.g. MCule vendor catalogue (mcule.com / database / ), and then transformed into synthons based on reaction compatibility with the intermediate 101. Each enumeration vector of the intermediate 101 has a functional group that facilitates chemistry. Each of these functional groups may be analysed to determine compatible chemistry according to a repository of encoded chemical reactions. These reactions may be of various types, but can be expressed as a machine readable representation of atom and bond rearrangements induced by a reaction, e.g. SMIRKS (which are known in the art). The first and second synthon sets 102a, 102b may be defined to include certain synthons based on the determined compatible chemistry.

[0134] In one example, a defined database of reaction SMIRKS may be used to identify compatible chemistry for a given enumeration vector. Compatible SMIRKS may then be transformed into synthon transformation SMIRKS, which may be regarded as defining a reactant’s transformation from a building block to a synthon, i.e. the transformation that removes the functional group(s) that is not present in the final molecule and places a linker atom in that position. Once the synthon transformation reactions have been made (by a suitable software tool / function), the synthon transformation SMIRKS may be exhaustively applied to all reactants that have the substructure needed to be considered a valid reactant for the reaction. The synthons may then be stored and deduplicated after all compatible chemistry is considered and all synthons have been produced.

[0135] Returning to Figure 1 , a first round or initial subset 103a, 103b of each respective synthon set 102a, 102b is selected or sampled. This may be performed by randomly sampling a defined number of synthons from each of the respective first and second synthon sets 102a, 102b. Alternatively, this may be performed in a manner that increases or maximises the diversity of the sampled subset of synthons (e.g. relative to random sampling) . Diversity may be determined with reference to the molecular or structural features of the synthons, e.g. the specific fragments present in the synthons, and so sampling to increase diversity may involve sampling a subset of synthons that increases the number of different types of fragments I structural features present in the synthon subset. Further alternatively, the initial subsets 103a, 103b may be obtained using hypercube sampling.

[0136] The number of synthons sampled for the initial subset 103a, 103b may be at least one - and typically more than one - order of magnitude fewer than the number of synthons in the respective synthon set 102a, 102b. The first and second initial subsets 103a, 103b may have the same number of synthons as one another.

[0137] The active learning method 10 then involves determining a set 104 of enumerated molecules that includes each molecule obtained by enumerating the intermediate fragment 101 with each of the synthons from the first and second initial subsets 103a, 103b. In one example, this is performed by enumerating the intermediate 101 with synthons from the first initial subset 103a and synthons from the second initial subset 103b. That is, in such an example the enumerated molecule set 104 will include a molecule having each different combination of synthons from the first and second initial subsets 103a, 103b, used to enumerate the intermediate 101. In a different example, molecules in the enumerated molecule set 104 may be obtained by enumerating the intermediate 101 with synthons from the first initial subset 103a and synthons from a first dummy subset of synthons. Such a dummy subset may be sampled randomly, e.g. from a dataset of possible synthons given the available chemistry of the intermediate. In such an example, the molecule set 104 also includes molecules obtained by enumerating the intermediate 101 with synthons from the second initial subset 103b and synthons from a second dummy subset of synthons. In some cases, the first and second dummy subsets may be the same. Furthermore, in some examples an enumerated molecule set may be obtained that includes both: molecules obtained by enumerating the intermediate with combinations of synthons from the two initial sets 103a, 13b; and molecules obtained by enumerating the intermediate with a synthon from one of the initial sets 103a, 103b and a dummy synthon subset.

[0138] The active learning method 10 then involves sampling or selecting molecules from the enumerated molecule set 104 to obtain a molecule subset 105. This is to reduce the amount of data for each synthon - which will be beneficial for the following, computationally-expensive scoring step - while still testing all of the sampled synthons. For instance, if each initial synthon subset 103a, 103b includes 103synthons, then that would provide 106enumerated molecules while only testing 2x103synthons. However, appropriate sampling from the 106molecule space still allows for 2x103synthons to be tested, but where an appropriate I desired number of data points for each synthon may be selected. For instance, a 2x103synthon space may need only 10 data points each, corresponding to 2x104enumerated molecules (i.e. orders of magnitude fewer than the original 106molecule space). The sampling to obtain the molecule subset 105 may be performed in any suitable manner, e.g. pseudo-randomly, but in a manner that ensures sufficient data points for each synthon are selected.

[0139] Each of the molecules in the molecule subset 105 is then scored according to a defined scoring function or objective function to obtain a plurality of molecule scores 106. The objective function may be for mathematically scoring one or more properties of a molecule relative to one or more respective desired properties of a candidate compound for the particular drug discovery project being undertaken. For instance, the objective function may score a molecule according to its predicted binding affinity after it has been docked with a defined target (protein). The objective function may alternatively or additionally score a molecule based on its shape similarity to the defined target, or based on a multiparameter optimisation (MPO) score. Indeed, the objective function may score a molecule based on a combination of properties such as one or more of: efficacy / potency against a desired target; selectivity against non-desired targets; low probability of toxicity; and, good drug metabolism and pharmacokinetic properties (ADME). The active learning method 10 then involves training surrogate or machine learning models to predict the molecule score as a function of fragments or structural features of synthons I molecules. A different surrogate model is trained for each different reaction vector that is being used to enumerate the intermediate, i.e. for each synthon set that is defined. That is, in the presently described example two surrogate models 107a, 107b are trained.

[0140] The first surrogate model 107a is trained using each of the enumerated molecules from the molecule subset 105 that includes a synthon from the first initial subset 103a (which may be all of the enumerated molecules in the molecule subset 105). In particular, the surrogate model is trained using an input-output pair of data for each respective enumerated molecule, where the input is (fragments I structural features of) the synthon from the first initial subset 103a included in said enumerated molecule, and the output is the molecule score of said enumerated molecule. Note that in the training data each synthon will have a plurality of molecule scores associated therewith, namely, the number of molecules in the initial subset 105 that include the respective synthon. The second surrogate model 107b is trained in a corresponding manner using each of the enumerated molecules from the molecule subset 105 that includes a synthon from the second initial subset 103b.

[0141] In more detail, each synthon includes a number of fragments or structural features that combine to form its chemical structure. Such structural features can be represented in any suitable manner. For instance, one way in which to describe the structure of a synthon or molecule is via fingerprinting. In particular, the fingerprint of a particular synthon may be represented as mathematical objects - e.g. a series of bits or list of integer numbers - that reflect which particular structural features or substructures are present or absent in the synthon.

[0142] There are several different classes of fingerprints, such as topological fingerprints, structural fingerprints, and circular fingerprints. A common circular fingerprinting method is Extended Connectivity Fingerprinting (ECFP). A number of ECFP methods are known, such as ECFP0, ECFP2, ECFP4 and ECFP6. As is known in the art, determining a fingerprint of a synthon will generally include assigning each atom in a compound with an identifier, updating these identifiers based on adjacent atoms, removing duplicates, and then forming a vector from the list of identifiers. In the described example, the input to the surrogate models is the fingerprint vector of the relevant synthon. In other examples, the input to the surrogate models may be a graph representation of a synthon, e.g. when a message passing neural network is used.

[0143] Once the surrogate models 107a, 107b are trained, they are used to infer or predict a molecule score 108a, 108b associated with certain synthons (or synthon fingerprints). The first trained surrogate model 107a is used to predict molecule score 108a for each of the synthons in the first synthon set 102a, and the second trained surrogate model 107b is used to predict molecule score 108b for each of the synthons in the second synthon set 102b. That is, the fingerprint vector or graph representation of a given synthon is input to the relevant trained surrogate model, and the trained surrogate model is executed to output predicted molecule score.

[0144] The use of synthons - rather than building blocks - in the present invention is important as synthons remove chemical moieties that are needed for the chemistry to occur and that will not be included in a product molecule (with the exception of a linker atom, that will be removed). As such, when using the surrogate models to predict the effectiveness of adding I including a particular synthon in an enumerated molecule, only atoms that will present in the final molecule will be considered in the model evaluation, thereby increasing the accuracy of the models. Also, as the linker atom is the same for all synthons for a single enumeration vector, then this does not have a significant signal in the prediction outcome for a given synthon. The use of synthons allows for the exploration of building blocks that come from multiple chemistries on the same footing, thereby increasing the diversity and scope of the explorable chemical space.

[0145] Returning to Figure 1 , the active learning method 10 then selects further subsets of synthons - one for each of the reaction vectors of the intermediate 101 being enumerated - based on the predicted molecule scores. These further subsets of synthons may be final subsets of synthons to be used to enumerate (final) molecules for synthesis and testing. However, more typically, the described steps of the active learning method will be repeated in an iterative manner a plurality of times to eventually obtain the final subsets of synthons. In such an example, first and second further synthon subsets 109a, 109b are selected, and the steps of enumerating 104, sampling 105, scoring 106, retraining / updating the surrogate models 107a, 107b and predicting 108a, 108b are repeated in a corresponding manner (to as described above) for the further synthon subsets 109a, 109b. The selection of the first and second further synthon subsets 109a, 109b may typically involve ranking the synthons in the respective first and second synthon sets 102a, 102b according to their predicted molecule score as determined by the respective first and second trained surrogate models 107a, 107b, and then selecting the highest ranking, i.e. highest scoring, synthons for inclusion in the respective first and second further synthon subsets 109a, 109b. As with the initial synthon subsets 103a, 103b, the further synthon subsets 109a, 109b will typically include a defined number of synthons, and so selecting the further subsets involves selecting said defined number of the highest ranking synthons for inclusion in the respective subsets.

[0146] In some examples, synthons that have previously been selected in an initial or further first or second subset may not be selected again in a subsequent respective further first or second subset. In such examples, selection of a further subset may involve selection of the highest ranked synthons that have not previously been selected. In different examples, synthons may be selected repeatedly across different iterations of the active learning loop.

[0147] The active learning loop 10 may be performed for any suitable number of iterations. For instance, the method may be performed for a prescribed number of iterations, e.g. 5, 10, 15, etc. Alternatively, the method may be performed until an improvement in the molecule scores of synthons in the selected (further) subsets slows or stops. This could be with reference to an average (mean or median) molecule score of synthons in the further subset, or it could be with reference to the top scoring synthon (or certain number of top scoring synthons) in the further subsets.

[0148] Once the active learning loop 10 has been performed for the required number of iterations, the first and second final subsets 110a, 110b are selected. This selection may typically involve selecting the highest scoring synthons according to the most up-to-date respective first and second trained surrogate models 107a, 107b, i.e. the retrained I updated surrogate models according to the most recent I final iteration of the active learning loop 10.

[0149] Figure 2 illustrates the steps of a method 20 in accordance with examples of the invention. The method 20 of Figure 2 may be regarded as incorporating the active learning loop 10 of Figure 1. Steps of the method 20 are performed computationally, i.e. implemented on one or more computer processors. The computer may be in the form of any suitable computing device, for instance one or more functional units or modules implemented on one or more computer processors. Such functional units may be provided by suitable software running on any suitable computing substrate using conventional or custom processors and memory. The one or more functional units may use a common computing substrate (for example, they may run on the same server) or separate substrates, or one or both may themselves be distributed between multiple computing devices. A computer memory may store instructions for performing the methods performed by the computer, and the processor(s) may execute the stored instructions to perform the described method.

[0150] At step 201 , the method 20 involves defining the intermediate molecule 101 and first and second synthons sets 102a, 102b. In the described example, the intermediate molecule 101 has a first enumeration vector and a second enumeration vector. However, in different examples the intermediate may have greater or fewer enumeration vectors. The number of defined synthon sets matches the number of enumeration vectors. The first and second synthons sets 102a, 102b include first and second synthons for enumerating the first and second enumeration vectors of the intermediate 101 , respectively. The intermediate molecule 101 and first and second synthons sets 102a, 102b may be defined as described above. The intermediate molecule 101 and first and second synthons sets 102a, 102b may be stored in, and retrieved from, a database accessible by the computer processor.

[0151] The sets 102a, 102b may include any suitable number of synthons, and the sets 102a, 102b may include the same or a similar number of synthons as one another. For instance, each set 102a, 102b may include at least 105synthons. In some examples, each set 102a, 102b may include at least 106synthons. Indeed, approximately 106synthons may be regarded as being a typical number of reagents per enumeration vector.

[0152] At step 202, the method 20 involves selecting the first and second synthon subsets 103a, 103b. These subsets 103a, 103b include some of the first and second synthons from the first and second synthon sets 102a, 102b, respectively. The subsets 103a, 103b may include a defined number of synthons, i.e. the defined number of synthons may be selected for inclusion in the respective subsets 103a, 103b from the respective sets 102a, 102b. The number of synthons in each subset 103a, 103b may be at least one order of magnitude less than the number of synthons in each respective synthon set 102a, 102b, and typically at least two orders of magnitude less. As a purely illustrative example, in a case in which each synthon set 102a, 102b includes a number of synthons of the order of 106, the number of synthons in each synthon subset 103a, 103b may be of the order of 103.

[0153] Selection of the initial subsets may be performed in different ways. In one example, the selection is performed according to a random selection of the first and second synthons from the synthon sets 102a, 102b. In another example, selection may be performed in a manner that prioritises or maximises the diversity of synthons in the subsets according to a suitable measure of diversity. For instance, synthons may be selected to maximise the number of different types of structural features I fragments (of synthons) present in the relevant subset. Any suitable diversity metric may be used for this purpose, e.g. a score I metric as described in WO 2022 / 084696 A1 .

[0154] In a further example, the synthon subsets 103a, 103b may be selected according to a hypercube sampling strategy. In such an example, the method involves enumerating a molecule (from the intermediate 101) for each different combination of first and second synthons in the respective first and second sets 102a, 102b. A hypercube sample of the obtained plurality of enumerated molecules is then determined. Each of the first and second synthons that are present in the hypercube sampled set of molecules is then included in the respective first and second synthon subsets 103a, 103b. Hypercube sampling is performed as known in the art. In particular, the sampling of molecules is based on the first and second synthons present therein, i.e. the first and second synthons may be regarded as the parameters of the hypercube. That is, in an illustrative example in which a first set includes elements [A1 , B1 , C1 ] and a second set includes elements [A2, B2, C2], the hypercube sample would include [[A1 , A2], [B1 , B2], [C1 , C2]].

[0155] Returning to Figure 2, step 203 involves determining the molecule set 104 comprising a plurality of product molecules each obtained by enumerating the first and second enumeration vectors of the intermediate molecule 101. The enumeration may be performed in different ways. In one example, the molecule set 104 includes each product molecule obtained by enumerating the intermediate 101 with each different combination of first and second synthons in the first and second subsets 103a, 103b. In such a case, if the number of synthons in each of the subsets 103a, 103b is of the order of 103then the number of enumerated product molecules in the molecule set 104 will be of the order of 106. In another example, the molecule set 104 includes each product molecule obtained by enumerating the intermediate 101 with each first synthon from the first subset 103a and each (dummy) synthon from a first dummy synthon set, and each product molecule obtained by enumerating the intermediate 101 with each second synthon from the second subset 103b and each (dummy) synthon from a second dummy synthon set. The first and second dummy sets may be the same, and they may include the same or a similar number of synthons as the first and second synthon subsets 103a, 103b. The first and second dummy sets may be randomly sampled, e.g. from the first and / or second synthons sets 102a, 102b.

[0156] At step 204, the method 20 involves selecting the molecule subset 105 comprising a subset of the product molecules in the molecule set 104. The number of molecules in the molecule subset 105 may be at least one order of magnitude less than the number of molecules in the molecule set 104, and typically at least two orders of magnitude less. The selection may involve selecting a defined number of molecules from the molecule set 104. The selection may be performed according to a random selection of molecules from the molecule set 104. Alternatively, the selection may be performed in a manner that ensures a sufficient spread of different first and second synthons from the first and second synthon subsets 103a, 103b is represented in the molecule subset 105. For instance, the molecule subset 105 may be selected such that at least a defined number of each of the first and second synthons from the first and second synthon subsets 103a, 103b is included in the molecules of the molecule subset 105. Within that constraint, the selection may be performed randomly (which may be regarded as a pseudo-random selection). The defined number may be any suitable number, e.g. 1 , 2, 3, 4, 5, 10, 15, 20, 25, 50 or any other suitable number.

[0157] It is noted that if a hypercube sampling strategy is performed as described above then a separate molecule subset selection step 204 is not needed. Indeed, the hypercube sampling strategy may be regarded as collectively performing steps 202, 203 and 204 of the method 20.

[0158] At step 205, the method 20 involves determining the molecule score for each of the product molecules in the molecule subset 105. As described above, the molecule scores are determined via execution of a defined mathematical scoring I objective function that scores a molecule according to how desirable one or more predicted properties of the molecule is, e.g. relative to a profile of one or more desirable properties of a candidate or ideal molecule (for a particular drug discovery project under consideration). That is, the scoring function can provide a score based on a single predicted property of a molecule, or it can provide a score that is an aggregate of scores for more than one property, e.g. multi - objective scoring. The closer or more proximal a predicted property of a molecule is to a defined, desired (value of the) property, the higher the molecule score determined by the scoring function may be.

[0159] The scoring function is typically computationally expensive to execute. As such, the greater the number of product molecules to score, the higher the computational cost. By first selecting only some of the synthons in the initial synthon sets for molecule enumeration, and then selecting only some of the enumerated molecules for scoring, the number of product molecules to be scored is only a fraction of the entire combinatorial I molecule space that is possible from the starting sets, thereby vastly reducing the computational cost while still providing an effective search of the combinatorial space for desirable molecules.

[0160] Any suitable biological, physical or chemical properties of a product molecule may be scored. Typically, desirable properties may include one or more of: high efficacy / potency against a desired target molecule; selectivity against non-desired targets; low probability of toxicity; and good drug metabolism and pharmacokinetic properties (ADME).

[0161] In one example, a docking score function may be used to determine molecule score. Such a function may provide a molecule score that is indicative of a prediction of a binding affinity of the product molecule when it is docked with a defined target molecule. The defined target molecule may be a protein with which it is desired that a drug molecule interacts in order to provide a therapeutic effect against a specific disease, i.e. the defined target may be known to be associated with the specific disease. An example of a docking function that may be used for these purposes is the OpenEye Hybrid function, as described here: htps: / / docs.evesopen.com / applications / oedockinQ / hvbrid / hybrid.html. However, it will be understood that any suitable docking function / application may be used.

[0162] In another example, a shape similarity function may be used to determine molecule score. Such a function may provide a molecule score that is indicative of a prediction of a shape similarity of the product molecule to the defined target molecule. An example of a shape similarity function that may be used for these purposes is the OpenEye ROCS function, as described here: htps: / / docs.eyesopen.com / applications / rocs / . In such an example, the molecule score may be a TanimotoCombo (ColorTanimoto + ShapeTanimoto) score as defined here: htps: / / docs.eyesopen.com / applications / rocs / rocs / rocs report file.html. Again, it will be understood that any suitable shape similarity function I application may be used.

[0163] Returning to Figure 2, at step 206 the method 20 involves training the surrogate models 107a, 107b for describing a relationship between synthons and predicted molecule scores. A different surrogate model is trained for each set 102a, 102b of synthons that is defined (or, equivalently, for each enumeration vector that is to be used for elaboration of the intermediate 101). That is, in the described example two surrogate models are trained. In particular, the first surrogate model 107a is configured to take first synthons as input and provides a molecule score as output. Correspondingly, the second surrogate model 107a, 107b is configured to take second synthons as input and provides a molecule score as output.

[0164] The surrogate models 107a, 107b are trained using the product molecules in the molecule subset 105 and their associated molecule scores. In particular, the first surrogate model is trained using each of the product molecules in the molecule subset 105 that includes one of the first synthons in the first synthon subset 103a, which in some examples will be all of the molecules in the molecule subset 105. The second surrogate model is trained in a corresponding manner using molecules that include one of the second synthons from the second synthon subset 103b. As described above, each surrogate model may take as input a fingerprint representation of a synthon and output a value, e.g. a one-dimensional value I number, indicative of molecule score. As such, the first surrogate model is trained using, as input-output pairs, first synthons and corresponding molecule scores. Similarly, the second surrogate model is trained using, as input-output pairs, second synthons and corresponding molecule scores.

[0165] The surrogate models 107a, 107b may typically be machine learning (ML) models. One example of such a model is a random forest (RF) model. The skilled person would readily understand how to implement an RF model for these purposes; however, one way to implement this is using the scikit learn RF model found here: learn.org / stable / modules / qenerated / skleam. ensemble. RandomForestClassifier.html.

[0166] Another example of an ML model that may be used is a feed-forward neural network (NN) model. Again, the skilled person would readily understand how to implement an NN model for these purposes; however, one way to implement this may be found here: htps: / / Qithub.eom / colevQroup / molpal / blob / main / molpal / models / nnmodels.py#L32.

[0167] A further ML model that may be used is a message-passing neural network (MPN) model. Once again, the skilled person would readily understand how to implement an MPN model for these purposes; however, one way to implement this may be found here: htps: / / Qithub.com / colevQroup / molpal / blob / main / molpal / models / mpnmodels.py.

[0168] Returning to Figure 2, after the surrogate models 107a, 107b have been trained, at step 207 the method 20 involves executing the trained first and second surrogate models to obtain (predicted) molecule scores 108a, 108b for different respective first and second synthons. As the surrogate models 107a, 107b are trained on fingerprints or graphs describing structural features I fragments of synthons then the surrogate models predict molecule score as a function of structural features present in synthons. This model execution step 207 may involve obtaining predicted molecule scores for each of the first and second synthons in the first and second synthon sets 102a, 102b (using the first and second surrogate models, respectively).

[0169] Referring again to Figure 2, the method 20 involves looping back to step 202 after step 207 to repeat steps 202-207 a plurality of times as part of an iterative process (as illustrated in Figure 1). In particular, when the process loops back to step 202, further first and second synthon subsets 109a, 109b are selected. These further subsets 109a, 109b may be selected in different ways. For instance, the further subsets 109a, 109b may be selected in the same manner as the initial subsets 103a, 103b, e.g. randomly or pseudo-randomly from the synthon sets 102a, 102b. Alternatively, the further subsets 109a, 109b may be selected based on the molecule scores 108a, 108b determined at step 207. For instance, the first and second synthons having the highest molecule scores associated therewith may be selected for inclusion in the further first and second subsets 109a, 109b, respectively, i.e. greedy acquisition of the synthons. Furthermore, in some examples synthons in the synthon sets 102a, 102b may be selected for inclusion in the subsets 103a, 103b, 109a, 109b on more than one occasion, i.e. the same synthon can be selected in a plurality of iterations of the method I loop. Such a strategy may be referred to as a ‘no drop’ or ‘with replacement’ strategy. Alternatively, in other examples, once a particular synthon has been selected for inclusion in a particular (initial or further) subset 103a, 103b, 109a, 109b it is removed from the pool of synthons that may be selected as part of subsets in subsequent iterations of the method I loop. Such a strategy may be referred to as a ‘drop’ or ‘without replacement’ strategy.

[0170] Steps 203-207 of the method 20 are then repeated in a similar manner described above for the further synthon subsets 109a, 109b selected at step 202. This process is repeated in an iterative manner until a stop condition is satisfied. For instance, the stop condition may be that the repeated steps 202-207 have been performed a defined number of times, e.g. 1 , 2, 3, 4, 5, 10, 20, 50, 100 or any suitable number of times. Note that in some examples, this means that the method steps 202-207 are performed only once.

[0171] The stop condition may alternatively, or additionally, include that a rate of improvement of the molecule scores of the top scoring synthons obtained from the surrogate models falls below a threshold rate of improvement. This may be determined as comparing an average I mean molecule score of a defined number of the top scoring synthons against a threshold molecule score.

[0172] Referring again to Figure 2, once the stop condition is satisfied the method 20 involves, at step 208, selecting the final first and second synthon subsets 110a, 110b. The final first and second subsets 110a, 110b may include a defined number of first and second synthons, respectively, e.g. a same or similar number of synthons selected in the initial and further synthon subsets 103a, 103b. The final subsets 110a, 110b are selected based on the molecule scores determined at a final I most recent I current iteration of the method 20 (or active learning loop 10). The final subsets 110a, 110b are typically selected to include synthons associated with the highest molecule scores (as determined at step 207 of a current iteration). For instance, step 208 may involve performing a greedy acquisition of (a defined number of) first synthons according to predicted molecule score using the first surrogate model 107a to obtain the final first subset 110a. Similarly, and in addition, step 208 may involve performing a greedy acquisition of (a defined number of) second synthons according to predicted molecule score using the second surrogate model 107b. The final subsets may be filtered to remove certain synthons, such as duplicates or synthons having undesirable properties (perhaps not scored by the scoring function), e.g. associated with high toxicity levels, so that the final subsets may not be the absolute highest-scoring molecules, but rather the highest-scoring molecules that are not otherwise removed I deselected for other reasons.

[0173] At step 209, the method 20 involves determining one or more high-scoring product molecules using the final first and second synthon subsets 110a, 110b. In particular, high- scoring product molecules are generated by enumerating the first and second enumeration vectors of the intermediate 101 with different combinations (e.g. all combinations) of the first and second synthons in the final first and second synthon subsets 110a, 110b. In other words, the best predicted synthons in each set are used to enumerate product molecules. As these high-scoring product molecules have synthons associated with high molecule scores then there is a higher probability of the molecules having desired properties, e.g. of a candidate molecule. Indeed, these enumerated molecules will typically then be scored with the objective function.

[0174] The method 20 may then involve taking forward some or all of the generated high-scoring product molecules forward for synthesis and testing, e.g. the molecules scoring highest against the objective function. Some of the generated high-scoring molecules may be filtered out of consideration for synthesis, e.g. because they have one or more undesirable features, such as they have previously been synthesised ad tested, that they are too similar to one or more previously synthesised or not-yet-synthesised molecules, associated with high toxicity, etc.

[0175] The method 20 may (further) score the high-scoring molecules using the defined scoring function or a trained machine learning model, e.g. a neural network, that scores product molecules according to predicted property scores. The top ranked subset of these molecules according to this scoring may then be selected for synthesis and testing.

[0176] Typically, only a relatively small number of molecules may be synthesised and tested at each cycle I iteration of a drug discovery process I project to identify candidate compounds. A certain number of the high-scoring molecules may be selected for testing, e.g. using one or more of the approaches described above, or randomly. The number of high-scoring molecules may be less than 100, less than 50, less than 25, less than 10, less than 5, or any other suitable number.

[0177] The method 20 may then involve synthesising the selected high-scoring product molecules, and testing the synthesised high-scoring product molecules to determine one or more biological properties of the high-scoring product molecules. Note that the steps of synthesising and testing the high-scoring molecules may not be computer implemented steps (unlike the other steps of the method 20). The biological properties tested may include the property(ies) predicted by the scoring functions. The way in which the high- scoring molecules being synthesised and tested are selected, i.e. according to the steps of the method 20 described above, means that the synthesised molecules are more likely to exhibit desired properties in line with those of a candidate compound for the particular drug discovery process under consideration.

[0178] The results obtained during synthesis and testing may be used to retrain one or more models used to select a next set of molecules to be synthesised and tested, as part of a feedback, iterative process. This further increases the likelihood of suitable candidate compounds being identified in fewer synthesis and design loops, thereby saving significant levels of cost and time.

[0179] In one example, the biological properties of molecules determined during testing are used to retrain the first and second surrogate models 107a, 107b ahead of a next run of the method 20. In another example, the biological properties of molecules determined during testing are used to retrain a separate machine learning model that predicts molecular property scores as a function of structural features of molecules, e.g. the machine learning model used to select which high-scoring molecules from the final sets 110a, 110b to synthesise. A next batch of selected molecules for synthesis then benefits from more accurate property predictions, thereby increasing the probability of identifying candidate compounds more quickly.

[0180] Specific examples of the described method are now described. As an initial evaluation, two vector positions were enumerated from an intermediate fragment. Deduplicated synthons were created and randomly sampled from the open-source MCule reagent database (mcule.com / database / ) given reaction compatibilities. Each synthon set contained a specific number, e.g. one thousand, of randomly selected synthons, giving an enumerated space of product molecules, e.g. one million product molecules when two sets of one thousand synthons are used.

[0181] The active learning cycle was tested and varied in three distinct experiments to obtain the most optimal configurations: the selection strategy for acquiring synthons for subsequent rounds and whether synthons should be considered for replacement; the number of synthons considered while selecting; and the model used for training. These methods were evaluated on two scoring functions to show applicability across both structure and ligandbased scoring methods.

[0182] Data acquisition for each round I iteration was 0.2% of the total space over five rounds, meaning the total space acquired at the end of the active learning cycle was 1 %. The first round of synthons was acquired through random acquisition. The term ‘acquisition pool’ refers to the set I pool of synthons that can be selected from during acquisition. In a ‘no replacement’ (or ‘drop’) strategy I method, once a synthon is selected from the acquisition pool, it is removed and is no longer available to make product molecules (in subsequent rounds I iterations). In a ‘replacement’ (or ‘no drop’) strategy, selected synthons are returned to the acquisition pool and can be selected in subsequent rounds I iterations.

[0183] To interpret the results from these studies, an enrichment factor (EF) is considered, where the EF is taken to be the percentage of top scoring molecules found in the model-guided enumeration over the percentage of top scoring molecules found in a strategy in which random acquisition with no prediction was performed. These scores are based entirely on the output of the enumerated space of the top synthons and will not include those molecules that are in the training data.

[0184] CDK2 (6GLIH) was used as the model protein system in these experiments. The ligand is Cc1 ncc(-c2ccnc(Nc3ccc(S(C)(=O)=O)cc3)n2)n1C(C)C . The core scaffold ( Nc1ncccn1 ) was used as the ‘fragment hit’ or intermediate in this experiment as it is ‘fragment-like’ and could be identified as a fragment that made favourable interactions with the protein target in the binding domain. Moreover, this ligand had two defined vectors for growth, as indicated in Figure 3. Enumeration vectors are off the aryl chloride and amine.

[0185] Once the intermediate building block was decided on, reaction compatibility for each vector position was identified. The reactions being considered were not ring-forming and all reactions are from a single position on the building block. There are some compatible reactions that enumerate off more than one vector; however, these reactions were not chosen during this experiment. Table 1 shows examples of compatible reactions for each of the two enumeration vectors in this experiment, and then which of those compatible reactions are single-vector reactions that may be used in the present case.

[0186] Table 1

[0187] This fragment substructure was searched in the database MCule building block catalogue and the building block Nc1 nccc(CI)n1 was identified as a suitable intermediate by medicinal chemists for elaborating because of the two distinct vectors for growth: the chlorine on the pyrimidine core and the primary amine. The positions of both these atoms were identified as vectors in a previous assessment. The complementary building blocks for these reactions to the intermediate were retrieved from MCule. The building blocks and intermediate were then processed to translate them into synthons. Vector 1 , the amine vector on the intermediate and its complementary building blocks, received a II atom at its connection point. Vector 2, the aryl-chloride vector on the intermediate and its complementary building blocks, received an Np atom. This is illustrated in Figure 4. This created two synthon sets for enumeration. The synthons were deduplicated (using SMILES) for each synthon set. Example synthons are illustrated in Figure 5. To form the product molecule, the intermediate synthon is simply combined with the complementary synthons at the atoms II and Np. One example would be as follows:

[0188] Intermediate synthon: [U]Nc1 nccc([Np])n1

[0189] Vector 1 synthon: COC(=O)[C@H]1CC[C@H]([U])CC1 (Chan-Lam Amine)

[0190] Vector 2 synthon: COC1 (C[Np])CCNCC1 (Negishi)

[0191] Product molecule: COC(=O)[C@@H]1 CC[C@@H](Nc2nccc(CC3(OC)CCNCC3)n2)CC1

[0192] As mentioned above, this experiment started by randomly selecting a specific number of synthons, e.g. 1000, for each vector, creating two synthon sets of the specific number of synthons each. These synthons sets are considered the acquisition pool, and if synthons are not replaced after selection, they will be dropped from this pool so that no other enumerated product molecule will have that synthon. The total enumerated space for this experiment when the synthons sets each have 1000 synthons is 1 ,000,000 product molecules.

[0193] In the first step of the active learning cycle, a specific number of synthons, e.g. 100, from each synthon set were acquired randomly. In the ‘no replacement’ method, once these synthons were selected for the next round, the synthons were not replaced back into the acquisition pool. These subsets of synthons were enumerated exhaustively to create a molecule set of a specific number of product molecules, e.g. 10,000 when 100 synthons were acquired from each set. A specific number of product molecules were then sampled randomly from this molecule set to obtain a molecule subset, e.g. 2000 molecules were sampled when 10,000 molecules were acquired.

[0194] The molecules in the molecule subset were scored with a given objective function or scoring function. Two scoring methods were used: docking and shape similarity. OpenEye Hybrid (htps: / / docs.evesopen.com / applications / oedockinQ / hvbrid / hybrid.html) was used as the docking function. In particular, product molecules went through a series of preparation and conformer generation before docking. The preparation step uses OpenEye tools to complete tautomer generation and stereo-chemistry control. A maximum of five conformers per molecule were created. The docking step then uses OpenEye Hybrid to dock five poses. The best score for each molecule is used as the docking score for that molecule.

[0195] OpenEye ROCS (htps: / / docs.eyesopen.com / applications / rocs / ) was used as the shape similarity function. Similarly to the docking procedure, OpenEye tools were used to prepare and create conformers for product molecules. A maximum of five conformers per molecule were created. ROCS was then used to complete shape similarity scoring. A Tanimoto ComboScore (docs.eyesopen.com / applications / rocs / rocs / rocs_report_file.html), which is ColorTanimoto + ShapeTanimoto, was used as the scoring metric for these experiments. The best score for each molecule is used as the shape similarity score for that molecule.

[0196] A surrogate model was trained for each synthon set, i.e. two models in total in this case, modelling synthon fingerprint to the score the product molecules received. The models were trained on all data points, and not just those acquired in the given round. Each model was used to predict the best synthons for its respective synthon set.

[0197] Two acquisition strategies were tested: greedy and random. In the greedy acquisition runs, a specific number, e.g. 100, of the top predicted synthons from each set would be selected for the next round. In the random acquisition runs, synthons would be randomly selected for the next round and would not consider the model predictions.

[0198] Two replacement strategies were tested: ‘replacement’ and ‘no replacement’, as defined above. In the replacement runs, prediction occurred on all synthons. In the no replacement runs, prediction occurred only on synthons that have not been selected yet (i.e. those still in the acquisition pool) and once synthons were selected for the next round, the synthons were dropped from the acquisition pool. This process then repeated at the enumeration step for five rounds. This gives a total of 10,000 product molecules scored in the example in which 2000 molecules are sampled for the molecule subset. Importantly, product molecules were never rescored.

[0199] Three different surrogate models were tested: random forest (RF), feed-forward neural network (NN) and message-passing neural network (MPN). The RF method used in these experiments was the ‘Random ForestRegressor’ found in the ‘SKLearn’ python package. The ‘n_estimators’ was set to 100 and ‘max-depth’ was set to none. All default parameters were used in these experiments. For fitting the model, the X values were the bits from the synthon fingerprint and the Y value was the Docking and Tanimoto ComboScore.

[0200] The NN used in this experiment was derived from the software package MolPAL (molecular pool-based active learning), which is for batched, Bayesian optimisation in a virtual screening environment. The underlying architecture is a ‘keras. Sequential’ model with a batch size of 4096 and two fully connected layers of size 100 each with an output size of 1 . The hidden layers used the ReLLI activation function. These were the default values in the MolPAL repository. The X values were an array of bits from the synthon fingerprint. The Y value was the Docking and Tanimoto ComboScore for their respective experiments.

[0201] The MPN used in these experiments was the directed message-passing neural network (D-MPNN) and was based on the implementation found in MolPal. The underlying model consists of both a PyTorch message-passing neural network architecture and the MoleculeModel class from the Chemprop library. The default values found in MolPAL were used in this work: a batch size of 50, a ReLLI activation function, a learned encoded representation of dimension 300, and the message-passing phase fully connected to an output layer of size 1 .

[0202] The synthons and product molecules were translated into feature fingerprints for X input into both the RF and NN models. 2048-bit morgan fingerprints with a radius of 2 using RDKit’s ‘GetMorganFingerprintAsBitVect’ function was used. The MPN used its own encoded molecular graph representation for learning molecular features.

[0203] Figures 6(a) and 6(b) illustrate an example of how well the models can retrieve synthons that, when enumerated together, produce top-scoring molecules. In this example, each of the two synthon sets has 1000 synthons, and the RF surrogate model was used. Each combination of the different replacement (replacement I no replacement) and acquisition strategies (greedy / random) was tested. Figure 6(a) shows results when the docking scoring function is used, and Figure 6(b) shows results when the shape similarity scoring function is used. At the end of each round when model training is complete, the models predict the top 250 synthons from each synthon set. These synthons are then enumerated exhaustively on the intermediate. This creates a product space of about 62,500 molecules, or 6.25% of the total enumerated space. The percentage of top-1000 scores found in this product space relative to the 1 ,000,000 molecule product space is then found.

[0204] Figures 6(a) and 6(b) then show plots of the percentage of top-1000 scoring molecules acquired against number of molecules in the training data for each of different strategy combinations, namely, ‘replacement, greedy’ 61a, 61b, ‘replacement, random’ 62a, 62b, ‘no replacement, greedy’ 63a, 63b, and ‘no replacement, random’ 64a, 64b. Also plotted in Figures 6(a) and 6(b) is the results of a ‘random acquisition, no prediction’ strategy (plots 65a, 65b) which shows the product space that is retrieved if no model was used to predict top synthons and synthons were instead randomly retrieved from the synthon pool. Error bars are included in Figures 6(a) and 6(b), and represent the standard error across 30 runs.

[0205] In the docking trials, the ‘no replacement, greedy’ and ‘replacement, greedy’ are the best performing acquisition functions, within error of each other, with EF values of 11.48 and 11.56, respectively. The ‘no replacement, random’ and ‘replacement, random’ functions perform slightly worse with EF values of 9.63 and 9.18, respectively. In the shape similarity trials, the ‘replacement, greedy’ function performs slightly, but significantly, better than the ‘no replacement, greedy’ function, with RF values of 16.01 and 15.63, respectively. The ‘no replacement, random’ and ‘replacement, random’ models perform worse with EF values of 14.54 and 14.26, respectively.

[0206] Figures 7(a) and 7(b) illustrate an example comparing different numbers of synthons being acquired at each acquisition step. There were two competing ideas here. With more synthons being acquired in each step, there would be fewer data points per synthon (given the same number of molecules scored), meaning synthons could be more easily misrepresented given an outlier, and the amount of synthon space explored would be larger. However, with fewer synthons, the robustness of data for a given synthon would be increased, but the amount of synthon space explored would be decreased. Thus, this experiment tests the optimal number of synthons for acquisition. The results illustrated in Figures 7(a) and 7(b) were obtained using the ‘no replacement, greedy’ strategy, given the statistically significant result obtained in the shape similarity plot in Figure 6(b). A random forest model was used as the model for training in this experiment. Moreover, hypercube enumeration is used for improved computational efficiency and stratified sampling efforts. The number of synthons chosen were 50 (plots 71a, 71 b), 100 (plots 72a, 72b), 150 (plots 73a, 73b) and 200 (plots 74a, 74b). As there are five acquisition rounds with 1000 synthons in each set, and the ‘no replacement’ method is being implemented, the maximum number of synthons selected for each round was 200.

[0207] Like in Figure 6, Figures 7(a) and 7(b) show the percentage of top-1000 scoring molecules found when enumerating the predicted top 250 synthons from set one and set two using the surrogate model trained on the molecules in the training data. Also like Figure 6, the efficacy of the models to predict the top scoring synthons was evaluated on docking and shape similarity functions for Figures 7(a) and 7(b), respectively. Again like Figure 6, the results obtained from the ‘random acquisition, no prediction’ strategy are shown in Figures 7(a) and 7(b) as plots 75a, 75b, respectively.

[0208] The general trend is that as the number of synthons that are explored increases, the models improve their ability to predict top scoring molecules. However, a surprising deviation from this trend happens as the number of synthons increases from 150 to 200. That is, the percentage of top scoring molecules found in the ‘150 synthons’ selection method in the docking trials is 70.1% (RF = 12.03), while for the ‘200 synthons’ selection method it is 66.2% (RF = 11.35). Similarly, in the shape similarity trials the percentage of top scoring molecules found for the ‘150 synthons’ selection method is 96.5% (RF = 16.55), while for the ‘200 synthons’ selection method it is 95.73% (RF = 16.42). Both the docking and shape similarity percentages are significantly different.

[0209] Thus, it seems that for this data the ‘150 synthons’ selection method is statistically the most performant for both docking and shape similarity. As mentioned above, this could be explained by the difference in number of data points for a given synthon that was tested. In the ‘200 synthons’ selection method about 10 scores are associated with each synthon. However, in the ‘150 synthons’ selection method there are about 3.3 more scores per data point. This difference in number of data points per synthon could be especially significant given the relatively small size of the synthon space, meaning any outliers for a given synthon could have a significant impact on how that synthon is interpreted by the models and how synthons are retrieved in future selection rounds.

[0210] Figures 8(a) and 8(b) illustrate an example of how different surrogate model architectures influence the effectiveness of the method to retrieve synthons that produce top scoring product molecules. Three different surrogate model architectures were tested: random forest (RF), feed-forward neural networks (NN), and message-passing neural networks (MPN). The NN and MPN models originate from MolPAL (with minor adjustments to handling fingerprints for the NN) which has been tuned for and shown efficacy in dealing with molecular learning problems.

[0211] All of the experiments in this example were performed using the ‘greedy, no replacement’ strategy with 150 synthons selected at each iteration I run, i.e. the best performing combination of methods from Figures 6 and 7. Like in Figures 6 and 7, Figures 8(a) and 8(b) show the percentage of top-1000 scoring molecules found when enumerating the predicted top 250 synthons from set one and set two using the surrogate model trained on the molecules in the training data. Also like Figures 6 and 7, the efficacy of the models to predict the top scoring synthons was evaluated on docking and shape similarity functions for Figures 8(a) and 8(b), respectively. The RF results are plots 81a, 81 b, the NN results are plots 82a, 82b, and the MPN results are plots 83a, 83b. Again like Figures 6 and 7, the results obtained from the ‘random acquisition, no prediction’ strategy are shown in Figures 8(a) and 8(b) as plots 84a, 84b, respectively.

[0212] Here, it is observed that the MPN performs significantly better than the NN and RF models, with the percentage of top scoring molecules found in docking being 76.8%, 72.1 % and 70.1 %, respectively, while for shape similarity the percentages were 97.7%, 97.2% and 96.5%, respectively. Perhaps most intriguing is the performance of the MPN model for the first acquisition batch, which achieves 89.4% of the top scoring molecules by only acquiring 0.2% of the total enumerated space for building the models. The MPN is the most performant surrogate model tested.

[0213] Throughout the testing to find ideal selection and modelling conditions, there was a significant improvement in the percentage of top-1000 scoring molecules acquired. That is, for docking, the best method, namely, the ‘greedy, no replacement’ method, acquired 66.9% of the top-1000 scoring molecules. In the last trial when testing the best surrogate model, the MPN model acquired an average of 76.8% of the top-1000 scoring molecules for docking. Similarly for shape similarity, there was an improvement from 93.3% acquisition of the top-1000 scoring molecules in the ‘greedy, no replacement’ method to 97.7% in the MPN method. This is a 14.8% and 4.7% improvement in acquisition efficiency in docking and shape similarity, respectively, which was accomplished by consecutive optimisation of selection and modelling methods throughout these three experiments.

[0214] In all three experiments, the model was able to find more top-scoring molecules in the shape similarity metric compared to the docking method. The number of unique synthons in top-1000 scoring molecules is greater for docking than it is for shape similarity. For instance, for the examples in which the enumerated molecule space is 1 ,000,000 molecules, there are 226 unique synthons in the top-1000 scoring molecules for set one and 130 unique synthons for set two. Conversely, for shape similarity, there are 106 unique synthons in set one and 103 unique synthons in set two. To assess the success of the models’ ability to prioritise synthons that produce high scoring molecules, the top 250 synthons were chosen to be exhaustively enumerated with each other. Thus, given that there are more unique synthons in the top-1000 scoring molecules in the docking experiments, the probability of acquiring all synthons in that top 250 is lower, even if the model is able to learn shape similarity and docking at the same rate.

[0215] Some further specific examples of the method in accordance with the invention are now described. As described above, the method finds synthons that, when enumerated together, create top scoring molecules without having to score the entire molecule space. Importantly, by reducing the number of synthons that need to be enumerated together, a decrease in the entire enumerated space by orders-of-magnitude is achieved, thereby decreasing computational costs by a similar degree.

[0216] Figures 9(a) and 9(b) illustrate an example of how different combinations of enumeration and selection strategies influence the effectiveness of the method to retrieve synthons that produce top scoring product molecules. Similarly to the above examples, Figures 9(a) and 9(b) show the percentage of top scoring enumerated molecules acquired by only selecting 250 synthons from each of two synthon sets (out of 1000 total sampled synthons in each set) from what the model deems to be the ‘best’ synthons, given a certain number of molecules that have already been learned upon. The model used here is a random forest model as described above, and Figures 9(a) and 9(b) score the molecules using a shape similarity function and a docking function, respectively, as described above. The x-axis shows the number of scored molecules added to the model training data. After each round, the model predicts the best 250 synthons in each synthon set, enumerates those synthons together (with the intermediate), scores the enumerated product, and finds the overlap in the top 1000 scored molecules between this subset of 62,500 (250 x 250) and the top 1000 scored molecules in the 1 ,000,000 (1000 x 1000) total enumerated molecules. This overlap is communicated as a percentage on the y-axis.

[0217] Two different enumeration strategies are tested. In the ‘random enumeration’ strategy, the selected synthons from one synthon set are enumerated with synthons from a randomly sampled group of synthons. In the ‘best synthons enumeration’ strategy, the selected synthons from one synthon set are enumerated with the selected synthons from the other synthon set. Two different selection strategies are selected, namely, the ‘drop’ (‘no replacement’) and ‘no drop’ (‘replacement’) strategies outlined above.

[0218] In Figures 9(a) and 9(b), there are shown plots of: a ‘no drop, random enumeration’ strategy 91a, 91b; a ‘no drop, best synthons enumeration’ strategy 92a, 92b; a ‘drop, best synthons enumeration’ strategy 93a, 93b; and a ‘drop, random enumeration’ strategy 94a, 94b. Also shown is a ‘random acquisition, no prediction’ strategy 95a, 95b, which shows results if no active learning loop was used and synthons were acquired randomly, and for which it may be seen that the percentage of top scoring molecules would be acquired when enumerated together (here -6.25%). Computationally, 92.75% fewer molecules were scored than in an exhaustive workflow (250 synthons x 250 synthons = 62,500 molecules, or 1000 synthons x 1000 synthons = 1 ,000,000 molecules, 6.25% of the total space scored plus the 1% scored in the active learning loop). For shape similarity (Tanimoto combo score), the best model achieves an Enrichment Factor (EF), i.e. the ratio of the percentage of top-k scores found by the model-guided search to the percentage of top-k scores found by a random search for the same number of objective function calculations, of 14.3.

[0219] Figures 10(a) and 10(b) illustrate an example of how different enumeration I sampling strategies influence the effectiveness of the method to retrieve synthons that produce top scoring product molecules. Again similarly to above, an RF model is used, and two synthon sets of 1000 synthons are defined, with 250 synthons being selected from each set. A ‘drop’ selection strategy was used throughout this experiment, and molecules were again scored based on shape similarity (Figure 10(a)) and docking (Figure 10(b)). Figures 10(a) and 10(b) show plots of: a ‘drop, hypercube best enumeration’ strategy 1001a, 1001b which uses hypercube sampling; a ‘drop, random selection’ strategy 1002a, 1002b which does not use any prediction method and randomly selects the synthons; a ‘drop, best synthons enumeration, hypercube’ strategy 1003a, 1003b which tests all synthons on a first run using hypercube sampling, and does not drop them, but then drops them in subsequent runs; a ‘drop, best synthons enumeration’ strategy 1004a, 1004b which is similar to the ‘drop, hypercube best enumeration’ strategy, but employs exhaustive enumeration and sampling; and a ‘drop, random enumeration' strategy 1005a, 1005b. Also shown is a ‘random acquisition, no prediction’ strategy 1006a, 1006b, which shows results if no active learning loop was used, and synthons were acquired randomly. The best synthons enumeration and hypercube best enumeration results are relatively similar, as expected, but the hypercube approach uses less computational time as enumerating every combination is not necessary.

[0220] Figures 11 (a) and 11 (b) illustrate an example of how different comparing different numbers of synthons being acquired at each acquisition step. An RF model is used along with a ‘drop, hypercube sampling, greedy acquisition’ strategy throughout, as well as shape similarity (Figure 11 (a)) and docking (Figure 11 (b)) scoring methods. Figures 11 (a) and 11 (b) show plots in which 50 (plots 1101a, 1101b), 100 (plots 1102a, 1102b), 150 (plots 1103a, 1103b) and 200 (plots 1104a, 1104b) synthons are tested at each round, with five rounds being executed in each case. Note that in Figures 9 and 10, 100 synthons were tested at each round. Figures 11 (a) and 11 (b) also show the ‘random acquisition, no prediction’ strategy 1105a, 1105b.

[0221] Figure 12 illustrates results of a further experiment using the active learning loop of Figure 1. Figure 12(a) illustrates the chemical space generated for the CDK2 intermediate used in this example. Figures 12(b) and 12(c) represent acquisition of the top 100 molecules by scoring 6% of the total space. The ‘SAL’ lines 1201a, 1201b represent drop (no replacement), greedy, 100 synthons being sampled in each round (with 2000 molecules being scored each round), and Random ForestRegressor as the model run. The ‘SAL optimized’ lines 1202a, 1202b represents the best model obtained through parameter optimization. It uses drop (no replacement), greedy, 150 synthons being sampled in each round, and the message-passing neural network. It also enumerates compounds in a piece-wise linear fashion, which essentially enumerates better scoring synthons than those that do not score as well. These two lines, SAL and SAL optimized, can be compared to TS 1203a, 1203b which stands for Thompson sampling, which is another acquisition technique that attempts to achieve the same goal (i.e. acquire the top-scoring molecules with scoring as little space as possible). Figures 12(b) and 12(c) show SAL performance on a benchmarking set of one-million molecules as measured by the top 100 molecules acquired as a function of the number of molecules in training data. The total enumerated molecules after training and final enumeration equals 6% of the total chemical space. The Figure 12(b) shows acquisition of SAL run with Hybrid Docking as its objective function, and Figure 12(c) shows acquisition of SAL run with ROCS TanimotoCombo. The SAL run represents best active learning model from first trial. SAL (optimized) represents best model found through several rounds of parameter optimization. Thompson sampling (TS) does not use the same training data technique as SAL, but results show the final percentage if top 100 molecules acquired after scoring 6% of benchmarking set.

[0222] Figures 12(d) and 12(e) show the performance of the best active-learning method on various sizes of chemical space for Docking and ROCS TanimotoCombo, respectively. The percentile scores for the 5000 top-ranking molecules are plotted. Each run was allocated one million objective function calculations. Thus, for the space of 106molecules, all molecules were evaluated exhaustively while the larger chemical space runs used the active-learning method. It is seen that the method is able to consistently find better molecules when exploring larger spaces. Figures 12(f) and 12(g) show the distributions of molecule scores for each chemical space found in the experiments described in Figures 12(d) and 12(e), respectively.

[0223] Experimentation to illustrate the effectiveness of the described active-learning method when the molecule scoring function is a multi-parameter optimisation function has also been performed. Indeed, it has been shown that the molecules that are produced from this method are "drug-like" (i.e. rule of 5 like principles) and meet specific ADMET and synthesizability metrics. Figures 13(a), 13(b), 13(c) show scatter plots illustrating multiparameter scores of molecules obtained using the active learning method with a suitably defined multi-parameter scoring function (indicated by 1301a, 1301b, 1301c), compared against the multi-parameter scores of molecules obtained by random acquisition (indicated by 1302a, 1302b, 1302c), over a large chemical space of molecules, e.g. 1012molecules. Figures 13(a), 13(b), 13(c) respectively show results across the three protein systems CDK2, BACE1 , D2. Specifically, Figures 13(a), 13(b), 13(c) show plots of molecule QED (Quantitative Estimate of Drug-Likeness) score against docking score (as defined above) for molecules obtained by the different approaches. QED score was determined as described here: www.nature.com / articles / nchem.1243. Figures 14(a), 14(b), 14(c) also show scatter plots illustrating multi-parameter scores of molecules obtained using the active learning method with a suitably defined multi-parameter scoring function (indicated by 1401a, 1401 b, 1401c), compared against the multi-parameter scores of molecules obtained by random acquisition (indicated by 1402a, 1402b, 1402c), over a large chemical space of molecules, e.g. 1012molecules. The same three protein systems were considered as in Figure 13. Figure 14 differs from Figure 13 in that the multi-parameter scoring function in Figure 14 used a combination of QED score and TanimotoCombo score (as defined above). It may be seen in Figures 13 and 14 that the described active learning method shows superior performance in all cases, as measured by the improved Pareto front formed between the relevant 3D scoring function and QED, relative to the baseline.

[0224] In the above-described examples, a starting material referred to as an intermediate molecule that has one or more defined enumeration vectors is provided / defined. The methods of the invention are used to select synthons to enumerate the one or more enumeration vectors with, in order to obtain product molecules that are more likely to exhibit desired properties with reference to a desired property profile of a drug to be designed. In other examples, the described methods of the invention may be used to select synthons or building blocks to be combined with intermediate molecules at each step of a defined synthetic route / pathway in order to obtain more desirable product molecules. Indeed, ‘more desirable’ refers to product molecules that are not only more likely to exhibit desired properties (binding affinity, selectivity, etc.), but that are also more likely to be readily synthesisable. This is described in greater detail below.

[0225] As outlined above, synthesizability is a common issue with generative methods given that many do not have an innate awareness of building blocks or chemical reactions that are available. Indeed, while computational approaches are often used to design and optimise hit molecules in drug discovery, they often undervalue or even ignore synthetic feasibility. This can result in the synthesisability of compounds being design in silico being relatively low.

[0226] In examples of the present invention, synthesisability is imposed on product molecules being obtained by enumerating molecules based on a known / defined synthetic route / pathway and known building blocks (or synthons). The combinatorial nature of combining / substituting different building blocks at (a plurality of) different chemical reaction steps means that enumerating and scoring all possible molecules becomes intractable. Using the methods of the invention as described above to select which building blocks I synthons to use at each chemical reaction step can therefore assist in identifying desired molecules without needing to enumerate all possible molecules.

[0227] In more detail, in these described examples molecules are generated efficiently based on known synthetic routes and optimised for defined scoring / reward functions (such as the scoring functions described above). In particular, using the active learning approach of the methods of the present invention to select the most promising building blocks I synthons means that top-scoring molecules are identified while only applying the (expensive) scoring function to a small fraction of the total combinatorial space. This approach can be used for both ligand-based and structure-based drug discovery. For a given synthetic route, it can be chosen to use the methods of the invention to select synthons at all, or only some of, the chemical reaction steps of the synthetic route. For instance, in a synthetic route having five chemical reaction steps it may be chosen to select synthons using the described method at only two or three of the steps.

[0228] As is known in the art, a synthetic route / pathway is a series / plurality of steps to be followed / performed in a defined order to make a product molecule from smaller, less complex structures. Each step in a synthetic route is a chemical reaction. A synthetic route may have / need any suitable number of chemical reaction steps to obtain a product molecule, e.g. 1-10 steps. A starting material on which a synthetic route is to be performed / implemented may be referred to herein as an initial intermediate molecule. Compounds formed at each step of the synthetic route may also be referred to as (further) intermediate molecules.

[0229] At each step of a synthetic route, a (current) intermediate molecule - i.e. the intermediate molecule obtained from a step of the synthetic route immediately prior to the current step - is enumerated / combined with a building block I synthon (reagent) using known and defined catalysts, reaction conditions, etc. to obtain a further intermediate molecule (or product molecule for the final step). For known synthetic routes, the reagents, catalysts, reaction conditions, etc., for each step are defined so as to optimise the chemical reaction being performed, e.g. maximise yield. That starting material or initial intermediate molecule may also be regarded / referred to herein a building block or synthon. In this described example, ‘building blocks’ and ‘synthons’ may be used interchangeably because only a single reaction type per synthon set (at each respective chemical reaction step of the synthetic route) is considered and so modification of the building blocks (to obtain the synthons) is not needed. However, it will be understood that in different examples of the invention more than one reaction type per synthon set may be considered such that modification of the building blocks to obtain synthons is needed.

[0230] Figure 15 illustrates an overview of a method 1500 according to the invention for obtaining desirable product molecules based on a known synthetic route. As part of a setup phase of the method 1500, at step 1501 an example molecule may be provided along with a known / defined synthetic route that can be implemented to obtain the molecule. The synthetic route has a defined number of chemical reaction steps, and catalysts, reaction conditions (e.g. temperature) are specified for each step. The known molecule and synthetic route may be retrieved as needed in any suitable manner, e.g. from computer memory.

[0231] At step 1502, the method 1500 involves performing atom mapping and template extraction. In particular, this involves identifying atoms of intermediates or building blocks that are involved in the reaction at each step of the defined synthetic route. This may be performed in any suitable manner, such as using one or more of the approaches described in “Atom- to-Atom Mapping: A Benchmarking Study of Popular Mapping Algorithms and Consensus Strategies", Madzhidov et al., ChemRxiv, 30 September 2020, Version 1. The identified atoms may be used to extract rules I templates relating to the reaction at a respective step. This may then be used, at step 1503, to define suitable substructures of building blocks that may be used in a given reaction step.

[0232] A database 1504 of building blocks I synthons, e.g. a publicly available database of known (e.g. commercially available) is searched for building blocks having similar / suitable substructures as the building blocks used to obtain the provided molecule using the provided synthetic route. Suitable building blocks are retrieved from the database 1504 at step 1505. Building blocks that are retrieved for a given step are regarded as possible alternative building blocks that can be used at the respective step to enumerate an intermediate molecule according to the specific chemical reaction performed at the respective step. The retrieval process may involve retrieving all building blocks in the database 1504 that have a specific substructure, or specific substructures, as defined at step 1503. Alternatively, or in addition, building blocks having similar substructures to the extracted / defined substructures may be retrieved from the database 1504, where the measure of similarity may be according to any suitable technique known to the skilled person.

[0233] The setup phase of the method 1500 can then comprise defining a plurality of sets of building blocks I synthons - specifically one synthon set per step of the synthetic route - each including a set of the synthons I building blocks retrieved as being suitable alternatives for use at the respective chemical reaction step.

[0234] The method 1500 includes an active learning phase. At step 1506, the method 1500 involves generating a set of product molecules by sampling the synthons in the plurality of synthon sets. In the described example, this is performed by hypercube enumeration / sampling as is described above. However, this may be performed using one of the other sampling methods described above, e.g. randomly selecting a subset of synthons in each synthon set and exhaustively enumerating all product molecules using the synthon subsets. Each of the enumerated product molecules are then scored at step 1507 using a defined scoring function, such as those described above, e.g. ROCS, docking score or MPO (multiparameter optimisation). At step 1508, the method involves training a surrogate model for each step of the synthetic route (for which corresponding building blocks are being selected). In a corresponding manner to the examples described above, the surrogate model for a respective step of the synthetic route is trained using input-output pairs of the building block used to enumerate the intermediate molecule at the respective chemical reaction step and the molecule score of the product molecule obtained at step 1507.

[0235] The trained surrogate models can then be used to predict molecule scores associated with all of the synthons in the respective sets or subsets. A greedy acquisition of building blocks for each chemical reaction step may then be performed at step 1510 according to the obtained molecule scores. Any suitable number of building blocks can be selected according to this greedy acquisition process. This process may be repeated as part of an iterative process in a corresponding manner to as described above. For instance, the process may be performed for a defined number of iterations and once it is determined that this defined number has been reached at step 1511 then the active learning phase may be regarded as completed.

[0236] In a final acquisition phase of the method 1500, an acquisition of building blocks from the entire available pool of building blocks may be performed at step 1512 based on the most recently trained version of the respective surrogate models. This provides the synthons for each chemical reaction step most likely to provide high-scoring product molecules while also being readily synthesisable. A final exhaustive enumeration is performed at step 1513 in a corresponding manner to the previous example to obtain final product molecules that can be scored with the scoring function at step 1514. As described before, the highest scoring molecules, or a subset thereof, may be taken forward for synthesis / testing.

[0237] Figure 16 schematically illustrates an example of a product molecule 1600 obtained by implementing a synthetic route having a number of chemical reaction steps. At each step a (current) intermediate molecule 1601a, 1601b, 1601c is enumerated with a synthon I building block. The building blocks 1602a, 1602b, 1602c of the product molecule 1600 can be analysed and sets of alternative building blocks for each step can be retrieved and analysed as described above in order to generate high-scoring product molecules.

[0238] Many modifications may be made to the described examples without departing from the scope of the appended claims.

Claims

CLAIMS1. A method for drug design, the method comprising computer-implemented steps of: defining an intermediate molecule having a first enumeration vector and a second enumeration vector; defining a first synthon set comprising a plurality of first synthons each for enumerating the first enumeration vector of the intermediate molecule; defining a second synthon set comprising a plurality of second synthons each for enumerating the second enumeration vector of the intermediate molecule; selecting a first synthon subset comprising a subset of the plurality of first synthons in the first synthon set; selecting a second synthon subset comprising a subset of the plurality of second synthons in the second synthon set; determining a molecule set comprising a plurality of product molecules each obtained by enumerating the first and second enumeration vectors of the defined intermediate molecule, wherein the molecule set comprises: for each of the first synthons in the first synthon subset, each product molecule obtained by enumerating the first enumeration vector with the respective first synthon and enumerating the second enumeration vector with each of a plurality of defined second proxy synthons; and for each of the second synthons in the second synthon subset, each product molecule obtained by enumerating the second enumeration vector with the respective second synthon and enumerating the first enumeration vector with each of a plurality of defined first proxy synthons; selecting a molecule subset comprising a subset of the product molecules in the molecule set; determining, by executing a defined scoring function, a molecule score for each of the product molecules in the molecule subset, wherein the molecule score is indicative of a proximity of a property of the respective product molecule to a desired property of a drug molecule to be designed; training a first surrogate model to output predicted molecule score as a function of structural features of first synthons, wherein the first surrogate model is trained using the first synthon and corresponding determined molecule score for each product molecule in the molecule subset;training a second surrogate model to output predicted molecule score as a function of structural features of second synthons, wherein the second surrogate model is trained using the second synthon and corresponding determined molecule score for each product molecule in the molecule subset; for each of the first synthons in the first synthon set, executing the trained first surrogate model to determine a first predicted molecule score as a function of structural features of the respective first synthon; for each of the second synthons in the second synthon set, executing the trained second surrogate model to determine a second predicted molecule score as a function of structural features of the respective second synthon; selecting a final first synthon subset comprising a subset of the plurality of first synthons in the first synthon set, wherein the selection is based on the determined first predicted molecule scores of the respective first synthons in the first synthon set; selecting a final second synthon subset comprising a subset of the plurality of second synthons in the second synthon set, wherein the selection is based on the determined second predicted molecule scores of the respective second synthons in the second synthon set; and determining one or more high-scoring product molecules, wherein each high- scoring product molecule is determined by enumerating the first enumeration vector of the intermediate molecule with one of the first synthons in the final first synthon subset and enumerating the second enumeration vector of the intermediate molecule with one of the second synthons in the final second synthon subset.

2. A method according to Claim 1 , wherein the first synthon subset comprises a defined number of the first synthons in the first synthon set, and wherein the second synthon subset comprises a defined number of the second synthons in the second synthon set.

3. A method according to Claim 2, wherein the defined number of synthons in the first and second synthon subsets is at least one order of magnitude less than a total number of synthons in the first and second synthon sets, respectively; optionally, wherein the defined number is at least two orders of magnitude less than the total number.

4. A method according to any previous claim, wherein the total number of first synthons in the first synthon set is at least 105synthons, and optionally at least 106synthons, andwherein the total number of second synthons in the second synthon set is at least 105synthons, and optionally at least 106synthons.

5. A method according to any previous claim, wherein the plurality of defined first proxy synthons is the first synthon subset, and wherein the plurality of defined second proxy synthons is the second synthon subset.

6. A method according to any of Claims 1 to 4, wherein the plurality of defined first proxy synthons is a randomly sampled subset of first synthons from the first synthon set, and wherein the plurality of defined second proxy synthons is a randomly sampled subset of second synthons from the second synthon set.

7. A method according to any previous claim, wherein selecting the first synthon subset comprises performing a random selection of first synthons in the first synthon set, and wherein selecting the second synthon subset comprises performing a random selection of second synthons in the second synthon set.

8. A method according to any of Claims 1 to 6, wherein selecting the first synthon subset comprises selecting first synthons to maximise a value of a diversity metric indicative of a diversity of structural features present in the first synthon subset, and wherein selecting the second synthon subset comprises selecting second synthons to maximise a value of the diversity metric indicative of a diversity of structural features present in the second synthon subset.

9. A method according to any of Claims 1 to 6, wherein the steps of selecting the first and second synthon subsets, determining the molecule set, and selecting the molecule subset collectively comprise: determining a master molecule set comprising a plurality of product molecules each obtained by enumerating the first and second enumeration vectors of the defined intermediate molecule, wherein the master molecule set comprises: for each of the first synthons in the first synthon set, each product molecule obtained by enumerating the first enumeration vector with the respective first synthon and enumerating the second enumeration vector with each of a plurality of defined second master proxy synthons; andfor each of the second synthons in the second synthon set, each product molecule obtained by enumerating the second enumeration vector with the respective second synthon and enumerating the first enumeration vector with each of a plurality of defined first master proxy synthons; and selecting the molecule subset by performing a hypercube sampling of molecules in the master molecule set, wherein the hypercube sampling comprises sampling molecules based on the first and second synthons present therein.

10. A method according to Claim 9, wherein the plurality of defined first master proxy synthons is the first synthon set, and wherein the plurality of defined second master proxy synthons is the second synthon set.11 . A method according to any previous claim, wherein a number of product molecules in the selected molecule subset is at least one order of magnitude less, and optionally at least two orders of magnitude less, than a number of product molecules in the molecule set.

12. A method according to any previous claim, wherein selecting the molecule subset comprises performing a random selection of the product molecules in the molecule set.

13. A method according to any previous claim, wherein: for each of the first synthons in the first synthon subset, the molecule subset is selected to include at least a defined number of product molecules obtained by enumerating the first enumeration vector with the respective first synthon; optionally, wherein the defined number is 1 , 2, 3, 4, 5, 10, 15, 20, 25 or 50; and for each of the second synthons in the second synthon subset, the molecule subset is selected to include at least a defined number of product molecules obtained by enumerating the second enumeration vector with the respective second synthon; optionally, wherein the defined number is 1 , 2, 3, 4, 5, 10, 15, 20, 25 or 50.

14. A method according to any previous claim, wherein the defined scoring function comprises a docking score function, wherein the molecule score is indicative of a prediction of a binding affinity of the respective product molecule when it is docked with a defined target molecule.

15. A method according to any previous claim, wherein the defined scoring function comprises a shape similarity function, wherein the molecule score is indicative of a prediction of a shape similarity of the respective product molecule to a defined target molecule.

16. A method according to any previous claim, wherein the first surrogate model is a first machine learning model, and wherein the second surrogate model is a second machine learning model.

17. A method according to Claim 16, wherein the first machine learning model is a first random forest model, and wherein the second machine learning model is a second random forest model.

18. A method according to Claim 16, wherein the first machine learning model is a first feed forward neural network model, and wherein the second machine learning model is a second feed forward neural network model.

19. A method according to Claim 16, wherein the first machine learning model is a first message passing neural network model, and wherein the second machine learning model is a second message passing neural network model.

20. A method according to any previous claim, the method comprising, prior to the steps of selecting the final first synthon subset and the final second synthon subset, iteratively repeating steps of: selecting a further first synthon subset comprising a subset of the plurality of first synthons in the first synthon set; selecting a further second synthon subset comprising a subset of the plurality of second synthons in the second synthon set; determining a further molecule set comprising a plurality of product molecules each obtained by enumerating the first and second enumeration vectors of the defined intermediate molecule, wherein the further molecule set comprises: for each of the first synthons in the further first synthon subset, each product molecule obtained by enumerating the first enumeration vector with the respective first synthon and enumerating the second enumeration vector with each of a plurality of defined second proxy synthons; andfor each of the second synthons in the further second synthon subset, each product molecule obtained by enumerating the second enumeration vector with the respective second synthon and enumerating the first enumeration vector with each of a plurality of defined first proxy synthons; selecting a further molecule subset comprising a subset of the molecules in the further molecule set; determining, using the defined scoring function, the molecule score for each of the molecules in the further molecule subset; retraining the first surrogate model using the first synthon and corresponding determined molecule score for each product molecule in the further molecule subset to obtain an updated first surrogate model; retraining the second surrogate model using the second synthon and corresponding determined molecule score for each product molecule in the further molecule subset to obtain an updated second surrogate model; for each of the first synthons in the first synthon set, executing the updated first surrogate model to determine the first predicted molecule score; and for each of the second synthons in the second synthon set, executing the updated second surrogate model to determine the second predicted molecule score, until a stop condition is satisfied.21 . A method according to Claim 20, wherein the further first synthon subset comprises a defined number of the first synthons in the first synthon set, and wherein the further second synthon subset comprises a defined number of the second synthons in the second synthon set; optionally, wherein the defined number of synthons in the further first and second synthon subsets is at least one order of magnitude less than a total number of synthons in the first and second synthon sets, respectively; further optionally, wherein the defined number is at least two orders of magnitude less than the total number.

22. A method according to Claim 20 or Claim 21 , wherein: selecting the further first synthon subset comprises performing a greedy selection of first synthons in the first synthon set according to determined first predicted molecule score; and selecting the further second synthon subset comprises performing a greedy selection of second synthons in the second synthon set according to determined second predicted molecule score.

23. A method according to Claim 20 or Claim 21 , wherein selecting the further first synthon subset comprises performing a random selection of first synthons in the first synthon set, and wherein selecting the further second synthon subset comprises performing a random selection of second synthons in the second synthon set.

24. A method according to any of Claims 20 to 23, wherein selecting the further first synthon subset comprises only selecting first synthons not selected in a previous iteration of the method, and wherein selecting the further second synthon subset comprises only selecting second synthons not selected in a previous iteration of the method.

25. A method according to any of Claims 20 to 24, wherein the stop condition is satisfied when a first rate of improvement of first predicted molecule scores of the first synthons in the further first synthon subset of a current iteration of the method relative to a previous iteration of the method is below a defined first threshold rate of improvement, and / or when a second rate of improvement of second predicted molecule scores of the second synthons in the further second synthon subset of the current iteration of the method relative to the previous iteration of the method is below a defined second threshold rate of improvement; optionally, wherein the defined first threshold rate of improvement is equal to the defined second threshold rate of improvement.

26. A method according to Claim 25, wherein the first rate of improvement is determined according to a calculation of average first predicted molecule score of the first synthons in the further first synthon subset of the current iteration of the method relative to the previous iteration of the method, and wherein the second rate of improvement is determined according to a calculation of average second predicted molecule score of the second synthons in the further second synthon subset of the current iteration of the method relative to the previous iteration of the method.

27. A method according to any of Claims 20 to 24, wherein the stop condition is satisfied when the repeated steps have been performed a defined number of times; optionally, wherein the defined number of times is any one of 1 , 2, 3, 4, 5, 10, 20, 50 or 100.

28. A method according to any previous claim, wherein the plurality of defined first proxy synthons is the further first synthon subset, and wherein the plurality of defined second proxy synthons is the further second synthon subset.

29. A method according to any of Claims 20 to 28, wherein a number of product molecules in the selected further molecule subset is at least one order of magnitude less, and optionally at least two orders of magnitude less, than a number of product molecules in the molecule set.

30. A method according to any previous claim, wherein selecting the further molecule subset comprises performing a random selection of the product molecules in the further molecule set.

31. A method according to any previous claim, wherein: for each of the first synthons in the further first synthon subset, the further molecule subset is selected to include at least a defined number of product molecules obtained by enumerating the first enumeration vector with the respective first synthon; optionally, wherein the defined number is 1 , 2, 3, 4, 5, 10, 15, 20, 25 or 50; and for each of the second synthons in the further second synthon subset, the further molecule subset is selected to include at least a defined number of product molecules obtained by enumerating the second enumeration vector with the respective second synthon; optionally, wherein the defined number is 1 , 2, 3, 4, 5, 10, 15, 20, 25 or 50.

32. A method according to any previous claim, wherein: selecting the final first synthon subset comprises ranking the first synthons in the first synthon set according to determined first predicted molecule score; defining a first ranked subset comprising a defined number of the top ranked first synthons; and selecting at least some of the top ranked first synthons in the first ranked subset to be included in the final first synthon subset; and selecting the final second synthon subset comprises ranking the second synthons in the second synthon set according to determined second predicted molecule score; defining a second ranked subset comprising a defined number of the top ranked second synthons; and selecting at least some of the top ranked second synthons in the second ranked subset to be included in the final second synthon subset.

33. A method according to any previous claim, wherein the defined number of synthons in the final first and second synthon subsets is at least one order of magnitude less than a total number of synthons in the first and second synthon sets, respectively; optionally, wherein the defined number is at least two orders of magnitude less than the total number.

34. A method according to any previous claim, wherein the determined high-scoring product molecules form a high-scoring molecule set, wherein the high-scoring molecule set comprises, for each of the first synthons in the final first synthon set, each high-scoring product molecule obtained by enumerating the first enumeration vector with the respective first synthon and enumerating the second enumeration vector with each of the plurality of second synthons in the final second synthon set.

35. A method according to any previous claim, the method comprising selecting at least some of the high-scoring product molecules for synthesis.

36. A method according to Claim 35, wherein selecting the high-scoring product molecules for synthesis comprises: determining, by executing the defined scoring function, or retrieving, from a database storing previously-determined molecule scores, the molecule score for each of the high-scoring product molecules; ranking the high-scoring product molecules according to determined molecule score; and, selecting a defined number of the top-ranked high-scoring product molecules for synthesis.

37. A method according to Claim 36, the method comprising, prior to selecting the defined number of the top-ranked high-scoring product molecules, performing one or more filters to remove one or more high-scoring product molecules from the ranking.

38. A method according to Claim 37, wherein the one or more filters are configured to remove high-scoring product molecules having one or more undesirable properties; optionally, wherein the undesirable properties include high toxicity or that the high-scoring molecule is a duplicate of a previously-tested molecule.

39. A method according to any of Claims 36 to 38, wherein the defined number of the topranked high-scoring product molecules is less than 100, less than 50, less than 25, less than 10, and / or less than 5.

40. A method according to any of Claims 35 to 39, the method comprising synthesising the selected high-scoring product molecules, and testing the synthesised high-scoring product molecules to determine one or more biological properties of the high-scoring product molecules.

41. A method according to Claim 40, wherein the determined one or more biological properties includes the property associated with the molecule score of the first and second surrogate models.

42. A method according to Claim 41 , the method comprising determining the molecule score of the respective tested high-scoring product molecules.

43. A method according to Claim 42, the method comprising: retraining the trained first surrogate model, or the updated first surrogate model, using the first synthon and corresponding determined molecule score for each tested high- scoring product molecule to obtain a tested first surrogate model; and retraining the trained second surrogate model, or the updated second surrogate model, using the second synthon and corresponding determined molecule score for each tested high-scoring product molecule to obtain a tested second surrogate model.

44. A method according to Claim 43, the method comprising: selecting a new first synthon subset comprising a subset of the plurality of first synthons in the first synthon set; selecting a new second synthon subset comprising a subset of the plurality of second synthons in the second synthon set; determining a new molecule set comprising a plurality of product molecules each obtained by enumerating the first and second enumeration vectors of the defined intermediate molecule, wherein the new molecule set comprises: for each of the first synthons in the new first synthon subset, each product molecule obtained by enumerating the first enumeration vector with the respectivefirst synthon and enumerating the second enumeration vector with each of the plurality of defined second proxy synthons; and for each of the second synthons in the new second synthon subset, each product molecule obtained by enumerating the second enumeration vector with the respective second synthon and enumerating the first enumeration vector with each of the plurality of defined first proxy synthons; selecting a new molecule subset comprising a subset of the product molecules in the new molecule set; determining, by executing the defined scoring function, the molecule score for each of the product molecules in the new molecule subset; retraining the tested first surrogate model using the first synthon and corresponding determined molecule score for each product molecule in the new molecule subset to obtain a new first surrogate model; retraining the tested second surrogate model using the second synthon and corresponding determined molecule score for each product molecule in the new molecule subset to obtain a new second surrogate model; for each of the first synthons in the first synthon set, executing the new first surrogate model to determine a first predicted molecule score as a function of structural features of the respective first synthon; for each of the second synthons in the second synthon set, executing the new second surrogate model to determine a second predicted molecule score as a function of structural features of the respective second synthon; selecting a new final first synthon subset comprising a subset of the plurality of first synthons in the new first synthon set, wherein the selection is based on the determined first predicted molecule scores of the respective first synthons in the first synthon set; selecting a new final second synthon subset comprising a subset of the plurality of second synthons in the second synthon set, wherein the selection is based on the determined second predicted molecule scores of the respective second synthons in the second synthon set; and determining one or more new high-scoring product molecules, wherein each high- scoring product molecule is determined by enumerating the first enumeration vector of the intermediate molecule with one of the first synthons in the new final first synthon subset and enumerating the second enumeration vector of the intermediate molecule with one of the second synthons in the new final second synthon subset.

45. A method according to Claim 44, the method comprising selecting at least some of the new high-scoring product molecules, and synthesising the selected new high-scoring product molecules; optionally, testing the synthesised new high-scoring product molecules to determine one or more biological properties of the new high-scoring product molecules.

46. A method according to any of Claims 40 to 45, the method comprising: defining a machine learning model for approximating one or more biological properties of molecules as a function of the one or more structural features of said molecules; and, training the machine learning model using the tested high-scoring product molecules.

47. A method according to Claim 46, wherein the machine learning model is at least one of: a Bayesian optimisation model; a regression model; a clustering model; a decision tree model; a random forest model; and, a neural network model.

48. A method according to Claim 46 or Claim 47, the method comprising executing the machine learning model, after the training step, to predict one or more molecules in a defined population of molecules having one or more desired biological properties.

49. A method according to Claim 48, the method further comprising synthesising at least one of the one or more predicted molecules.

50. A method according to any previous claim, wherein the intermediate molecule is defined in dependence on a defined target molecule with which it is desired that determined product molecules interact; optionally, wherein the defined target molecule is associated with a defined disease of interest.

51. A method according to any previous claim, wherein one or more of the high-scoring product molecules, or predicted molecules, are a candidate drug or therapeutic molecule having a desired biological, biochemical, physiological and / or pharmacological activity against a defined target molecule.

52. A method according to Claim 51 , wherein the predetermined target molecule is an in vitro and / or in vivo therapeutic, diagnostic or experimental assay target.

53. A method according to Claim 51 or Claim 52, wherein the candidate drug or therapeutic molecule is for use in medicine; for example, in a method for the treatment of an animal, such as a human or non-human animal.

54. A method according to any previous claim, wherein the structural features of the first and second synthons correspond to fragments present in the respective first and second synthons.

55. A method according to Claim 54, wherein the fragments present in each of the first and second synthons are represented as a respective molecular fingerprint; optionally, wherein the molecular fingerprint is an Extended Connectivity Fingerprint, ECFP; further optionally, ECFPO, ECFP2, ECFP4, ECFP6, ECFP8, ECFP10 or ECFP12.

56. A non-transitory, computer-readable storage medium storing instructions thereon that when executed by a computer processor causes the computer processor to perform the method of any of Claims 1 to 39.

57. A computing device comprising a processor that is configured to perform a method according to any of Claims 1 to 39.

58. A method for drug design, the method comprising computer-implemented steps of: defining a synthetic route having a plurality of chemical reaction steps including a first chemical reaction step and a second chemical reaction step to be performed after the first chemical reaction step; defining a first synthon set comprising a plurality of first synthons each for combining with an initial intermediate molecule at the first chemical reaction step to obtain a further intermediate molecule; defining a second synthon set comprising a plurality of second synthons each for combining with the further intermediate molecule at the second chemical reaction step to obtain a product molecule; selecting a first synthon subset comprising a subset of the plurality of first synthons in the first synthon set; selecting a second synthon subset comprising a subset of the plurality of second synthons in the second synthon set;determining a molecule set comprising a plurality of product molecules each obtained by implementing the defined synthetic route starting at the initial intermediate molecule, wherein the molecule set comprises: for each of the first synthons in the first synthon subset, each product molecule obtained by combining the initial intermediate molecule with the respective first synthon at the first chemical reaction step and combining the further intermediate molecule, obtained from the first chemical reaction step, with each of a plurality of defined second proxy synthons at the second chemical reaction step; and for each of the second synthons in the second synthon subset, each product molecule obtained by combining the initial intermediate molecule with each of a plurality of defined first proxy synthons at the first chemical reaction step and combining the further intermediate molecule, obtained from the first chemical reaction step, with the respective second synthon at the second chemical reaction step; selecting a molecule subset comprising a subset of the product molecules in the molecule set; determining, by executing a defined scoring function, a molecule score for each of the product molecules in the molecule subset, wherein the molecule score is indicative of a proximity of a property of the respective product molecule to a desired property of a drug molecule to be designed; training a first surrogate model to output predicted molecule score as a function of structural features of first synthons, wherein the first surrogate model is trained using the first synthon and corresponding determined molecule score for each product molecule in the molecule subset; training a second surrogate model to output predicted molecule score as a function of structural features of second synthons, wherein the second surrogate model is trained using the second synthon and corresponding determined molecule score for each product molecule in the molecule subset; for each of the first synthons in the first synthon set, executing the trained first surrogate model to determine a first predicted molecule score as a function of structural features of the respective first synthon; for each of the second synthons in the second synthon set, executing the trained second surrogate model to determine a second predicted molecule score as a function of structural features of the respective second synthon;selecting a final first synthon subset comprising a subset of the plurality of first synthons in the first synthon set, wherein the selection is based on the determined first predicted molecule scores of the respective first synthons in the first synthon set; selecting a final second synthon subset comprising a subset of the plurality of second synthons in the second synthon set, wherein the selection is based on the determined second predicted molecule scores of the respective second synthons in the second synthon set; and determining one or more high-scoring product molecules, wherein each high- scoring product molecule is obtained by implementing the defined synthetic route starting at the initial intermediate molecule, comprising combining the initial intermediate molecule with one of the first synthons in the final first synthon subset at the first chemical reaction step and combining the further intermediate molecule, obtained from the first chemical reaction step, with one of the second synthons in the final second synthon subset at the second chemical reaction step.

59. A method according to Claim 58, wherein the first synthon subset comprises a defined number of the first synthons in the first synthon set, and wherein the second synthon subset comprises a defined number of the second synthons in the second synthon set.

60. A method according to Claim 59, wherein the defined number of synthons in the first and second synthon subsets is at least one order of magnitude less than a total number of synthons in the first and second synthon sets, respectively; optionally, wherein the defined number is at least two orders of magnitude less than the total number.

61. A method according to any of Claims 58 to 60, wherein the total number of first synthons in the first synthon set is at least 105synthons, and optionally at least 106synthons, and wherein the total number of second synthons in the second synthon set is at least 105synthons, and optionally at least 106synthons.

62. A method according to any of Claims 58 to 61 , wherein the plurality of defined first proxy synthons is the first synthon subset, and wherein the plurality of defined second proxy synthons is the second synthon subset.

63. A method according to any of Claims 58 to 62, wherein the plurality of defined first proxy synthons is a randomly sampled subset of first synthons from the first synthon set,and wherein the plurality of defined second proxy synthons is a randomly sampled subset of second synthons from the second synthon set.

64. A method according to any of Claims 58 to 63, wherein selecting the first synthon subset comprises performing a random selection of first synthons in the first synthon set, and wherein selecting the second synthon subset comprises performing a random selection of second synthons in the second synthon set.

65. A method according to any of Claims 58 to 63, wherein selecting the first synthon subset comprises selecting first synthons to maximise a value of a diversity metric indicative of a diversity of structural features present in the first synthon subset, and wherein selecting the second synthon subset comprises selecting second synthons to maximise a value of the diversity metric indicative of a diversity of structural features present in the second synthon subset.

66. A method according to any of Claims 58 to 63, wherein the steps of selecting the first and second synthon subsets, determining the molecule set, and selecting the molecule subset collectively comprise: determining a master molecule set comprising a plurality of product molecules each obtained by implementing the defined synthetic route starting at the initial intermediate molecule, wherein the master molecule set comprises: for each of the first synthons in the first synthon set, each product molecule obtained by combining the initial intermediate molecule with the respective first synthon at the first chemical reaction step and combining the further intermediate molecule, obtained from the first chemical reaction step, with each of a plurality of defined second master proxy synthons at the second chemical reaction step; and for each of the second synthons in the second synthon set, each product molecule obtained by combining the initial intermediate molecule with each of a plurality of defined first master proxy synthons at the first chemical reaction step and combining the further intermediate molecule, obtained from the first chemical reaction step, with the respective second synthon at the second chemical reaction step; and selecting the molecule subset by performing a hypercube sampling of molecules in the master molecule set, wherein the hypercube sampling comprises sampling molecules based on the first and second synthons present therein.

67. A method according to Claim 66, wherein the plurality of defined first master proxy synthons is the first synthon set, and wherein the plurality of defined second master proxy synthons is the second synthon set.

68. A method according to any of Claims 58 to 67, wherein a number of product molecules in the selected molecule subset is at least one order of magnitude less, and optionally at least two orders of magnitude less, than a number of product molecules in the molecule set.

69. A method according to any of Claims 58 to 68, wherein selecting the molecule subset comprises performing a random selection of the product molecules in the molecule set.

70. A method according to any of Claims 58 to 69, wherein: for each of the first synthons in the first synthon subset, the molecule subset is selected to include at least a defined number of product molecules obtained by combining the initial intermediate molecule with the respective first synthon at the first chemical reaction step; optionally, wherein the defined number is 1 , 2, 3, 4, 5, 10, 15, 20, 25 or 50; and for each of the second synthons in the second synthon subset, the molecule subset is selected to include at least a defined number of product molecules obtained by combining the further intermediate molecule, obtained from the first chemical reaction step, with the respective second synthon at the second chemical reaction step; optionally, wherein the defined number is 1 , 2, 3, 4, 5, 10, 15, 20, 25 or 50.

71. A method according to any of Claims 58 to 70, wherein the defined scoring function comprises a docking score function, wherein the molecule score is indicative of a prediction of a binding affinity of the respective product molecule when it is docked with a defined target molecule.

72. A method according to any of Claims 58 to 71 , wherein the defined scoring function comprises a shape similarity function, wherein the molecule score is indicative of a prediction of a shape similarity of the respective product molecule to a defined target molecule.

73. A method according to any of Claims 58 to 72, wherein the first surrogate model is a first machine learning model, and wherein the second surrogate model is a second machine learning model.

74. A method according to Claim 73, wherein the first machine learning model is a first random forest model, and wherein the second machine learning model is a second random forest model.

75. A method according to Claim 73, wherein the first machine learning model is a first feed forward neural network model, and wherein the second machine learning model is a second feed forward neural network model.

76. A method according to Claim 73, wherein the first machine learning model is a first message passing neural network model, and wherein the second machine learning model is a second message passing neural network model.

77. A method according to any of Claims 58 to 76, the method comprising, prior to the steps of selecting the final first synthon subset and the final second synthon subset, iteratively repeating steps of: selecting a further first synthon subset comprising a subset of the plurality of first synthons in the first synthon set; selecting a further second synthon subset comprising a subset of the plurality of second synthons in the second synthon set; determining a further molecule set comprising a plurality of product molecules each obtained by implementing the defined synthetic route starting at the initial intermediate molecule, wherein the further molecule set comprises: for each of the first synthons in the further first synthon subset, each product molecule obtained by combining the initial intermediate molecule with the respective first synthon at the first chemical reaction step and combining the further intermediate molecule, obtained from the first chemical reaction step, with each of a plurality of defined second proxy synthons at the second chemical reaction step; and for each of the second synthons in the further second synthon subset, each product molecule obtained by combining the initial intermediate molecule with each of a plurality of defined first proxy synthons at the first chemical reaction step andcombining the further intermediate molecule, obtained from the first chemical reaction step, with the respective second synthon at the second chemical reaction step; selecting a further molecule subset comprising a subset of the molecules in the further molecule set; determining, using the defined scoring function, the molecule score for each of the molecules in the further molecule subset; retraining the first surrogate model using the first synthon and corresponding determined molecule score for each product molecule in the further molecule subset to obtain an updated first surrogate model; retraining the second surrogate model using the second synthon and corresponding determined molecule score for each product molecule in the further molecule subset to obtain an updated second surrogate model; for each of the first synthons in the first synthon set, executing the updated first surrogate model to determine the first predicted molecule score; and for each of the second synthons in the second synthon set, executing the updated second surrogate model to determine the second predicted molecule score, until a stop condition is satisfied.

78. A method according to Claim 77, wherein the further first synthon subset comprises a defined number of the first synthons in the first synthon set, and wherein the further second synthon subset comprises a defined number of the second synthons in the second synthon set; optionally, wherein the defined number of synthons in the further first and second synthon subsets is at least one order of magnitude less than a total number of synthons in the first and second synthon sets, respectively; further optionally, wherein the defined number is at least two orders of magnitude less than the total number.

79. A method according to Claim 77 or Claim 78, wherein: selecting the further first synthon subset comprises performing a greedy selection of first synthons in the first synthon set according to determined first predicted molecule score; and selecting the further second synthon subset comprises performing a greedy selection of second synthons in the second synthon set according to determined second predicted molecule score.

80. A method according to Claim 77 or Claim 78, wherein selecting the further first synthon subset comprises performing a random selection of first synthons in the first synthon set, and wherein selecting the further second synthon subset comprises performing a random selection of second synthons in the second synthon set.

81. A method according to any of Claims 77 to 80, wherein selecting the further first synthon subset comprises only selecting first synthons not selected in a previous iteration of the method, and wherein selecting the further second synthon subset comprises only selecting second synthons not selected in a previous iteration of the method.

82. A method according to any of Claims 77 to 81 , wherein the stop condition is satisfied when a first rate of improvement of first predicted molecule scores of the first synthons in the further first synthon subset of a current iteration of the method relative to a previous iteration of the method is below a defined first threshold rate of improvement, and / or when a second rate of improvement of second predicted molecule scores of the second synthons in the further second synthon subset of the current iteration of the method relative to the previous iteration of the method is below a defined second threshold rate of improvement; optionally, wherein the defined first threshold rate of improvement is equal to the defined second threshold rate of improvement.

83. A method according to Claim 82, wherein the first rate of improvement is determined according to a calculation of average first predicted molecule score of the first synthons in the further first synthon subset of the current iteration of the method relative to the previous iteration of the method, and wherein the second rate of improvement is determined according to a calculation of average second predicted molecule score of the second synthons in the further second synthon subset of the current iteration of the method relative to the previous iteration of the method.

84. A method according to any of Claims 77 to 82, wherein the stop condition is satisfied when the repeated steps have been performed a defined number of times; optionally, wherein the defined number of times is any one of 1 , 2, 3, 4, 5, 10, 20, 50 or 100.

85. A method according to any of Claims 58 to 84, wherein the plurality of defined first proxy synthons is the further first synthon subset, and wherein the plurality of defined second proxy synthons is the further second synthon subset.

86. A method according to any of Claims 77 to 85, wherein a number of product molecules in the selected further molecule subset is at least one order of magnitude less, and optionally at least two orders of magnitude less, than a number of product molecules in the molecule set.

87. A method according to any of Claims 58 to 86, wherein selecting the further molecule subset comprises performing a random selection of the product molecules in the further molecule set.

88. A method according to any of Claims 58 to 87, wherein: for each of the first synthons in the further first synthon subset, the further molecule subset is selected to include at least a defined number of product molecules obtained by combining the initial intermediate molecule with the respective first synthon at the first chemical reaction step; optionally, wherein the defined number is 1 , 2, 3, 4, 5, 10, 15, 20, 25 or 50; and for each of the second synthons in the further second synthon subset, the further molecule subset is selected to include at least a defined number of product molecules obtained by combining the further intermediate molecule, obtained from the first chemical reaction step, with the respective second synthon at the second chemical reaction step; optionally, wherein the defined number is 1 , 2, 3, 4, 5, 10, 15, 20, 25 or 50.

89. A method according to any of Claims 58 to 88, wherein: selecting the final first synthon subset comprises ranking the first synthons in the first synthon set according to determined first predicted molecule score; defining a first ranked subset comprising a defined number of the top ranked first synthons; and selecting at least some of the top ranked first synthons in the first ranked subset to be included in the final first synthon subset; and selecting the final second synthon subset comprises ranking the second synthons in the second synthon set according to determined second predicted molecule score; defining a second ranked subset comprising a defined number of the top ranked second synthons; and selecting at least some of the top ranked second synthons in the second ranked subset to be included in the final second synthon subset.

90. A method according to any of Claims 58 to 89, wherein the defined number of synthons in the final first and second synthon subsets is at least one order of magnitude less than a total number of synthons in the first and second synthon sets, respectively; optionally, wherein the defined number is at least two orders of magnitude less than the total number.

91. A method according to any of Claims 58 to 90, wherein the determined high-scoring product molecules form a high-scoring molecule set, wherein the high-scoring molecule set comprises, for each of the first synthons in the final first synthon set, each high-scoring product molecule obtained by combining the respective first synthon with the initial intermediate molecule at the first chemical reaction step and combining the further intermediate molecule, obtained from the first chemical reaction step, with each of the plurality of second synthons in the final second synthon set in the second chemical reaction step.

92. A method according to any of Claims 58 to 91 , the method comprising selecting at least some of the high-scoring product molecules for synthesis.

93. A method according to Claim 92, wherein selecting the high-scoring product molecules for synthesis comprises: determining, by executing the defined scoring function, or retrieving, from a database storing previously-determined molecule scores, the molecule score for each of the high-scoring product molecules; ranking the high-scoring product molecules according to determined molecule score; and, selecting a defined number of the top-ranked high-scoring product molecules for synthesis.

94. A method according to Claim 93, the method comprising, prior to selecting the defined number of the top-ranked high-scoring product molecules, performing one or more filters to remove one or more high-scoring product molecules from the ranking.

95. A method according to Claim 94, wherein the one or more filters are configured to remove high-scoring product molecules having one or more undesirable properties; optionally, wherein the undesirable properties include high toxicity or that the high-scoring molecule is a duplicate of a previously-tested molecule.

96. A method according to any of Claims 93 to 95, wherein the defined number of the topranked high-scoring product molecules is less than 100, less than 50, less than 25, less than 10, and / or less than 5.

97. A method according to any of Claims 93 to 96, the method comprising synthesising the selected high-scoring product molecules, and testing the synthesised high-scoring product molecules to determine one or more biological properties of the high-scoring product molecules.

98. A method according to Claim 97, wherein the determined one or more biological properties includes the property associated with the molecule score of the first and second surrogate models.

99. A method according to Claim 98, the method comprising determining the molecule score of the respective tested high-scoring product molecules.

100. A method according to Claim 99, the method comprising: retraining the trained first surrogate model, or the updated first surrogate model, using the first synthon and corresponding determined molecule score for each tested high- scoring product molecule to obtain a tested first surrogate model; and retraining the trained second surrogate model, or the updated second surrogate model, using the second synthon and corresponding determined molecule score for each tested high-scoring product molecule to obtain a tested second surrogate model.

101. A method according to Claim 100, the method comprising: selecting a new first synthon subset comprising a subset of the plurality of first synthons in the first synthon set; selecting a new second synthon subset comprising a subset of the plurality of second synthons in the second synthon set; determining a new molecule set comprising a plurality of product molecules each obtained by implementing the defined synthetic route starting at the initial intermediate molecule, wherein the new molecule set comprises: for each of the first synthons in the new first synthon subset, each product molecule obtained by combining the initial intermediate molecule with therespective first synthon at the first chemical reaction step and combining the further intermediate molecule, obtained from the first chemical reaction step, with each of the plurality of defined second proxy synthons at the second chemical reaction step; and for each of the second synthons in the new second synthon subset, each product molecule obtained by combining the initial intermediate molecule with each of the plurality of defined first proxy synthons at the first chemical reaction step and combining the further intermediate molecule, obtained from the first chemical reaction step, with the respective second synthon at the second chemical reaction step; selecting a new molecule subset comprising a subset of the product molecules in the new molecule set; determining, by executing the defined scoring function, the molecule score for each of the product molecules in the new molecule subset; retraining the tested first surrogate model using the first synthon and corresponding determined molecule score for each product molecule in the new molecule subset to obtain a new first surrogate model; retraining the tested second surrogate model using the second synthon and corresponding determined molecule score for each product molecule in the new molecule subset to obtain a new second surrogate model; for each of the first synthons in the first synthon set, executing the new first surrogate model to determine a first predicted molecule score as a function of structural features of the respective first synthon; for each of the second synthons in the second synthon set, executing the new second surrogate model to determine a second predicted molecule score as a function of structural features of the respective second synthon; selecting a new final first synthon subset comprising a subset of the plurality of first synthons in the new first synthon set, wherein the selection is based on the determined first predicted molecule scores of the respective first synthons in the first synthon set; selecting a new final second synthon subset comprising a subset of the plurality of second synthons in the second synthon set, wherein the selection is based on the determined second predicted molecule scores of the respective second synthons in the second synthon set; and determining one or more new high-scoring product molecules, wherein each high- scoring product molecule is obtained by implementing the defined synthetic route startingat the initial intermediate molecule, comprising combining the initial intermediate molecule with one of the first synthons in the new final first synthon subset at the first chemical reaction step and combining the further intermediate molecule, obtained from the first chemical reaction step, with one of the second synthons in the new final second synthon subset at the second chemical reaction step.

102. A method according to Claim 101 , the method comprising selecting at least some of the new high-scoring product molecules, and synthesising the selected new high-scoring product molecules; optionally, testing the synthesised new high-scoring product molecules to determine one or more biological properties of the new high-scoring product molecules.

103. A method according to any of Claims 97 to 102, the method comprising: defining a machine learning model for approximating one or more biological properties of molecules as a function of the one or more structural features of said molecules; and, training the machine learning model using the tested high-scoring product molecules.

104. A method according to Claim 103, wherein the machine learning model is at least one of: a Bayesian optimisation model; a regression model; a clustering model; a decision tree model; a random forest model; and, a neural network model.

105. A method according to Claim 103 or Claim 104, the method comprising executing the machine learning model, after the training step, to predict one or more molecules in a defined population of molecules having one or more desired biological properties.

106. A method according to Claim 105, the method further comprising synthesising at least one of the one or more predicted molecules.

107. A method according to any of Claims 58 to 106, wherein the intermediate molecule is defined in dependence on a defined target molecule with which it is desired that determined product molecules interact; optionally, wherein the defined target molecule is associated with a defined disease of interest.

108. A method according to any of Claims 58 to 107, wherein one or more of the high- scoring product molecules, or predicted molecules, are a candidate drug or therapeutic molecule having a desired biological, biochemical, physiological and / or pharmacological activity against a defined target molecule.

109. A method according to Claim 108, wherein the predetermined target molecule is an in vitro and / or in vivo therapeutic, diagnostic or experimental assay target.

110. A method according to Claim 108 or Claim 109, wherein the candidate drug or therapeutic molecule is for use in medicine; for example, in a method for the treatment of an animal, such as a human or non-human animal.

111. A method according to any of Claims 58 to 110, wherein the structural features of the first and second synthons correspond to fragments present in the respective first and second synthons.

112. A method according to Claim 111 , wherein the fragments present in each of the first and second synthons are represented as a respective molecular fingerprint; optionally, wherein the molecular fingerprint is an Extended Connectivity Fingerprint, ECFP; further optionally, ECFP0, ECFP2, ECFP4, ECFP6, ECFP8, ECFPIO or ECFP12.

113. A method according to any of Claims 58 to 112, wherein the initial intermediate molecule is an initial synthon from an initial synthon subset comprising a plurality of initial synthons, wherein the initial synthon subset is selected from an initial synthon set comprising a plurality of initial synthons.

114. A method according to Claim 113, wherein the molecule set includes product molecules including different initial synthons from the initial synthon subset.

115. A method according to Claim 114, comprising training an initial surrogate model to output predicted molecule score as a function of structural features of initial synthons, wherein the initial surrogate model is trained using the first synthon and corresponding determined molecule score for each product molecule in the molecule subset;for each of the initial synthons in the first synthon set, executing the trained initial surrogate model to determine a first predicted molecule score as a function of structural features of the respective initial synthon; selecting a final initial synthon subset comprising a subset of the plurality of initial synthons in the initial synthon set, wherein the selection is based on the determined initial predicted molecule scores of the respective initial synthons in the initial synthon set, wherein each of the one or more high-scoring product molecules is determined using one of the initial synthons from the final initial synthon subset.

116. A non-transitory, computer-readable storage medium storing instructions thereon that when executed by a computer processor causes the computer processor to perform the method of any of Claims 58 to 115.

117. A computing device comprising a processor that is configured to perform a method according to any of Claims 58 to 115.

118. A method for drug design, the method comprising computer-implemented steps of: defining a synthetic route having a plurality of defined chemical reaction steps to be performed in a defined order to obtain a product molecule from an initial intermediate molecule; for each of the defined chemical reaction steps of the defined synthetic route: defining a synthon set comprising a plurality of synthons each for combining, at the respective chemical reaction step, with an intermediate molecule, obtained from a previous chemical reaction step of the defined synthetic route; selecting a synthon subset comprising a subset of the plurality of synthons in the synthon set; determining a molecule set comprising a plurality of product molecules each obtained by implementing the defined synthetic route starting at the initial intermediate molecule and using one of the synthons from the respective synthon subset at each respective chemical reaction step of the synthetic route; selecting a molecule subset comprising a subset of the product molecules in the molecule set; determining, by executing a defined scoring function, a molecule score for each of the product molecules in the molecule subset, wherein the molecule score is indicative ofa proximity of a property of the respective product molecule to a desired property of a drug molecule to be designed; for each of the defined chemical reaction steps of the defined synthetic route: training a respective surrogate model to output predicted molecule score as a function of structural features of synthons, wherein the respective surrogate model is trained using the synthon, from the respective synthon subset, and corresponding determined molecule score for each product molecule in the molecule subset; for each of the synthons in the respective synthon set, executing the respective trained surrogate model to determine a predicted molecule score as a function of structural features of the respective synthon; selecting a respective final synthon subset comprising a subset of the plurality of synthons in the respective synthon set, wherein the selection is based on the determined predicted molecule scores of the respective synthons in the respective synthon set; and determining one or more high-scoring product molecules, wherein each high- scoring product molecule is obtained by implementing the defined synthetic route starting at the initial intermediate molecule and using one of the synthons from the respective final synthon subset at each respective chemical reaction step of the synthetic route.

119. A method according to Claim 118, wherein the plurality of synthons sets includes a synthon set comprising a plurality of synthons for use as the initial intermediate molecule.

120. A method according to Claim 118 or Claim 119, wherein defining the first and second synthon sets comprise: providing an example product molecule and performing an atom mapping process, based on the example product molecule, to identify atoms involved in each of the first and second reaction steps; and determining molecular substructures, based on the identified atoms, for performing each of the first and second reaction steps, wherein the synthons in the first and second synthon sets are selected based on the determined molecular substructures for the respective first and second reaction steps.121 . A method according to Claim 120, wherein defining the first and second synthon sets comprise retrieving synthons, having the determined molecular substructures for the respective first and second reaction steps, from a defined database.

122. A method according to Claim 121 , wherein defining the first and second synthon sets comprise retrieving synthons, having molecular substructures similar to the determined molecular substructures for the respective first and second reaction steps, from the defined database, wherein molecular similarity is determined according to a defined molecular similarity metric.

123. A method for drug design, the method comprising computer-implemented steps of: defining a plurality of synthon sets each comprising a respective plurality of synthons; for each of the plurality of synthon sets, selecting a synthon subset comprising a subset of the plurality of synthons in the respective synthon set; determining a molecule set comprising a plurality of product molecules each obtained by combining together one of the synthons from each of the respective synthon subsets; selecting a molecule subset comprising a subset of the product molecules in the molecule set; determining, by executing a defined scoring function, a molecule score for each of the product molecules in the molecule subset, wherein the molecule score is indicative of a proximity of a property of the respective product molecule to a desired property of a drug molecule to be designed; for each of the plurality of synthon sets: training a surrogate model to output predicted molecule score as a function of structural features of synthons, wherein the surrogate model is trained using the synthon, from the synthon subset of the respective synthon set, and corresponding determined molecule score for each product molecule in the molecule subset; for each of the synthons in the respective synthon set, executing the respective trained surrogate model to determine a predicted molecule score as a function of structural features of the respective synthon; selecting a final synthon subset comprising a subset of the plurality of synthons in the respective synthon set, wherein the selection is based on thedetermined predicted molecule scores of the synthons in the respective synthon set; and determining one or more high-scoring product molecules, wherein each high-scoring product molecule is obtained by combining together one of the synthons from each of the respective final synthon subsets.

124. A method according to Claim 123, wherein combining together one of the synthons from each of the respective synthon subsets comprises one of: enumerating a plurality of enumeration vectors of an intermediate molecule with a respective one of the synthons from the respective synthon subsets; and implementing a defined synthetic route having a plurality of chemical reaction steps performed in order, wherein each chemical reaction step comprises enumerating a respective intermediate molecule with a respective one of the synthons from the respective synthon subsets.

Citation Information

Patent Citations

  • Drug optimisation by active learning

    WO2022084696A1