Virtual co-compound screening model
A machine learning model using physical and experimental data with active learning optimizes co-crystal screening, addressing inefficiencies in existing methods by improving accuracy and scalability for diverse APIs and co-formers.
Patent Information
- Application Number
- PCT/EP2025/051906
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-30
- Filing Date
- 2025-01-27
- Publication Date
- 2025-08-07
AI Technical Summary
Existing co-crystal screening methods are time-consuming, costly, and lack accuracy due to inconsistent experimental data and limited applicability domains, particularly in predicting co-crystal formation for a broad chemical variety of APIs and co-formers.
A machine learning model trained with a combination of software-based physical models and experimental data, using features like excess enthalpy, free energy of mixing, and shape fitting parameters, to predict co-crystal formation, employing active learning for iterative experiment selection and data clustering to cover a broad chemical space.
Enhances the accuracy and scalability of co-crystal screening by optimizing experimental data selection and covering a wide range of chemical varieties, reducing the need for extensive experimental trials.
Smart Images

Figure IMGF000014_0001 
Figure IMGF000015_0001 
Figure 00000021_0000
Abstract
Description
[0001] VIRTUAL CO-COMPOUND SCREENING MODEL
[0002] The hereby described invention discloses a system and a method to determine suitable co-compound combinations via a combined software based approach of first principle calculations and a machine learning.
[0003] Technical Field
[0004] The invention deals with the technological area of a software-supported formulation development.
[0005] Background and description of the prior art
[0006] The inactive ingredient database of the FDA contains thousands of entries, which may act as suitable co-forming excipients. However, the high number of potential co-formers and the API specific crystal engineering requirements render the screening for co-crystals a challenging task (DOI: 10.1016 / j.drudis.2023.103763). The successful synthesis of co-crystals usually involves many experimental trials, which are time-consuming and costly. In the practical translation to market, the development cost is a primary factor which directly relies on the number of experimental trials required to find a suitable co-former.
[0007] Different experimental screening techniques exist such as solvent assisted solid grinding, crystallization from solution, slow solvent evaporation from solution, antisolvent addition, slurry conversion etc. However, all those techniques entail the risk of crystallizing single components, solvate formation or the induction of polymorphic transitions. In addition, solution based methods require equal or similar solubility of the compounds in the solvent or solvent mixture.
[0008] Consequently, different prediction approaches have been developed in order to support the experimental screening process. Basically, it is possible to distinguish between physics-based models and data-driven, in particular machine learning, models. Among physics-based models the conductor like screening model for realistic solvents (COSMO-RS) is most commonly used for the calculation of thermodynamic state functions describing for instance the mixing of two components. Those state functions can be exploited as a measure for successful co-crystal formation.
[0009] In contrast, data driven models rely on training machine learning algorithms on input data. In case of co-crystal formation this is typically based on experimental data containing the information about successful or unsuccessful co-crystal formation, respectively. Multiple approaches using different algorithms and descriptors representing molecular features are published in literature, like the table on page 4 of the above mentioned review paper and an Yuan et al., CrystEngComm, 2021 ,23, 6039-6044. Yuan et al. describes a combination of COSMO-RS and machine learning for the purpose of predicting whether an API and a co-former will form a co-crystal. It is suggested to compute Hex, the excess enthalpy of the two compounds together, from COSMO-RS theory and then train a separate machine learning model (random forest) using various standard molecular descriptors from RDKit, PaDEL and Mordred. The final model provides then a ranking scheme for how to combine the results of COSMO-RS and their ML model: Based on the output of the two, thresholds are defined for the ML probability and Hex and then either the prediction by COSMO-RS or their model is followed (page 3 of the supporting information). This approach still does not deliver optimal results as simply the better output is chosen but no data from one path is used the optimize the results of the respective other path.
[0010] With respect to data driven models substantial data amount is needed to train a reliable algorithm. In the majority of cases, published data used for this purpose often lacks negative data points, which are essential for any machine learning approach. In addition, many published models rely on data assembled from different sources leading to inconsistencies with respect to experimental procedures such as co-crystallization techniques, analytics and interpretation. Furthermore, public models often lack a comprehensive applicability domain, meaning they cannot be applied to any structural problem due to limited chemical variety of the input training data. Consequently, co-crystallization cannot be simulated at large scale nor for many systems in an automated way. The task of this patent application is therefore to provide an improved machine learning model approach to facilitate co-compound screening for a broad chemical variety of APIs that is more accurate than the known approaches from the prior art.
[0011] Summary of the invention
[0012] This task has been solved by a method to create a machine learning model for conducting a virtual co-compound screening, wherein the model is trained with training data created both by calculating features of at least one API and at least one co-former with a software-based physical model and with using experimental data containing information about co-compound formation of the at least one API and the at least one co-former, wherein calculating the features with the softwarebased model comprise of calculating at least one single molecule feature of the at least one API and I or the at least one co-former and at least one interaction feature. The core of the invention is hereby the point that the machine learning model which is used for performing the virtual co-compound screening is created and trained by using both training data which derives from experimental data gained by performing real experiments and training data which is calculated by applying a software which uses a physical model, meaning a set of equations, to calculate relevant features of the desired co-compound.
[0013] According to the present invention, the term “co-compound” comprises co-crystals, their salts and mixtures thereof. Co-crystals are crystals consisting of at least two different molecular entities crystallizing in the same crystal lattice connected via non- covalent interactions such as hydrogen bonds, Van-der-Waals interactions etc. In case of pharmaceutical co-crystals at least one molecular entity is an active pharmaceutical ingredient, while the co-former needs to be classified as pharmaceutically acceptable. Salts are similar to co-crystals but their interactions are dominated by ionic forces, meaning for many pharmaceutical applications they involve a proton transfer. In addition, the starting co-former does not need to be a solid. Advantageous and therefore preferred further developments of this invention emerge from the associated subclaims and from the description and the associated drawings.
[0014] One of those preferred further developments of the disclosed method comprise that the at least one interaction feature is selected from a list comprising of excess enthalpy (Hex), free energy of mixing (Gmix), shape fitting parameters and / or solubility of API in co-former. These are the most preferred interaction features relevant for co-compound formation but also others might be suitable.
[0015] Another one of those preferred further developments of the disclosed method comprise that the experimental data result from experiments that are performed by applying a software-based model driven iterative experiment selection. For gaining the optimal experiment training data for the final machine learning model it is preferable to optimize the selection of the performed experiments, since there is usually only a limited amount of experiments possible due to time and budget constraints. For this it can be helpful to use a software which applies at least one model, e.g. machine-learning models, to optimize the experiments selection process. This results in an iterative software / model based selection process.
[0016] Another one of those preferred further developments of the disclosed method comprise that the iterative experiments are run in batches, after each batch the newly obtained experimental data is added to a dataset and an ensemble of additional machine learning models is trained to make predictions for possible combinations of at least one API and at least one co-former that have not yet been tested, and those combinations which have the lowest agreement will then be used as the next experiment. This way, the information contained in the dataset can be drastically increased while avoiding experiments that do not provide enough new information.
[0017] Another one of those preferred further developments of the disclosed method comprise that on the experimental data resulting from the model driven iterative experiment selection a group cross validation split for hyperparameter tuning is used, so that all data points with the same at least one API are either used in a training or in a test split. This ensures the data validity and avoids overfitting.
[0018] Another one of those preferred further developments of the disclosed method comprise that the experimental data result from experiments with multiple APIs and multiple co-formers and wherein the selection of APIs and co-formers is based on single-molecule features to cover a broad chemical space. This allows the balance between the requirement that for the optimal final machine learning model the training data needs to cover the greatest possible chemical diversity, while at the same time chemical diversity is too vast to experimentally determine the result for all possible combinations of all molecules.
[0019] Another one of those preferred further developments of the disclosed method comprise that the multiple APIs and multiple co-formers used in the experiments are selected by featurizing and then clustering them in feature space, so that the final selection yields APIs and co-formers across their whole respective chemical spaces. This means the dataset was designed to create a basis for a generalizable model which allows to cover a broad chemical space on the contrary to most of the literature datasets have a narrow focus, e.g. only amino acids as co-formers, or contain co-compounds of non-API / non-co-former structures like CCSD. Feature space itself means that the dataset of all considered APIs represented by their computed molecular features - typically a NxM numerical matrix. This is also repeated with the co-formers separately from the APIs.
[0020] Another one of those preferred further developments of the disclosed method comprise that the at least one single molecule feature is selected from a list comprising of molecular mass, sigma moments, melting and boiling points, vapor pressure, viscosity, solubilities and free enthalpies of solvation in different solvents and surface area. These single molecule features are calculated by the softwarebased physical model and are used to train the final machine learning model which enhances its effectiveness compared to being trained using only experiment data.
[0021] Another one of those preferred further developments of the disclosed method comprise that the machine learning model is trained to calculate a success score for each given at least one API and at least one co-former. By doing so the user can e.g. create a list of suitable co-compounds sorted by the success rate provided by trained final machine learning model and then choose the co-compounds with the best success rate.
[0022] Another one of those preferred further developments of the disclosed method comprise that the machine learning model performs a binary classification of either co-compound or no co-compound based on the calculated success score and the imbalance in the binary classification is accounted for by weighing of the data points. As a further development the computer, meaning a software which uses the machine learning model, can be configured via applying a threshold or other another decision value if the scored combination of at least one API and at least one coformer is a suitable co-compound or not. This makes it easier for the user who otherwise would have needed to do so himself.
[0023] A further solution of the given task comprises a system to create a machine learning model for conducting a virtual co-compound screening, comprising of a computer which runs a software to perform the machine learning model, a database connected to the computer for storing data and laboratory devices for performing the respective iterative co-compound experiments, which are configured to perform the method steps of the previously described methods.
[0024] Another solution of the given task is a method for selecting at least one co-former for a given API to form a co-compound via a computer (6) comprising the following steps: providing single features of the new API and the at least one co-former, and interaction features of the new API and the at least one co-former to a machine learning model trained by a method as previously described, wherein the machine learning model calculates a success score for the new API and at least one coformer, and the at least one co-former matching a predefined selection criterion regarding the calculated success score is selected.
[0025] One preferred further development of this disclosed method comprise that the predefined selection criterion is the at least one co-former with the highest success score. Using the highest success score is for most use cases the logical choice. But there might be specific scenarios where also other at least one co-formers with lower success scores might be a better choice, e.g. if the co-former with the highest success score is not available due to legal or logistic reasons or if more than one co-former is required e.g. if multiple experiments are supposed to be performed.
[0026] Summarized can be said, that the present co-compound prediction model is based on consistent and dedicated screening data consisting of a broad variety of APIs and co-forming excipients. The present approach is based on two major premises. First a chemical space clustering for comprehensive applicability domain and second a model driven iterative experiment selection, also known as active learning. Usage of parameters derived from first principles for co-forming compounds as descriptors in the machine learning algorithm. This allows the algorithm to consider information about the interaction of both molecules from more valuable parameters than solely molecular descriptors. The invented solution tackles multiple problems in a way that has not been used like this before: First, the dataset was designed to create a basis for a generalizable model, e.g. APIs and co-formers that were used for experiments in the dataset were selected by featurizing them and then clustering them in feature space, so that the final selection yielded APIs and co-formers across their whole respective chemical spaces. In contrast, most of the literature datasets have a narrow focus (e.g. only amino acids as co-formers) or contain co-compounds of non-API / non-co-former structures (e.g. Cambridge Structural Database, CSD). There are additional problems with the literature datasets: they have a very skewed distribution of positive to negative examples, they combine sources of very different quality, and different experimental procedures, therefore providing inhomogeneous data sets that are less suitable to train a quality prediction tool. In contrast, the experimental procedure according to the invention was kept constant across all experiments.
[0027] Second the selection of experiments was performed with the help of active learning and therefore contains more information than a dataset where the API / co-former combinations had just been randomly sampled. In active learning, experiments are run in batches. After each batch, the newly obtained data is added to the dataset and an ensemble of models is trained to make predictions for API / co-former combinations that have not yet been tested. Those combinations for which ensemble has the lowest agreement (sometimes called uncertainty, entropy, etc.) will then be used as the next experiment. This way, the information contained in the dataset can be drastically increased while avoiding experiments that add less new information.
[0028] The final model is trained on single molecule features of both the API and the coformer, as well as features that describe interactions of the two, like Hex. These interactions features have not been used yet in the literature as inputs to a machine learning model for co-compound prediction. All examples are based on some featurization of both API and the co-former independently. Additionally, the model is trained with using experimental data containing information about co-compound formation of the at least one API and the at least one co-former, as explained above. In contrast, the model in Yuan et al. an interaction feature is used independently of a machine learning model, and then a ranking is proposed to use either the machine learning model’s results or follow the prediction based on a single interaction feature.
[0029] Detailed description of the invention
[0030] The methods and and used system according to the invention and functionally advantageous developments of those are described in more detail below with reference to the associated drawings using at least one preferred exemplary embodiment. In the drawings, elements that correspond to one another are provided with the same reference numerals.
[0031] The drawings show:
[0032] Figure 1 : A schematic overview about the workflow of the invented method.
[0033] Figure 2: A summary of the involved system components.
[0034] Figure 3: The schematic structure of the model architecture.
[0035] The workflow of the invented method according to a preferred exemplary embodiment is described schematically in Figure 1. Figure 2 shows an overview about the participating hardware of the invented Screening Tool 1. Apart from the necessary laboratory equipment 9, the hardware consists mainly of a suitable computer 2 hosting the software 3 that operates the used Machine Learning Models 4. Every kind of computer 2 that is suitable to be used with the respective software 3 can be used, e.g. a standard personal computer or an industrial pc. The data 5 used for training the models is stored at a database 10 is preferably realized as an integrated memory in the computer, but an external database I memory, e.g. in from of a cloud space memory, which is accessable via a network connection is also a further valid preferred embodiment.
[0036] Regarding the preferred exemplary embodiment the following working example will provide a more detailed explanation of the general workflow disclosed by Figure 1.
[0037] Laboratory experiments
[0038] For the laboratory experiments a set of 30 APIs and 40 co-formers was used (details see section Chemical space coverage). Not every possible combination was tested but for the first iteration 428 co-compound experiments were chosen with the help of active learning in an iterative manner (details see section Model driven experiment selection).
[0039] Co-compound preparation
[0040] For all experiments a solvent assisted solid grinding has been used: API (50 mg, or equivalent amount for liquid APIs) and 1 mol eq co-former were added into an HPLC vial. The mixture was wetted with acetone (25 pl) and 2 x 3 mm stainless steel grinding balls added. The samples were ground on a planetary mill at 500 rpm for 2 hours (for Screen A). After grinding the samples were uncapped to evaporate the acetone for a minimum of 4 hours.
[0041] Screen A
[0042] The resulting solids from the co-compound preparation were analyzed by XRPD.
[0043] XRPD diffractograms were collected on a Bruker D8 diffractometer using Cu Ka radiation (40 kV, 40 mA) in reflection geometry and a 0-20 goniometer fitted with a Ge monochromator. The incident beam passes through a 2.0 mm divergence slit followed by a 0.2 mm anti-scatter slit and knife edge. The diffracted beam passes through an 8.0 mm receiving slit with 2.5° Soller slits followed by the Lynxeye Detector. The software used for data collection was Diffrac Plus XRPD Commander and data analysis was HighScore Plus.
[0044] Samples were run under ambient conditions as flat plate specimens using powder as received. The sample was prepared on a polished, zero-background (510) silicon wafer by gently pressing onto the flat surface or packed into a cut cavity. The sample was rotated in its own plane.
[0045] The details of the standard Pharmorphix data collection method are:
[0046] • Angular range: 2 to 42° 20
[0047] • Step size: 0.05° 20
[0048] • Collection time: 0.5 s / step (total collection time: 6.40 min)
[0049] When required other methods for data collection are used with details as follows. The details of the screening data collection method are:
[0050] • Angular range: 2 to 31° 20
[0051] • Step size: 0.06° 20
[0052] • Collection time: 0.5 s / step (total collection time: 4 min)
[0053] All samples that showed XRPD pattern of either the single API or co-former were continued to Screen B.
[0054] Screen B
[0055] The samples from the co-compound preparation were wetted with ethyl acetate (25 pl) and re-ground as mentioned above. After grinding the samples were uncapped to evaporate the ethyl acetate for a minimum of 4 hours. The resulting solids were analyzed by XRPD as described above.
[0056] All samples that showed XRPD pattern of either the single API or co-former were continued to Screen C. Screen C
[0057] Methanol (1 ml) and the samples from the co-compound preparation were added into the vial and briefly shaken by hand. If the samples dissolved, they were uncapped to evaporate at RT. If the samples were undissolved after shaking, they were heated to 50 °C without stirring for 1 hour before being uncapped to evaporate at RT. After evaporation the resulting solids were analyzed by XRPD as describe above.
[0058] Samples at any stage of screening with an XRPD diffractogram different to a physical mixture of components were analyzed by thermogravimetric analysis (TGA), differential scanning calorimetry (DSC) and nuclear magnetic resonance (NMR). Additional characterization was carried out by ion chromatography (IC) if the co-former was not detectable by 1 H NMR or by HPLC if sample degradation was suspected.
[0059] Nuclear Magnetic Resonance (NMR)
[0060] 1H NMR spectra were collected on a Bruker 400 MHz instrument equipped with an auto-sampler and controlled by a Avance NEO nanobay console. Samples were prepared in DMSO-cfe solvent, unless otherwise stated. Automated experiments were acquired using ICON-NMR configuration within Topspin software, using standard Bruker-loaded experiments (1H). Off-line analysis was performed using ACD Spectrus Processor.
[0061] Differential Scanning Calorimetry (DSC)
[0062] TA Instruments Q2000
[0063] DSC data were collected on a TA Instruments Q2000 equipped with a 50 position auto-sampler. Typically, 0.5 - 3 mg of each sample, in a pin-holed aluminium pan, was heated at 10 °C / min from 25 °C to 300 °C. A purge of dry nitrogen at 50 ml / min was maintained over the sample.
[0064] The instrument control software was Advantage for Q Series and Thermal Advantage and the data were analysed using Universal Analysis or TRIOS. TA Instruments Discovery DSC
[0065] DSC data were collected on a TA Instruments Discovery DSC equipped with a 50 position auto-sampler. Typically, 0.5 - 3 mg of each sample, in a pin-holed aluminium pan, was heated at 10 °C / min from 25 °C to 300 °C. A purge of dry nitrogen at 50 ml / min was maintained over the sample.
[0066] The instrument control software was TRIOS and the data were analysed using TRIOS or Universal Analysis.
[0067] Thermo-Gravimetric Analysis (TGA)
[0068] TA Instruments Q500
[0069] TGA data were collected on a TA Instruments Q500 TGA, equipped with a 16 position autosampler. Typically, 5 - 10 mg of each sample was loaded onto a pre-tared aluminium DSC pan and heated at 10 °C / min from ambient temperature to 350 °C. A nitrogen purge at 60 ml / min was maintained over the sample.
[0070] The instrument control software was Advantage for Q Series and Thermal Advantage and the data were analysed using Universal Analysis or TRIOS.
[0071] TA Instruments Discovery TGA
[0072] TGA data were collected on a TA Instruments Discovery TGA, equipped with a 25 position auto-sampler. Typically, 5 - 10 mg of each sample was loaded onto a pre-tared aluminium DSC pan and heated at 10 °C / min from ambient temperature to 350 °C. A nitrogen purge at 25 ml / min was maintained over the sample.
[0073] The instrument control software was TRIOS and the data were analysed using TRIOS or Universal Analysis.
[0074] Chemical Purity Determination by High Performance Liquid Chromatography (HPLC)
[0075] Purity analysis was performed on an Agilent HP1100 / 1 nfinity II 1260 series system equipped with a diode array detector and using OpenLAB software. The full method details are provided below: Table 1 HPLC method for chemical purity determinations
[0076] Ion Chromatography (IC)
[0077] Data were collected on a Metrohm 930 Compact IC Flex with 858 Professional autosampler and 800 Dosino dosage unit monitor, using IC MagicNet software. Accurately weighed samples were prepared as stock solutions in a suitable solvent. Quantification was achieved by comparison with standard solutions of known concentration of the ion being analysed. Analyses were performed in duplicate, and an average of the values is given unless otherwise stated. Table 2 IC method for anion chromatography
[0078] Chemical space coverage
[0079] For a predictive model the training data needs to cover a certain chemical diversity. At the same time chemical diversity is too vast to experimentally determine the result for all possible combinations of all molecules. To alleviate this, a set of 83 well known pharmaceutical APIs was chosen. Based on a selection of single-molecule features described in the section ‘Model architecture’ they are grouped into 30 clusters using an unsupervised clustering algorithm. From each cluster one representative API is selected, also considering price and availability. This approach is repeated for coformers, considering an initial list of 1600 co-former from different sources, e.g. GRAS, EAFLIS, FDA list for iner excipients and the CSD, which were grouped into 40 clusters as described above. From each cluster one representative API is selected, also considering toxicity and availability. The combinatorial set of 30 APIs and 40 co-formers spans now the set of possible experiments, guaranteeing the coverage of a certain chemical diversity by design. Model driven experiment selection (active learning)
[0080] The aforementioned data set of 30 APIs and 40 co-formers spans is still too large to be measured in its entirety. To achieve the optimal model within a maximum of N performable measurements, active learning for planning the experiments is used. That means data points are measured in a weekly iterative manner. To this end, at each iteration an ensemble of models is trained on all measurements that are ready at that iteration. Each model is trained on a slightly different bootstrap training data sample. Then the ensemble predictions for each unmeasured data point are obtained. For the measurements in the next iteration, the datapoints that had the largest disagreement in their ensemble predictions are chosen. This choice allows the overall model to improve as fast as possible in its predictive performance and confidence, as opposed to selecting experiments randomly. This iterative process can be repeated until the maximum of N performable measurements is reached.
[0081] The next point is about the model architecture. Figure 3 shows its schematic structure. The final machine learning model uses a set of features derived from COSMO-RS theory as implemented in the COSMOtherm software which uses a set of equations meaning a kind of physical model. Note, that while the disclosed approach achieves its goals efficiently, due to the physical meaningfulness of the features, also other features could potentially be used in further preferred embodiments, including finger prints or machine learned representations for in former autoencoders, graph neural networks etc. As is, the current final machine learning model according to the preferred embodiment uses calculated features from COSMO-RS or simple molecular properties which can be divided into two groups: (1) single molecule features like molecular mass, sigma moments, melting and boiling points, vapor pressure, viscosity, solubilities and free enthalpies of solvation in different solvents, surface area, etc. and (2) interaction features where properties of the interaction of both the API and the conformer are computed, including the excess enthalpy, shape / fit parameters, etc. In the preferred embodiment, both the single molecule features of API and 15onformer as well as the interaction features of both are used as input. However, in further preferred embodiments other suitable features can be derived from these features, e.g. the sum / difference / ratio of the single molecule features of API and co-former or features from applying dimensionality reduction to the original features. The final machine learning model is then trained on the data collected in the active learning phase above using a group cross validation split for hyperparameter tuning so that all data points with the same API are either in the training or test split. The machine learning model is trained to perform a binary classification (co-compound I no co-compound) based its original success score output. Imbalance in the two classes is accounted for by weighing of the data points. A gradient boosting model minimizing the crossentropy as loss function is used to achieve this, but other models could potentially be used as well in further preferred embodiments: logistic regression, random forests or other ensemble methods, neural networks, etc. After hyperparameter optimization, the final machine learning model is trained on the full dataset without cross validation.
[0082] The used machine learning model in the invention is therefore different from the known approaches in the described state of the art, in particular the XtalPi paper. In the hereby disclosed invention multiple features from COSMO-RS are used, which describe a joint property of the API-co-former combination,, like Hex, Gmix, shape fitting parameters, and others, and then these are used together with single molecule features as input for a machine learning model. The final model does not rank between different, separate approaches, but basically uses the prediction from COSMO-RS together with other properties as the input.
[0083] List of references
[0084] 1 Co-Compound screening tool
[0085] 2 Computer
[0086] 3 Software
[0087] 4 Machine Learning Model
[0088] 5 Model Data
[0089] 6 User
[0090] 7 Display
[0091] 8 User Interface (GUI)
[0092] 9 Lab I experiment equipment
[0093] 10 Database
[0094] 11 Calculated Features
[0095] 12 Physical Model
[0096] 13 Experiment Data
[0097] 14 Experiment Results
[0098] 15 Predicted Co-Compund Success Rate
[0099] 16 Binary Classification
Claims
Patent claims1. A method to create a machine learning model for conducting a virtual cocompound screening via a computer, wherein the model is trained on the computer with training data created both by calculating features of at least one API and at least one co-former with a software-based physical model and with using experimental data containing information about co-compound formation of the at least one API and the at least one co-former, wherein calculating the features with the software-based model comprise of calculating at least one single molecule feature of the at least one API and I or the at least one co-former and at least one interaction feature.
2. A method according to claim 1 , wherein the at least one interaction feature is selected from a list comprising excess enthalpy (Hex), free energy of mixing (Gmix), shape fitting parameters and / or solubility of API in co-former.
3. A method according to any of the previous claims, wherein the experimental data result from experiments that are performed by applying a software-based model driven iterative experiment selection.
4. A method according to claim 3, wherein the iterative experiments are run in batches, after each batch the newly obtained experimental data is added to a dataset and an ensemble of additional machine learning models is trained to make predictions for possible combinations of at least one API and at least one co-former that have not yet been tested, and those combinations which have the lowest agreement will then be used as the next experiment.
5. A method according to claim 3 or claim 4, wherein on the experimental data resulting from the model driven iterative experiment selection a group cross validation split for hyperparameter tuning is used, so that all data points with the same at least one API are either used in a training or in a test split.
6. A method according to any of the previous claims, wherein the experimental data result from experiments with multiple APIs and multiple co-formers andwherein the selection of APIs and co-formers is based on single-molecule features to cover a broad chemical space.
7. A method according to claim 6, wherein the multiple APIs and multiple coformers used in the experiments are selected by featurizing and then clustering them in feature space, so that the final selection yields APIs and co-formers across their whole respective chemical spaces8. A method according to any of the previous claims, wherein the at least one single molecule feature is selected from a list comprising of molecular mass, sigma moments, melting and boiling points, vapor pressure, viscosity, solubilities and free enthalpies of solvation in different solvents and surface area.
9. A method according to any of the previous claims, wherein the machine learning model is trained to calculate a success score for each given at least one API and at least one co-former.
10. A method according to claim 9, wherein the machine learning model performs a binary classification of either co-compound or no co-compound based on the calculated success score and the imbalance in the binary classification is accounted for by weighing of the data points.1 1. A system to create a machine learning model for conducting a virtual cocompound screening, comprising of a computer which runs a software to perform the machine learning model, a database connected to the computer for storing data and laboratory devices for performing the respective iterative cocompound experiments, which are configured to perform the method steps of claims 1 to 10.
12. Method for selecting at least one co-former for a unknown API to form a cocompound via a computer (6) comprising the following steps: providing single features of the unknown API and the at least one co-former, and interaction features of the unknown API and the at least one co-former to a machine learning model trained by a method according to any of the previous claims,wherein the machine learning model calculates a success score for the unknown API and at least one co-former, and the at least one co-former matching a predefined selection criterion regarding the calculated success score is selected.
13. A method according to claim 12, wherein the predefined selection criterion is the at least one co-former with the highest success score.