Digital Selection of Viscosity-Reducing Excipients for Protein Formulations

JP2024519756A5Pending Publication Date: 2025-05-16MERCK PATENT GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023569631
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-05-10
Filing Date
2022-05-09
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing methods for selecting viscosity-reducing excipients for protein formulations are time-consuming and require prior knowledge of the protein's structural details, which are often confidential and not readily available, limiting the ability to optimize formulations with excipient combinations for synergistic viscosity reduction and protein stability.

Method used

A computer-mediated method using a machine learning model that predicts the viscosity of protein formulations with unknown proteins by analyzing excipient representations and patterns in a dataset, allowing for the selection of optimal excipient combinations without requiring detailed protein information, and guiding experimental design for verification.

Benefits of technology

Significantly reduces the number of laboratory tests needed to determine the viscosity of new formulations, enabling efficient identification of excipients that reduce viscosity and improve protein stability, while avoiding the need for costly and complex protein descriptor collection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000018_0000
    Figure 00000018_0000
  • Figure 00000018_0001
    Figure 00000018_0001
  • Figure 00000018_0002
    Figure 00000018_0002
Patent Text Reader

Abstract

A method for selecting, via a computer (6), at least one viscosity-altering excipient (2) for a formulation (8) containing at least one unknown protein (11), comprising the steps of: providing a dataset (1) from a database describing the viscosity of several known formulations containing at least one protein and, optionally, at least one viscosity-altering excipient (2); generating, via in silico simulation, a representation of at least one excipient (2) from the list of excipients by the computer (6); and using a machine learning model (5) performed on the computer (6) to recognize patterns in the dataset (1) using the generated representation of the at least one excipient (2). evaluating the viscosity-altering effect of at least one viscosity-altering excipient (2) selected from a list of excipients on a new formulation (8) containing at least one unknown protein (11) and at least one viscosity-altering excipient (2) by applying the recognized pattern to the provided data of the protein (11); depending on the evaluation results, selecting at least one excipient from the list according to an acquisition criterion and applying it to the unknown protein (11), wherein the provided data of the at least one unknown protein (11) are data describing the viscosity of a protein composition containing at least one unknown protein (11) and optionally at least one viscosity-altering excipient (2) together.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a computer-assisted method for selecting the viscosity of a viscosity-reducing excipient for a protein composition. [Background technology]

[0002] Description of Background and Prior Art Monoclonal antibodies (mABs) and other protein therapeutics are typically administered parenterally. Subcutaneous injections are particularly popular for the delivery of protein therapeutics due to their potential to simplify patient administration (fast, low-volume injections) and reduce treatment costs (shorter medical support). To ensure patient compliance, subcutaneous dosage forms are desirably isotonic and capable of being injected in small volumes (<2.0 ml per injection site). To reduce injection volume, proteins are often administered at concentrations of 1 mg / ml to 150 mg / ml.

[0003] At the same time, mAB-based therapies typically require administration of several mg / kg. Thus, the combination of high therapeutic doses and low injection volumes creates a need for highly concentrated formulations of therapeutic antibodies. However, antibodies are large proteins and therefore have a complex conformation as well as numerous functional groups, which makes their formulation difficult, especially when high concentrations are required.

[0004] One of the main problems with highly concentrated protein solutions is viscosity. Proteins tend to form highly viscous solutions, primarily due to non-native self-association. Furthermore, proteins exhibit increased rates of aggregation and particle formation at such high concentrations.

[0005] These problems relate to both the manufacturing process and administration to patients, where highly viscous, highly concentrated protein formulations pose particular challenges to ultrafiltration and sterile filtration. Additionally, tangential flow filtration is often used for buffer exchange and increasing protein concentration. However, viscous solutions exhibit increased backpressure and shear stress during injection and filtration, which can destabilize therapeutic proteins and / or increase processing times. The increased shear stress frequently results in product loss. Both aspects negatively impact the economics of the process.

[0006] At the same time, high viscosity severely limits the injectability of the protein and is therefore unacceptable for administration.

[0007] Certain excipients and excipient combinations have been identified to reduce the viscosity of protein formulations, however, using screening approaches to identify the best excipients or excipient combinations is time consuming. Particularly in light of competitive timelines and limited amounts of test material (protein), a data-driven approach that reduces the number of experiments and accelerates excipient selection would be highly beneficial.

[0008] The state of the art is the discovery of several approaches for protein formulation design that incorporate machine learning models. For example, the paper "Machine learning models of antibody-excipient preferential interactions for use in computational formulation design" (Mol. Pharmaceutics 2020, 17, 3589-3599, DOI. 10.1021 / acs.molpharmaceut.0c00629) describes the ability to describe the interactions between formulation excipients and proteins in solution using computer simulations. This allows formulation design to begin early in the development of new antibody therapeutics. To do so, it discloses a feature set that numerically describes local regions of the antibody surface for use in machine learning applications.

[0009] Another approach is summarized in the review "Prediction Machines: Applied Machine Learning for Therapeutic Protein Design and Development," published in the Journal of Pharmaceutical Sciences on December 2, 2020. It describes the application of machine learning models to better understand the nonlinear concentration-dependent viscosity of protein solutions, predict protein oxidation and deamidation rates, classify subvisible particles, and compare protein physical stability. Additionally, various machine learning approaches are used to provide improved modeling results for regression and classification of previously published data.

[0010] Another approach is known from International Patent Application WO 2021 / 0413S4A1, which discloses a method for predicting a property of a potential protein formulation, wherein a set of formulation descriptors is classified as belonging to a particular one of a plurality of predetermined groups, each corresponding to a different value range for the protein formulation property; classifying the set of descriptors includes applying at least a first portion of the set of descriptors as input to a first machine learning model. The method also includes selecting a second machine learning model from a plurality of models corresponding to the different groups based on the classification. The method also includes predicting a value of the protein formulation property corresponding to the set of descriptors by applying at least a second portion of the set of formulation descriptors as input to the selected model. The method also includes displaying the value of the protein formulation property to a user and / or storing it in memory.

[0011] However, while these publications demonstrate methods for simulating the viscosity of protein solutions, they still suffer from several important drawbacks. First, they all require prior knowledge of the target protein, either in the form of data, descriptors, properties, or structural details. This is a significant drawback because protein developers are reluctant to share detailed information about protein candidates, or the information is simply unavailable. Experiments to obtain these parameters can be tedious, time-consuming, and therefore costly. Furthermore, collecting protein descriptors is complex and may not fully reflect the protein's properties. Regarding the task of formulation design, the most serious drawback of existing approaches is that they are limited to predicting the concentration-dependent viscosity of fixed protein formulations and do not provide experimental designs to help users explore other formulations. Therefore, the current state-of-the-art does not provide a way to directly use the insights gained from the predictions to optimize formulations with other excipients, let alone other proteins. Another issue is that these documents only consider the use of a single excipient. Combining excipients is beneficial because different excipients may exhibit synergistic viscosity reduction and / or improved protein stability compared to single excipients with similar viscosity-reducing effects. Excipient combinations are not covered by the models in the above references.

[0012] The objective of this patent application is to solve the above problems. A further objective is to find a more efficient machine learning-based approach to determine and identify the best excipient or excipient combination for changing the viscosity of a protein formulation. A further objective is to find a machine learning approach to identify the best experimental design for finding the optimal excipient or excipient combination for changing the viscosity of a protein formulation. Summary of the Invention

[0013] Summary of the Invention This issue is 1. A method for selecting, via a computer, at least one viscosity-altering excipient for a formulation containing at least one unknown protein, the method comprising the steps of: providing a dataset from a database describing the viscosity of several known formulations containing at least one protein and, optionally, at least one viscosity-altering excipient; generating, by a computer, a representation of at least one excipient from the list of excipients via in silico simulation; using a machine learning model implemented on a computer to recognize patterns in the dataset using the generated representation of the at least one excipient, and evaluating the viscosity-altering effect of the at least one viscosity-altering excipient selected from the list of excipients on a new formulation containing the at least one unknown protein and the at least one viscosity-altering excipient by applying the recognized pattern to the provided data of the at least one unknown protein; and, depending on the evaluation results, selecting at least one excipient from the list according to a capture criterion and applying it to the unknown protein, wherein the provided data of the at least one unknown protein is data describing the viscosity of a protein composition together containing the at least one unknown protein and, optionally, at least one viscosity-altering excipient. was resolved by

[0014] This procedure provides a more efficient method for exploring the viscosity of protein-excipient formulations compared to state-of-the-art approaches. A key advantage is that the use of a trained machine learning model significantly reduces the number of actual laboratory tests that need to be performed to determine the resulting viscosity of each excipient used. To do so, the model evaluates the viscosity-altering effect of at least one viscosity-altering excipient added from the list of excipients in the dataset used to create the representation. One evaluation option is for the model to predict the resulting viscosity for all possible excipients (excipients used in the dataset) so that the most suitable combination of unknown protein and excipient can be selected. However, the model can also predict the resulting viscosity for excipients selected from a list, or other evaluation methods can be used. The higher the model's predictive accuracy, the fewer time- and resource-intensive actual tests that need to be performed. The model's predictive accuracy typically increases with the number of viscosity measurements provided. When no or very few measurements are available, the machine learning model primarily uses the excipient representation and data from other known formulations, preferably those similar to the new formulation.

[0015] After the machine learning model predicts the viscosity of a new formulation and the most suitable excipients are selected, the ground truth viscosity of the formulation can be measured, and the respective data can be fed back into the model to improve the accuracy of subsequent predictions. Any standard personal or industrial computer equipped with a processor and respective working and storage memory can be used to run the machine learning model. While it is possible to use the same computer to run the in silico simulation and the machine learning model, it is often more efficient to use two different computers specially configured to run each application. The sufficiency criteria can suggest, for example, the excipient predicted to most reduce viscosity or the excipient that will provide the greatest information gain. The solution to the task also includes a software product stored on a computer-readable storage medium and including instructions that, when executed by a computer, cause the computer to perform the method steps disclosed in the previous chapter.

[0016] As defined herein, "unknown protein" refers to a protein to be tested by the described method. The properties and / or characteristics of these proteins, such as specific protein descriptors, are not necessarily known at the time the disclosed method is performed. In particular, unknown proteins are proteins not included in the databases of the aforementioned methods. More specifically, comprehensive viscosity measurements, with or without viscosity-altering excipients, are not available, except for data describing the viscosity of protein compositions containing at least one unknown protein and, optionally, at least one viscosity-altering excipient, which is required as data provided for the above-described methods. One advantage of the disclosed methods over known prior art is that specific information about the protein used is not required, thereby allowing the use of unknown proteins. At the other end of the spectrum is a "known formulation containing at least one protein and, optionally, at least one viscosity-altering excipient," which refers to a protein composition in which the protein itself and some or all of its properties and / or characteristics are known. In particular, these proteins are included in databases, and more specifically, viscosity measurements can be performed with or without viscosity-altering excipients. Optionally, the formulation can include one or more known viscosity-altering excipients.

[0017] As defined herein, "a new formulation containing at least one unknown protein and at least one viscosity-altering excipient" refers to a protein composition containing an unknown protein and at least one viscosity-altering excipient as defined above. The formulation contains one or more known viscosity-altering excipients. According to the present invention, the viscosity of a new formulation as defined above is predicted. Preferably, the viscosity of multiple formulations is predicted, for example, a formulation containing at least one unknown protein and at least one viscosity-altering excipient A and a formulation containing at least one unknown protein and at least one viscosity-altering excipient B. More preferably, the viscosity of all possible formulations is predicted. In this context, "all possible combinations" refers to all combinations of at least one viscosity-altering excipient selected from the list of excipients used to generate the dataset and at least one unknown protein. In a further embodiment, at least one group of viscosity-altering excipients is selected from the list of excipients.

[0018] As defined herein, "data for at least one unknown protein provided" refers to data describing the viscosity of a protein composition containing at least one unknown protein without or with at least one viscosity-altering excipient. In this context, "data describing viscosity" refers to data generated by at least one viscosity measurement of a protein composition. The protein composition contains at least one unknown protein, where the protein itself and its properties and characteristics are unknown. Optionally, the protein composition can contain one or more known viscosity-altering excipients. The one or more known viscosity-altering excipients are viscosity-altering excipients that were also used to generate the dataset.

[0019] The data provided for at least one unknown protein does not address protein descriptors or properties. Collections of protein descriptors, as used in other methods, are complex and may not adequately reflect protein properties.

[0020] Furthermore, protein developers do not need to share confidential information about their protein candidates, and costly MD simulations and homology modeling can be avoided. Furthermore, structural details of the protein, as would otherwise be required, may not be available. In preferred embodiments of the present invention, only a limited set of viscosity measurements, only one viscosity measurement, or no viscosity measurements are required.

[0021] Advantageous and therefore preferred further developments of the invention emerge from the associated subclaims, the description and the associated drawings.

[0022] One of the preferred further developments of the disclosed method includes that a dataset is generated by experimental measurement and stored in a database via a computer. Actual laboratory testing is also a recommended method for generating a dataset that will later be used by a machine learning model. The more accurate and representative this dataset is, the better the results of the machine learning model will be. This applies to both datasets containing the viscosity and excipients of known formulations and datasets of new formulations.

[0023] Another preferred further development of the disclosed method includes using a combination of two or more excipients from the list as at least one excipient from the list that most significantly changes the viscosity of the new formulation (8). Combinations of two or more excipients can be beneficial as different excipients may exhibit synergistic viscosity reduction and / or improved protein stability compared to single excipients with similar viscosity-reducing effects.

[0024] Another preferred further development of the disclosed method includes that specific experimental measurements are proposed to the formulation specialist, who performs these respective experiments in the laboratory to verify the predicted viscosity and adds the verified results to the provided dataset in the database via a computer, thereby training the machine learning model. Further, the viscosity values ​​predicted from the machine learning model can also be proposed to the formulation specialist. Measurement of the resulting viscosity of the mentioned new formulation is preferably performed by the formulation specialist. It is also possible for the specialist to perform the measurement with the support of a robotic machine and software. With suitable hardware and software, the measurement can also be performed fully automatically.

[0025] Another of these preferred further developments of the disclosed method includes that initial data describing the viscosity of a new formulation without excipients and / or already validated excipients is used as input data for the new formulation. More specifically, if there is already known data about the new formulation, e.g., from previous measurements or other sources, this data is input to the machine learning model, thereby further reducing the amount of testing or measurements required to achieve an accurate prediction.

[0026] Another preferred further development of the disclosed method includes creating and training a machine learning model by combining a dataset describing the viscosity of at least one prototype protein formulation with a representation of at least one excipient or combination thereof. The dataset initially used to create the machine learning model includes known formulations and their excipients. If the machine learning model is used to predict the viscosity of a new, possibly unknown, formulation, the machine learning model is further trained with known properties of the formulation, with or without excipients (if available), and / or experimental measurement data obtained as a result of confirmatory laboratory testing. If the known properties are not available in the required digital representation format, they must be converted, respectively.

[0027] Another preferred further development of the disclosed method involves modeling the viscosity values ​​of a given protein formulation in the form of a Gaussian process, and using the model predictions to guide a formulation specialist through a Bayesian optimal experimental design, which then allows the formulation specialist to perform the necessary measurements on the excipients or combinations thereof suggested by the machine learning model.

[0028] Another preferred further development of the disclosed method includes training the machine learning model on a computer by performing at least one of the following steps: optimizing machine learning model parameters with training data from the dataset by maximizing the marginal likelihood of the training data; evaluating the posterior distribution of viscosity values ​​of untested excipients or combinations thereof based on the machine learning model, thereby predicting viscosity; selecting a new set of excipients or combinations thereof by optimizing the obtained scores obtained from the calculated posterior distribution; proposing the new excipients or combinations thereof to a formulation specialist, who then performs respective experiments in a lab and measures the obtained viscosities; and adding the obtained measurements to the training data.

[0029] Another preferred further development of the disclosed method includes that the viscosity prediction obtained from the posterior distribution of viscosity values ​​is based on a pH-dependent feature vector characterizing the excipients used in the considered formulation and the excipient concentration levels used. This represents the most preferred method for a machine learning model to predict viscosity. However, the machine learning model is not limited to this method. If there are other methods for predicting viscosity values, they can be incorporated into the machine learning model and implemented.

[0030] In another preferred further development of the disclosed method, the acquisition criteria include assessing which viscosity-altering excipient is expected to reduce viscosity the most. Alternative embodiments may include other acquisition criteria that suggest, for example, experiments expected to yield the greatest information gain, experiments that yield the greatest model change, experiments that provide the greatest probability of improving formulation viscosity beyond the level of the observed best settings, experiments that yield the greatest expected improvement over the current optimal formulation, or any other systematic trade-off between exploring the formulation search space and utilizing previously gathered knowledge.

[0031] Another preferred further development of the disclosed method includes measuring viscosity in a protein formulation containing at least a protein, at least one viscosity-changing agent, at least one buffer, at least one stabilizer and at least one surfactant in an aqueous solution.This combination of components is the most common and therefore is preferably used.However, if there are other combinations that are necessary and / or more suitable for the claimed method, they can also be used.

[0032] Another preferred further development of the disclosed method involves computationally generating a representation of the excipient in the form of physical parameters, similar to a molecular fingerprint. The physical parameters describe the excipient and its properties so that a machine learning model can process the parameters and use them to predict the viscosity they will cause in a particular protein formulation. Possible parameters include, but are not limited to, charge distribution, dipole moment, quadrupole moment trace and anisotropy, polarizability, molecular London dispersion coefficient (C6), logP water / hexane partition coefficient, solvent-accessible surface area, molecular orbital energy HOMO-LUMO gap, etc.

[0033] Another solution to the problem of this patent application is a machine learning model implemented on a computer, which is created and trained as described in the previous section.

[0034] Another preferred further development of the disclosed machine learning models includes replacing Gaussian processes with any other model architecture that serves the same purpose, in particular other types of stochastic processes, generalized linear models, neural networks, support vector machines, tree-based models, ensemble models, etc.

[0035] Detailed Description of the Invention The methods, machine learning models and software products according to the invention, as well as their advantageous functional developments, are described in more detail below by means of at least one preferred exemplary embodiment and with reference to the associated drawings, in which corresponding elements are provided with the same reference numerals. The drawings show: [Brief explanation of the drawings]

[0036] [Figure 1] FIG. 1 shows an overview of the method of the present invention. [Figure 2] Figure 2 shows an overview of the system components involved. [Figure 3] Figure 3 shows the training of the machine learning model used. [Figure 4] FIG. 4 is a result chart showing the execution performance of the method of the present invention.

[0037] The solution to this problem is a software tool that empowers users to make data-driven decisions and solve formulation challenges. The tool consists of three components: 1. Experimental Data 10: The viscosity of various prototype protein formulations was measured to generate Data Set 1 of 600 data points. 2. Excipient 2 is represented in the form of relevant physical parameters and molecular fingerprints, which were generated through in silico simulations and experimentally cross-validated. 3. A machine learning model 5 that uses representation 2 from step 2 to recognize patterns in the data from step 1 and predict viscosity 3 of a new protein excipient formulation 8.

[0038] The intended interaction with the developed software tool 7 is described schematically in Figure 1. Figure 2 shows an overview of the participating hardware. Apart from the necessary experimental equipment, the hardware mainly consists of a suitable computer 6 hosting the software 7 operating the used machine learning model 5. Any kind of computer 6 suitable for use with the respective software 7 can be used, for example a standard personal computer or an industrial PC. Dataset 1 is generated by measuring the viscosity of a solution / formulation 8 containing a protein and a solution containing the same protein solution and further containing at least one viscosity-reducing excipient 2. Preferably, the at least one viscosity-reducing excipient 2 is a single viscosity-reducing excipient or a combination of two viscosity-reducing excipients.

[0039] Viscosity reduction is measured by comparing the viscosity of a protein composition that does not contain viscosity-reducing excipient 2 or a combination of viscosity-reducing excipients to the viscosity of a protein composition that contains the viscosity-reducing excipient or a combination of viscosity-reducing excipients. Measurements are carried out using different proteins at defined concentrations, different viscosity-reducing excipients or combinations of viscosity-reducing excipients at defined concentrations.

[0040] Typically, the protein composition is a liquid composition and further contains at least one buffer and at least one stabilizer. The buffer and pH are selected depending on the protein, and the pH is usually adjusted using NaOH or HCl. The composition may further contain a pharmaceutically acceptable diluent, solvent, carrier, adhesive, binder, preservative, solubilizer, stabilizer, surfactant, penetration enhancer, emulsifier, or bioavailability enhancer. Those skilled in the art know how to select appropriate additives and parameters for a liquid composition.

[0041] In a preferred embodiment, the composition according to the invention is a liquid formulation 8 and the protein is a therapeutic protein. Therapeutic proteins include antibody drugs, Fc fusion proteins, anticoagulants, blood factors, bone morphogenetic proteins, artificial protein scaffolds, enzymes, growth factors, hormones, interferons, interleukins, antibody-drug conjugates (ADCs), thrombolytic agents, etc. Therapeutic proteins can be naturally occurring or recombinant proteins. Their sequences can be natural or engineered. In a particularly preferred embodiment, the protein in the compositions and formulations according to the invention is an antibody, in particular a therapeutic antibody. In a further particularly preferred embodiment, the protein in the compositions and formulations according to the invention is a plasma-derived protein, in particular IgG or hyperIgG. Some pharmaceutical formulations containing plasma proteins are composed of a mixture of different plasma proteins.

[0042] As used herein, the term "plasma derived proteins" refers to proteins obtained from donor plasma by plasma fractionation. The donor can be human or non-human. An example of a plasma protein is an immunoglobulin. As used herein, the term "IgG" refers to immunoglobulin type G. As used herein, the term "IgM" refers to immunoglobulin type M. As used herein, the term "IgA" refers to immunoglobulin type A. As used herein, the term "hyper IgG" refers to a preparation of IgG purified from a donor infected with or vaccinated against a particular disease, said donor may be human or non-human. As used herein, the term "antibody" refers to monoclonal antibodies (including full-length or intact monoclonal antibodies), polyclonal antibodies, multivalent antibodies, multispecific antibodies (e.g., bispecific antibodies), and antibody fragments.

[0043] An antibody fragment comprises only a portion of an intact antibody, generally including the antigen-binding site of the intact antibody, and thus retains antigen-binding ability. Examples of antibody fragments encompassed by this definition include the following: Fab fragments, Fab' fragments, Fd fragments, Fd' fragments, Fv fragments, dAb fragments, isolated CDR regions, F(ab')2 fragments, as well as single-chain antibody molecules, diabodies, and linear antibodies. In one embodiment, the protein is a biosimilar. A "biosimilar" is defined herein as a biological drug that is highly similar to another already approved biological drug. In a preferred embodiment, the biosimilar is a monoclonal antibody. In one embodiment, compositions and formulations according to the invention comprise one or more protein species. The present invention is not limited to proteins within a particular molecular weight range. Preferably, the protein molecular weight is between 120 kDa and 250 kDa, more preferably between 130 kDa and 180 kDa.

[0044] To test viscosity reduction by viscosity-reducing excipient 2, one or more protein concentrations are selected that increase the viscosity of solution 8. The resulting viscosity of solution 8 is at least 20-25 mPas. -1 The viscosity should be

[0045] In a preferred embodiment, the protein concentration in the compositions and formulations according to the invention is at least 1 mg / ml, at least 50 mg / ml, preferably at least 75 mg / ml, and more preferably at least 100 mg / ml. In another preferred embodiment, the protein concentration is between 90 mg / ml and 300 mg / ml, more preferably between 100 and 250 mg / ml, and even more preferably between 120 and 210 mg / ml. The present invention is particularly useful for these high-concentration protein compositions.

[0046] There is no limitation on the choice of proteins to generate the dataset. For example, the following proteins can be used to set up a dataset: cetuximab, evolocumab, infliximab, reslizumab, etanercept (fusion protein).

[0047] As defined herein, "viscosity" refers to the resistance to flow of a substance (typically a liquid). Viscosity is related to the concept of shear stress; it can be understood as the effect of different layers of a fluid exerting shear stress against each other or against other surfaces as they move against each other. Several viscosity standards exist. The unit of viscosity is Ns / m, also known as Pascal-second (Pa-s). 2 Viscosity may be "kinematic" or "absolute." Kinematic viscosity is a measure of the rate at which momentum is transmitted through a fluid. It is measured in Stokes (St). Kinematic viscosity is a measure of the resistance of a fluid to flow under the influence of gravity. When two fluids of equal volume and different viscosities are placed in the same capillary viscometer and allowed to flow by gravity, the more viscous fluid will take longer to flow through the capillary than the less viscous fluid. If, for example, one fluid takes 200 seconds (s) to complete its flow and another takes 400 s, the second fluid is twice as viscous as the first on the kinematic viscosity scale. The dimension of kinematic viscosity is the length 2 / hour. Kinematic viscosity is generally expressed in centistokes (cSt). The SI unit of kinematic viscosity is mm 2 / s, which is equal to 1 cSt. "Absolute viscosity" (sometimes called "dynamic viscosity" or "simple viscosity") is the product of kinematic viscosity and fluid density. Absolute viscosity is expressed in units of centipoise (cP). The SI unit of absolute viscosity is millipascal-second (mPa-s), where 1 cP = 1 mPas.

[0048] Viscosity can be measured at a given shear rate or multiple shear rates, for example, using a viscometer. "Zero-shear extrapolated" viscosity can be determined by constructing a best-fit line through the four highest shear points on a plot of absolute viscosity versus shear rate and linearly extrapolating the viscosity to zero shear. Alternatively, for Newtonian fluids, viscosity can be determined by averaging viscosity values ​​at multiple shear rates. Viscosity can also be measured at single or multiple shear rates (also called flow rates) using a microfluidic viscometer, where absolute viscosity is obtained from the change in pressure as the liquid flows through the channel. Viscosity is equal to shear stress versus shear rate. In some embodiments, viscosity measured with a microfluidic viscometer can be directly compared to extrapolated zero-shear viscosity, for example, viscosity extrapolated from viscosity measured at multiple shear rates using a cone and plate viscometer. According to the present invention, the viscosity of the composition and formulation 8 is reduced if at least one of the above methods exhibits a stabilizing effect. Preferably, the viscosity is measured using mVROC™ Technology at 20° C. More preferably, the viscosity is measured using mVROC™ Technology at 20° C. Most preferably, the viscosity is measured using mVROC™ Technology with a 500 μl syringe, 3000 s -1 or 2000s -1 The viscosity is measured at 20°C using a shear rate of 100 rpm and a volume of 200 μl. Those skilled in the art will be familiar with viscosity measurements using mVROC™ technology, particularly the selection of the above parameters. Detailed specifications, methods, and settings are described in 901003.5.1-mVROC_User's_Manual.

[0049] As used herein, "shear rate" refers to the rate of change of velocity as one layer of fluid passes over an adjacent layer. The velocity gradient is the rate of change of velocity with distance from the plate. This simple case shows a uniform velocity gradient of shear rate (v1-v2) / h, with units of (cm / sec) / (cm) = 1 / sec. Therefore, the units of shear rate are reciprocal seconds, or more commonly, reciprocal time. In microfluidic viscometers, changes in pressure and flow rate are related to shear rate. "Shear rate" refers to the rate at which a material deforms. Formulations 8 containing proteins and viscosity-reducing agents typically exhibit a shear rate of approximately 0.5 s when measured using a cone and plate viscometer and a spindle appropriately selected by one skilled in the art to accurately measure the viscosity of the sample in the viscosity range of interest. -1 From about 200s -1 (i.e., a 20 cP sample is most accurately measured with a CPE40 spindle attached to a DV2T viscometer (Brookfield)); when measured using a microfluidic viscometer, approximately 20 s -1 ~about 3,000s -1 Greater than.

[0050] For classical "Newtonian" fluids, as that term is generally used herein, viscosity is essentially independent of shear rate. However, for "non-Newtonian fluids," viscosity either decreases or increases with increasing shear rate; by way of example, the fluids are "shear-thinning" or "shear-thickening," respectively. In the case of concentrated (i.e., highly concentrated) protein solutions, this can manifest as pseudoplastic shear-thinning behavior (i.e., a decrease in velocity with shear rate).

[0051] In certain embodiments, the compositions and formulations of the present invention exhibit a reduction in viscosity of at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70% or 75% compared to the same composition without the at least one first excipient. In certain embodiments, the compositions and formulations of the present invention exhibit a reduction in viscosity of at least 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70% or 75% compared to the same composition without the at least one first excipient and the at least one second excipient 2.

[0052] The present invention further provides a pharmaceutical formulation 8 according to the present invention, which has a viscosity of 1 mPas to 60 mPas, preferably 1 mPas to 50 mPas, more preferably 1 mPas to 30 mPas, and most preferably 1 mPas to 20 mPas. The compositions typically have a pH between 4 and 8, preferably between 5 and 7.2. In one embodiment, compositions and formulations 8 have a pH of exactly 5 or exactly 7.2. The pH is selected depending on the protein, and the pH is typically adjusted using NaOH or HCl. Those skilled in the art know how to select the pH of a protein composition.

[0053] At least one stabilizer is a compound suitable for enhancing protein stability. Suitable stabilizers are known in the art and include suitable sugars and / or surfactants. Suitable sugars as stabilizers are known in the literature, such as sucrose or trehalose. In a preferred embodiment, the sugar is sucrose. Suitable surfactants are known in the literature, such as polysorbate 20, polysorbate 80, or poloxamer 188. In another preferred embodiment, the surfactant is polysorbate 80. The addition of an additional stabilizer further amplifies the effectiveness of formulation 8 of the composition according to the present invention. Preferably, the sugar has a sucrose concentration of 50-100 mg / ml, more preferably 50 mg / ml. Preferably, the surfactant has a polysorbate 80 concentration of 0.01-0.2 mg / ml, more preferably 0.05 mg / ml. A buffer solution is prepared by adding at least one suitable buffer to the protein solution. Suitable buffers are known in the art, such as acetate-citrate or phosphate buffers. The concentration of the buffer is usually 1 to 50 mM.

[0054] According to the present invention, a "viscosity changing excipient 2" is a compound that can affect the viscosity of a liquid formulation. This definition includes a "viscosity reducing excipient 2" that is suitable for reducing the viscosity of liquid formulation 8 when added to formulation 8 in the concentration ranges defined below. Preferably, liquid formulation 8 is a protein solution.

[0055] There are no limitations on the choice of viscosity-reducing excipient 2 for creating dataset 1. A dataset can be created using the following viscosity-reducing excipients 2: guanidine hydrochloride, L-arginine, L-carnitine hydrochloride, L-ornithine hydrochloride, L-serine, lysine, meglumine, quinine hydrochloride, thiamine hydrochloride, ascorbic acid, benzenesulfonic acid, camphorsulfonic acid, thiamine pyrophosphate, disodium succinate sodium tartrate, folic acid, gluconic acid, glucuronic acid, pyridoxine, sodium p-toluenesulfonate, thiamine monophosphate, urea, aminocaproic acid, caffeine, cyanocobalamin, glycine, isoleucine, leucine, nicotinamide, phenylalanine, proline, sodium chloride, and valine.

[0056] Such viscosity-reducing agents are used at a concentration suitable for reducing the viscosity of the protein solution 8. In a preferred embodiment, the protein concentration in compositions and formulations according to the present invention is at least 1 mg / ml, at least 50 mg / ml, preferably at least 75 mg / ml, and more preferably at least 100 mg / ml. In another preferred embodiment, the protein concentration is between 90 mg / ml and 300 mg / ml, more preferably the protein concentration is between 100 and 250 mg / ml, and even more preferably between 120 and 210 mg / ml. Preferably, one viscosity-altering excipient 2 is used at a concentration of up to 200 mM, more preferably up to 150 mM, and most preferably 75 mM or 150 mM. When two viscosity-altering excipients 2 are used, each is preferably at a concentration of up to 150 mM, more preferably at a concentration of up to 100 mM, and most preferably at a concentration of 75 mM. As a result, concentration levels above 300 mM are undesirable, so the concentrations of both excipients 2 should not exceed 150 mM. If, for some reason, there is an unequal distribution between the two preferred excipients 2, the respective ratios will change. The same rule applies when more than one excipient 2 is used. These concentration levels are most effective for viscosity reduction and are therefore preferred. However, the method is not limited to these specific values.

[0057] The excipient dataset 1 is based on a simplified molecular-input line-entry system (SMILES) representation. All viscosity measurements are performed in defined, but distinct, environments characterized by pH values. To incorporate changes in excipient protonation, a pH-dependent microspecies distribution is generated using ChemAxons predictions. The pH dependence is pH 4–8. Each microspecies is converted to a three-dimensional structure using Marvin molconverter. From the three-dimensional trial structures, the ensemble of all conformers in aqueous solution at room temperature is calculated. For this purpose, the CREST algorithm is employed, which is a global search based on metamechanics, structure intersection, and simulated annealing run on an extended tight-binding level quantum mechanical potential energy surface encompassing the generalized born and surface-accessible region implicit solvation model (GFN2-xTB+GBSA(water)).

[0058] The zero-point and thermodynamic contributions are encompassed by the rigid-rotor-harmonic-oscillator (RRHO) model. Individual geometries are further refined with the density functional approximation B97-3c within the Conductor-Like Screening Model for Real Solvents (COSMO-RS) within the 2019.0.4 parameterization. The final single point used for the Boltzmann ensemble of conformers is composed of the electronic energy, RRHO contribution, and solvation free energy. Structures with contributions below 1% are ignored.

[0059] These microspecies ensembles are the basis for quantum chemical calculations at the density functional theory level, simulating molecular observables such as charge distribution, dipole moment, quadrupole moment trajectory and anisotropy, polarizability, molecular London dispersion coefficient (C6), logP water / hexane partition coefficient, solvent-accessible surface area, and molecular orbital energy HOMO-LUMO gap. These quantum mechanical features are complemented by a set of topological molecular fingerprints. This expanded set of 200 standardized fingerprints is generated based on the same microspecies ensemble using the RDKit. Together, this results in a high-dimensional, pH-dependent feature vector for each excipient.

[0060] The developed machine learning model 5 combines experimental data obtained in the lab with calculated in silico excipient characteristics to construct a predictive model for formulation viscosity. Based on this model 5, an optimal experiment schedule is provided. A formulation specialist 9 uses these suggestions to perform the recommended experiments 10, and then feeds the newly obtained viscosity data into the system, as exemplarily shown in Figure 3. Through this process, experiments are performed focusing on formulations with the greatest potential for viscosity reduction.

[0061] For a given protein, viscosity values ​​are modeled in the form of a Gaussian Process (GP), and the predictions of the model 5 are used in a preferred embodiment to guide the formulation specialist 9 through Bayesian optimal experimental design. This guidance involves several steps: 1. In the case where the set of viscosity measurements for a particular excipient / excipient combination (=training data) is potentially empty, the GP model parameters are optimized by maximizing the marginal likelihood of the training data. 2. The posterior distribution of viscosity values ​​for untested excipients / excipient combinations is estimated based on the GP model. 3. A new set of excipient / excipient combinations is selected by optimizing the obtained scores obtained from the calculated posterior distribution. 4. A new set of excipient / excipient combinations will be proposed to the Formulation Specialist 9, who will then conduct experiments on each in the lab to determine the resulting viscosity. 5. The obtained measurements are added to the training data and the process is repeated from step 1. Includes.

[0062] The predictions in step 2 are based on the pH-dependent feature vectors characterizing excipient 2 and the excipient concentration levels used in formulation 8 under consideration. The challenge with steps 1–5 of this procedure is that the measured viscosity depends not only on the selected excipient combination but also on the ground-truth protein concentration, which may vary from measurement to measurement. To account for deviations from the target concentration, GP Model 5 is designed to predict the relative change in viscosity rather than absolute viscosity values. More precisely, it predicts the relative viscosity reduction relative to the theoretical viscosity level achieved at the actual protein concentration without excipients. The required theoretical value is obtained from an exponential regression model calculated from concentration-dependent viscosity measurements of the unformulated protein solution.

[0063] Although the experimental design considered in steps 1-5 follows a typical optimization procedure, existing black-box GP models5 provided by major software suites such as GPyTorch, BoTorch, and GPflow are not applicable to the given scenario due to the specific data characteristics that need to be encoded. Therefore, specialized GP kernel structures are designed given the following domain- and problem-specific characteristics. These properties are as follows: There is no natural order to the combination of excipients 2 contained in a given formulation, i.e., adding excipient A + excipient B is equivalent to adding excipient B + excipient A. The kernel used is designed to be permutation invariant with respect to the excipients 2 added. A given formulation 8 can contain a variable number of excipients 2. The kernel is constructed to be flexible in its ability to handle a flexible number of excipients. The induced viscosity-reducing effect of formulation 8 depends on both the given protein concentration and the applied excipient concentrations. The dependence on these concentrations is clearly reflected in the structure of the kernel used. Combining excipients can result in synergistic viscosity-reducing effects that cannot be described through the properties of each excipient. A general, generic kernel structure based on automatic association detection of individual feature dimensions cannot adequately capture these multivariate relationships. To generalize from measured data to untested excipient combinations, the kernel used uses a linear subspace projection that is optimized during the parameter fitting process.

[0064] Particularly challenging is generalizing the predicted viscosity 3 to new proteins due to the lack of chemical information characterizing the protein's global and local interactions. Therefore, a further extended preferred embodiment is recommended. This includes a database containing viscosity measurements of various formulations 8 that exhibit typical interaction patterns between proteins and excipients 2. These interaction patterns can be used as prior information for the predicted viscosity 3 of new proteins in the form of an additional kernel component that biases the prediction of matching protein-excipient patterns. One approach to achieve this, although others are possible, is to capture the influence of proteins via a multitasking kernel model, such as the intrinsic co-regionalization model or its variants. In alternative embodiments of the present invention, the Gaussian process can be replaced by other machine learning models, and other appropriate model components can take over the task of generalizing across proteins.

[0065] In the following, a specific example is disclosed to demonstrate the advantages of using the software tool 7 compared to performing a randomly selected, uninformed search in the following experiment. The goal is to reduce the viscosity of the protein solution 8 below a specified threshold. Given the vast array of excipients 2 available on the market, finding a suitable excipient combination is challenging. To avoid an exhaustive screening study testing all candidate formulations, an informed, data-driven search is performed with the help of the proposed software tool 7.

[0066] To achieve this goal, follow these steps: 1) The specific formulation conditions, especially the pH value, and which excipients 2 can be considered as potential candidates are defined. 2) A small number of concentration-dependent viscosity measurements are performed on solutions containing the new protein without at least one viscosity-reducing agent. This data is input into a software tool 7, which estimates a base viscosity curve for the unformulated protein based on the predicted viscosity 3. By consulting the software tool 7 after each measurement, the user is instructed as to which protein concentration to consider next, and this information will be available once a sufficient amount of data has been collected. 3) The software tool 7 then recommends the first excipient 2 or excipient combination to test. The user performs each experiment 10 in the lab and reports the measured viscosity to the tool. In an iterative process, the user is prompted to run further experiments depending on the latest measurements reported to the tool 7 until a formulation 8 with a sufficiently low viscosity is found. is carried out.

[0067] If measurements have already been taken before using Tool 7, e.g. for excipient 2 that is not on the candidate list, the user can report the corresponding viscosity before starting the process, allowing Tool 7 to provide improved recommendations from the start.

[0068] In an alternative embodiment, the user can run several experiments at once before consulting software tool 7 after each iteration. In this so-called "batch mode," the user can input the desired number of experiments to run in parallel during the next iteration, e.g., for purposes of scheduling laboratory resources. Software tool 7 then optimizes its recommendations in such a way as to optimize the expected information gain from running the experiments simultaneously.

[0069] Figure 4 shows the viscosity reduction achieved with both search strategies depending on the number of experiments performed by the user. In the given example, a total of 629 experiments covering 6 proteins and 33 excipients were considered. To average the results across all proteins tested, the measured viscosity reduction is reported relative to the maximum reduction observed per protein, and the number of experimental steps is shown relative to the total number of experiments (10) performed per protein. Mean values ​​(solid lines) and standard deviations (shaded areas) obtained from several experimental repetitions are shown. These repetitions are obtained by considering different sets of initial measurements provided to the software tool and different random experimental paths for the random baseline strategy.

[0070] Consistent with the theoretical number of steps required, a random strategy would find the expected optimal excipient combination after performing 50% of all possible experiments. Using the invented approach, this number can be reduced by half on average.

[0071] A further embodiment of the present invention is a new formulation 8 containing at least one viscosity-altering excipient 2 selected by the method provided above.

[0072] A further embodiment of the present invention is a pharmaceutical formulation comprising the novel formulation 8 and at least one viscosity-altering excipient 2 selected by the method described above.

[0073] A further embodiment of the present invention is a pharmaceutical formulation comprising the novel formulation 8 and at least one viscosity-altering excipient 2 selected by the method described above.

[0074] example Example 1. Experimental data generation / viscosity measurement General Concepts of Experiments To generate the experimental data for Dataset 1, various protein compositions were prepared and various viscosity-reducing excipients were tested for viscosity reduction. The following commercially available proteins were used: cetuximab, evolocumab, infliximab, reslizumab, etanercept. The following commercially available viscosity-reducing excipients 2 were used: guanidine hydrochloride, L-arginine, L-carnitine hydrochloride, L-ornithine hydrochloride, L-serine, lysine, meglumine, quinine hydrochloride, thiamine hydrochloride, ascorbic acid, benzenesulfonic acid, camphorsulfonic acid, thiamine pyrophosphate, disodium succinate, disodium tartrate, folic acid, gluconic acid, glucuronic acid, pyridoxine, sodium p-toluenesulfonate, thiamine monophosphate, urea, aminocaproic acid, caffeine, cyanocobalamin, glycine, isoleucine, leucine, nicotinamide, phenylalanine, proline, sodium chloride, valine, and combinations thereof. Below is exemplified the measurement of viscosity reduction of valine as viscosity-reducing excipient 2 in an infliximab solution. The general concepts of this particular example can be carried over to all other proteins and viscosity-reducing agents used.

[0075] When a single viscosity-reducing agent was used, the viscosity was typically measured at a concentration of 150 mM. When a combination of two excipients 2 was used, the viscosity was typically measured at a concentration of 75 mM for each excipient. In some cases, the concentration of viscosity-reducing excipient 2 was adjusted depending on the solubility of excipient 2. The buffer, pH, protein concentration, and optional stabilizers and / or surfactants were selected depending on the protein used. Typically, the buffer, pH, stabilizers, and / or surfactants of commercial protein-containing products were used. The protein solution was concentrated to a viscosity of at least 20 mPas. -1In some cases, viscosity was measured at multiple protein concentrations.

[0076] Viscosity measurement Buffer adjustment 5 mM phosphate buffer was prepared by appropriately mixing sodium dihydrogen phosphate and disodium hydrogen phosphate to produce a pH of 7.2 and dissolving the mixture in ultrapure water. The ratio was determined using the Henderson-Hasselbalch equation. The pH was adjusted using HCl and NaOH, if necessary. 50 mg / ml sucrose and 0.05 mg / ml polysorbate 80 were added as stabilizers.

[0077] Sample preparation Each excipient solution was prepared at 150 mM valine in phosphate buffer pH 7.2. The pH was adjusted as needed using HCl or NaOH. Concentrated infliximab solutions containing the desired excipients were prepared using centrifugal filters (Amicon, 30 kDa MWCO), exchanging the original buffer with a buffer containing the respective excipient, and reducing the volume of Solution 8. The protein was subsequently diluted to 122 mg / ml and 143 mg / ml, respectively. In a similar manner, identical protein solutions were prepared except that they did not contain valine.

[0078] Measurement of protein concentration Protein concentrations were determined using absorption spectroscopy applying the Beer-Lambert law. If the excipient itself had strong absorbance at 280 nm, the Bradford assay was used. The concentrated protein solution was diluted so that the expected concentration at the time of measurement was between 0.3 and 1.0 mg / mL. For absorption spectroscopy, a BioSpectrometer® kinetic (Eppendorf, Hamburg, Germany) was used to determine the absorbance at 280 nm with a protein extinction coefficient A of 0.1% at 280 nm = 1.428. Because some excipients 2 themselves exhibit strong absorption at 280 nm, a Bradford assay was required for concentration determination. The Bradford assay used a kit from Thermo Scientific™ (Thermo Fisher, Waltham, Massachusetts, USA) and a bovine gamma globulin standard. Absorbance was measured at 595 nm using a Thermo Scientific™ (Thermo Fisher, Waltham, Massachusetts, USA). Protein concentrations were determined by linear regression of a standard curve ranging from 125 to 1500 μg / ml.

[0079] Viscosity measurement The viscosity was measured using mVROC™ Technology (manufactured by Rheo Sense, San Ramon, California, USA). Measurements were performed using a 500 μl syringe and 3000 s -1 The measurements were performed at 20°C using a shear rate of 0.25 s. A volume of 200 μl was used. All samples were measured in triplicate. Viscosity reduction was calculated by comparing the absolute viscosity of the protein composition with and without valine.

[0080] 2. Measurement of concentration-dependent viscosity of raw protein solution Below we take the measurement of the concentration-dependent viscosity of infliximab as an example, and the general concepts of this particular example can be taken for all other proteins. Buffer and sample preparation was performed as described above. Protein was subsequently diluted to 13, 30, 42, 68, 79, 80, 103, 110, 117.30, 121, and 148.2 mg / ml, respectively. Protein concentration measurements using absorption spectroscopy applying the Beer-Lambert law were performed as described above. Viscosity measurements of different infliximab concentrations were performed using mVROC™ Technology (RheoSense, San Ramon, CA, USA).

[0081] Reference List 1. Dataset 2. Excipients (expressions) 3 Predicted viscosity 4. Selected excipients (combinations) 5. Machine Learning Models 6. Computer used 7 Software Tools 8 New formulations 9. User (Formulation Specialist) 10 Experimental measurement data 11 Unknown Protein

Claims

1. 1. A method for selecting, via a computer (6), at least one viscosity-altering excipient (2) for a formulation (8) containing at least one unknown protein (11), comprising the steps of: Providing a data set (1) from a database describing the viscosity of several known formulations containing at least one protein and optionally at least one viscosity-altering excipient (2); generating, by a computer (6) via in silico simulation, a representation of at least one excipient (2) from a list of excipients; - evaluating the viscosity altering effect of at least one viscosity altering excipient (2) selected from the list of excipients on a new formulation (8) containing at least one unknown protein (11) and at least one viscosity altering excipient (2) by applying the recognized pattern to provided data of the at least one unknown protein (11) using a machine learning model (5) performed on a computer (6) that recognizes a pattern in the dataset (1) using the generated representation of the at least one excipient (2); Depending on the evaluation results, selecting at least one excipient from the list according to the acquisition criteria and applying it to the unknown protein (11); Including, wherein the provided data of at least one unknown protein (11) is data describing the viscosity of a protein composition comprising at least one unknown protein (11) and, optionally, at least one viscosity-altering excipient (2), The method.

2. 2. The method of claim 1, wherein the evaluation of the viscosity-altering effect of the at least one viscosity-altering excipient (2) is performed by predicting the viscosity (3) of a new formulation (8) containing at least one unknown protein (11) and at least one viscosity-altering excipient (2).

3. 2. The method of claim 1, wherein the data set (1) is generated by experimental measurements (10) and stored in a database via a computer (6).

4. 2. The method according to claim 1, wherein a combination of two or more excipients from the list is used as at least one excipient from the list that most significantly changes the viscosity of the new formulation (8).

5. 2. The method of claim 1, wherein at least one specific experimental measurement (10) is proposed to a formulation specialist (9), who performs at least one respective experiment (10) in a laboratory to validate the predicted viscosity (3), and trains the machine learning model (5) by adding the results of the validation to the provided dataset (1) in a database via a computer (6).

6. 2. The method of claim 1, wherein the machine learning model (5) is created and trained by combining a dataset (1) describing the viscosity of at least one prototype protein formulation (8) with a representation of at least one viscosity-altering excipient (2) or a combination thereof.

7. 7. The method of claim 6, wherein viscosity values ​​of a given formulation (8) are modeled via a machine learning model (5) in the form of a Gaussian process, and model predictions (3) are used to guide a formulation specialist (9) via a Bayesian optimal experimental design.

8. Training the machine learning model (5) on a computer (6) comprises the following steps: Optimizing machine learning model parameters on training data from dataset (1) by maximizing the marginal likelihood of the training data; - assessing the posterior distribution of viscosity values ​​for an untested excipient (2) or combination thereof based on a machine learning model (5), thereby predicting viscosity (3); selecting a new set of excipients (2) or combinations thereof by optimizing the obtained scores obtained from the calculated posterior distribution; - propose new excipients (2) or combinations thereof to the formulation specialist (9), who then performs the respective experiments (10) in the laboratory and measures the resulting viscosities; Add the obtained measurements (10) to the training data: The method of claim 6, which is carried out by performing the steps at least once.

9. 9. The method of claim 8, wherein the prediction of the viscosity (3) obtained from the posterior distribution of the viscosity values ​​(3) is based on a pH-dependent feature vector characterizing the excipients (2) used in the formulation (8) considered and on the excipient concentration levels used.

10. 2. The method of claim 1, wherein the obtaining criterion is which viscosity-changing excipient (2) reduces the viscosity the most.

11. 2. The method according to claim 1, wherein the representation of the excipient (2) is generated by a computer (6) in the form of physical parameters and a molecular fingerprint.

12. 2. The method of claim 1, wherein the generated representations of excipients (2) are experimentally cross-validated.

13. 2. The method of claim 1, wherein the generated representation of the excipient (2) comprises quantum mechanical features and is optionally complemented with a set of topological molecular fingerprints.

14. A computer-implemented machine learning model created and trained in accordance with claims 5 to 13.

15. 15. The machine learning model of claim 14, wherein the Gaussian process is replaced by other model architectures that fulfill the same objective, in particular other types of stochastic processes, deep Bayesian networks, generalized linearity models, neural networks, support vector machines, tree-based models, ensemble models, etc.