Protein solution

A neural network trained with diverse data types predicts protein solution viscosity accurately, addressing the challenge of high-concentration viscosity prediction for therapeutic proteins, ensuring effective pharmaceutical formulations and administration.

JP2025538138APending Publication Date: 2025-11-26LONZA AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025525613
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-25
Filing Date
2023-11-03
Publication Date
2025-11-26

AI Technical Summary

Technical Problem

Existing methods struggle to accurately predict the concentration-dependent viscosity of protein solutions, particularly at high concentrations relevant for subcutaneous administration of therapeutic proteins like monoclonal antibodies, which is crucial for assessing drug-related CMC issues and ensuring effective pharmaceutical formulations.

Method used

A computer-implemented neural network is trained using a combination of experimental, computational, and in silico data to predict the concentration-dependent viscosity of protein solutions, incorporating input parameters such as hydrophobicity, diffusion interaction parameter, isoelectric point, and charged patch size, achieving high accuracy in viscosity prediction.

Benefits of technology

The neural network achieves prediction accuracy of greater than 0.95, enabling early evaluation of protein solutions' suitability for pharmaceutical use by accurately determining viscosity within the desired range, thus optimizing formulation and administration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025538138000001_ABST
    Figure 2025538138000001_ABST
Patent Text Reader

Abstract

The present invention relates to a method for providing a computer-implemented neural network configured to predict the concentration-dependent viscosity of a protein solution, and further to a method for predicting the concentration-dependent viscosity of a protein solution, determining the concentration-dependent viscosity of a protein solution, and providing a pharmaceutical product comprising the protein solution by using such a computer-implemented neural network.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for providing a computer-implemented neural network configured to predict the concentration-dependent viscosity of a protein solution, and further to a method for predicting the concentration-dependent viscosity of a protein solution, determining the concentration-dependent viscosity of a protein solution, and providing a pharmaceutical product comprising the protein solution by using such a computer-implemented neural network. [Background technology]

[0002] Therapeutic proteins, such as monoclonal antibodies (mAbs), have become important factors in the treatment of a wide variety of diseases, including cancer, immune-mediated disorders, and infectious diseases. The most common route of administration for protein-based pharmaceuticals (DPs) is intravenous, which requires the patient to be hospitalized and the pharmaceutical to be administered by a medical professional. For patients with chronic diseases, the need for repeated drug administration, and therefore hospitalization, can cause significant burden and stress and jeopardize the success of the intended therapy. Subcutaneous (sc) injection allows patients to self-administer protein-based pharmaceuticals, such as monoclonal antibody-based pharmaceuticals, through the use of prefilled syringes, autoinjectors, or other delivery devices, which often improves the quality of life and compliance of patients with chronic conditions.

[0003] However, subcutaneous administration has certain limitations. Typically, the volume of a single injection is limited to less than about 2 mL, which is determined by the available subcutaneous space and the patient's tolerable pain perception. This volume limitation necessitates the use of pharmaceuticals with high drug substance concentrations.

[0004] Generally, mAbs have high specificity but also require significant therapeutic dosages, which therefore result in highly concentrated pharmaceutical preparations (DPs) containing solutions of proteins such as mAbs, often exceeding 100 mg / mL of protein in solution for subcutaneous administration.

[0005] However, as concentration increases, intermolecular distance decreases, allowing proteins to interact with each other. These protein-protein interactions (PPIs) then exponentially affect and determine the solubility, aggregation, and viscosity of proteins (e.g., mAbs). PPIs are determined, among other things, by the primary amino acid sequence of a protein (e.g., mAb) and the resulting three-dimensional structure containing charged or hydrophobic patches. Solution conditions can also affect protein-protein interactions, for example, by adjusting the size of charged patches via pH or by shielding charged patches via short-range electrostatic interactions using salts, buffers, amino acids, or other charged excipients. Arginine is a common excipient tested for viscosity reduction, and its dual mode of action is to shield both charged and hydrophobic patches. 17 of 34 FDA-approved drugs containing high concentrations of mAbs use salts or amino acids as excipients, presumably to reduce protein-protein interactions and thus the viscosity of the mAb solution. More exploratory excipients that have demonstrated viscosity-reducing potential but have not been applied in marketed pharmaceuticals, primarily because they are not approved as parenteral excipients or because of concerns regarding toxicity, include poly-l-glutamic acid, caffeine, hydrophobic salts, or amino acid derivatives.

[0006] Highly viscous solutions can be a significant obstacle in the development of protein-based pharmaceuticals. Disadvantages include high costs due to significant purification losses and low recovery rates, difficulties in manufacturing or filling, and poor administration due to the need for strong injection forces, which can result in long and painful administration. Generally, solutions with dynamic viscosities greater than about 15-30 mPa*s, or even greater than about 15-20 mPa*s, are considered problematic. The desire to develop high-concentration formulations is not a priori obvious for new molecules and may also arise as a result of the need for unexpectedly high doses or changes in the target administration route. Clinical trials typically begin in Phase 1 with formulations at lower concentrations (<50 mg / mL protein in solution), while higher protein concentrations (≥100 mg / mL protein in solution) are typically explored in later phases once safety and efficacious dose levels have been established.

[0007] Predicting the viscosity of new proteins (especially mAbs) at high protein concentrations early in development is essential to assess drug-related CMC (chemistry, manufacturing, and quality control) issues once dose ranges are established later in development.

[0008] Highly concentrated mAb solutions have been extensively studied over the past two decades to understand the factors that lead to high viscosity.

[0009] In early development, multiple candidates are often available in small quantities and need to be tested in preformulation studies for their stability and solubility. A considerable amount of research has been done over the past few years to predict the viscosity of high concentration solutions using experimental data from low concentration experiments.

[0010] Various techniques are known to predict the viscosity of protein solutions used as pharmaceuticals in clinical therapy, which often rely on experimental parameters such as the diffusion interaction parameter (kD) or computational tools that utilize information derived from the primary sequence of the protein.

[0011] Until Roberts's study (Woldeyes, MA, Qi, W., Razinkov, VI, Furst, EM, & Roberts, CJ. How Well Do Low- and High-Concentration Protein Interactions Predict Solution Viscosities of Monoclonal Antibodies? J Pharm Sci 108, 142-154 (2019)), experimental colloidal interaction data, primarily the diffusion interaction parameter (kD) and the second virial coefficient (A2), had been found to be able to at least qualitatively predict potentially high-viscosity mAbs (problematic mAbs). However, Roberts highlighted numerous cases where the predictions were unsuccessful. This calls into question the validity of using low-concentration experimental data as a predictive tool. Most publications also reference a fairly small number of samples and linearly describe the relationship between experimental data and viscosity. The complexity of the origin of solution viscosity at high mAb concentrations suggests that a nonlinear modeling approach may be more appropriate. In recent years, experimental approaches to predict the concentration-dependent viscosity of antibodies have been complemented by in silico methods that aim to identify crucial molecular descriptors, such as solvent exposure, local charge, hydrophobic effects, and surface patches. One model for predicting viscosity is known from Agrawal et al. (Agrawal, NJ, et al., Computational tool for the early screening of monoclonal antibodies for their viscosities, MAbs 8, 43-48 (2016)), which uses space charge maps (SCMs) to identify mAbs with high viscosity. SCMs apply molecular dynamics (MD) simulations to calculate a viscosity score for screening antibodies at high concentrations. However, MD simulations are computationally expensive and require structural information, which is a major bottleneck in their application.The principles of SCM were also used in a machine learning approach to predict the viscosity of mAb solutions (Lai, P.-K. DeepSCM: An efficient convolutional neural network surrogate model for the screening of therapeutic antibody viscosity. Comput Struct Biotechnol J, 20, 2143-2152 (2022)). The deep learning algorithm established by Lai used preprocessed antibody sequences as input for model training and SCM scores obtained from MD simulations as output. A DeepSCM surrogate model for SCM scores was developed based on a convolutional neural network (CNN) architecture. The SCM scores were then used to predict viscosity at high concentrations. Lai reported an accuracy of 0.9. Summary of the Invention

[0012] The present invention refers to a method for providing a computer-implemented neural network configured to predict the concentration-dependent viscosity of a protein solution based on a plurality of input parameters associated with the protein solution. Furthermore, the present invention, in particular the method, may include providing and using a computer-implemented neural network, in particular a trained computer-implemented neural network. In particular, the providing and using may be performed to provide a trained computer-implemented neural network, in particular configured to predict the concentration-dependent viscosity of a protein solution based on a plurality of input parameters associated with the protein solution. The method, in particular the method using a computer-implemented neural network, in particular the method using a trained computer-implemented neural network, may include: - providing a plurality of training data sets, each of the plurality of training data sets being associated with a particular protein solution and including a plurality of input parameters indicative of the particular protein solution and at least one associated output parameter indicative of the concentration-dependent viscosity of the particular protein solution; - training the neural network based on the training data set to provide a computer-implemented neural network configured to predict concentration-dependent viscosity, wherein the input parameters are: i) experimental data, preferably selected from apparent surface hydrophobicity as measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), second virial coefficient (A2), and combinations thereof; ii) computational data, preferably selected from the isoelectric point (pI) of the protein, the fragment variable (Fv) charge (Fv charge), and combinations thereof; and iii) Preferably, the in silico data includes at least one of hydrophobicity and charged patch size, in particular selected from score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total.

[0013] Further embodiments of the present invention include computer-implemented methods for predicting the concentration-dependent viscosity of a protein solution by using such a computer-implemented neural network, in particular a trained computer-implemented neural network, and more particularly, the provision and use of a computer-implemented neural network trained as described above. The computer-implemented methods for predicting the concentration-dependent viscosity of a protein solution of the present invention have high accuracy, which can be reflected by an R value of greater than 0.95, and even greater than 0.99.

[0014] Yet further embodiments of the present invention include methods for determining the concentration-dependent viscosity of a protein solution by using such a computer-implemented neural network, and in particular such a trained computer-implemented neural network.

[0015] A still further embodiment of the present invention comprises a method for providing a pharmaceutical product comprising a protein solution, which allows for an effective and efficient evaluation of the suitability of said protein solution for use as a pharmaceutical product already at an early stage of development.

[0016] Preferably, any of the methods disclosed herein that use a computer-implemented method for predicting or determining the concentration-dependent viscosity of a protein solution uses a computer-implemented neural network that was initially configured (hereinafter sometimes referred to as "trained") to predict the concentration-dependent viscosity of a protein solution as described herein, and includes all embodiments thereof.

[0017] In the context of the present disclosure, the task of a computer-implemented neural network to calculate concentration-dependent viscosity may be referred to as determining or predicting concentration-dependent viscosity. The terms "determining" and "predicting" can be used interchangeably in connection with the calculations of a computer-implemented neural network.

[0018] In the context of the present disclosure, the suitability of a protein to be used as a pharmaceutical in the form of a solution is determined by the viscosity of the protein solution at high concentrations, in particular at concentrations above 100 mg / mL of protein in solution, which must not exceed 15-30 mPa*s, preferably 15-20 mPa*s, in particular 15 mPa*s.

[0019] These objects are solved by the subject matter of the independent patent claims. Preferred embodiments are set out in the description, the drawings and the dependent claims.

[0020] Accordingly, the present invention also includes a method for providing a computer-implemented neural network (hereinafter sometimes referred to as a "neural network"), preferably an artificial neural network (ANN), also referred to as a trained ANN. In particular, the present invention includes the provision and use of such a computer-implemented neural network. The neural network, in particular a trained neuron network, is configured to predict or determine the concentration-dependent viscosity of a protein solution depending on a plurality of input parameters associated with said protein solution. The method includes the steps of providing a plurality of training datasets, each of the plurality of training datasets being associated with a particular protein solution and including a plurality of input parameters indicative of the particular protein solution and at least one associated output parameter indicative of the concentration-dependent viscosity of the particular protein solution; and training the neural network, i.e., a trained neural network or an untrained neural network, based on the training datasets to provide a computer-implemented neural network configured to predict concentration-dependent viscosity, i.e., providing a trained ANN.

[0021] The term "protein solution" refers to an aqueous solution of a protein, preferably a therapeutic protein (often referred to as an active pharmaceutical ingredient, "API"). Thus, a protein solution refers to any solution that includes at least one protein dissolved or substantially dissolved in an aqueous medium, such as water, a buffer, or a cell culture medium.

[0022] The term "protein" refers to a polypeptide consisting of a sequence of amino acids. A protein can be, for example, a therapeutic protein used in the diagnosis, treatment, and / or prevention of a disease or disorder. A protein can be a naturally occurring protein, i.e., a protein produced by a non-recombinant cell occurring in nature, or it can be produced by genetic engineering or a recombinant cell and can include molecules having the amino acid sequence of a naturally occurring protein, or molecules having deletions, additions, and / or substitutions of one or more amino acids of the naturally occurring sequence, or molecules having the amino acid sequence of a protein unrelated to the naturally occurring protein. The term also includes amino acid polymers in which one or more amino acids are chemical analogues of the corresponding naturally occurring amino acid polymer.

[0023] The therapeutic proteins contained in the protein solution may be antibodies, which may include monoclonal and polyclonal antibodies, whole antibodies, antibody-drug conjugates, chimeric antibodies, humanized antibodies, human antibodies, or hybrid antibodies with dual or multiple antigen or epitope specificities, antibody fragments and antibody subfragments (e.g., Fab, Fab', F(ab')2, fragments, etc.), hybrid fragments of any immunoglobulin or hybrid fragment of any natural, synthetic, or genetically engineered protein that act like antibodies by binding to a specific antigen to form a complex, in one embodiment, a monoclonal antibody. The term "epitope" refers to a protein determinant capable of specific binding to an antibody. Epitopes usually consist of chemically active surface groupings of molecules, such as amino acids or sugar side chains, and typically have specific three-dimensional structural characteristics as well as specific charge characteristics. Conformational and nonconformational epitopes are distinguished in that the binding to the former, but not the latter, is lost in the presence of denaturing solvents.

[0024] A preferred therapeutic protein is a monoclonal antibody. As used herein, the term "monoclonal antibody" (mAb) refers to an antibody obtained from a population of substantially homogeneous antibodies, i.e., the individual antibodies comprising the population are identical except for possible naturally occurring mutations that may be present in minor amounts. Monoclonal antibodies are highly specific, being directed against a single antigenic site. Furthermore, in contrast to polyclonal antibody preparations, which include different antibodies directed against different determinants (epitopes), each monoclonal antibody is directed against a single antigenic site on the antigen. The modifier "monoclonal" indicates the character of the antibody as being obtained from a substantially homogeneous antibody population and should not be construed as requiring production of the antibody by any particular method. Preferred monoclonal antibodies are of the immunoglobulin G (IgG) class, such as IgG1, IgG2, IgG3, and IgG4.

[0025] Suitable aqueous media for protein solutions are known to those skilled in the art. Preferably, the aqueous medium is or comprises water or an aqueous buffer such as a histidine-HCl buffer.

[0026] Other buffers may be used and buffers commonly used in pharmaceutical formulations such as acetate, phosphate, succinate or citrate buffers will be known to those skilled in the art.

[0027] Preferably, the aqueous medium is a buffer solution having a pH of 5 to 8, more preferably 5 to 7.5, and even more preferably 5.5 to 6.5.

[0028] The aqueous medium may also have other pH values ​​suitable for providing a protein solution, and typical pH values ​​of aqueous media for pharmaceutical preparations are known to those skilled in the art.

[0029] Preferably, any input parameter and at least one associated output parameter (the latter provided for configuring a computer-implemented neural network) are also determined in a similar aqueous medium used for the intended pharmaceutical product containing the protein solution, or more preferably, in the same aqueous medium having the same pH, same buffer, same viscosity, particularly the same buffer viscosity, etc.

[0030] At least one pharmaceutically acceptable excipient may also be included in the protein solution. Suitable excipients are known to those skilled in the art, such as stabilizers, pH adjusters such as one or more buffers, etc. In some embodiments, polysorbates such as polysorbate 20, 40, 60, or 80 may be used as stabilizers, which can improve the stability of proteins in aqueous solutions. Other stabilizers may be, for example, sugars or sugar alcohols, surfactants, salts, or antioxidants (e.g., L-methionine).

[0031] As described above, a method for providing a computer-implemented neural network, particularly a trained computer-implemented neural network, also referred to as a trained ANN, includes providing a plurality of training datasets. Furthermore, using a computer-implemented neural network, particularly a trained neural network, may include providing a plurality of training datasets. Each of the training datasets is associated with a specific protein solution. Specifically, at least two of the plurality of training datasets are associated with different protein solutions. More specifically, each of the training datasets may be associated with a specific protein, i.e., a protein contained in the protein solution. Preferably, the plurality of training datasets may be associated with different proteins, particularly constituting a set of proteins. In other words, each protein contained in the set of proteins is associated with at least one training dataset.

[0032] Preferably, the proteins underlying each training dataset are different proteins but similar to each other in at least one aspect. The proteins in the set of proteins can be any of the types of proteins described above. More preferably, all proteins in the set of proteins of the training dataset are antibodies, and even more preferably, all proteins in the set of proteins are monoclonal antibodies.

[0033] To improve the training of computer-implemented neural networks, particularly ANNs, more specifically the computer-implemented neural networks to be trained, the number of training datasets, and therefore the number of proteins, may be greater than 2. The accuracy of the viscosity prediction (or determination) can be expressed, for example, by an accuracy value comparable to the value of the coefficient of determination R2, which ranges from 0 (reflecting no accuracy of the determination, which may also be called the accuracy of the prediction) to 1 (reflecting a 100% accurate determination).

[0034] For example, the number of training datasets (and therefore proteins in the protein set) may be at least 15, preferably at least 20, more preferably at least 25. Generally, more training datasets result in higher accuracy. As the number of proteins in the protein set increases, the accuracy of the determination no longer increases linearly and approaches an upper limit of 1. For example, when 25 training datasets are provided, the accuracy of the determination is usually quite high.

[0035] In the context of the present disclosure, the term "concentration-dependent viscosity" refers to a parameter that indicates the viscosity of a protein solution at one or more protein concentrations of the protein solution, particularly the dynamic viscosity of the protein solution. Thus, concentration-dependent viscosity relates at least one protein concentration of the protein solution to the corresponding viscosity of the protein solution. For example, concentration-dependent viscosity is the viscosity value η of a protein solution at a selected protein concentration, particularly a single protein concentration. csIn other words, the concentration-dependent viscosity can be or can be indicative of a viscosity in the form of a viscosity value η of a protein solution at a selected protein concentration. cs can be expressed and / or determined as

[0036] Alternatively, concentration-dependent viscosity may refer to the viscosity of a protein solution at two or more protein concentrations. Specifically, concentration-dependent viscosity refers to a function f that relates the viscosity of a protein solution to the protein concentration of the protein solution. η , in particular, can be represented by a mathematical function. For example, concentration-dependent viscosity can be represented by a curve diagram relating the viscosity value of a protein solution to the protein concentration.

[0037] Furthermore, the concentration-dependent viscosity can be expressed by the function given by equation (1) below: f η (c)=A×e B×c (1) In the formula, f η refers to a function relating the viscosity of a protein solution to the protein concentration, c refers to the protein concentration, and A and B refer to constants.

[0038] Equation (1) can be linearized to obtain the intercept A and slope B using natural logarithms, as in equation (2). ln η rel =ln A+B×c(2) Equation (2): Linearized viscosity behavior, where η rel refers to relative viscosity, A refers to intercept, B refers to slope, and c refers to protein concentration.

[0039] In a further development, the concentration-dependent viscosity of the protein solution may be represented by at least one, preferably both, of the constants A and B in the above equation (1). In other words, the computer-implemented neural network, in particular the trained computer-implemented neural network, more particularly the trained ANN and / or the computer-implemented neural network to be trained may be configured to determine at least one, preferably both, of the constants A and B in the above equation (1).

[0040] The term computer-implemented neural network (also referred to herein as "neural network") may be an artificial neural network. Thus, the term neural network as used in this disclosure may be used interchangeably with the term "artificial neural network" (ANN), and in particular, with the terms training an ANN or being trained. In general, an artificial neural network refers to a computing system using interconnected nodes, and in particular, a functional unit that utilizes machine learning. An artificial neural network is a subset of machine learning, also known as deep learning, and uses unstructured datasets and hidden layers. As with regular machine learning, the dataset for establishing a neural network is typically divided into a training dataset and a validation dataset.

[0041] In an ANN, interconnected nodes, which may also be referred to as artificial neurons, can receive signals, process these signals, and send signals to their connected nodes based on the received and processed signals. Typically, the signals sent to a node are referred to as inputs and usually represent real numbers. The output of each node is typically calculated by a nonlinear function, also referred to as an activation function, based on the sum of its inputs. Nodes and their connections typically have weights, which are adjusted during the training phase based on a training dataset. A validation dataset is then used to validate the trained ANN, allowing the accuracy of the trained ANN to be determined. Typically, in an ANN, nodes are organized into layers, as described below. Typically, signals travel from the first layer (i.e., input layer), through hidden layers, to the last layer (i.e., output layer), possibly after passing through the layers multiple times.

[0042] The general functionality and operation of such ANNs is well known to those skilled in the art and will therefore not be described further. Rather, the technical features of ANNs that are interrelated with the present invention are highlighted below.

[0043] Typically, to provide an ANN suitable for predicting concentration-dependent viscosity, a trained ANN, particularly an untrained ANN, may first be provided, which may then be appropriately trained based on a training dataset. Thus, the method may further include providing a neural network to be trained (preferably an untrained ANN).

[0044] Specifically, the neural network, particularly the neural network to be trained and / or the trained neural network, may include an input layer having a plurality of input nodes, each of which receives one input parameter. In other words, each of the plurality of input nodes preferably receives one input parameter. The received input parameters preferably differ between the input nodes.

[0045] Alternatively or additionally, the neural network, particularly the neural network to be trained and / or the trained neural network, may include at least one hidden layer, particularly a single hidden layer with multiple hidden nodes, particularly four or more hidden nodes. According to one configuration, the neural network, particularly the neural network to be trained and / or the trained neural network, may include a single hidden layer with four hidden nodes. The hidden nodes may use tan h as an activation function, which transforms values ​​between -1 and 1. Alternatively, a sigmoid may be used as the activation function for the hidden nodes.

[0046] Alternatively or additionally, the neural network, particularly the neural network to be trained and / or the trained neural network, may include an output layer having at least one output node for providing an output parameter. Specifically, the output layer may have multiple output nodes, each of which determines and outputs a single output parameter. The output parameter provided by the output layer may be used to determine the concentration-dependent viscosity. Alternatively, the output parameter provided by the output layer may indicate or represent the concentration-dependent viscosity. For example, the output layer may be configured to provide at least one output parameter indicating at least one of the constants A and B in the above equation (1). In other words, the output layer may be configured to provide at least one value as the output parameter used as the constant A or the constant B in the above equation (1). For example, the output layer may be configured to provide a first output parameter via a first output node used as the constant A and a second output parameter via a second output node used as the constant B to determine the concentration-dependent viscosity represented by the function specified in the above equation (1).

[0047] Generally, a neural network, in particular a neural network to be trained and / or a trained neural network, receives a plurality of input parameters based on which at least one output parameter is calculated. Thus, a training data set includes or consists of the input parameters and at least one output parameter.

[0048] In the step of training the neural network, preferably a training data set, in particular, the weights of the nodes of the neural network and their connections are adapted depending on the training data set, which is used to implement or establish the neural network. In this way, the trained neural network becomes a trained neural network, i.e., the neural network provided by the above method. In the following, the training data set will be further described. This provides an ANN that allows training and therefore prediction of the viscosity of a protein solution with high accuracy.

[0049] Each set of training data sets preferably has the same number and type of parameters. Furthermore, as described above, each training data set is associated with a specific protein solution, in particular a specific protein. Specifically, the training data sets can be subdivided into input parameters, i.e., parameters representing the input to the neural network, and output parameters, i.e., parameters representing the output of the neural network. The input parameters indicate the specific protein solution with which the corresponding training data set is associated. Thus, the input parameters may indicate properties of the specific protein solution or of the protein contained therein. The output parameters indicate the concentration-dependent viscosity of the specific protein solution. Thus, the output parameters represent or can indicate the concentration-dependent viscosity of the specific protein solution.

[0050] Preferably, the input parameters included in the training data set are indicative of protein-protein interactions in a relevant protein solution.

[0051] Intermolecular distance decreases with increasing concentration, and protein-protein interactions then exponentially affect and determine the solubility, aggregation, and viscosity of protein solutions. Therefore, by providing input parameters indicative of the protein-protein interactions of a protein solution to a neural network, the proposed method takes into account these properties, which substantially affect the viscosity of a protein solution at increasing protein concentrations. By doing so, effective and efficient prediction of the viscosity of a protein solution can be achieved.

[0052] Generally, protein-protein interactions are influenced by the amino acid sequence of the protein and, consequently, the three-dimensional structure with charged or hydrophobic patches. Furthermore, solution conditions can affect protein-protein interactions by adjusting the size of the charged patches through pH, ​​or by shielding the charged patches through short-range electrostatic interactions using salts, buffer substances, amino acids, or other charged excipients.

[0053] As mentioned above, multiple input parameters can be i) experimental data obtained by detecting parameters from the provided protein solution, preferably the experimental data representing parameters selected from protein hydrophobicity, diffusion interaction parameter (kD), net protein charge, zeta potential, second virial coefficient (A2), third virial coefficient (A3), apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), and combinations thereof; ii) computational data, preferably calculated from the primary amino acid sequence of the protein at a pH of 5.0 to 7.0, preferably representing parameters selected from the protein's isoelectric point (pI), fragment variable (Fv) charge (Fv charge), Fv symmetry parameter, Fv hydrophobicity, Vh charge, Vl charge, hinge charge, and hydrophobic solvent accessible surface area, and combinations thereof; and iii) in silico data, preferably including at least one of in silico data selected from hydrophobicity and charged patch size, and combinations thereof.

[0054] In other words, the plurality of input parameters may include at least one selected from the group consisting of i) experimental data, ii) computational data, and iii) in silico data.

[0055] In particular, the calculated data and in silico data may be data calculated depending on the primary amino acid sequences of the proteins contained in the protein solution.

[0056] The i) experimental data is preferably a parameter selected from at least one of the apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), the diffusion interaction parameter (kD), the second virial coefficient (A2), and a combination thereof. In other words, the i) experimental data may include at least one parameter selected from the group consisting of the apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), the diffusion interaction parameter (kD), and the second virial coefficient (A2). More preferably, the i) experimental data consists of the parameters of the apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), the diffusion interaction parameter (kD), and the second virial coefficient (A2).

[0057] The ii) calculated data is preferably a parameter selected from at least one of the isoelectric point (pI) of the protein, the variable fragment (Fv) charge (Fv charge), and a combination thereof. In other words, the ii) calculated data may include at least one parameter selected from at least one of the isoelectric point (pI) of the protein and the variable fragment (Fv) charge (Fv charge). More preferably, the ii) calculated data consists of the parameters isoelectric point (pI) of the protein and the variable fragment (Fv) charge (Fv charge).

[0058] The iii) in silico data is preferably a parameter selected from at least one of hydrophobicity and charged patch size, such as score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total. More preferably, the iii) in silico data consists of score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total. In other words, the iii) in silico data includes at least one parameter selected from the group consisting of score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total.

[0059] The use of computational data and in silico modeling has the advantage of being relatively easy to access and does not require materials or laboratory work.

[0060] Methods for determining the diffusion interaction parameter (kD) are known to those skilled in the art. A suitable method may illustratively be via dynamic light scattering (DLS), for example, according to the procedure further outlined below.

[0061] Virial coefficients can be determined in a variety of ways, such as by static light scattering (SLS), in which case the virial coefficients are generally designated A2 for the second virial coefficient and A3 for the third virial coefficient, or by measuring osmotic pressure, in which case the virial coefficients are generally designated B2 for the second virial coefficient and B3 for the third virial coefficient.

[0062] Methods for determining the second (A2) or third (A3) virial coefficients are known to those skilled in the art. A suitable method may illustratively be via static light scattering (SLS), for example, according to the procedure further outlined below.

[0063] Methods for determining apparent surface hydrophobicity are known to those skilled in the art. A suitable method may illustratively be via hydrophobic interaction chromatography (HIC), for example, according to the procedure further outlined below.

[0064] Methods for determining the calculated data are known to those skilled in the art. Suitably, the respective parameters, such as the Fv charge, are calculated at a pH of 5-8, or 5-7.5, or 5.5-6.5, or about 5.5, or about 6.0, or about 6.5, preferably pH 6.0. Those skilled in the art can follow the procedure further outlined below and return to commonly known platforms such as Prot pi|Protein Tool, https: / / www.protpi.ch / Calculator / ProteinToo, site operator Roland Josuran Prot pi, 8820 Wädenswil, Switzerland.

[0065] The isoelectric point (pI) can be calculated "manually" from the pKa values ​​of the amino acid residues in the primary sequence.

[0066] Methods for determining in silico data are known to those skilled in the art. Suitably, they may be determined using the software BioLuminate (version 3.80, Schroedinger, LLC, New York, NY), for example, according to the procedure further outlined below.

[0067] Specifically, the neural network, in particular the neural network to be trained and / or the trained neural network, may receive input parameters selected from all three of the above data groups i) to iii), but reliable prediction of concentration-dependent viscosity is also possible when the neural network receives input parameters from only two of the three data groups, such as receiving input parameters from the above data groups i) and ii), or from data groups ii) and iii), or from data groups i) and iii), preferably from data groups ii) and iii).

[0068] As described above, the neural network, particularly the neural network to be trained and / or the trained neural network, receives a plurality of input parameters based on which at least one output parameter is calculated. The output parameter may represent or be indicative of the concentration-dependent viscosity of a particular protein solution. Thus, the output parameter may be a viscosity value η that relates the viscosity of the protein solution to a particular protein concentration in the protein solution. cs Alternatively or additionally, the at least one output parameter may be indicative of the viscosity of the protein solution at a plurality of protein concentrations. Specifically, the at least one output parameter may be indicative of the function f η More specifically, the at least one output parameter may represent at least one, and preferably both, of the constant A and the constant B in equation (1) above.

[0069] The following provides an example of a method for determining at least one output parameter to be included in a training dataset. First, different protein solutions, referred to as protein solution samples, may be provided, which contain the same type of protein but at different concentrations. Based on these samples, concentration-dependent viscosity data may be generated, for example, by using a VROC viscometer (Rheosense) for each concentration. In this way, multiple value pairs may be provided, each of which relates the viscosity of the protein solution to the protein concentration in the protein solution. These measured viscosity data, particularly experimentally measured viscosity data, may then be processed in a next step based on the following equations (3) to (5):

[0070] For a substance, equation (3) can be used to calculate the relative viscosity of different protein solution samples. To do so, the viscosity of the buffer solution contained in the protein solution can first be determined, preferably at a defined temperature. Once the viscosity of the buffer solution is known, the relative viscosity of the protein solution sample can then be calculated based on the viscosity data measured according to equation (3). By doing so, a pair of concentration and relative viscosity values ​​is determined for each protein solution sample.

number

[0071] Equation (4) can then be used to describe the exponential concentration-dependent viscosity of a protein solution. This equation can be linearized to obtain the intercept A and slope B using natural logarithms, as in equation (5). η rel =A×e B×cs (4) Equation (4): Exponential viscosity behavior; η rel = relative viscosity; A = intercept; B = slope; cs = selected mAb concentration ln η rel =ln A+B×cs(5) Equation (5): Linearized viscosity behavior; η rel = relative viscosity; A = intercept; B = slope; cs = selected mAb concentration

[0072] Based on the determined pair of values ​​of concentration and relative viscosity of the protein solution sample, the viscosity descriptors A and B for each protein can then be determined based on equations (4) and / or (5) by applying a curve fitting technique, such as, for example, a least squares fit.

[0073] For each protein in the set of proteins of the training and validation datasets, a set of numbers of protein solutions having different protein concentrations (num-conc) is provided, and the sample viscosity η of each of these is determined. num-conc is at least two, preferably 2 to 12, e.g., 4, 5, 6, 7, 8, 9, 10, 11, or 12, with a typical value of 6. These pairs of values ​​of protein concentration and sample viscosity serve to determine at least one associated output parameter of the training and validation datasets, as described above.

[0074] The different concentrations in the set of protein solutions may range, for example, from a minimum concentration (min-conc) of 10 mg / mL to a maximum concentration (max-conc) of 320 mg / mL, or from 10 to 280 mg / mL, or from 10 to 240 mg / mL, or from 10 to 210 mg / mL, with a typical value for the minimum concentration being 30 mg / mL and a typical value for the maximum concentration being 180 mg / mL. Preferably, when the set of protein solutions includes three or more protein solutions each having a different protein concentration, the difference between every two adjacent concentrations, i.e., the difference between two subsequent concentrations, is the same, so that (max-conc-min-conc) / (num-conc-1)=diff-conc, and conc n+1 -conc n =diff-conc. [Table 1]

[0075] In one embodiment, the set of protein solutions for each protein in the set of proteins includes six protein solutions having six different concentrations: 30 mg / mL, 60 mg / mL, 90 mg / mL, 120 mg / mL, 150 mg / mL, and 180 mg / mL, such that [Table 2]

[0076] The method according to the present invention may further include a step of validating the trained neural network after the step of configuring the neural network (also referred to herein as training) has been performed. In this "validation" step, multiple validation data sets may be provided. In other words, the method may further include validating the trained neural network based on multiple validation data sets, each of which includes multiple input parameters and at least one associated output parameter. The validation data sets may have the same or different number and types of parameters as the training data set, but it is preferable that the validation data set uses the same types of parameters as the training data set. Thus, each validation data set may be divided into input parameters and associated output parameters, similar to the training data set. The validation data sets may then be used to verify the proper implementation or training of the neural network. To do so, the trained neural network may receive the input parameters of the validation data set to calculate at least one corresponding associated output parameter. The calculated output parameter may then be compared with at least one associated output parameter included in the validation data set. This may be performed for each set of validation data sets to determine whether a proper configuration of the neural network has been established.

[0077] In a further development, the trained neural network may be stored on a data medium, for example a non-transitory storage medium configured to store digital data of a type that will be apparent to those skilled in the art in view of this disclosure.

[0078] A further aspect of the present invention comprises a computer-implemented method for determining or predicting the concentration-dependent viscosity of a protein solution by using a computer-implemented neural network, in particular a trained neural network, more particularly a neural network trained as described above. In particular, a further aspect of the present invention may comprise the provision and use of a computer-implemented method for determining or predicting the concentration-dependent viscosity of a protein solution by using a computer-implemented neural network. The neural network, in particular the trained neural network, is configured to predict the concentration-dependent viscosity of a protein solution depending on a plurality of input parameters associated with said protein solution, preferably the input parameters being: i) experimental data, preferably selected from apparent surface hydrophobicity as measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), second virial coefficient (A2), and combinations thereof; ii) computational data, preferably selected from the isoelectric point (pI) of the protein, the fragment variable (Fv) charge (Fv charge), and combinations thereof; and iii) Preferably, the in silico data includes at least one of hydrophobicity and charged patch size, in particular selected from score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total.

[0079] In other words, the input parameters may include at least one selected from the group consisting of i) experimental data, ii) computational data, and iii) in silico data.

[0080] Preferably, the computer-implemented method of the present invention utilizes a neural network, in particular the above-described trained neural network (i.e., the neural network provided by the above-described method). Accordingly, the technical features described herein in relation to the method for providing a computer-implemented neural network, in particular in relation to the provision and use of a computer-implemented neural network, may therefore also apply and refer to the computer-implemented method for predicting the concentration-dependent viscosity of a protein solution, and vice versa. This applies in particular to the above-described features of the neural network to be trained or the trained neural network, the protein solution, any input parameters, and any output parameters.

[0081] The method of the present invention can be used to predict the viscosity of any suitable protein solution. As described above, the protein solution may contain a therapeutic protein, preferably an antibody, more preferably a monoclonal antibody (mAb), such as a monoclonal antibody of the IgG1 or IgG2 subtype. Thus, the protein solution may be an aqueous medium containing water, preferably a buffer, more preferably a histidine-HCl buffer.

[0082] The neural networks, particularly trained neural networks, used in the methods of the present invention are configured to determine or predict concentration-dependent viscosity in dependence on multiple input parameters.

[0083] As mentioned above, the concentration-dependent viscosity is calculated by the viscosity value η of a protein solution at a selected protein concentration. cs Alternatively or additionally, the concentration-dependent viscosity can be expressed and / or determined as a function f that relates the viscosity of a protein solution to the protein concentration. η Specifically, the concentration-dependent viscosity can be expressed and / or determined as a function f η More specifically, a neural network may be represented by or be expressed by a function fη The control unit 10 may be configured to calculate at least one of the constant A and the constant B.

[0084] In the method according to the present invention, a neural network, in particular a trained neural network, receives input parameters and determines the concentration-dependent viscosity of a protein solution based thereon. The input parameters are preferably associated with the protein solution whose viscosity is to be predicted. Specifically, the input parameters preferably refer to proteins contained in the protein solution. More preferably, the input parameters indicate protein-protein interactions in the protein solution. Furthermore, the input parameters may correspond to those input parameters contained in the training dataset, as described above. Thus, the above technical features related to the parameters contained in the training dataset may also apply and similarly refer to the input parameters received by the neural network to calculate the concentration-dependent viscosity, and vice versa.

[0085] More specifically, in this method, the neural network, in particular the trained neural network, receives input parameters as described above. Thus, the input parameters are: i) Preferably, the experimental data includes at least one of the following: apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), second virial coefficient (A2), and combinations thereof. More preferably, the experimental data i) consists of the following parameters: apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), and second virial coefficient (A2). ii) calculated data, preferably selected from the isoelectric point (pI) of the protein, the variable fragment (Fv) charge (Fv charge), and combinations thereof. More preferably, ii) the calculated data consists of the parameters isoelectric point (pI) of the protein and the variable fragment (Fv) charge (Fv charge), iii) In silico data, preferably selected from hydrophobic and charged patch sizes, in particular score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total.

[0086] In other words, the input parameters may include at least one selected from the group consisting of i) experimental data, ii) computational data, and iii) in silico data.

[0087] Specifically, as described above, input parameters can be selected from all three of the above data groups i) to iii), but reliable prediction of concentration-dependent viscosity is also possible when the neural network receives input parameters from only two of the three data groups, such as receiving input parameters from the above data groups i) and ii), or from data groups ii) and iii), or from data groups i) and iii), preferably from data groups ii) and iii).

[0088] In a further aspect of the present invention, there is provided a method for determining the concentration-dependent viscosity of a protein solution, the method comprising predicting the concentration-dependent viscosity of the protein solution by using a computer-implemented neural network, and comparing the concentration-dependent viscosity to a target viscosity, in particular to determine the viscosity characteristics of the protein solution.

[0089] The proposed method may utilize the above neural network and the above computer-implemented method for predicting the concentration-dependent viscosity of a protein solution. Thus, the above technical features, particularly in relation to the neural network and its use for predicting the concentration-dependent viscosity of a protein solution, may also be applied and referred to in the method for determining the concentration-dependent viscosity of a protein solution, and vice versa.

[0090] The proposed method can be used to evaluate whether a protein solution is subject to high or too high viscosity, particularly at a given protein concentration. In other words, comparing the concentration-dependent viscosity with a target viscosity can be performed to determine whether a protein solution is subject to high or too high viscosity, particularly exceeding the target viscosity (particularly at a given protein concentration). For the reasons cited herein above, a protein solution can be considered to have too high a viscosity if it becomes unsuitable for pharmaceutical use, i.e., if its production results in high costs due to high purification losses and low recovery rates, difficulties in manufacturing or filling, and poor administration due to the need for strong injection forces, long administration times, and potential pain. Generally, solutions with a dynamic viscosity greater than 15-30 mPa*s, or even greater than 15-20 mPa*s, are considered problematic and therefore "too high."

[0091] In a further development, the method may be used to determine the pharmaceutical suitability of a protein solution, in particular a protein solution used or prepared as a pharmaceutical product, To do so, a step of comparing the concentration-dependent viscosity with a target viscosity is carried out to determine the pharmaceutical suitability of the protein solution.

[0092] By providing a method in which the pharmaceutical suitability of a protein solution is evaluated based on the concentration-dependent viscosity determined by a computer-implemented neural network, particularly a trained neural network, the proposed method makes it possible to easily and reliably take into account the concentration-dependent viscosity of a protein solution, which depends on the concentration of the protein in the protein solution, particularly at an early stage of the development process. In this way, improved validation of protein solutions for use as pharmaceuticals is possible at an early stage, which has an impact on the resulting pharmaceutical product.

[0093] The proposed method can be used to evaluate the suitability of a protein solution for use in any pharmaceutical product that contains or is constituted by a protein solution.

[0094] As described above, the method includes comparing the determined concentration-dependent viscosity with a target viscosity. In the context of the present disclosure, the term "target viscosity" refers to a desired or predetermined viscosity of a protein solution, or the target viscosity may represent an upper threshold value for the viscosity of the protein solution. For example, the target viscosity may indicate a maximum viscosity value of the protein solution at a given concentration. Specifically, the target viscosity may indicate a maximum viscosity value of the protein solution of 15 mPa*s to 30 mPa*s, or 15 mPa*s to 25 mPa*s, or 15 mPa*s to 20 mPa*s, e.g., 15 mPa*s, 18 mPa*s, 20 mPa*s, or 25 mPa*s. Viscosities exceeding the maximum value may be considered problematic when using the protein solution as a pharmaceutical.

[0095] Therefore, this step can be performed to determine the suitability of the protein solution as a pharmaceutical. In other words, by comparing the concentration-dependent viscosity with the target viscosity, it is possible to evaluate whether the protein solution has a preferable concentration-dependent viscosity when used as a pharmaceutical. The suitability of the protein solution can be determined if the concentration-dependent viscosity meets or is less than the target viscosity, i.e., if the viscosity is within the desired range even when the protein is used at a high concentration, for example, >50 mg / mL, or >60 mg / mL, >70 mg / mL, or >80 mg / mL, >90 mg / mL, or >100 mg / mL. By doing so, the application-specific viscosity of the protein solution can be predicted, and whether the protein solution is suitable for use as a pharmaceutical in clinical therapy can be evaluated.

[0096] Preferably, the step of comparing the concentration-dependent viscosity to the target viscosity takes into account the protein concentration of the protein solution, in other words, the concentration-dependent viscosity and the target viscosity may be compared at a particular protein concentration or within a range of protein concentrations.

[0097] When comparing the concentration-dependent viscosity and the target viscosity at a specific protein concentration, the protein concentration of the protein solution can be determined first. For example, the maximum protein concentration can be determined. This can refer to the protein concentration that is not expected to exceed when the protein solution is used as a pharmaceutical. The maximum protein concentration can be in the range of 100 mg / mL to 220 mg / mL, or 120 mg / mL to 180 mg / mL, for example, 120 mg / mL, 150 mg / mL, or 180 mg / mL. Then, based on the determined concentration-dependent viscosity, the viscosity of the protein solution at the maximum protein concentration can be determined and compared with the target viscosity.

[0098] When comparing the concentration-dependent viscosity and the target viscosity within a protein concentration range, the realistic protein concentration range of the protein in the protein solution for clinical therapy can be determined first. For example, the protein concentration range can range from 20 mg / mL, 50 mg / mL, or 100 mg / mL to the maximum protein concentration. Then, it can be determined whether the concentration-dependent viscosity of the protein solution exceeds the target viscosity within the protein concentration range.

[0099] In a further aspect of the present invention, there is provided a method for providing (or preparing) a pharmaceutical product comprising a protein solution, the method comprising the steps of predicting (or determining) the concentration-dependent viscosity of the protein solution by using a computer-implemented neural network, comparing the concentration-dependent viscosity with a target viscosity to determine the suitability of the protein solution as a pharmaceutical product, and the optional step of preparing the pharmaceutical product if the suitability of the protein solution has been determined.

[0100] The proposed method may utilize the above neural network, the above computer-implemented method for predicting the concentration-dependent viscosity of a protein solution, and the above method for determining the concentration-dependent viscosity. Therefore, the above technical features may also be applied to and refer to a method for providing a pharmaceutical product, and optionally a method for preparing a pharmaceutical product, and vice versa.

[0101] In this method, if the suitability of the protein solution as a pharmaceutical product is determined as a result of comparing the concentration-dependent viscosity with the target viscosity, then a step of preparing a pharmaceutical product can be performed. If the suitability of the protein solution as a pharmaceutical product is not determined, the protein solution can be adapted or modified, and the steps of determining the concentration-dependent viscosity of the protein solution and comparing the determined concentration-dependent viscosity with the target viscosity can be performed again based on the adapted or modified protein solution.

[0102] Further provided is a method for determining a suitable concentration of a protein in a protein solution for a pharmaceutical product, the method comprising: determining a concentration-dependent viscosity of the protein solution by using a computer-implemented neural network; and comparing the determined concentration-dependent viscosity with a target viscosity to determine an upper limit for the concentration of the protein in the protein solution that is still acceptable for the pharmaceutical product.

[0103] Further provided is a method for identifying an upper limit on the concentration of a protein in a protein solution that must not be exceeded to avoid an unacceptably high viscosity of the protein solution, the method comprising: determining a concentration-dependent viscosity of the protein solution by using a computer-implemented neural network; and comparing the determined concentration-dependent viscosity with a target viscosity to determine whether a protein in the protein solution induces a viscosity of the protein solution that is unacceptable for pharmaceutical use above a certain concentration.

[0104] Further provided is a method for identifying proteins in solution that have an unacceptable concentration-dependent viscosity, the method comprising: determining the concentration-dependent viscosity of a protein solution by using a computer-implemented neural network; and comparing the determined concentration-dependent viscosity to a target viscosity to determine whether a protein in the protein solution, above a certain concentration, induces a viscosity of the protein solution that is unacceptable for pharmaceutical use.

[0105] The present invention further provides the use of a computer-implemented neural network to facilitate the formulation of a pharmaceutical product, wherein the neural network is configured to predict the concentration-dependent viscosity of a protein solution for use in determining the suitability of the protein solution for preparation as a pharmaceutical product.

[0106] Furthermore, the present invention provides a method for determining the pharmaceutical suitability of a protein solution, comprising the steps of: experimentally detecting a parameter from the provided protein solution, preferably wherein the obtained experimental data represents a parameter selected from protein hydrophobicity, diffusion interaction parameter (kD), net protein charge, zeta potential, second virial coefficient (A2), third virial coefficient (A3), apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), and combinations thereof; determining the concentration-dependent viscosity of the protein solution by using a computer-implemented neural network; and comparing the determined concentration-dependent viscosity to a target viscosity to determine the pharmaceutical suitability of the protein solution.

[0107] Additionally, the present invention provides a method for providing a trained artificial neural network (ANN) for determining the concentration-dependent viscosity of a protein solution by training the ANN with input parameters describing proteins included in a set of proteins.

[0108] Furthermore, the present invention provides a method for determining (or predicting) the concentration-dependent viscosity of a solution of a protein by providing a trained ANN with input parameters indicative of the protein or input parameters of the protein and having the ANN calculate an output parameter indicative of or which is the concentration-dependent viscosity.

[0109] The present invention further provides an apparatus for predicting the concentration-dependent viscosity of an experimental protein solution, comprising: an experimental protein solution; a computer component including a neural network configured to predict the viscosity of an experimental protein solution; A neural network is trained using a plurality of predetermined data sets; each of the plurality of predetermined data sets comprises (i) a plurality of input parameters indicative of at least one physical characteristic of the predetermined protein solution, and (ii) at least one output parameter indicative of the concentration-dependent viscosity of the predetermined protein solution; An apparatus is provided, wherein the plurality of input parameters comprises at least one of experimentally derived data, computationally derived data, and in silico derived data.

[0110] The proposed device may utilize the above neural network and the above computer-implemented method for predicting the concentration-dependent viscosity of a protein solution. Accordingly, the above technical features, particularly in relation to neural networks, more particularly in relation to the above trained neural network, and its use for predicting the concentration-dependent viscosity of a protein solution, may therefore also be applied and referred to in the proposed device, and vice versa.

[0111] The present invention further provides a system for determining the concentration-dependent viscosity of a protein solution, the system comprising: A protein solution; and a neural network configured to determine the viscosity of the protein solution, wherein determining the viscosity of the protein solution is dependent on a plurality of input parameters (IPs) associated with the protein solution, and further wherein the input parameters (IPs) include at least one from the group consisting of: (i) experimental data, (ii) computational data, and (iii) in silico data.

[0112] The proposed system may utilize the above neural network and the above computer-implemented method for predicting the concentration-dependent viscosity of a protein solution. Accordingly, the above technical features, particularly related to neural networks, more particularly related to the above trained neural network, and its use for predicting the concentration-dependent viscosity of a protein solution, may therefore also be applied and referred to in the proposed system, and vice versa.

[0113] The present disclosure will be more readily understood by reference to the following detailed description when considered in conjunction with the accompanying drawings, in which: [Brief explanation of the drawings]

[0114] [Figure 1a] 1 illustrates exemplary computing components that may be used to implement various features of embodiments of the present invention. [Figure 1b] FIG. 1 is a flow diagram illustrating a method for providing a pharmaceutical product, according to an embodiment of the present invention. [Figure 2] 2 shows a schematic diagram of a computer-implemented neural network used in the method shown in FIG. 1; [Figure 3] 2 shows schematically a further computer-implemented neural network for use in the method shown in FIG. 1; [Figure 4] 2 shows the training and validation data sets used to train and implement the neural network used in the method shown in FIG. 1; [Figure 5] FIG. 10 shows a diagram illustrating a comparison between viscosity values ​​calculated by the proposed neural network and measured viscosity values. [Figure 6-1]Examples of predicted viscosity curves are shown (predicted vs. measured viscosity for selected mAbs; x indicates measured value, black lines are drawn using predicted values ​​of slope (B) and intercept (A) inserted into the formula X (y = A x e(B x x)). Dashed lines indicate viscosity thresholds of problematic mAbs at 15 mPa*s. A) is mAb13, B) is mAb2, C) is mAb17, D) is mAb24, E) is mAb26, and F) is mAb27). [Figure 6-2] Same as above. [Figure 6-3] Same as above. [Figure 7] The difference in predicted viscosity values ​​calculated from the predicted viscosity curve using the predicted values ​​of intercept and slope relative to the viscosity values ​​measured at the same concentration is shown (the percentage difference is presented as a bar, the absolute difference is presented as a dot). [Figure 8] The input variables and settings of the illustrated artificial neural network are shown (inputs are combined into one layer with four hidden nodes with a tan h activation function and a target decision viscosity descriptor, intercept A or slope B). [Figure 9a] Examples of training data and validation data are shown below (training data and validation data for model xx-A). [Figure 9b] Examples of training data and validation data are shown below (training data and validation data for model xx-B). DETAILED DESCRIPTION OF THE INVENTION

[0115] The invention will now be described in more detail with reference to the accompanying drawings, in which like elements are designated by the same reference numerals and repeated description may be omitted to avoid redundancy.

[0116] 1a is a schematic diagram illustrating an exemplary computing component 2 that may be used to implement various features of embodiments of the present invention. More specifically, computing component 2 generally includes a bus 3 and is connected to (i) a processor 4, (ii) memory 5 (e.g., random access memory (RAM) or other dynamic memory, read-only memory (ROM) or any other static storage device for storing static information and instructions for processor 4, etc.) for storing information and instructions executed by processor 4, (iii) storage device 6 (e.g., hard disk drive, solid-state disk drive, optical storage, etc.) comprising a non-transitory medium for long-term storage of digital information (e.g., computer software, digital data, etc.), and (iv) a communication interface 7 for allowing software and data to be transferred between computing component 2 and external devices via a communication channel 8.

[0117] Computing component 2 may be part of or constitute an apparatus for predicting the concentration-dependent viscosity of an experimental protein solution.

[0118] FIG. 1b shows a method for providing and optionally preparing a pharmaceutical product comprising a protein solution.

[0119] In step S0 of the method, a computer-implemented neural network 10 is provided, which is configured to predict the concentration-dependent viscosity of a protein solution depending on a plurality of input parameters associated with the protein solution. Step S0 represents a sub-method, i.e., a method as included in a method for preparing a pharmaceutical product. It will be understood that the computer-implemented neural network 10 (which is preferably a trained neural network) may be implemented using the computing component 2 described above, or any suitable computing system that will be apparent to those skilled in the art in view of the present disclosure.

[0120] The protein solution comprises a protein, preferably a therapeutic protein such as an antibody, which may be a monoclonal antibody, a polyclonal antibody, a whole antibody, an antibody-drug conjugate, a chimeric antibody, a humanized antibody, a human antibody, or a hybrid antibody with dual or multiple antigen or epitope specificity, which may include antibody fragments and antibody subfragments (e.g., Fab, Fab', F(ab')2, fragments, etc.), hybrid fragments of any immunoglobulin or hybrid fragment of any natural, synthetic, or genetically engineered protein that acts like an antibody by binding to a specific antigen to form a complex. In the embodiment shown, the therapeutic protein is a monoclonal antibody (mAb), preferably of the IgG1 or IgG2 subtype.

[0121] Additionally, the protein solution comprises an aqueous medium comprising water, preferably a buffer, more preferably a histidine-HCl buffer.

[0122] In a first substep S0.1, a neural network 10 to be trained, particularly an untrained neural network, is provided. The neural network 10 to be trained may be configured to receive a plurality of input parameters and calculate at least one output parameter based thereon. In the illustrated configuration, the neural network 10 is configured to calculate a concentration-dependent viscosity of the protein solution depending on the input parameters. In general, the concentration-dependent viscosity relates at least one protein concentration of the protein solution to a corresponding viscosity (particularly, dynamic viscosity) of the protein solution.

[0123] According to one embodiment of the present invention shown in FIG. 2, the neural network 10, in particular the neural network to be trained and the trained neural network, is configured to detect a selected protein concentration c S The method is configured to calculate the viscosity value of a protein solution at a selected protein concentration, c Smay refer to the maximum protein concentration expected to be relevant for a pharmaceutical product. The maximum protein concentration may be in the range of 100 mg / mL to 220 mg / mL, or 120 mg / mL to 180 mg / mL, for example, 120 mg / mL, 150 mg / mL, or 180 mg / mL. Thus, in this embodiment, the concentration-dependent viscosity is determined by the selected protein concentration c S η means the viscosity value of the protein solution at cs It is expressed as:

[0124] According to another embodiment of the present invention shown in Figure 3, the neural network 10, in particular the neural network to be trained and the trained neural network, is configured to calculate a mathematical function relating the viscosity value of the protein solution to the protein concentration of the protein solution. Specifically, in this embodiment, the concentration-dependent viscosity is calculated using the function f as specified in equation (1) above. η Thus, in this embodiment, the concentration-dependent viscosity is expressed by the function f η To do so, the constants A and B in the above equation (1) are calculated by the neural network 10.

[0125] In the following, the neural network 10 to be trained will be explained in more detail with reference to the embodiment shown in FIGS.

[0126] 2 shows an embodiment of a neural network 10 used in the method shown in FIG. 1, in particular the neural network to be trained and the trained neural network. The neural network 10 includes the inputs IN1 to IN i The input layer 10 includes an input layer 12 having a plurality of input nodes IN1 to IN2, where the index "i" refers to a positive integer. In other words, the input layer 10 includes i different input nodes IN1 to IN2. As can be seen in FIG. 2, each input node IN1 to IN2 n are the corresponding input parameters IP1 to IP i Receive.

[0127] Furthermore, the neural network 10 includes a plurality of hidden nodes HN1 to HN j where the index "j" refers to a positive integer, preferably j is 4, and the neural network has four hidden nodes HN1 to HN4, each of which is connected to a respective input node IN1 to IN i are connected to hidden nodes HN1 to HN j uses tan h as the activation function. Alternatively, a sigmoid can be used as the activation function.

[0128] Furthermore, the neural network 10 includes all hidden nodes HN1 to HN j The output layer 16 has one single output node ON connected to the selected protein concentration c s The parameter η, which is the viscosity of the protein solution at cs In this embodiment, the parameter η cs constitutes the concentration dependent viscosity.

[0129] 3 shows a further embodiment of a neural network 10 according to the present invention, in particular a neural network to be trained and a trained neural network. Compared to the configuration shown in FIG. 2, the output layer 16 of the neural network 10 is provided with two output nodes ON1 and ON2 providing different output parameters. In particular, the first output node ON1 provides an output parameter OP1 which is A, and the second output node ON2 provides an output parameter OP2 which is B. As mentioned above, these two parameters are determined by the function f η In this embodiment, the function f represented by the constants A and B is η constitutes the concentration dependent viscosity.

[0130] In the next substep S0.2, multiple training datasets S Tis provided, each of which is associated with a particular protein solution and includes a plurality of input parameters IP indicative of the particular protein solution and at least one associated output parameter OP indicative of the concentration-dependent viscosity of the particular protein solution. Further, a validation data set S having the same number and types of input parameters and output parameters is provided. V FIG. 4 shows the training data set S used to properly implement the neural network 10. T and validation dataset S V Here is a table that shows the general structure of each training dataset S T and each validation dataset S V is the input node IN1 to IN i The input parameters IP1 to IP i and two output parameters OP1 and OP2 representing values ​​provided by two output nodes ON1 and ON2.

[0131] Specifically, the computational model underlying the neural network 10 is adapted to m different training data sets S during the training phase in the next sub-step S0.3. T In this context, the parameter "m" denotes a positive integer.

[0132] Specifically, in substep S0.3, the training dataset S T Based on this, a neural network 10, particularly a trained neural network, is trained to provide a computer-implemented neural network 10, particularly a trained neural network, that is suitable for and configured to predict concentration-dependent viscosity. To do so, a training data set S T All input parameters IP and output parameters OP of the neural network 10 are fed to the neural network 10 to adapt the underlying computational model of the neural network 10, in particular to adapt the weights of the nodes and their connections.

[0133] In a further optional sub-step, several n different validation data sets S V is used to validate the neural network 10, in particular the trained neural network. In this context, the parameter "n" denotes a positive integer. Preferably, the validation data set S V Each of these is a training dataset S T In order to validate the provided neural network 10, particularly the trained neural network, a validation data set S V The input parameters associated with each validation data set S are fed to the neural network 10, based on which V Then, the calculated output parameters are used to calculate the output parameters of the validation data set S V is compared with the output parameters contained in

[0134] For example, to train such a neural network, 20 or more (e.g., 24) different training datasets S T Furthermore, two or more (e.g., three) different validation data sets S V can be used.

[0135] As mentioned above, the neural model may include a number of i different input nodes and therefore may receive i different input parameters IP per dataset. The input parameters IP preferably represent protein-protein interactions in the associated protein solution. Furthermore, the input parameters IP may include different types of input parameters that can be categorized into i) experimental data, ii) computational data, and iii) in silico data.

[0136] The experimental data can be obtained by detecting parameters from a provided protein solution. Preferably, i) the experimental data represents a parameter selected from the group consisting of apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), a diffusion interaction parameter (kD), a second virial coefficient (A2), and a combination thereof. More preferably, i) the experimental data consists of the parameters of apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), a diffusion interaction parameter (kD), and a second virial coefficient (A2). The calculated data can be calculated from the primary sequence of the protein at a pH of 5.0 to 7.0.

[0137] Preferably, the calculated data represent a parameter selected from the isoelectric point (pI) of the protein, the variable fragment (Fv) charge (Fv charge), and a combination thereof. More preferably, ii) the calculated data consists of the parameters isoelectric point (pI) of the protein and the variable fragment (Fv) charge (Fv charge).

[0138] The in silico data is preferably selected from hydrophobic patch size and charged patch size, and combinations thereof, more preferably the size of the patch (Å 2 units), score (sum of all contributing patch scores associated with this patch), and type (pos=positively charged, neg=negatively charged, hyd=hydrophobic). Thus, the descriptors derived from the modeling are scorepos / neg / hyd Fvtotal, sizepos / neg / hyd Fvtotal, countpos / neg / hyd Fvtotal, and any combination thereof. More preferably, iii) the in silico data consists of scorepos Fvtotal, sizepos Fvtotal, countpos Fvtotal, scoreneg Fvtotal, sizeneg Fvtotal, countneg Fvtotal, scorehyd Fvtotal, sizehyd Fvtotal, and counthyd Fvtotal.

[0139] Below, we will look at 14 different input nodes IN1 to IN 14and 14 different input parameters IP1 to IP 14 1. One embodiment of a neural network 10, and in particular a trained neural network, comprising input parameters IP1 to IP2 is described. It will be apparent to those skilled in the art that these input parameters represent only examples of multiple possibilities. Therefore, the input parameters described below should not be understood to form a limitation. Thus, more or fewer input parameters may be used. In a further embodiment, the neural network comprises input parameters IP1 to IP3. 14 The input parameters may include at least one of:

[0140] Specifically, the input parameters IP1 to IP3 are i) experimental data, the input parameters IP4 and IP5 are ii) calculated data, and the input parameters IP6 to IP 14 iii) in silico data. IP1 refers to HIC RT [min]; IP2 refers to kD [mL / g], IP3 is A2 x E -4 [mol*mL / g2], IP4 refers to pI, IP5 refers to the Fv charge at pH 6, IP6 refers to the score pos Fv total, IP7 refers to the size pos Fv total, IP8 refers to the count pos Fv total, IP9 refers to the score neg Fv total. IP 10 refers to the size neg Fv total, IP 11 refers to the count neg Fv total. IP 12 refers to the score hyd Fv total, IP 13 refers to the size hyd Fv total, IP 14 refers to the count hyd Fv total.

[0141] In one embodiment, the input parameters IP1 to IP3 are used as experimental data, IP4 and IP5 are used as input parameters, and IP6 to IP7 are used as input parameters. 14 A combination of in silico data consisting of:

[0142] In another embodiment, ii) calculation data consisting of input parameters IP4 and IP5, and iii) input parameters IP6 to IP 14 In this embodiment, the neural network 10 is configured with 11 different input nodes IN4 to IN5. 14 and 11 different input parameters IP4 to IP 14 Equipped with.

[0143] According to the previous embodiment, a neural network 10, particularly a trained neural network, was used to model the viscosity of a solution of a monoclonal antibody. Figure 5 shows a diagram in which the output values ​​provided by the neural network 10 for modeling the viscosity of a solution of a monoclonal antibody are compared with measured viscosity values. As can be seen, the neural network 10, particularly the trained neural network, effectively and reliably predicts the concentration-dependent viscosity of the protein solution, particularly by calculating a viscosity curve.

[0144] In the next step S1 of the method, a protein solution is identified, based on which a pharmaceutical product can be produced.

[0145] In step S2, concentration-dependent viscosity, in particular protein concentration-dependent viscosity, is predicted by using a computer-implemented neural network 10, in particular by using a trained neural network. Step S2 therefore represents a sub-method, i.e., a computer-implemented method for predicting the concentration-dependent viscosity of a protein solution.

[0146] In this step, the concentration-dependent viscosity of the protein solution is calculated using the neural network 10 provided in step S0, in particular the trained neural network. To do so, the input parameters IP, in particular IP1 to IP2 associated with the protein solution identified in step S1, are used. 14 are determined and provided to the neural network 10, which calculates at least one output parameter OP indicative of the concentration-dependent viscosity.

[0147] In the next step S3, the concentration-dependent viscosity determined in step S2 is compared with a target viscosity to determine the pharmaceutical suitability of the protein solution. Steps S2 and S3 together represent a sub-method, i.e. a method for determining the concentration-dependent viscosity of a protein solution.

[0148] In this context, the target viscosity indicates the maximum viscosity of the protein solution. Specifically, the target viscosity indicates a maximum value of 15 mPa*s to 30 mPa*s, or 15 mPa*s to 25 mPa*s, or 15 mPa*s to 20 mPa*s, for example, 15 mPa*s, or 18 mPa*s, or 20 mPa*s, or 25 mPa*s. In this step, the suitability of the protein solution is determined by the particularly selected protein concentration c S or a selected protein concentration range (e.g., 20 mg / mL to a selected protein concentration c S ), it is determined if the concentration-dependent viscosity does not exceed the threshold or maximum value of the target viscosity (i.e., 15 mPa*s or 20 mPa*s). However, if the viscosity of the protein solution exceeds the target viscosity, the suitability of the protein solution is not determined (i.e., the protein solution is determined to be unsuitable as a pharmaceutical product, or the protein solution is determined to be potentially unsuitable as a pharmaceutical product).

[0149] According to the embodiment shown in FIG. 2, in step S3, the parameter η cs is compared to the target viscosity. The parameter η csdoes not exceed the target viscosity, the suitability of the protein solution as a pharmaceutical product is determined (i.e., the protein solution is determined to be suitable as a pharmaceutical product, or the protein solution is determined to be potentially suitable as a pharmaceutical product).

[0150] According to the embodiment shown in FIG. 3, in step S3, the derived function f η Use the VI to determine whether the viscosity of the protein solution is consistent across a selected protein concentration range (e.g., 20 mg / mL to a selected protein concentration c S Alternatively, this step may first involve measuring the viscosity of a sample at a selected protein concentration, c S Viscosity of the protein solution at η cs is the viscosity η cs (i.e., the selected protein concentration c S Before determining whether the viscosity at η Therefore, if the target viscosity is not exceeded, the suitability of the protein solution as a pharmaceutical product is determined (i.e., the protein solution is determined to be suitable as a pharmaceutical product, or the protein solution is determined to be potentially suitable as a pharmaceutical product).

[0151] 1, if the suitability of the protein solution is determined in step S3, the method may then proceed to optional step S5, in which a pharmaceutical product is prepared (or manufactured) based on the protein solution identified in step S1, in particular at a desired protein concentration. If the suitability of the protein solution is not determined in step S3, the method may proceed to step S1, in which a new protein solution is identified, in particular by adapting or modifying the initial protein solution, and then steps S2 to S4 are performed again.

[0152] It will thus be appreciated that by screening potential protein solutions using the above-described method (i.e., by screening potential protein solutions using the neural network 10 provided in step S0), it is possible to identify candidate protein solutions for use as pharmaceuticals without actually preparing and empirically testing the potential protein solutions. The above-described in silico method for screening potential protein solutions makes it possible to screen a large number of candidate protein solutions without requiring empirical measurement of the quality of every candidate protein solution. As a result, the present invention solves a long-unmet need for a rapid, high-throughput method for screening candidate protein solutions for use as pharmaceuticals that avoids the time and effort inherent in empirical (i.e., laboratory-based) screening.

[0153] It will be apparent to those skilled in the art that these embodiments and items merely show examples of multiple possibilities. Therefore, the embodiments shown below should not be understood as forming limitations on these features and configurations. Any possible combination and configuration of the described features can be selected according to the scope of the present invention.

[0154] [Table 3]

[0155] Materials and Methods Experimental setup To evaluate the influence of various input parameters on the predictive ability of the artificial neural network, three different models (Tables 4 and 5 below) were created using data from mAbs 1–25. Either experimental input (retention time in HIC; kD, A2) alone, in silico-derived input (patch size) alone, or both inputs were entered into the modeling. All three models included Fv charge and pI, regardless of the input selected. Each model was created by dividing the input data into a training set and a validation set, which contained information from randomized mAbs for each new artificial neural network model. The training set included data, input variables, and viscosity descriptors for 18–20 mAbs. The validation set included only the input variables for the remaining 7–5 mAbs for the purpose of predicting their viscosity descriptors. These viscosity descriptors were derived from the linearization of the viscosity-concentration curves for each mAb, intercept A, and slope B.

[0156] R for the created model 2 The values ​​show the interdependence of the validation and training sets (Table 1; example graphs are shown in Figures 9a and 9b). In some cases, the quality of one set is superior (R 2 >0.99), the second set is not as good (R 2 <0.95). Artificial neural networks using both variables (experimental and in silico) achieved the highest quality models, with even the lowest R 2 >0.92 (intercept A of the training set). Artificial neural networks using only in silico derived inputs had a minimally low "worst" R of >0.90 (slope B of the training set). 2 Finally, artificial neural networks containing only experimental data achieve a minimum R of >0.75. 2 This is comparable to published results using a linear correlation of kD and viscosity, and is superior to the linear correlation of A2 and kD for the intercept and slope parameters performed on this data set.

[0157] The slope parameter B describes the steepness of the viscosity curve, which exponentially increases with protein concentration. This is important in describing potentially "problematic" mAbs. Such "problematic" mAbs may exhibit moderate viscosity at low protein concentrations, but a significant increase in viscosity above values ​​typically considered acceptable for pharmaceuticals (i.e., values ​​above approximately 15-25 mPa*s). Consequently, an artificial neural network provided the best prediction for slope B, and therefore was used with all the following input variables:

[0158] Of all the available data, mAbs 26 and 27 were not used to create the artificial neural network model but were kept separate for additional validation studies. [Table 4]

[0159] Table 1 shows a comparison of models created using different inputs that predict either the intercept (A) or slope (B) of the concentration-dependent viscosity curve of a mAb. An "x" indicates that a particular set of inputs was used for the model, and a "-" indicates that a set of input parameters was not used. The experimental inputs refer to HIC retention time, kD, and A2, while the in silico inputs include data from surface patch analysis, which provides information on the size, score, and count of positive, negative, and hydrophobic surface patches. Fv charge and pI were included regardless of the input selected. The model quality parameters shown were obtained from plotting the A or B values ​​calculated from the measured viscosity against the predicted values ​​of the respective models. R 2 = coefficient of determination; SSE = standard squared error; RMSE = root mean squared error

[0160] Category Classification Using experimental and in silico data as inputs, we developed models for category classification by training them on whether a mAb exhibited a viscosity above a threshold of 15 mPa*s at a specific concentration. This was done for concentrations of 120, 150, and 180 mg / mL. Of the 25 mAbs used in model development, three mAbs at 120 mg / mL, six at 150 mg / mL, and 15 at 180 mg / mL exhibited viscosities above 15 mPa*s. The confusion matrices for these models are shown in Table 2 below. From this, it is clear that both the training and validation sets contained problematic (defined in this experiment as having a viscosity ≥ 15 mPa*s) and non-problematic (defined in this experiment as having a viscosity < 15 mPa*s) mAbs. The developed models had excellent predictive ability regardless of the input used, and all of them exhibited a misclassification rate of 0 (Table 8). To further evaluate the predictive ability of these models, two mAbs not used in the training and validation sets were selected for validation. mAb26 exhibits benign behavior at 120 and 150 mg / mL, but exceeds 15 mPa*s at 180 mg / mL, which was correctly predicted by the model disclosed herein (No / No / Yes). mAb27 does not exhibit problematic behavior at any of the concentrations, again correctly predicted by the model presented herein (No / No / No). [Table 5]

[0161] Viscosity curve prediction To obtain more comprehensive information about the viscosity of the mAbs, we predicted the complete concentration-dependent viscosity curves, more specifically, the intercept (A) and slope (B) of the linearized exponential function. Using the predicted values ​​of A and B, a theoretical viscosity curve can be constructed and compared to the actual measured values. An example of such a comparison is shown in Figure 6A-D. Despite not matching the actual values ​​exactly, the predicted values ​​are very close, and the predicted curve reflects the actual concentration-dependent viscosity in a similar manner. The predicted percentage and absolute difference compared to the measured viscosity of all mAbs are presented in Figure 7. The calculated average difference between the predicted and measured viscosity values ​​across all 27 mAbs is shown in Table 3. The average absolute difference (in mPa*s) between the calculated and measured values ​​ranged from 0.1 to 4.1 mPa*s, with the difference gradually increasing as the protein concentration increased. The relative difference % ranged from 8.2 to 26.2%, with the largest difference occurring at concentrations between 90 and 120 mg / mL, and the difference decreasing as the protein concentration decreased and increased. [Table 6]

[0162] To validate the model again, mAbs 26 and 27 were used. Figures 6E and 6F show the predicted viscosity curves for both mAbs. The predicted curve course for mAb 27 (Figure 6F) appears very similar to the measured values, while the predicted curve for mAb 26 (Figure 6E) does not accurately reflect the steepness of the measured curve above 150 mg / mL.

[0163] The use of computational data and in silico modeling has the advantage of being relatively easily accessible and does not require materials and laboratory work.

[0164] The present invention provides an artificial neural network and method for predicting and determining the viscosity of proteins, such as mAbs, in solution. The model can be used to predict viscosity classifications above or below a given threshold, e.g., 15 mPa*s, and can be used to predict viscosity curves. While already good at categorizing, the model for viscosity curve prediction exhibits high power and good viscosity curve prediction. The use of multiple input variables derived from experimental data as well as from computational data and in silico modeling is advantageous.

[0165] Monoclonal antibodies mAbs were obtained. Double gene vectors containing heavy and light chains were transfected into CHOK1SV GS-KO cells and grown as stable pooled cultures under selective conditions. Clarified supernatants were obtained by centrifugation followed by filter sterilization using a 0.22 μm filter. Protein A chromatography was used to purify the mAbs. All proteins were concentrated to a final concentration of 10 mg / mL and buffer exchanged into formulation buffer (protein solution) (20 mM histidine-HCl, pH 6.0) by tangential flow filtration. The mAbs were of different subtypes, IgG1 or IgG2 (see Tables 4 and 5 below).

[0166] buffer solution All described lab experiments (HIC, DLS, SLS, viscosity) were performed in a buffer (20 mM histidine-HCl, pH 6.0) that served as the solution for proteins (and later pharmaceuticals).

[0167] Protein concentration Concentration determination was performed using an Agilent Cary 60 UV-spectrophotometer with a variable pathlength extension SoloVPE. For each measurement, 30 μL of sample was loaded into a cuvette and measured at 280 nm using the appropriate specific extinction coefficient.

[0168] Hydrophobic interaction chromatography The hydrophobic surface characteristics of all mAbs were determined by hydrophobic interaction chromatography (HIC). Proteins were analyzed at 10 mg / mL in formulation buffer, and 5 μL was injected onto a ProPac Hic-10 column (ThermoScientific) and separated using a Waters HPLC system. Starting conditions of 95% mobile phase A (1 M ammonium sulfate in 20 mM sodium phosphate, pH 7.0) were linearly decreased to 95% mobile phase B (20 mM sodium phosphate, pH 7.0) over 39 min. The flow rate was set at 1 mL / min with a column temperature of 24 °C.

[0169] Dynamic and static light scattering Dynamic light scattering (DLS) and static light scattering (SLS) measurements were performed using a DynaPro plate reader III (Dynamics software, Wyatt Technologies). Antibody stock solutions were filtered through 0.22 μm PVDF filters (Millex GV), and serial dilutions of seven concentrations of protein from 10 mg / mL to 2 mg / mL were prepared in formulation buffer. Samples were transferred in triplicate to a 384-well plate (Aurora), and the plate was centrifuged at 750*g for 2 min to remove air bubbles. The measurement temperature was set to 25°C. The laser power was set to 20%, the attenuation level was set to 0%, and 20 acquisitions of 5 s duration were performed for each well. DLS was used to evaluate the diffusion interaction parameter kD (mL / g). The mutual diffusion coefficient Dm (m2 / s) was plotted against the protein concentration (g / mL), and kD was obtained from the slope of the linear fit. The second virial coefficient, A2 (mol*mL / g), was obtained from SLS measurements. Plate calibration was performed using dextran (Sigma) with a defined molecular weight of 36.9 ± 0.1 kDa. Solvent offset was measured in triplicate for the formulation buffer. Reciprocal molecular weight (mol / g) was plotted against protein concentration (g / mL), and A2 was obtained from the slope of the linear fit.

[0170] In silico modeling of mAb Fv patches Modeling of the mAb was performed using the software BioLuminate (version 3.80, Schroedinger, LLC, New York, NY). Homology modeling of the Fv region was performed by using an antibody prediction tool. Framework templates of IgG1 or IgG2 isotypes were selected from the pdb database based on the highest composite score. Optimal CDR loop clusters were automatically selected. For modeling, the standard pre-settings of the software were maintained, except that the pH was set to 6.0 to represent the experimental setting. The surface of the modeled mAb Fv region was analyzed using the protein surface analysis tool in BioLuminate software to determine the patch size (Å). 2 The modeling yielded the following descriptors: score (pos / neg / hyd Fv total), score (sum of all contributing patch scores related to this patch), and type (pos = positively charged, neg = negatively charged, hyd = hydrophobic). The descriptors derived from the modeling are therefore score pos / neg / hyd Fv total, size pos / neg / hyd Fv total, and count pos / neg / hyd Fv total.

[0171] Calculating Fv charge To calculate the Fv charge, the variable heavy and variable light chains of each antibody were analyzed using the protpi protein tool (https: / / www.protpi.ch / Calculator / ProteinTool). Each of the two chains was defined as a subunit of the entire protein. The set modifier for post-translational modifications was an overall disulfide bridge for cysteine ​​residues in the Fv. Charges were calculated at pH 6.0.

[0172] pI calculation The pI was calculated "manually" from the pKa values ​​of the amino acid residues in the primary sequence.

[0173] Rheometry and Viscosity Descriptors The mAb was concentrated to approximately 180 mg / mL using a spin filter with a 30 kDa molecular weight cutoff and then diluted to six concentrations ranging from 180 to 30 mg / mL. Concentration-dependent viscosity data were generated using a VROC viscometer (Rheosense) for each concentration. To obtain descriptive information on the concentration-dependent viscosity of the mAb samples, the experimentally measured viscosity data were processed based on equations (3) to (5) above.

[0174] The relative viscosity of the mAb samples was calculated using equation (3) above. To do so, the viscosity of the buffer solution was first measured, which was 0.92 mPa*s at 25°C. Knowing the viscosity of the buffer solution, the relative viscosity of the mAb samples was then calculated based on the measured viscosity values ​​according to equation (3). Thus, a pair of concentration and relative viscosity values ​​was determined for each mAb sample.

[0175] Equation (4) above was used to describe the exponential concentration-dependent viscosity of a mAb solution. This equation can be linearized to obtain the intercept A and slope B using natural logarithms, as in equation (5).

[0176] Based on the determined paired values ​​of concentration and relative viscosity of the mAB samples, the viscosity descriptors A and B of each mAb were then determined based on equations (4) and (5) above by applying a least-squares fit.

[0177] Data Analysis and Artificial Neural Network Modeling The ANN model creation followed the approach shown in Figure 8. All input parameters (experimental data, calculated values ​​from the sequence, in silico-derived data, and viscosity descriptors) are listed in Tables 4 and 5. Each model was trained on a categorical response, aiming to identify mAbs that exhibit viscosity values ​​above a threshold of 15 mPa*s at different protein concentrations. ANNs were generated using the software JMP v.16.0.0 (SAS Institute Inc.). The activation function used for all nodes was a tan h function, which transformed values ​​from -1 to 1. For all models, one hidden layer with four nodes was sufficient. To prevent the network from overfitting the model and losing predictive ability, the data were divided into a training set and a validation set. The method used was K-fold, where 25 mAbs were selected and their datasets were divided into K sets. Each K set contained the data set of all 25 mAbs, but each K set had two different subsets, the training set and the validation set. The mAbs were randomly distributed across the two subsets in each K set. Each K set was used to verify the fit of the model to the remaining data, and the total K set was fitted. The value of K was set to 5 for each model. The distribution of mAbs across the K sets was also uncorrelated among the three models but was random, so each model had a different K set. The models reported by the JMP software were based on the best log-likelihood. The quality of the ANN was determined using the coefficient of determination (R2), standard squared error (SSE), and root mean squared error (RMSE) of the training and validation data sets.

[0178] [Table 7]

[0179] [Table 8-1] [Table 8-2]

[0180] [Table 9]

[0181] [Table 10]

[0182] As described above, the viscosity descriptors, intercept A and slope B, are calculated from the values ​​shown in Tables 6 and 7. To do so, first, the viscosity of the buffer is determined (viscosity was measured to be 0.92 mPa*s). The viscosity of the buffer is then used to calculate the relative viscosity of each mAb for each concentration and viscosity value pair in Tables 6 and 7, based on equation (3) above. A least-squares fit (or any other mathematical procedure for finding a relationship, particularly a best-fit curve, for a given set of points) is then performed using the six value pairs for concentration and viscosity (relative viscosity) of each mAb, thereby fitting the six value pairs to the exponential function of equation (4) above. The calculated function is then linearized as shown in equation (5) above to provide the A and B values ​​for each mAb included in Table 5.

[0183] [Table 11]

Claims

1. 1. A method for providing a computer-implemented neural network (10) configured to predict the concentration-dependent viscosity of a protein solution depending on a plurality of input parameters (IP) associated with said protein solution, comprising: - Multiple training datasets (S T ), wherein the plurality of training data sets (S T ) each of which is associated with a particular protein solution and includes a plurality of input parameters (IP) indicative of said particular protein solution and at least one associated output parameter (OP) indicative of the concentration-dependent viscosity of said particular protein solution; - the training data set (S T training a neural network based on the concentration-dependent viscosity of the sample to provide the computer-implemented neural network (10) configured to predict the concentration-dependent viscosity; The input parameters (IP) are: i) experimental data; ii) calculation data, and iii) in silico data.

2. The method of claim 1 , wherein the protein solution comprises a therapeutic protein, preferably an antibody.

3. 3. The method of claim 2, wherein the antibody is a monoclonal antibody (mAb), preferably of the IgG1 or IgG2 subtype.

4. The method of any one of claims 1 to 3, wherein the protein solution comprises an aqueous medium comprising water, preferably a buffer, more preferably a histidine-HCl buffer.

5. The concentration-dependent viscosity is a value (η) that relates the viscosity of the protein solution to a particular protein concentration in the protein solution. cs 5. The method according to claim 1, wherein

6. The concentration-dependent viscosity is a function (f η 5. The method according to claim 1, wherein the arithmetic unit represents a function, in particular a mathematical function.

7. The concentration dependent viscosity is a function of: f η (c)=A×e Bc (In the formula, f η 7. The method of claim 6, wherein ∑ ∑ a ∑ b ...

8. 8. The method of claim 7, wherein the concentration-dependent viscosity of the protein solution is represented by at least one of a constant A and a constant B.

9. The method further includes providing a neural network to be trained, the neural network to be trained comprising: an input layer (12) having a plurality of input nodes (IN), each of said plurality of input nodes (IN) receiving one input parameter (IP); at least one hidden layer (14), in particular only one hidden layer (14), having a plurality of hidden nodes (HN), in particular four or more hidden nodes (HN); and an output layer (16) having at least one output node (ON) for providing at least one output parameter (OP) indicative of said concentration-dependent viscosity.

10. The method of claim 9 , wherein the hidden nodes (HN) use tan h as an activation function.

11. The at least one output parameter provided by the at least one output node (ON) is A viscosity value (η) relating the viscosity of the protein solution to a particular protein concentration of the protein solution cs ) or The following functions: f η (c)=A×e Bc (In the formula, f η 11. The method of claim 9 or 10, wherein A denotes a function relating the viscosity of the protein solution to the protein concentration, c denotes the protein concentration, and A and B denote constants, and / or a constant A and a constant B of the function.

12. The method according to any one of claims 1 to 11, wherein said input parameters (IP) describe protein-protein interactions.

13. The method according to any one of claims 1 to 12, wherein the neural network is an artificial neural network.

14. Multiple validation datasets (S V ), further comprising the step of validating the neural network (10) trained on the plurality of validation data sets (S V 14. The method of any one of claims 1 to 13, wherein each of the input parameters (IP) comprises a number of input parameters (IP) and at least one associated output parameter (OP).

15. 1. A computer-implemented method for predicting the concentration-dependent viscosity of a protein solution by using a neural network, comprising: the neural network is configured to predict a concentration-dependent viscosity of the protein solution depending on a plurality of input parameters associated with the protein solution; The input parameters (IP) are: i) experimental data; ii) calculation data, and iii) in silico data.

16. said input parameters (IP) comprising at least one of i) experimental data, ii) computational data, or iii) in silico data; the experimental data i) is selected from apparent surface hydrophobicity as measured by hydrophobic interaction chromatography (HIC), a diffusion interaction parameter (kD), a second virial coefficient (A2), and combinations thereof; the calculated data ii) is selected from the isoelectric point (pI) of the protein, the variable fragment (Fv) charge (Fv charge), and a combination thereof; 16. The method of any one of claims 1 to 15, wherein the in silico data iii) is selected from hydrophobicity and charged patch sizes, in particular score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total.

17. 17. The method of any one of claims 1 to 14 or 16, wherein said input parameters (IP) are selected from said ii) computational data and said iii) in silico data.

18. 1. A method for determining the concentration-dependent viscosity of a protein solution, comprising: predicting the concentration-dependent viscosity of said protein solution by using a computer-implemented neural network (10); and comparing the concentration-dependent viscosity to a target viscosity.

19. 20. The method of claim 18, wherein the step of comparing the concentration-dependent viscosity with the target viscosity is performed to determine pharmaceutical suitability of the protein solution.

20. 1. A method for providing a pharmaceutical product comprising a protein solution, comprising: Predicting the concentration-dependent viscosity of the protein solution (S2) by using a computer-implemented neural network (10); (S3) comparing the concentration-dependent viscosity with a target viscosity to determine the pharmaceutical suitability of the protein solution; Optionally, if the suitability of said protein solution is determined, preparing said pharmaceutical product (S5).

21. 1. An apparatus for predicting the concentration-dependent viscosity of an experimental protein solution, comprising: an experimental protein solution; a computer component including a neural network configured to predict the viscosity of the experimental protein solution; the neural network is trained using a plurality of predetermined data sets; each of the plurality of predetermined data sets comprises: (i) a plurality of input parameters indicative of at least one physical characteristic of a predetermined protein solution; and (ii) at least one output parameter indicative of a concentration-dependent viscosity of the predetermined protein solution; The apparatus, wherein the plurality of input parameters comprises at least one of experimentally derived data, computationally derived data, and in silico derived data.

22. 1. A system for determining the concentration-dependent viscosity of a protein solution, comprising: A protein solution; and a neural network configured to determine the viscosity of the protein solution, wherein determining the viscosity of the protein solution is dependent on a plurality of input parameters (IPs) associated with the protein solution, and further wherein the input parameters (IPs) comprise at least one from the group consisting of: (i) experimental data, (ii) computational data, and (iii) in silico data.