Protein solutions
By using computer-implemented neural network methods, combining experimental and computational data, predicting the concentration-dependent viscosity of protein solutions, the problem of difficult prediction in the prior art is solved, and high-accuracy prediction is achieved, supporting the early development evaluation of drug products.
Patent Information
- Application Number
- CN202380076044.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-25
- Filing Date
- 2023-11-03
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art is difficult to effectively predict the concentration-dependent viscosity of protein solutions, especially at high protein concentrations, which affects the development and use of drug products.
The neural network method is adopted to train the neural network through the training data set, and the concentration-dependent viscosity of protein solution is predicted using experimental data, computational data and computer simulation data.
High accuracy prediction of protein solution concentration-dependent viscosity is achieved, reflecting R2 values to achieve accuracy of 0.95 or higher, supporting the early stage evaluation of drug products.
Smart Images

Figure CN120129943A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for providing a computer-implemented neural network configured to predict the concentration-dependent viscosity of a protein solution. Further, the present invention relates to methods for predicting the concentration-dependent viscosity of a protein solution, determining the concentration-dependent viscosity of a protein solution, and providing a pharmaceutical product comprising a protein solution by using such a computer-implemented neural network. Background Art
[0002] Therapeutic proteins, such as monoclonal antibodies (mAbs), have become an important factor in the treatment of various diseases, such as cancer, immune-mediated disorders, and infectious diseases. The most common route of administration of protein-based pharmaceutical products (DPs) is the intravenous route, but this requires the patient to be hospitalized and the pharmaceutical product to be administered by a healthcare professional. For patients with chronic diseases, the need for repeated drug administration and thus hospitalization poses a significant burden and stress that can jeopardize the success of the intended treatment. Subcutaneous (s.c.) injection allows patients to self-administer protein-based pharmaceutical products, such as monoclonal antibody-based pharmaceutical products, by using prefilled syringes, autoinjectors, or other delivery devices, and this generally improves the quality of life and compliance of patients with chronic conditions.
[0003] However, subcutaneous administration has certain limitations. Generally, the single injection volume is limited to below about 2 mL, depending on the available subcutaneous space of the patient and the tolerable pain. Due to this volume limitation, pharmaceutical products with a high drug substance concentration are required.
[0004] Generally, mAbs have high specificity, but they also require a relatively large therapeutic dose. Thus, this results in a pharmaceutical product (DP) containing a protein solution, such as an mAb, having too high a concentration, usually more than 100 mg / mL of protein in the solution (for subcutaneous administration).
[0005] However, as the concentration increases, the intermolecular distance shortens and proteins may interact. These protein-protein interactions (PPIs) then multiply and determine the solubility, aggregation, and viscosity of proteins (such as mAbs). PPIs are particularly determined by the primary amino acid sequence of the protein (such as mAb) and the resulting three-dimensional structure with charged or hydrophobic patches. Solution conditions can also affect protein-protein interactions, for example, by regulating the size of charged patches via pH, or by shielding charged patches via short-range electrostatic interactions using salts, buffers, amino acids, or other charged excipients. Arginine is a common excipient for testing viscosity reduction, with a dual mode of action of shielding both charged and hydrophobic patches. Seventeen out of 34 FDA-approved drug products with high mAb concentrations use salts or amino acids as excipients, presumably to reduce protein-protein interactions and thereby lower the viscosity of mAb solutions. More exploratory excipients with proven viscosity reduction potential but not applied in commercial drug products are poly-l-glutamic acid, caffeine, hydrophobic salts, or amino acid derivatives, mainly due to not being approved as excipients for parenteral administration or concerns about toxicity.
[0006] High-viscosity solutions can be a major obstacle in the development of protein-based drug products. Disadvantages may include high costs due to high losses and low recoveries during purification, difficulty in manufacturing or filling, and ultimately low administrability due to the need for high injection forces and slow administration, which can be potentially painful. Generally, solutions with a dynamic viscosity higher than about 15 to 30 mPa*s or even higher than about 15 to 20 mPa*s are considered problematic. For new molecules, the need to develop high-concentration formulations may not be obvious initially, and it may also be due to unexpectedly high doses or a change in the target administration route. Clinical trials typically start in the first phase with formulations at lower concentrations (protein in solution < 50 mg / mL), and once safety and effective dose levels are established, high protein concentrations (protein in solution ≥ 100 mg / mL) are usually explored in later phases.
[0007] When the dose range is established, for CMC (chemistry, manufacturing, and control) issues related to drug products in later development, it is important to predict the viscosity of new proteins (especially mAbs) at high protein concentrations during the early development stage.
[0008] For high-concentration mAb solutions, a great deal of work has been done in the past two decades to understand the factors contributing to high viscosity.
[0009] In early development, small amounts of multiple candidate proteins are typically obtained, and their stability and solubility need to be tested in pre-formulation studies. Considerable work has been done in the past few years to predict the viscosity of high-concentration solutions using experimental data from low-concentration experiments.
[0010] To predict the viscosity of protein solutions used as drug products in clinical treatment, different techniques are known, which typically rely on experimental parameters (such as the diffusion interaction parameter (kD)) or on computational tools that utilize information derived from the protein's primary sequence.
[0011] Until the work of Roberts (Woldeyes, M.A., Qi, W., Razinkov, V.I., Furst, E.M. & Roberts, C.J. How Well Do Low-and High-Concentration Protein Interactions Predict Solution Viscosities of Monoclonal Antibodies?J Pharm Sci 108, 142–154 (2019)), experimental data on colloidal interactions (primarily the diffusion interaction parameter (kD) and the second virial coefficient (A2)) were found to at least qualitatively predict mAbs with potentially high viscosities (problematic mAbs). Although Roberts highlighted many cases of predictive failure, this questioned the validity of using low-concentration experimental data as a predictive tool. Most publications also mentioned a rather small number of samples and described the relationship between experimental data and viscosity in a linear manner. The complexity of the sources of solution viscosity at high mAb concentrations suggests that a non-linear modeling approach might be more appropriate. In recent years, experimental methods for predicting the concentration-dependent viscosity of antibodies have been complemented by computational simulation methods, which aim to identify decisive molecular descriptors such as solvent exposure, local charge, hydrophobic effects, and surface patches. A model for predicting viscosity is known from Agrawal et al. (Agrawal, N.J. et al. Computational tool for the early screening of monoclonal antibodies for their viscosities, MAbs 8, 43–48 (2016)), which uses a spatial charge map (SCM) to identify mAbs with high viscosities. The SCM applies molecular dynamics (MD) simulations to calculate scores for screening antibody viscosities at high concentrations. However, the computational cost of molecular dynamics simulations is high and requires structural information, which is a major application bottleneck. The principle of the SCM has also been used to test machine learning methods for predicting mAb solution viscosity (Lai, P.-K. DeepSCM: An efficient convolutional neural network surrogate model for the screening of therapeutic antibody viscosity. Comput Struct Biotechnol J, 20, 2143–2152 (2022)).The deep learning algorithm established by Lai uses preprocessed antibody sequences as input and SCM scores obtained from MD simulations as output for model training; and a DeepSCM surrogate model for SCM scores is developed based on the convolutional neural network (CNN) architecture. Then, the SCM scores are used to predict the high-concentration viscosity. The accuracy reported by Lai is 0.9. Summary of the Invention
[0012] The present invention relates to a method for providing a computer-implemented neural network configured to predict the concentration-dependent viscosity of a protein solution based on a plurality of input parameters associated with the protein solution. Further, the present invention, particularly the method, may include providing and using a computer-implemented neural network, particularly a computer-implemented neural network to be trained. Specifically, the providing and using may be performed to provide a trained computer-implemented neural network, particularly one configured to predict the concentration-dependent viscosity of a protein solution based on a plurality of input parameters associated with the protein solution. The method specifically uses a computer-implemented neural network, particularly a computer-implemented neural network to be trained, and the method includes the following steps:
[0013] - Providing a plurality of training data sets, each of the plurality of training data sets being associated with a specific protein solution and including a plurality of input parameters indicative of the specific protein solution and at least one associated output parameter indicative of the concentration-dependent viscosity of the specific protein solution; and
[0014] - Training a neural network based on the training data sets to provide a computer-implemented neural network configured to predict the concentration-dependent viscosity. The input parameters include at least one of the following:
[0015] i) Experimental data, preferably selected from the apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), the diffusion interaction parameter (kD), the second virial coefficient (A2), and combinations thereof;
[0016] ii) Computational data, preferably selected from the isoelectric point (pI) of the protein, the variable fragment (Fv)-charge (Fv-charge), and combinations thereof; and
[0017] iii) Computer simulation data, preferably selected from the hydrophobic and charged patch sizes, particularly the positive Fv total score, positive Fv total size, positive Fv total count, negative Fv total score, negative Fv total size, negative Fv total count, hydrophobic Fv total score, hydrophobic Fv total size, and hydrophobic Fv total count.
[0018] Another embodiment of the present invention includes a computer-implemented method for predicting the concentration-dependent viscosity of a protein solution by using such a computer-implemented neural network, in particular a trained computer-implemented neural network. More specifically, this method is provided when providing and using the computer-implemented neural network to be trained as described above. The computer-implemented method of the present invention for predicting the concentration-dependent viscosity of a protein solution has high accuracy, which can be reflected by an R2 value greater than 0.95 or even greater than 0.99.
[0019] Another embodiment of the present invention includes a method for determining the concentration-dependent viscosity of a protein solution by using such a computer-implemented neural network, in particular such a trained computer-implemented neural network.
[0020] Another embodiment of the present invention includes a method for providing a pharmaceutical product containing a protein solution, which can effectively and efficiently evaluate the suitability of the protein solution as a pharmaceutical product in the early stage of development.
[0021] Preferably, any method for predicting or determining the concentration-dependent viscosity of a protein solution using a computer-implemented method disclosed herein uses a computer-implemented neural network initially configured (sometimes hereinafter referred to as "trained") to predict the concentration-dependent viscosity of the protein solution, as described herein, including all its embodiments.
[0022] In the context of the present disclosure, the task of a computer-implemented neural network to calculate the concentration-dependent viscosity can be referred to as determining or predicting the concentration-dependent viscosity; the terms "determine" and "predict" related to the calculation of the computer-implemented neural network can be used interchangeably.
[0023] In the context of the present disclosure, the suitability of a protein as a pharmaceutical product in solution form is determined by the viscosity of the protein solution at high concentrations. In particular, when the concentration of the protein in the solution exceeds 100 mg / mL, the viscosity should not exceed 15 mPa*s to 30 mPa*s, preferably 15 mPa*s to 20 mPa*s, especially 15 mPa*s.
[0024] These objectives are solved by the subject matter of the independent claims. Preferred embodiments are set forth in the present specification, the drawings and the dependent claims.
[0025] Accordingly, the present invention also includes a method for providing a computer-implemented neural network (sometimes hereinafter referred to as "neural network"), preferably an artificial neural network (ANN), also referred to as a trained ANN. Specifically, the present invention includes providing and using such a computer-implemented neural network. The neural network, particularly the trained neural network, is configured to predict or determine the concentration-dependent viscosity of a protein solution based on a plurality of input parameters associated with the protein solution. The method includes the steps of providing a plurality of training data sets, each of the plurality of training data sets being associated with a specific protein solution and including a plurality of input parameters indicative of the specific protein solution and at least one associated output parameter indicative of the concentration-dependent viscosity of the specific protein solution; and training the neural network (i.e., the neural network to be trained or the untrained neural network) based on the training data sets to provide a computer-implemented neural network (i.e., to provide a trained ANN), the computer-implemented neural network being configured to predict the concentration-dependent viscosity.
[0026] The term "protein solution" refers to an aqueous solution of a protein, preferably a therapeutic protein (commonly referred to as the active pharmaceutical ingredient "API"). Thus, a protein solution refers to any solution containing at least one protein dissolved or substantially dissolved in an aqueous medium such as water, buffer, cell culture medium, etc.
[0027] The term "protein" means a polypeptide consisting of an amino acid sequence. The protein can be a therapeutic protein, e.g., a therapeutic protein for diagnosing, treating, and / or preventing a disease or disorder. The protein can be a natural protein, i.e., a protein produced by naturally occurring and non-recombinant cells; or it can be produced by genetically engineered or recombinant cells and can include a molecule having the amino acid sequence of a natural protein, or a molecule having a deletion, addition, and / or substitution of one or more amino acids of the natural sequence, or a molecule having an amino acid sequence of a protein unrelated to the natural protein. The term also encompasses amino acid polymers in which one or more amino acids are chemical analogs of the corresponding naturally occurring amino acid polymers.
[0028] The therapeutic protein contained in the protein solution can be an antibody, which can include monoclonal antibodies and polyclonal antibodies, intact antibodies, antibody-drug conjugates, chimeric antibodies, humanized antibodies, human antibodies, or hybrid antibodies with dual or multiple antigen or epitope specificities, antibody fragments and antibody sub-fragments, such as Fab, Fab', F(ab')2 fragments, etc., including any immunoglobulin or hybrid fragment of any natural, synthetic, or genetically engineered protein that acts like an antibody by binding to a specific antigen to form a complex in one embodiment of a monoclonal antibody. The term "epitope" refers to a protein determinant capable of specifically binding an antibody. Epitopes are usually composed of chemically active surface groups of molecules such as amino acids or sugar side chains, and epitopes usually have specific three-dimensional structural characteristics as well as specific charge characteristics. The difference between conformational epitopes and non-conformational epitopes is that in the presence of a denaturing solvent, binding to the former rather than the latter is lost.
[0029] The preferred therapeutic protein is a monoclonal antibody. As used herein, the term'monoclonal antibody' (mAb) refers to an antibody obtained from a substantially homogeneous antibody population, i.e., the individual antibodies comprising the population are identical, except for possible naturally occurring mutations that may be present in minor amounts. Monoclonal antibodies are highly specific, targeting a single antigenic site. In addition, in contrast to polyclonal antibody preparations that include different antibodies directed against different antigenic points (determinants or epitopes), each monoclonal antibody targets a single antigenic point on the antigen. The modifier "monoclonal" indicates the characteristic of the antibody being obtained from a substantially homogeneous antibody population and should not be construed as requiring the antibody to be produced by any particular method. Preferred monoclonal antibodies are antibodies of the immunoglobulin G (IgG) class, such as IgG1, IgG2, IgG3, and IgG4.
[0030] Suitable aqueous media for the protein solution are known to those skilled in the art. Preferably, the aqueous medium is or contains water or an aqueous buffer, such as a histidine-HCl buffer.
[0031] Other buffers can be used, buffers known to those skilled in the art that are commonly used in formulating pharmaceutical products, such as acetate buffers, phosphate buffers, succinate buffers, or citrate buffers.
[0032] Preferably, the aqueous medium is a buffer with a pH of 5 to 8, more preferably 5 to 7.5, and even more preferably 5.5 to 6.5.
[0033] The aqueous medium can also have other pH values suitable for providing the protein solution, common pH values that aqueous media of pharmaceutical products are known to have by those skilled in the art.
[0034] Preferably, any one of any input parameter and at least one associated output parameter (the latter being provided for configuring a computer-implemented neural network) is determined in a similar aqueous medium which is also used in the intended pharmaceutical product containing the protein solution, or even more preferably in the same aqueous medium, such as having the same pH, the same buffer, the same viscosity, especially the same buffer viscosity, etc.
[0035] The protein solution may further comprise at least one pharmaceutically acceptable excipient. Suitable excipients are known to those skilled in the art, such as stabilizers, pH modifiers (such as one or more buffers), etc. In some embodiments, a polysorbate such as polysorbate 20, 40, 60 or 80 may be used as a stabilizer, which can improve the stability of the protein in an aqueous solution. Other stabilizers may be, for example, sugars or sugar alcohols, surfactants, salts or antioxidants, such as L-methionine.
[0036] As described above, the method for providing a computer-implemented neural network (especially a trained computer-implemented neural network (also referred to as a trained ANN)) includes the step of providing a plurality of training data sets. Further, using a computer-implemented neural network, especially a neural network to be trained, may include the step of providing a plurality of training data sets. Each of the training data sets is associated with a specific protein solution. Specifically, at least two of the plurality of training data sets are associated with different protein solutions. More specifically, each of the training data sets may be associated with a specific protein (i.e., the protein contained in the protein solution). Preferably, the plurality of training data sets may be associated with different proteins, which particularly constitute a proteome. In other words, each protein included in the proteome is associated with at least one training data set.
[0037] Preferably, the proteins on which each training data set is based are different proteins but are similar to each other in at least one aspect. The proteins in the proteome may be any of the above types of proteins. More preferably, all the proteins in the proteome of the training data sets are antibodies, and even more preferably all the proteins in the proteome are monoclonal antibodies.
[0038] To improve the training of a computer-implemented neural network (especially an ANN, more specifically a computer-implemented neural network to be trained), the number of training data sets and thus the number of proteins may be more than two. The prediction (or determination) accuracy of viscosity can be represented, for example, by an accuracy value comparable to the value of the coefficient of determination R2; these values range from 0 (reflecting the determination accuracy, which can also be called the prediction accuracy) to 1 (reflecting 100% accurate determination).
[0039] For example, the number of training data sets (and thus the number of proteins in the proteome) can be at least 15, preferably at least 20, more preferably at least 25. Generally, the larger the number of training data sets, the higher the accuracy. As the number of proteins in the proteome increases, the determination accuracy no longer increases linearly but approaches an upper limit of 1. For example, when 25 training data sets are provided, the determination accuracy is usually quite high.
[0040] In the context of the present disclosure, the term "concentration-dependent viscosity" refers to a parameter that indicates the viscosity of a protein solution, particularly the dynamic viscosity of the protein solution, at one or more protein concentrations of the protein solution. Thus, the concentration-dependent viscosity correlates at least one protein concentration of the protein solution with the corresponding viscosity of the protein solution. For example, the concentration-dependent viscosity can be or can indicate the viscosity of the protein solution in the form of a viscosity value η cs at a selected protein concentration, particularly at a single protein concentration. In other words, the concentration-dependent viscosity can be expressed and / or determined as η cs as the viscosity value of the protein solution at the selected protein concentration.
[0041] Alternatively, the concentration-dependent viscosity can indicate the viscosity of the protein solution at more than one protein concentration. Specifically, the concentration-dependent viscosity can be represented by a function f η (particularly a mathematical function) that correlates the viscosity of the protein solution with the protein concentration of the protein solution. For example, the concentration-dependent viscosity can be represented by a graph that correlates the viscosity value with the protein concentration of the protein solution.
[0042] Furthermore, the concentration-dependent viscosity can be represented by a function represented by Equation (1):
[0043] f η (c) = A * e B*c (1)
[0044] where f η refers to the function that correlates the viscosity of the protein solution with the protein concentration; c refers to the protein concentration; A and B refer to constants.
[0045] This Equation (1) can be linearized using the natural logarithm as in Equation (2) to obtain the intercept A and the slope B:
[0046] lnη rel = lnA + B * c (2)
[0047] Equation (2): Linearized viscosity behavior; where η rel refers to the relative viscosity; A refers to the intercept; B refers to the slope; c refers to the protein concentration.
[0048] In further development, the concentration-dependent viscosity of the protein solution can be represented by at least one (preferably both) of the constants A and B in Equation (1) above. In other words, a computer-implemented neural network, particularly a trained computer-implemented neural network, more particularly a trained ANN, and / or a computer-implemented neural network to be trained, can be configured to determine at least one, preferably both, of the constants A and B in Equation (1) above.
[0049] The term computer-implemented neural network (also referred to herein as "neural network") can be an artificial neural network. Thus, the term neural network used in the present disclosure can be used interchangeably with the term "artificial neural network" (ANN) (particularly with the term ANN to be trained or trained ANN). Generally speaking, an artificial neural network refers to a computing system that employs interconnected nodes, which are particularly functional units that utilize machine learning. An artificial neural network is a subset of machine learning (also known as deep learning) and uses unstructured data sets and hidden layers. Similar to ordinary machine learning, the data set used to build a neural network is typically divided into a training data set and a validation data set.
[0050] In an ANN, the interconnected nodes (also known as artificial neurons) receive signals, process these signals, and send signals to the nodes connected to them based on the received and processed signals. Generally, the signals transmitted to a node are called inputs and typically represent real numbers. The output of each node is calculated through a non-linear function (also known as an activation function), usually based on the sum of its inputs. The nodes and their connections usually have weights that are adjusted during the training phase based on the training data set. Then, in order to validate the trained ANN, a validation data set is used to allow determination of the accuracy of the trained ANN. Generally, as will be described below, the nodes in an ANN are aggregated into layers. Usually, the signals may be transmitted from the first layer (i.e., the input layer) via the hidden layers to the last layer (i.e., the output layer) after traversing the layers multiple times.
[0051] The general functions and operations of such an ANN are well known to those skilled in the art and are therefore not further specified. Instead, the technical features of the ANN related to the present invention are set forth below.
[0052] Generally, in order to provide an ANN suitable for predicting concentration-dependent viscosity, an ANN to be trained, particularly an untrained ANN, can first be provided and then appropriately trained based on the training data set. Thus, the method can further include the step of providing a neural network to be trained, which is preferably an untrained ANN.
[0053] Specifically, a neural network, particularly a neural network to be trained and / or a trained neural network, may include an input layer having a plurality of input nodes, each of the plurality of input nodes receiving an input parameter. In other words, each of the plurality of input nodes preferably receives an input parameter. The input parameters received between the input nodes are preferably different from each other.
[0054] Alternatively or additionally, a neural network, particularly a neural network to be trained and / or a trained neural network, may include at least one hidden layer, particularly only one hidden layer, the at least one hidden layer having a plurality of hidden nodes, particularly four or more hidden nodes. According to one configuration, a neural network, particularly a neural network to be trained and / or a trained neural network, may include a single hidden layer having four hidden nodes. The hidden nodes may use tan h as the activation function. Tan h converts the value to between -1 and 1. Alternatively, sigmoid may be used as the activation function of the hidden nodes.
[0055] Alternatively or additionally, a neural network, particularly a neural network to be trained and / or a trained neural network, may include an output layer having at least one output node for providing an output parameter. Specifically, the output layer may have a plurality of output nodes, each of the plurality of output nodes determining and outputting a single output parameter. The output parameter provided by the output layer can be used to determine the concentration-dependent viscosity. Alternatively, the output parameter provided by the output layer may indicate or represent the concentration-dependent viscosity. For example, the output layer may be configured to provide at least one output parameter that indicates at least one of the constant A and the constant B in the above equation (1). In other words, the output layer may be configured to provide at least one value as the output parameter, which is used as the constant A or the constant B in the above equation (1). For example, the output layer may be configured to provide a first output parameter through a first output node, the first output parameter being used as the constant A, and provide a second output parameter through a second output node, the second output parameter being used as the constant B, to determine the concentration-dependent viscosity represented by the function specified in the above equation (1).
[0056] Generally, a neural network, particularly a neural network to be trained and / or a trained neural network, receives a plurality of input parameters and calculates at least one output parameter based on these input parameters. Therefore, the training dataset includes or consists of input parameters and at least one output parameter.
[0057] In the step of training a neural network, a training data set is preferably used to implement or establish the neural network, in particular by adjusting the weights of the neural network nodes and their connections according to the training data set. Thus, the neural network to be trained becomes a trained neural network, i.e., the neural network provided by the above method. Hereinafter, the training data set is further specified, which allows training and thus provides an ANN that enables the prediction of the viscosity of a protein solution with high accuracy.
[0058] Each set of training data preferably has the same number and type of parameters. Further, as described above, each training data set is associated with a specific protein solution, in particular with a specific protein. Specifically, the training data set can be subdivided into input parameters (i.e., parameters representing the neural network input) and output parameters (i.e., parameters representing the neural network output). The input parameters indicate the specific protein solution associated with the corresponding training data set. Thus, the input parameters can indicate the characteristics of a specific protein solution or the protein contained therein. The output parameters indicate the concentration-dependent viscosity of the specific protein solution. Thus, the output parameters can represent or indicate the concentration-dependent viscosity of a specific protein solution.
[0059] Preferably, the input parameters contained in the training data set indicate the protein-protein interactions in the associated protein solution.
[0060] The intermolecular distance decreases with increasing concentration, and the protein-protein interactions then affect and determine the solubility, aggregation, and viscosity of the protein solution exponentially. Thus, by providing the neural network with input parameters indicating the protein-protein interactions in the protein solution, the proposed method takes into account those characteristics that essentially affect the viscosity of the protein solution at increasing protein concentrations. In this way, an effective and efficient prediction of the viscosity of the protein solution can be achieved.
[0061] Generally, protein-protein interactions are affected by the amino acid sequence of the protein and the resulting three-dimensional structure with charged or hydrophobic patches. Further, the solution conditions can affect protein-protein interactions by adjusting the size of the charged patches via pH and by shielding the charged patches through short-range electrostatic interactions using salts, buffer substances, amino acids, or other charged excipients.
[0062] As described above, the plurality of input parameters includes at least one of the following:
[0063] i) Experimental data obtained by detecting parameters from the provided protein solution, preferably, wherein the experimental data represents a parameter selected from protein hydrophobicity, diffusion interaction parameter (kD), net protein charge, Zeta potential, second virial coefficient (A2), third virial coefficient (A3), apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), and combinations thereof;
[0064] ii) Calculated data, preferably calculated from the primary amino acid sequence of the protein at pH 5.0 to 7.0, preferably, wherein the calculated data represents a parameter selected from protein isoelectric point (pI), variable fragment (Fv)-charge (Fv-charge), Fv symmetry parameter, Fv hydrophobicity, Vh-charge, Vl-charge, hinge-charge, and hydrophobic solvent accessible surface area; and combinations thereof; and
[0065] iii) Computer simulation data, preferably selected from hydrophobicity and charged patch size and combinations thereof.
[0066] In other words, the plurality of input parameters may include at least one selected from i) experimental data, ii) calculated data, and iii) computer simulation data.
[0067] Specifically, the calculated data and computer simulation data may be data calculated based on the primary amino acid sequence of the protein contained in the protein solution.
[0068] i) The experimental data is preferably a parameter selected from at least one of apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), second virial coefficient (A2), and combinations thereof. In other words, i) the experimental data may include at least one parameter selected from the group consisting of apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), and second virial coefficient (A2). More preferably, i) the experimental data consists of an apparent surface hydrophobicity parameter measured by hydrophobic interaction chromatography (HIC), a diffusion interaction (kD) parameter, and a second virial coefficient (A2) parameter.
[0069] ii) The calculated data is preferably a parameter selected from at least one of the isoelectric point (pI) of the protein, variable fragment (Fv)-charge (Fv-charge), and combinations thereof. In other words, ii) the calculated data may include at least one parameter selected from at least one of the isoelectric point (pI) of the protein and variable fragment (Fv)-charge (Fv-charge). More preferably, ii) the calculated data consists of an isoelectric point (pI) parameter of the protein and a variable fragment (Fv)-charge (Fv-charge) parameter.
[0070] iii) The computer simulation data is preferably a parameter selected from at least one of the hydrophobic and charged patch sizes, such as the positive Fv total score, positive Fv total size, positive Fv total count, negative Fv total score, negative Fv total size, negative Fv total count, hydrophobic Fv total score, hydrophobic Fv total size, and hydrophobic Fv total count. More preferably, iii) the computer simulation data consists of the positive Fv total score, positive Fv total size, positive Fv total count, negative Fv total score, negative Fv total size, negative Fv total count, hydrophobic Fv total score, hydrophobic Fv total size, and hydrophobic Fv total count. In other words, iii) the computer simulation data includes at least one parameter selected from the group consisting of the positive Fv total score, positive Fv total size, positive Fv total count, negative Fv total score, negative Fv total size, negative Fv total count, hydrophobic Fv total score, hydrophobic Fv total size, and hydrophobic Fv total count.
[0071] Using computational data and computer simulation modeling has the advantage of being relatively easy to obtain and does not require materials and laboratory work.
[0072] Methods for determining the diffusion interaction parameter (kD) are known to those skilled in the art. Suitable methods can be, by way of example, via dynamic light scattering (DLS), for example following the procedure outlined further below.
[0073] The virial coefficients can be determined in various ways, such as by static light scattering (SLS), in which case, for the second virial coefficient, the virial coefficient is usually denoted by A2, and for the third virial coefficient, by A3; or by measuring the osmotic pressure, in which case, for the second virial coefficient, the virial coefficient is usually denoted by B2, and for the third virial coefficient, by B3.
[0074] Methods for determining the second virial coefficient (A2) or the third virial coefficient (A3) are known to those skilled in the art. Suitable methods can be, by way of example, via static light scattering (SLS), for example following the procedure outlined further below.
[0075] Methods for determining the apparent surface hydrophobicity are known to those skilled in the art. Suitable methods can be, by way of example, via hydrophobic interaction chromatography (HIC), for example following the procedure outlined further below.
[0076] The method for determining the calculated data is known to those skilled in the art. Appropriately, the corresponding parameters such as Fv charge are calculated at a pH value of 5 to 8, or 5 to 7.5, or 5.5 to 6.5, or about 5.5, or about 6.0, or about 6.5, preferably at a pH value of 6.0. Those skilled in the art can follow the procedures outlined further below and use well-known platforms such as Prot pi|Protein Tool, https: / / www.protpi.ch / Calculator / ProteinToo, site operator Roland Josuran Prot pi, 8820 Switzerland。
[0077] The isoelectric point (pI) can be calculated "manually" from the pKa values of the amino acid residues of the primary sequence.
[0078] The method for determining the computer simulation data is known to those skilled in the art. Appropriately, software BioLuminate (version 3.80; LLC, New York, NY) can be used to determine such methods, for example, by following the procedures outlined further below.
[0079] Specifically, a neural network, particularly a neural network to be trained and / or a trained neural network, can receive input parameters selected from all three of the above data groups i) to iii). However, if the neural network receives input parameters from only two of the three data groups (such as the neural network receives input parameters from data groups i) and ii), or from data groups ii) and iii), or from data groups i) and iii) above, preferably from data groups ii) and iii)), the concentration-dependent viscosity can also be predicted reliably.
[0080] As described above, a neural network, particularly a neural network to be trained and / or a trained neural network, receives a plurality of input parameters and calculates at least one output parameter based on these input parameters. The output parameter can represent or indicate the concentration-dependent viscosity of a specific protein solution. Thus, the output parameter can indicate the cs viscosity value that correlates the viscosity of the protein solution η with the specific protein concentration of the protein solution. Alternatively or additionally, at least one output parameter can indicate the viscosity of the protein solution at more than one protein concentration. Specifically, at least one output parameter can indicate the η function represented by the above equation (1). More specifically, at least one output parameter can indicate at least one of the constants A and B in the above equation (1), preferably both.
[0081] In the following, examples are provided on how to determine at least one output parameter included in a training dataset. First, different protein solutions (referred to as protein solution samples) can be provided, which contain the same type of protein at different concentrations. Based on these samples, for example, for each concentration by using a VROC viscometer (Rheosense), concentration-dependent viscosity data can be generated. In this way, multiple value pairs can be provided, where each of the multiple value pairs correlates the viscosity of the protein solution with the protein concentration in the protein solution. In the next step, these measured viscosity data, particularly the experimentally measured viscosity data, can be processed based on the following equations (3) to (5).
[0082] For a substance, Equation (3) can be used to calculate the relative viscosity of different protein solution samples. To do this, first, the viscosity of the buffer contained in the protein solution can be determined (preferably at a defined temperature). After knowing the viscosity of the buffer, the relative viscosity of the protein solution sample can be calculated based on the measured viscosity data according to Equation (3). In this way, a value pair of the concentration and relative viscosity of each protein solution sample can be determined.
[0083]
[0084] Equation (3): Relative viscosity; η rel = Relative viscosity; η 0 = Buffer viscosity; η = Sample viscosity
[0085] Then Equation (4) can be used to describe the exponential concentration-dependent viscosity of the protein solution. The equation can be linearized using the natural logarithm as in Equation (5) to obtain an intercept A and a slope B.
[0086] η rel = A*e B*cs (4)
[0087] Equation (4): Exponential viscosity behavior; η rel = Relative viscosity; A = Intercept; B = Slope; cs = Selected mAb concentration
[0088] lnη rel = lnA + B*cs (5)
[0089] Equation (5): Linearized viscosity behavior; η rel = Relative viscosity; A = Intercept; B = Slope; cs = Selected mAb concentration
[0090] Based on the determined value pairs of the concentration and relative viscosity of the protein solution samples, then the viscosity descriptors A and B for each protein can be determined based on Equation (4) and / or (5) by applying, for example, curve fitting techniques such as least squares fitting.
[0091] For each protein in the proteome of the training and validation datasets, a plurality (num-conc) of protein solutions with different protein concentrations are provided, and their respective sample viscosities η are determined; num-conc is at least 2, preferably from 2 to 12, such as 4, 5, 6, 7, 8, 9, 10, 11 or 12, and a typical value is 6; the values of these protein concentrations and sample viscosities are used to determine at least one associated output parameter of the training and validation datasets as described above.
[0092] The different concentrations in the set of protein solutions can be, for example, from a minimum concentration (mini-conc) of 10 mg / mL to a maximum concentration (max-conc) of 320 mg / mL, or from 10 mg / mL to 280 mg / mL, or from 10 mg / mL to 240 mg / mL, or from 10 mg / mL to 210 mg / mL; a typical value of min-conc is 30 mg / mL, and a typical value of max-conc is 180 mg / mL. Preferably, if the set of protein solutions contains 3 or more protein solutions with different protein concentrations each, the difference between every two adjacent concentrations, i.e., the difference between two consecutive concentrations, is the same, and thus
[0093] (max-conc - min-conc) / (num-conc - 1) = diff-conc; and
[0094] conc n+1 - conc n = diff-conc.
[0095] max-conc is the maximum concentration in the set of protein solutions
[0096] min-conc is the minimum concentration in the set of protein solutions
[0097] num-conc is the number of protein solutions with different concentrations in the set of protein solutions
[0098] diff-conc is the difference between two consecutive concentrations
[0099] n is an integer from 0 to (num-conc - 1)
[0100] conc n is the concentration of the (n + 1)th protein solution in the set of protein solutions
[0101] In one embodiment, the protein solution set for each protein in the proteome comprises six protein solutions having six different concentrations, these concentrations being 30 mg / mL, 60 mg / mL, 90 mg / mL, 120 mg / mL, 150 mg / mL, and 180 mg / mL, so that
[0102] the max conc is 180 mg / mL,
[0103] the min conc is 30 mg / mL,
[0104] the num conc is 6,
[0105] the diff conc is 30 mg / mL,
[0106] and n is an integer from 0 to 5.
[0107] The method according to the invention may further comprise the step of validating the trained ANN, i.e., after performing the step of configuring the neural network (also referred to herein as training). In this "validation" step, a plurality of validation data sets may be provided. In other words, the method may further comprise validating the trained neural network based on a plurality of validation data sets, each of the plurality of validation data sets comprising a plurality of input parameters and at least one associated output parameter. The validation data sets may have the same or different numbers and types of parameters as the training data sets, but preferably, the validation data sets use the same type of parameters as the training data sets. Thus, each validation data set may be divided into input parameters and associated output parameters, similar to the training data sets. Then, the validation data sets may be used to validate the correct implementation or training of the neural network. For this purpose, the trained neural network may receive the input parameters of the validation data set to calculate the corresponding at least one associated output parameter. Then the calculated output parameter may be compared with the at least one associated output parameter included in the validation data set. This may be done for each set of validation data sets to determine whether the correct configuration of the neural network has been established.
[0108] In a further development, the trained neural network may be stored on a data medium (such as a non-transitory storage medium) which is configured to store digital data that would be apparent to a person skilled in the art based on the present disclosure.
[0109] Another aspect of the present invention includes a computer-implemented method for determining or predicting the concentration-dependent viscosity of a protein solution by using a computer-implemented neural network, particularly a trained neural network, more particularly the trained neural network as described above. Specifically, another aspect of the present invention may include providing and using a computer-implemented method for determining or predicting the concentration-dependent viscosity of a protein solution by using a computer-implemented neural network. The neural network, particularly a trained neural network, is configured to predict the concentration-dependent viscosity of the protein solution based on a plurality of input parameters associated with the protein solution, preferably wherein the input parameters include at least one of the following:
[0110] i) Experimental data, preferably selected from the apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), the diffusion interaction parameter (kD), the second virial coefficient (A2), and combinations thereof;
[0111] ii) Computational data, preferably selected from the isoelectric point (pI) of the protein, the variable fragment (Fv)-charge (Fv-charge), and combinations thereof; and
[0112] iii) Computer simulation data, preferably selected from the hydrophobic and charged patch sizes, particularly the total positive Fv score, the total positive Fv size, the total positive Fv count, the total negative Fv score, the total negative Fv size, the total negative Fv count, the total hydrophobic Fv score, the total hydrophobic Fv size, and the total hydrophobic Fv count.
[0113] In other words, the input parameters may include at least one selected from i) experimental data, ii) computational data, and iii) computer simulation data.
[0114] Preferably, the computer-implemented method of the present invention utilizes a neural network, particularly the trained neural network as described above, i.e., the neural network provided by the above method. Therefore, the technical features related to the method of providing a computer-implemented neural network described herein, particularly the technical features related to the provision and use of a computer-implemented neural network, can thus also be applied to and referred to for the computer-implemented method for predicting the concentration-dependent viscosity of a protein solution, and vice versa. This particularly applies to the above features of the neural network, particularly the features of the neural network to be trained or the trained neural network, the protein solution, any input parameter, and any output parameter.
[0115] The method of the present invention can be used to predict the viscosity of any suitable protein solution. As described above, the protein solution can contain a therapeutic protein, preferably an antibody, more preferably a monoclonal antibody (mAb), such as a monoclonal antibody of the IgG1 or IgG2 subtype. Thus, the protein solution can be an aqueous medium containing water, preferably a buffer, more preferably a histidine-HCl buffer.
[0116] The neural network used in the method of the present invention, in particular the trained neural network, is configured to determine or predict the concentration-dependent viscosity based on a plurality of input parameters.
[0117] As described above, the concentration-dependent viscosity can be expressed and / or determined as the viscosity value η of the protein solution at a selected protein concentration. cs . Alternatively or additionally, the concentration-dependent viscosity can be expressed and / or determined as a function f that correlates the viscosity of the protein solution with the protein concentration. η . Specifically, the concentration-dependent viscosity can indicate or can be represented by the function depicted by the above equation f η . More specifically, the neural network can be configured to calculate at least one of the constants A and B of the function f η .
[0118] In the method according to the present invention, the neural network, in particular the trained neural network, receives input parameters and determines the concentration-dependent viscosity of the protein solution based thereon. The input parameters are preferably associated with the protein solution whose viscosity is to be predicted. Specifically, the input parameters preferably refer to the protein contained in the protein solution. More preferably, the input parameters indicate the protein-protein interactions in the protein solution. Further, the input parameters can correspond to those contained in the training dataset as described above. Thus, the above technical features related to the parameters contained in the training dataset can equally apply to and be referenced for the input parameters to be received by the neural network for calculating the concentration-dependent viscosity, and vice versa.
[0119] More specifically, in the method, the neural network, in particular the trained neural network, receives the input parameters as described above. Thus, the input parameters include at least one of the following:
[0120] i) Experimental data, preferably selected from the apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), the diffusion interaction parameter (kD), the second virial coefficient (A2), and combinations thereof. More preferably, i) the experimental data consists of the apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), the diffusion interaction parameter (kD), and the second virial coefficient (A2) parameters;
[0121] ii) Calculated data, preferably selected from the isoelectric point (pI) of a protein, variable fragment (Fv)-charge (Fv-charge), and combinations thereof. More preferably, ii) the calculated data consists of the isoelectric point (pI) of the parameter protein and variable fragment (Fv)-charge (Fv-charge); and
[0122] iii) Computer simulation data, preferably selected from hydrophobicity and charged patch size, in particular positive Fv total score, positive Fv total size, positive Fv total count, negative Fv total score, negative Fv total size, negative Fv total count, hydrophobic Fv total score, hydrophobic Fv total size, and hydrophobic Fv total count.
[0123] In other words, the input parameters can include at least one selected from i) experimental data, ii) calculated data, and iii) computer simulation data.
[0124] Specifically, as described above, the input parameters can be selected from all three of the above data groups i) to iii), but if the neural network receives input parameters from only two of the three data groups (such as the neural network receiving input parameters from data groups i) and ii), or from data groups ii) and iii), or from data groups i) and iii) above, preferably from data groups ii) and iii)), then the concentration-dependent viscosity can also be reliably predicted.
[0125] In another aspect of the present invention, a method for determining the concentration-dependent viscosity of a protein solution is provided. The method includes the steps of predicting the concentration-dependent viscosity of the protein solution by using a computer-implemented neural network; and comparing the concentration-dependent viscosity with a target viscosity, in particular for determining the viscosity characteristics of the protein solution.
[0126] The proposed method can utilize the above neural network and the above computer-implemented method to predict the concentration-dependent viscosity of the protein solution. Therefore, the above technical features, in particular the technical features related to the neural network and its use for predicting the concentration-dependent viscosity of the protein solution, can also be applied to and referred to for the method for determining the concentration-dependent viscosity of the protein solution, and vice versa.
[0127] The proposed method can be used to evaluate whether a protein solution has a high or excessive viscosity, especially at a predefined protein concentration. In other words, a step of comparing the concentration-dependent viscosity with a target viscosity can be performed to determine whether a protein solution has a high or excessive viscosity, especially at a predefined protein concentration, and in particular whether it exceeds the target viscosity. If the viscosity of the protein solution makes the protein solution unsuitable as a pharmaceutical product for the above reasons, i.e., if its production cost is high because of high losses, low recovery rates, difficulties in manufacturing or filling during the purification process, and ultimately low administrability due to the need for high injection force and slow administration and possible pain, the viscosity of the protein solution can be considered excessive. Generally, solutions with a dynamic viscosity higher than 15 mPa·s to 30 mPa·s or even higher than 15 mPa·s to 20 mPa·s are considered problematic and thus "excessive".
[0128] In further development, the method can be used to determine the suitability of a protein solution as a pharmaceutical product, especially a protein solution to be used as or prepared as a pharmaceutical product. For this purpose, a step of comparing the concentration-dependent viscosity with a target viscosity is performed to determine the suitability of the protein solution as a pharmaceutical product.
[0129] By providing a method for evaluating the suitability of a protein solution as a pharmaceutical product based on the concentration-dependent viscosity determined by a computer-implemented neural network, especially a trained neural network, the proposed method enables the concentration-dependent viscosity that depends on the protein concentration in the protein solution to be easily and reliably taken into account, especially in the early stages of the development process. In this way, better validation of protein solutions used as pharmaceutical products can be carried out at an early stage, which affects the pharmaceutical products to be produced.
[0130] The proposed method can be used to evaluate whether a protein solution is suitable for any pharmaceutical product containing or consisting of the protein solution.
[0131] As described above, the method includes a step of comparing the determined concentration-dependent viscosity with a target viscosity. In the context of the present disclosure, the term "target viscosity" refers to the desired or predefined viscosity of the protein solution; or the target viscosity can represent an upper limit of the viscosity of the protein solution. For example, the target viscosity can indicate the maximum value of the viscosity of the protein solution at a given concentration. Specifically, the target viscosity can indicate that the maximum value of the viscosity of the protein solution is between 15 mPa·s and 30 mPa·s, or between 15 mPa·s and 25 mPa·s, or between 15 mPa·s and 20 mPa·s, such as 15 mPa·s or 18 mPa·s or 20 mPa·s or 25 mPa·s. When a protein solution is used as a pharmaceutical product, a viscosity higher than the maximum value may be considered problematic.
[0132] Therefore, this step can be performed to determine the suitability of the protein solution as a drug product. In other words, by comparing the concentration-dependent viscosity with the target viscosity, it can be evaluated whether the protein solution has a concentration-dependent viscosity that is favorable when used as a drug product. When the concentration-dependent viscosity meets or is lower than the target viscosity, that is, even when using the protein at high concentrations, such as >50 mg / mL, or >60 mg / mL, >70 mg / mL, or >80 mg / mL, or >90 mg / mL, or >100 mg / mL, the suitability of the protein solution can be determined. Thus, the application-specific viscosity of the protein solution can be predicted to evaluate whether the protein solution is suitable for use as a drug product in clinical treatment.
[0133] Preferably, in the step of comparing the concentration-dependent viscosity with the target viscosity, the protein concentration of the protein solution is considered. In other words, the concentration-dependent viscosity and the target viscosity can be compared at a specific protein concentration or within a range of protein concentrations.
[0134] When comparing the concentration-dependent viscosity and the target viscosity at a specific protein concentration, first, the protein concentration of the protein solution can be determined. For example, the maximum protein concentration can be determined, which can refer to the protein concentration that the protein solution is expected not to exceed when used as a drug product. The maximum protein concentration can be in the range from 100 mg / mL to 220 mg / mL, or from 120 mg / mL to 180 mg / mL, such as 120 mg / mL or 150 mg / mL or 180 mg / mL. Then, based on the determined concentration-dependent viscosity, the viscosity of the protein solution at the maximum protein concentration can be determined and compared with the target viscosity.
[0135] When comparing the concentration-dependent viscosity and the target viscosity within a range of protein concentrations, first, the actual range of protein concentrations for clinical treatment of the protein in the protein solution can be determined. For example, the protein concentration range can be from 20 mg / mL or 50 mg / mL or 100 mg / mL to the maximum protein concentration. Then, it can be determined whether the concentration-dependent viscosity of the protein solution within this protein concentration range exceeds the target viscosity.
[0136] In another aspect of the present invention, a method for providing (or preparing) a drug product comprising a protein solution is given. The method includes the steps of predicting (or determining) the concentration-dependent viscosity of the protein solution by using a computer-implemented neural network; comparing the concentration-dependent viscosity with the target viscosity to determine the suitability of the protein solution as a drug product; and optionally preparing the drug product if the suitability of the protein solution is determined.
[0137] The proposed method can utilize the above neural network, the above computer-implemented method for predicting the concentration-dependent viscosity of a protein solution, and the above method for determining the concentration-dependent viscosity. Therefore, the above technical features can also be applied to and referred to in methods for providing and optionally preparing pharmaceutical products, and vice versa.
[0138] If, in this method, the suitability of the protein solution as a pharmaceutical product is determined by comparing the concentration-dependent viscosity with a target viscosity, the steps for preparing the pharmaceutical product can be carried out. If the suitability of the protein solution as a pharmaceutical product cannot be determined, the protein solution can be adjusted or changed, and the steps for determining the concentration-dependent viscosity of the protein solution and comparing the determined concentration-dependent viscosity with the target viscosity can be carried out again based on the adjusted or changed protein solution.
[0139] Furthermore, a method for determining the appropriate concentration of a protein in a protein solution for a pharmaceutical product is also provided; the method includes the step of determining the concentration-dependent viscosity of the protein solution by using a computer-implemented neural network; and the step of comparing the determined concentration-dependent viscosity with a target viscosity to determine the upper limit of the protein concentration in the protein solution, which is still acceptable for the pharmaceutical product.
[0140] Furthermore, a method for identifying the upper limit of the protein concentration in a protein solution is also provided, which should not be exceeded to avoid the protein solution having an unacceptably high viscosity; the method includes the step of determining the concentration-dependent viscosity of the protein solution by using a computer-implemented neural network; and the step of comparing the determined concentration-dependent viscosity with a target viscosity to determine whether the protein in the protein solution induces a viscosity of the protein solution at a concentration higher than a certain level, which is unacceptable for the pharmaceutical product.
[0141] Furthermore, a method for identifying a protein in a solution with an unacceptably concentration-dependent viscosity is also provided; the method includes the step of determining the concentration-dependent viscosity of the protein solution by using a computer-implemented neural network; and the step of comparing the determined concentration-dependent viscosity with a target viscosity to determine whether the protein in the protein solution induces a viscosity of the protein solution at a concentration higher than a certain level, which is unacceptable for the pharmaceutical product.
[0142] Furthermore, the present invention provides the use of a computer-implemented neural network for facilitating the preparation of a pharmaceutical product, wherein the neural network is configured to predict the concentration-dependent viscosity of a protein solution, which is used to determine the suitability of preparing the protein solution as a pharmaceutical product.
[0143] Furthermore, the present invention also provides a method for determining the suitability of a protein solution as a pharmaceutical product, comprising:
[0144] Experimentally detecting parameters from the provided protein solution, preferably, wherein the obtained experimental data represents parameters selected from protein hydrophobicity, diffusion interaction parameter (kD), net protein charge, Zeta potential, second virial coefficient (A2), third virial coefficient (A3), apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), and combinations thereof;
[0145] Determining the concentration-dependent viscosity of the protein solution by using a computer-implemented neural network; and
[0146] Comparing the determined concentration-dependent viscosity with a target viscosity to determine the suitability of the protein solution as a pharmaceutical product.
[0147] Furthermore, the present invention provides a method for providing a trained artificial neural network (ANN), which determines the concentration-dependent viscosity of a protein solution by training the ANN with input parameters describing the proteins contained in a set of proteins.
[0148] Furthermore, the present invention provides a method for determining (or predicting) the concentration-dependent viscosity of a protein solution, which comprises providing input parameters of or indicative of a protein to a trained ANN and having the ANN calculate an output parameter indicative of or being the concentration-dependent viscosity.
[0149] Furthermore, the present invention provides an apparatus for predicting the concentration-dependent viscosity of an experimental protein solution, the apparatus comprising:
[0150] An experimental protein solution;
[0151] A computer component, the computer component comprising a neural network configured to predict the viscosity of the experimental protein solution:
[0152] Wherein the neural network is trained using a plurality of predetermined data sets;
[0153] Wherein each of the plurality of predetermined data sets comprises (i) a plurality of input parameters indicative of at least one physical property of a predefined protein solution, and (ii) at least one output parameter indicative of the concentration-dependent viscosity of the predefined protein solution; and
[0154] Wherein the plurality of input parameters includes at least one of experimentally obtained data, computationally obtained data, and data obtained by computer simulation.
[0155] The proposed device can utilize the above neural network and the above computer-implemented method to predict the concentration-dependent viscosity of a protein solution. Therefore, the above technical features, particularly those related to the neural network, more specifically those related to the above-trained neural network, and its use for predicting the concentration-dependent viscosity of a protein solution, can thus also be applied to and referred to in relation to the proposed device, and vice versa.
[0156] Furthermore, the present invention provides a system for determining the concentration-dependent viscosity of a protein solution, the system comprising:
[0157] a protein solution; and
[0158] a neural network configured to determine the viscosity of the protein solution, wherein determining the viscosity of the protein solution depends on a plurality of input parameters (IP) associated with the protein solution, and further wherein the input parameters (IP) include at least one selected from the group consisting of: (i) experimental data; (ii) computational data; and (iii) computer simulation data.
[0159] The proposed system can utilize the above neural network and the above computer-implemented method to predict the concentration-dependent viscosity of a protein solution. Therefore, the above technical features, particularly those related to the neural network, more specifically those related to the above-trained neural network, and its use for predicting the concentration-dependent viscosity of a protein solution, can thus also be applied to and referred to in relation to the proposed system, and vice versa. BRIEF DESCRIPTION OF THE DRAWINGS
[0160] The present disclosure will be more readily understood by reference to the following detailed description when considered in conjunction with the accompanying drawings, in which:
[0161] Figure 1a exemplary computing components that can be used to implement various features of embodiments of the present invention are shown;
[0162] Figure 1b a flowchart showing a method for providing a pharmaceutical product according to an embodiment of the present invention is shown;
[0163] Figure 2 a computer-implemented neural network used in the method depicted in FIG. 1 is schematically shown;
[0164] Figure 3 another computer-implemented neural network used in the method shown in FIG. 1 is schematically shown;
[0165] Figure 4Shows the training dataset and validation dataset for training and implementing the neural network used in the method depicted in Figure 1;
[0166] Figure 5 Depicts a graph showing a comparison between the viscosity values calculated by the proposed neural network and the measured viscosity values;
[0167] Figure 6 depicts an example of a predicted viscosity curve (predicted viscosity vs. measured viscosity for selected mAbs; crosses represent measured values, the black line is plotted using the slope (B) and intercept (A) predicted values in the interpolation equation X(y = A*e(B*x)), and the dashed line represents the viscosity threshold at 15 mPa*s for the problematic mAb). A) is mAb 13, B) is mAb 2, C) is mAb 17, D) is mAb 24, E) is mAb 26, F) is mAb 27);
[0168] Figure 7 Shows the difference between the predicted viscosity values calculated from the predicted viscosity curve using the predicted values of the intercept and slope and the measured viscosity values at the same concentration (the percentage difference is represented by bars and the absolute difference is represented by dots);
[0169] Figure 8 Shows the input variables and settings of the artificial neural network as exemplified (the input combinations are in one layer, which has 4 hidden nodes, and these 4 hidden nodes have a tanh activation function and a target determination viscosity descriptor, intercept A or slope B);
[0170] Figure 9a and 9b Depicts examples of training and validation data (9a: model xx - A training and validation data, and 9b: model xx - B training and validation data). Detailed Description
[0171] Hereinafter, the present invention will be explained in more detail with reference to the accompanying drawings. In the drawings, similar elements are denoted by the same reference numerals, and for the sake of avoiding redundancy, their repeated description may be omitted.
[0172] Figure 1aFIG. 0 is a schematic diagram showing an exemplary computing component 2 that can be used to implement various features of embodiments of the present invention. More specifically, computing component 2 generally includes a bus 3 that is connected to (i) a processor 4, (ii) a memory 5 for storing information and instructions to be executed by processor 4 (e.g., random access memory (RAM) or other dynamic memory, read-only memory (ROM), or any other static storage device for storing static information and instructions of processor 4, etc.), (iii) a storage device 6 including a non-transitory medium (e.g., hard disk drive, solid state disk drive, optical memory, etc.) for long-term storage of digital information (e.g., computer software, digital data, etc.), and (iv) a communication interface 7 for allowing software and data to be transmitted between computing component 2 and external devices via a communication channel 8, which will be apparent to those skilled in the art from the present disclosure.
[0173] Computing component 2 can be part of or can constitute a device for predicting the concentration-dependent viscosity of an experimental protein solution.
[0174] Figure 1b FIG. 7 depicts a method for providing and optionally preparing a pharmaceutical product comprising a protein solution.
[0175] In step S0 of the method, a computer-implemented neural network 10 is provided, which is configured to predict the concentration-dependent viscosity of a protein solution based on a plurality of input parameters associated with the protein solution. Step S0 represents a sub-method, i.e., a method included in a method for preparing a pharmaceutical product. It should be understood that the computer-implemented neural network 10 (which is preferably a trained neural network) can be implemented using the above-described computing component 2 or using any suitable computing system that will be apparent to those skilled in the art from the present disclosure.
[0176] The protein solution contains a protein, preferably a therapeutic protein (such as an antibody), which can include monoclonal antibodies and polyclonal antibodies, intact antibodies, antibody-drug conjugates, chimeric antibodies, humanized antibodies, human antibodies, or hybrid antibodies having dual or multiple antigen or epitope specificities, antibody fragments and antibody sub-fragments (such as Fab, Fab', F(ab')2 fragments), including any immunoglobulin or hybrid fragment of any natural, synthetic, or genetically engineered protein that acts like an antibody by binding to a specific antigen to form a complex. In the illustrated embodiment, the therapeutic protein is a monoclonal antibody (mAb), preferably a monoclonal antibody of the IgG1 or IgG2 subtype.
[0177] Further, the protein solution contains an aqueous water medium, preferably a buffer, more preferably a histidine-HCl buffer.
[0178] In a first sub-step S0.1, a neural network 10 to be trained is provided, in particular an untrained neural network. The neural network 10 to be trained can be configured to receive a plurality of input parameters and, based thereon, calculate at least one output parameter. In the configuration shown, the neural network 10 is configured to calculate the concentration-dependent viscosity of a protein solution. Generally speaking, the concentration-dependent viscosity associates at least one protein concentration of the protein solution with the corresponding viscosity of the protein solution (in particular the dynamic viscosity).
[0179] According to Figure 2 one embodiment of the invention depicted in, the neural network 10 (in particular the neural network to be trained and the trained neural network) is configured to calculate the viscosity value c of a protein solution at a selected protein concentration s . The selected protein concentration c s can refer to the maximum protein concentration expected to be associated with a drug product. The maximum protein concentration can be in the range from 100 mg / mL to 220 mg / mL, or in the range from 120 mg / mL to 180 mg / mL, for example 120 mg / mL or 150 mg / mL or 180 mg / mL. Thus, in this embodiment, the concentration-dependent viscosity is denoted as η cs , i.e., the viscosity value c of the protein solution at the selected protein concentration s .
[0180] According to Figure 3 another embodiment of the invention depicted in, the neural network 10 (in particular the neural network to be trained and the trained neural network) is configured to calculate a mathematical function that associates the viscosity value of a protein solution with the protein concentration of the protein solution. Specifically, in this embodiment, the concentration-dependent viscosity is represented by the function f as specified in equation (1) above η . Thus, in this embodiment, the concentration-dependent viscosity is provided in the form of the function f η . For this purpose, the neural network 10 calculates the constants A and B in the above equation (1).
[0181] Hereinafter, the neural network 10 to be trained will be described in more detail with reference to Figure 2 and Figure 3 the embodiments depicted in.
[0182] Figure 2 depicts an embodiment of the neural network 10, in particular the neural network to be trained and the trained neural network, which is used in the method depicted in FIG. 1. The neural network 10 includes an input layer 12 having a plurality of input nodes IN 1 -IN i , where the index "i" is a positive integer. In other words, the input layer 10 includes i different input nodes IN. AsFigure 2 depicted in, each input node IN 1 -IN n receives a corresponding input parameter IP 1 -IP i .
[0183] Furthermore, the neural network 10 includes a hidden layer 14 having a plurality of hidden nodes HN 1 -HN j where the index "j" is a positive integer, preferably j is 4, and the neural network has four hidden nodes HN 1 -HN 4 , where each hidden node is connected to each input node IN 1 -IN i . The hidden node HN 1 -HN j uses tan h as the activation function. Alternatively, sigmoid can be used as the activation function.
[0184] Still further, the neural network 10 includes an output layer 16 having a single output node ON, and the output node is connected to all hidden nodes HN 1 -HN j . The output node ON provides a single output parameter in the form of parameter η cs , and the output parameter is the viscosity of the protein solution at the selected protein concentration c s . In this embodiment, the parameter η cs constitutes the concentration-dependent viscosity.
[0185] Figure 3 depicts another embodiment of the neural network 10 according to the present invention, particularly the neural network to be trained and the trained neural network. Compared with Figure 2 the configuration depicted in, the output layer 16 of the neural network 10 has two output nodes ON 1 and ON 2 , which provide different output parameters. Specifically, the first output node ON 1 provides an output parameter OP 1 , which is A, while the second output node ON 2 provides an output parameter OP 2 , which is B. As described above, these two parameters represent the constants of the function f η . In this embodiment, the constants A and B of the function f η constitute the concentration-dependent viscosity.
[0186] In the next sub-step S0.2, a plurality of training data sets S are provided T, each of the plurality of training datasets is associated with a specific protein solution and includes a plurality of input parameters IP indicative of the specific protein solution and at least one associated output parameter OP indicative of the concentration-dependent viscosity of the specific protein solution. Further, a validation dataset S having the same number and type of input and output parameters is also provided V 。 Figure 4 depicts a table that generally shows the training dataset S for the proper implementation of the neural network 10 T and the validation dataset S V 。 Each training dataset S T and each validation dataset S V includes input parameters IP representing the values that the input nodes IN 1 -IN i will receive and input parameters IP 1 -IP i and two output parameters OP representing the values that the two output nodes ON 1 and ON 2 will provide 1 and OP 2 。
[0187] Specifically, m different training datasets S are provided T Based on these training datasets, the computational model on which the neural network 10 is based is adjusted during the training phase in the next sub-step S0.3. In this context, the parameter "m" represents a positive integer.
[0188] Specifically, in sub-step S0.3, based on the training dataset S T the neural network 10 (specifically the neural network to be trained) is trained to provide a computer-implemented neural network 10, specifically a trained neural network, that is suitable and configured to predict concentration-dependent viscosity. To this end, all input parameters IP and output parameters OP of the training dataset S T are fed into the neural network 10 to adjust the computational model on which the neural network 10 is based, specifically to adjust the weights of the nodes and their connections.
[0189] In another optional sub-step, n different validation datasets S are used V to validate the neural network 10, specifically the trained neural network. In this context, the parameter "n" represents a positive integer. Preferably, each of the validation datasets S V is different from each of the training datasets S T To validate the provided neural network 10 (specifically the trained neural network), the input parameters associated with the validation dataset S V are fed into the neural network 10, and the neural network calculates each validation dataset S based on thisV The corresponding output parameters. Then compare the calculated output parameters with the output parameters included in the validation dataset S V .
[0190] For example, to train such a neural network, 20 or more (e.g., 24) different training datasets S T can be used. Further, two or more (e.g., three) different validation datasets S V can be used.
[0191] As described above, the neural model can include i different input nodes and can thus receive i different input parameters IP. The input parameter IP preferably indicates the protein-protein interaction in the protein solution associated therewith. Further, the input parameter IP can include different types of input parameters, which can be classified into i) experimental data, ii) computational data, and iii) computer simulation data.
[0192] The experimental data can be obtained from the protein solution by detecting the provided parameters. Preferably, i) the experimental data represents parameters selected from the apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), the diffusion interaction parameter (kD), the second virial coefficient (A2), and combinations thereof. More preferably, i) the experimental data consists of the apparent surface hydrophobicity parameter measured by hydrophobic interaction chromatography (HIC), the diffusion interaction (kD) parameter, and the second virial coefficient (A2) parameter. The computational data can be calculated from the primary sequence of the protein at pH 5.0 to 7.0.
[0193] Preferably, the computational data represents parameters selected from the isoelectric point (pI) of the protein, the variable fragment (Fv)-charge (Fv-charge), and combinations thereof. More preferably, ii) the computational data consists of the isoelectric point (pI) parameter of the protein and the variable fragment (Fv)-charge (Fv-charge) parameter.
[0194] The computer simulation data is preferably selected from hydrophobicity and charged patch size and combinations thereof; more preferably from the size of the patch (in (in units), score (the sum of the scores of all contributing patches associated with the patch), and type (positive (pos) = positively charged; negative (neg) = negatively charged; hydrophobic (hyd) = hydrophobic). Thus, the descriptors derived from the modeling are the positive / negative / hydrophobic Fv total score, positive / negative / hydrophobic Fv total size, positive / negative / hydrophobic Fv total count; and any combination thereof. More preferably, iii) the computer simulation data consists of the positive Fv total score, positive Fv total size, positive Fv total count, negative Fv total score, negative Fv total size, negative Fv total count, hydrophobic Fv total score, hydrophobic Fv total size, and hydrophobic Fv total count.
[0195] In the following, an embodiment of the neural network 10 (in particular, a trained neural network) is described, which is equipped with 14 different input nodes IN 1 -IN 14 and 14 different input parameters IP 1 -IP 14 . It will be apparent to those skilled in the art that these input parameters only depict examples of a variety of possibilities. Therefore, the input parameters described below should not be construed as forming a limitation. Thus, more or fewer input parameters can be used. In another embodiment, the neural network can utilize at least one of the input parameters IP 1 to IP 14 .
[0196] Specifically, the input parameters IP 1 -IP 3 are i) experimental data, the input parameters IP 4 and IP 5 are ii) calculated data, and the input parameters IP 6 to IP 14 are iii) computer simulation data. Regarding the substance:
[0197] IP 1 refers to HIC RT [min];
[0198] IP 2 refers to kD [mL / g];
[0199] IP 3 refers to A2*E -4 [mol*mL / g2];
[0200] IP 4 refers to pI;
[0201] IP 5 refers to the Fv charge at pH 6;
[0202] IP6 refers to the total positive Fv score;
[0203] IP 7 refers to the total positive Fv size;
[0204] IP 8 refers to the total positive Fv count;
[0205] IP 9 refers to the total negative Fv score;
[0206] IP 10 refers to the total negative Fv size;
[0207] IP 11 refers to the total negative Fv count;
[0208] IP 12 refers to the total hydrophobic Fv score;
[0209] IP 13 refers to the total hydrophobic Fv size; and
[0210] IP 14 refers to the total hydrophobic Fv count.
[0211] In one embodiment, a combination of i) experimental data consisting of input parameters IP 1 -IP 3 , ii) computational data consisting of input parameters IP 4 and IP 5 , and iii) computer simulation data consisting of input parameters IP 6 to IP 14 is used.
[0212] In another embodiment, only a combination of ii) computational data consisting of input parameters IP 4 and IP 5 and iii) computer simulation data consisting of input parameters IP 6 to IP 14 is used; in this embodiment, the neural network 10 is equipped with 11 different input nodes IN 4 -IN 14 and 11 different input parameters IP 4 -IP 14 .
[0213] According to the previous embodiment, the neural network 10, particularly the trained neural network, is used to model the viscosity of a monoclonal antibody solution. Figure 5A graph depicting the comparison of the output values provided by the neural network 10 for modeling the viscosity of a monoclonal antibody solution with the measured viscosity values is shown. It can be seen that the neural network 10, especially the trained neural network, effectively and reliably (especially by calculating the viscosity curve) predicts the concentration-dependent viscosity of the protein solution.
[0214] In the next step S1 of the method, a protein solution from which a drug product can be produced is identified.
[0215] In step S2, the concentration-dependent viscosity, especially the protein concentration-dependent viscosity, is predicted by using a computer-implemented neural network 10, especially by using a trained neural network. Thus, step S2 represents a sub-method, namely a computer-implemented method for predicting the concentration-dependent viscosity of a protein solution.
[0216] In this step, the neural network 10 provided in step S0, especially the trained neural network, is used to calculate the concentration-dependent viscosity of the protein solution. For this purpose, input parameters IP, especially IP 1 to IP 14 are determined, at least one of which, preferably all, is provided to the neural network 10 to calculate at least one output parameter OP indicating the concentration-dependent viscosity based on this.
[0217] In the next step S3, the concentration-dependent viscosity determined in step S2 is compared with a target viscosity to determine the suitability of the protein solution as a drug product. Steps S2 and S3 together represent a sub-method, namely a method for determining the concentration-dependent viscosity of a protein solution.
[0218] In this context, the target viscosity represents the maximum value of the viscosity of the protein solution. Specifically, the target viscosity represents the maximum value between 15 mPa·s and 30 mPa·s, or between 15 mPa·s and 25 mPa·s, or between 15 mPa·s and 20 mPa·s, such as 15 mPa·s or 18 mPa·s or 20 mPa·s or 25 mPa·s. In this step, if the concentration-dependent viscosity (especially at the selected protein concentration c s or for the selected protein concentration range, such as between 20 mg / mL and the selected protein concentration c s does not exceed the target viscosity (i.e., the threshold or maximum value 15 mPa·s or 20 mPa·s), the suitability of the protein solution is determined. However, if the viscosity of the protein solution exceeds the target viscosity, the suitability of the protein solution is not determined (i.e., it is determined that the protein solution is not suitable as a drug product, or it is determined that the protein solution may not be suitable as a drug product).
[0219] According to Figure 2 In the embodiment depicted in Figure 2 , in step S3, the parameter η cs is compared with the target viscosity. If the parameter η cs does not exceed the target viscosity, the suitability of the protein solution as a drug product is determined (i.e., it is determined that the protein solution is suitable as a drug product, or it is determined that the protein solution may be suitable as a drug product).
[0220] According to Figure 3 In the embodiment depicted in Figure 3 , in step S3, the derived function f η is used to determine whether the viscosity of the protein solution exceeds the target viscosity within a selected protein concentration range, for example, between 20 mg / mL and the selected protein concentration c s . Alternatively, in this step, the viscosity η η of the protein solution at the selected protein concentration c s can first be determined based on the function f cs , and then it is determined whether the viscosity η cs (i.e., the viscosity at the selected protein concentration c s ) exceeds the target viscosity. Thus, if it does not exceed the target viscosity, the suitability of the protein solution as a drug product is determined (i.e., it is determined that the protein solution is suitable as a drug product, or it is determined that the protein solution may be suitable as a drug product).
[0221] As indicated in step S4 of FIG. 1, if the suitability of the protein solution is determined in step S3, the method can proceed to the optional step S5, where a drug product is prepared (or produced) based on the protein solution identified in step S1 (particularly at the required protein concentration). If the suitability of the protein solution is not determined in step S3, the method can proceed to step S1, where a new protein solution is identified (particularly by adjusting or changing the initial protein solution), and then steps S2 to S4 are performed again.
[0222] Therefore, it can be seen that by using the above method to screen potential protein solutions (i.e., by using the neural network 10 provided in step S0 to screen potential protein solutions), candidate protein solutions for use as drug products can be identified without actually preparing and empirically testing the potential protein solutions. The above computer simulation method for screening potential protein solutions allows for screening a large number of candidate protein solutions without empirically measuring the quality of each candidate protein solution. Thus, the present invention addresses the long - unmet need for a method for rapid, high - throughput screening of candidate protein solutions for use as drug products, thereby avoiding the time and labor inherent in empirical (i.e., laboratory - based) screening.
[0223] It is obvious to those skilled in the art that these examples and items only depict examples of various possibilities. Therefore, the examples shown herein should not be construed as forming limitations on these features and configurations. Any possible combination and configuration of the described features can be selected according to the scope of the present invention.
[0224] Abbreviation
[0225]
[0226]
[0227] Materials and Methods
[0228] Experimental Setup
[0229] To evaluate the impact of various input parameters on the prediction ability of artificial neural networks, three different models were created, including data of mAb 1 - 25 (Tables 4 and 5 below herein), where only experimental inputs (retention time in HIC; kD, A2), or only computer - simulated inputs (patch size), or both inputs were fed into the modeling; in all three models, regardless of which input was selected, Fv - charge and pI were included. Each model was created by dividing the input data into a training set and a validation set, which included information on randomized mAbs from each new artificial neural network model. The training set contained data, input variables, and viscosity descriptors of 18 - 20 mAbs. The validation set only contained the input variables of the remaining 7 - 5 mAbs as data for the purpose of predicting their viscosity descriptors. These viscosity descriptors were derived from the linearization of the viscosity - concentration curve of each mAb, the intercept A, and the slope B.
[0230] The R 2 value of the created models shows the interdependence of the validation and training sets (Table 1; Figure 9a and 9b depicts an exemplary graph), because in several cases, the quality of one set was very good (R 2 > 0.99), while the quality of the second set was not so good (R 2 < 0.95). The artificial neural network using both variables (experimental and computer - simulated) obtained the highest - quality model, where the lowest value was R 2 > 0.92 (intercept A of the training set). The artificial neural network using only computer - simulated inputs achieved the minimization of the lower "worst" R 2 > 0.90 (slope B of the training set). Finally, the artificial neural network containing only experimental data achieved the lowest R 2> 0.75, which is comparable to the results published for the linear correlation of kD with viscosity and is superior to the linear correlation of A2 and kD with intercept and slope parameters done for this dataset.
[0231] The slope parameter B describes the steepness of the viscosity curve, i.e., the viscosity increases exponentially with protein concentration, which is very important for describing potentially "problematic" mAbs. Such "problematic" mAbs may show moderate viscosity at low protein concentrations, but the viscosity increases significantly, above values typically considered acceptable for drug products (i.e., above values of about 15 mPa*s - 25 mPa*s). Thus, in the following, an artificial neural network is used with all input variables because it provides the best prediction for slope B.
[0232] Among all the available data, mAbs 26 and 27 were not used for artificial neural network model creation but were saved separately for other validation tests.
[0233] Table 1:
[0234]
[0235] Table 1 shows a comparison of models created using different inputs that predict the intercept (A) or slope (B) of the concentration-dependent viscosity curve of mAbs. "x" indicates that a particular set of inputs was used for the model, while "-" indicates that the set of input parameters was not used. Experimental inputs refer to HIC retention time, kD, and A2, while computer simulation inputs include data from surface patch analysis, i.e., the size, score, and count information of positive, negative, and hydrophobic surface patches. Regardless of the input selected, Fv charge and pI are included. The quality parameters of the shown models were obtained by plotting the A or B values calculated from the measured viscosity against the predicted values of the corresponding models. R 2 = coefficient of determination; SSE = standard squared error; RMSE = root mean square error
[0236] Classification
[0237] Using experimental and computer simulation data as input, a classification model was created by training a model to determine whether an mAb shows a viscosity above the 15 mPa*s threshold at a specific concentration. This was done for concentrations of 120, 150, and 180 mg / mL, where out of 25 mAbs used in model creation, 3 mAbs had a viscosity above 15 mPa*s at 120 mg / mL, 6 mAbs at 150 mg / mL, and 15 mAbs at 180 mg / mL. The confusion matrices for these models are shown in Table 2 below, and clearly both the training set and the validation set contain problem (defined in this experiment as viscosity ≥ 15 mPa*s) and non-problem (defined in this experiment as viscosity < 15 mPa*s) mAbs. Regardless of the input used, the created models have excellent predictive power, with a misclassification rate of 0 for all models (Table 8). To further evaluate the predictive power of these models, two mAbs not used in the training set and validation set were selected for validation. mAb26 showed non-problematic behavior at 120 mg / mL and 150 mg / mL but exceeded 15 mPa*s at 180 mg / mL, which was correctly predicted by the models disclosed herein (no / no / yes). mAb 27 did not show problematic behavior at any concentration, which was again correctly predicted by the models proposed herein (no / no / no).
[0238] Table 2: Confusion matrix of the classification model using experimental and computer simulation data as input. Results for the validation set are in parentheses.
[0239]
[0240] Viscosity Curve Prediction
[0241] To obtain more detailed information about the viscosity of the mAb, the full concentration-dependent viscosity curve was predicted, or more specifically, the intercept (A) and slope (B) of a linear exponential function were predicted. Using the predicted values of A and B, a theoretical viscosity curve can be constructed and compared to the actual measured values, and an example of such a comparison is shown in Figures 6A - 6D. Although not an exact match to the actual values, the predicted values are very close, and the predicted curve reflects the actual concentration-dependent viscosity in a similar manner. Figure 7The percentage and absolute differences between the predicted and measured viscosities of all mAbs are shown. Table 3 lists the average calculated differences between the predicted and measured viscosity values for all 27 mAbs. The average absolute difference between the calculated and measured values (in mPa*s) is between 0.1 mPa*s - 4.1 mPa*s, and the difference gradually increases with increasing protein concentration. The relative % difference is between 8.2% - 26.2%, but follows a curvilinear function, with the largest difference at concentrations of 90 mg / mL - 120 mg / mL, and the difference decreasing with decreasing and increasing protein concentration.
[0242] Table 3: Comparison of Predicted and Measured Viscosities (including data for all mAbs in this study)
[0243]
[0244] To further validate the model, mAbs 26 and 27 were used, and the predicted viscosity curves for the two mAbs were plotted in Figures 6E and 6F. Although the trend of the predicted curve for mAb 27 (Figure 6F) is very similar to the measured values, the predicted curve for mAb 26 (Figure 6E) does not correctly reflect the steepness of the measured curve above 150 mg / mL.
[0245] Using computational data and computer simulation modeling has the advantage of relatively easy access and does not require materials and laboratory work.
[0246] The present invention provides artificial neural networks and methods for predicting and determining the viscosity of proteins (such as mAbs) in solution. These models can be used to predict viscosity classifications above or below a given threshold (e.g., 15 mPa*s) and to predict viscosity curves. Although the classification is already good, the viscosity curve prediction model shows high power and good viscosity curve prediction. It is advantageous to use a large number of input variables from experimental data, computational data, and computer simulation modeling.
[0247] Monoclonal Antibody
[0248] mAbs were obtained. A bicistronic vector containing the heavy and light chains was transfected into CHOK1SV GS-KO cells and cultured into a stable mixed culture under selection conditions. The clarified supernatant was obtained by centrifugation and then filtered and sterilized using a 0.22 μm filter. mAb purification was performed using Protein A chromatography. All proteins were concentrated to a final concentration of 10 mg / mL, and the buffer was exchanged to the formulation buffer (protein solution) (20 mM histidine-HCl, pH 6.0) by tangential flow filtration. The mAbs belong to different subtypes IgG1 or IgG2 (see Tables 4 and 5 below).
[0249] Buffer
[0250] All the described laboratory experiments (HIC, DLS, SLS, viscosity) were carried out in a buffer (20 mM histidine-HCl, pH 6.0) as a protein solution (later as a drug product).
[0251] Protein Concentration
[0252] For concentration determination, an Agilent Cary 60 UV-visible spectrophotometer with a variable path length extender SoloVPE was used. For each measurement, 30 μL of the sample was loaded into a cuvette and measured at 280 nm using the appropriate specific extinction coefficient.
[0253] Hydrophobic Interaction Chromatography
[0254] The hydrophobic surface properties of all mAbs were determined by hydrophobic interaction chromatography (HIC). The protein was analyzed at 10 mg / mL in the formulation buffer, 5 μL was injected onto a ProPac Hic-10 column (Thermo Scientific) and separated using a Waters HPLC system. The starting condition of 95% mobile phase A (1 M ammonium sulfate in 20 mM sodium phosphate pH 7.0) was linearly reduced to 95% mobile phase B (20 mM sodium phosphate pH 7.0) within 39 minutes. The flow rate was set at 1 mL / min and the column temperature was 24 °C.
[0255] Dynamic and Static Light Scattering
[0256] Dynamic light scattering (DLS) and static light scattering (SLS) measurements were performed on a DynaPro PlateReader III (with software Dynamics; Wyatt Technologies). The stock solution of the antibody was filtered through a 0.22 μm PVDF filter (MillexGV), and serial dilutions of seven concentrations with 10 - 2 mg / mL protein were prepared using the formulation buffer. The samples were transferred in triplicate to a 384-well plate (Aurora), and then the plate was centrifuged at 750*g for 2 minutes to remove air bubbles. The measurement temperature was set at 25 °C. The laser power was set at 20%, and the attenuation level was set at 0%. Twenty acquisitions of 5-second length were performed for each well. The diffusion interaction parameter kD (mL / g) was evaluated via DLS. The mutual diffusion coefficient Dm (m2 / s) was plotted against the protein concentration (g / mL), and kD was obtained from the slope of the linear fit. The second virial coefficient A2 (mol*mL / g) was obtained from the SLS measurement. The plate was calibrated using dextran (Sigma) with a predetermined molecular weight of 36.9 ± 0.1 kDa. The solvent offset of the formulation buffer was measured in triplicate. The reciprocal of the molecular weight (mol / g) was plotted against the protein concentration (g / mL), and A2 was obtained from the slope of the linear fit.
[0257] Computer Simulation Modeling Patch of mAb Fv
[0258] Modeling of the mAb was performed using software BioLuminate (version 3.80; LLC, New York, NY). Homology modeling of the Fv region was done using an antibody prediction tool. The framework template of isotype IgG1 or IgG2 was selected based on the highest combined score in the pdb database. The best CDR loop clusters were automatically selected. For the modeling, the standard presets of the software were retained, except that the pH was set to 6.0 to represent the experimental setup. The surface of the modeled mAb Fv region was analyzed using the protein surface analyzer tool in the BioLuminate software to obtain the size (in units), score (the sum of all contributing patch scores associated with the patch), and type (positive (pos) = positively charged; negative (neg) = negatively charged; hydrophobic (hyd) = hydrophobic). Thus, the descriptors derived from the modeling were the total positive / negative / hydrophobic Fv score, the total positive / negative / hydrophobic Fv size, and the total positive / negative / hydrophobic Fv count.
[0259] Fv Charge Calculation
[0260] To calculate the Fv charge, the variable heavy and variable light chains of each antibody were analyzed using the prot pi protein tool (https: / / www.protpi.ch / Calculator / ProteinTool). Each chain was defined as a subunit of the whole protein. The collective modifier for post-translational modification was the global disulfide bond of cysteine residues in the Fv. The charge was calculated at pH 6.0.
[0261] pI Calculation
[0262] The pI was calculated "manually" based on the pKa values of the amino acid residues in the primary sequence.
[0263] Rheological Determination and Viscosity Descriptor
[0264] The mAb was concentrated to approximately 180 mg / mL using a spin filter with a molecular weight cut-off of 30 kDa and then diluted to six concentrations between 180 - 30 mg / mL. Concentration-dependent viscosity data were generated for each concentration using a VROC viscometer (Rheosense). To obtain descriptive information on the concentration-dependent viscosity of the mAb samples, the experimentally measured viscosity data were processed based on the above equations (3) to (5).
[0265] The above equation (3) was used to calculate the relative viscosity of the mAb samples. For this purpose, the viscosity of the buffer was first measured at 25 °C to be 0.92 mPa*s. After knowing the viscosity of the buffer, the relative viscosity of the mAb samples was calculated based on the measured viscosity values according to equation (3). Thus, the concentration and relative viscosity value pairs for each mAb sample were determined.
[0266] The above equation (4) was used to describe the exponential concentration-dependent viscosity of the mAb solution. This equation can be linearized using the natural logarithm as in equation (5) to obtain the intercept A and slope B.
[0267] Based on the determined concentration and relative viscosity value pairs of the mAB samples, the viscosity descriptors A and B for each mAb were then determined by applying least squares fitting based on the above equations (4) and (5).
[0268] Data Analysis and Artificial Neural Network Modeling
[0269] The creation of the ANN model follows Figure 8The method shown in . All input parameters (experimental data, calculated values from sequences, computer simulation data, and viscosity descriptors) are listed in Tables 4 and 5. Each model was trained with a classification response aimed at identifying mAbs that showed viscosity values above the threshold of 15 mPa*s at different protein concentrations. The ANN was generated using the software JMP v.16.0.0 (SAS Institute Inc.). The activation function used for all nodes was the tanh function, which converts values to values between -1 and 1. For all models, one hidden layer with four nodes was sufficient. To prevent the network from overfitting the model and thus losing its predictive ability, the data was divided into a training set and a validation set. The method used was K-fold, where 25 mAbs and their datasets were divided into K sets. Each set in the K sets contained the datasets of all 25 mAbs, but in each K set, the two subsets, the training set and the validation set, were different, and the mAbs in each K set were randomly distributed to the two subsets, the training set and the validation set. Each set in the K sets was used to validate the model fit to the remaining data, and a total of K sets were fitted. The value of K for each model was set to 5. The distribution of mAbs among the K sets was also uncorrelated between the three models but was random, so each model had different K sets. The models reported by the JMP software were based on the best log-likelihood. The quality of the ANN was determined using the coefficient of determination (R2), the standard squared error (SSE), and the root mean square error (RMSE) for the training and validation datasets.
[0270] Table 4: Data included in the artificial neural network modeling. Experimental inputs (retention time in hydrophobic interaction chromatography (HIC); diffusion interaction parameter kD; second virial coefficient A2) and calculated inputs derived from the amino acid sequence (isoelectric point (pI) and Fv charge).
[0271]
[0272]
[0273] Table 5; Data included in the artificial neural network modeling. Inputs derived from computer simulations (scores and areas of positively charged, negatively charged, and hydrophobic patches). Viscosity descriptors intercept A and slope B were derived from the linearization of viscosity data (raw viscosity data in Tables 6 and 7). The data was fed into the artificial neural network following the Figure 8 scheme in .
[0274]
[0275]
[0276] Table 6: Viscosity raw data. For each mAb, a series of dilutions were prepared in 20 mM histidine-HCl pH 6 buffer with target concentrations (target c) of 180, 150, 120 mg / mL. The table contains the dynamic viscosity values (mPa*s ± SD) from 10 measurements and the measured protein concentration (mg / mL).
[0277]
[0278]
[0279] Table 7: Viscosity raw data. For each mAb, a series of dilutions were prepared in 20 mM histidine-HCl pH 6 buffer with target concentrations (target c) of 90, 60, 30 mg / mL. The table contains the dynamic viscosity values (mPa*s ± SD) from 10 measurements and the measured protein concentration (mg / mL).
[0280]
[0281]
[0282] As described above, the viscosity descriptors intercept A and slope B were calculated from the values shown in Tables 6 and 7. To do this, the viscosity of the buffer was first determined, and the measured viscosity was 0.92 mPa*s. Then, based on Equation (3) above, the relative viscosity of each mAb and of each pair of concentration and viscosity values in Tables 6 and 7 was calculated using the buffer viscosity. Subsequently, a least squares fit (or any other mathematical procedure for finding the relationship for a given set of points, particularly for finding the best fit curve) was performed on the 6 pairs of concentration and viscosity (relative viscosity) values for each mAb, thus fitting the 6 pairs of values to the exponential function of Equation (4) above. Then, as shown in Equation (5) above, the calculated function was linearized to provide the A and B values for each mAb included in Table 5.
[0283] Table 8: Regardless of the level of the input variables, i.e., only experiments, only computer simulations, or both types of variables (including Fv charge and pI regardless of the input selected in all three models), the models derived from artificial neural networks are able to correctly classify whether a certain mAb will exhibit problematic behavior, such as a viscosity higher than 20 mPa*s, above a certain protein concentration threshold.
[0284]
Claims
1. A method for providing a computer-implemented neural network (10), the computer-implemented neural network being configured to predict the concentration-dependent viscosity of a protein solution based on a plurality of input parameters (IP) associated with the protein solution, the method comprises the following steps: - Provide a plurality of training data sets (S T ), each of the plurality of training data sets being associated with a specific protein solution and including a plurality of input parameters (IP) indicative of the specific protein solution and at least one associated output parameter (OP) indicative of the concentration-dependent viscosity of the specific protein solution; and -Based on the training data set (S T ), training a neural network to provide the computer-implemented neural network (10), the computer-implemented neural network being configured to predict the concentration-dependent viscosity, wherein The input parameters (IP) include at least one of the following: i) Experimental data; ii) Computational data; and iii) Computer simulation data.
2. The method according to claim 1, wherein the protein solution comprises a therapeutic protein, preferably an antibody.
3. The method according to claim 2, wherein the antibody is a monoclonal antibody (mAb), preferably a monoclonal antibody of the IgG1 or IgG2 subtype.
4. The method according to any one of claims 1 to 3, wherein the protein solution comprises an aqueous aqueous medium, preferably a buffer, more preferably a histidine-HCl buffer.
5. The method according to any one of claims 1 to 4, wherein the concentration-dependent viscosity represents a value (η cs ) that correlates the viscosity of the protein solution with a specific protein concentration of the protein solution.
6. The method according to any one of claims 1 to 4, wherein the concentration-dependent viscosity represents a function (f η ) that correlates the viscosity of the protein solution with the protein concentration of the protein solution, in particular a mathematical function.
7. The method according to claim 6, wherein the concentration-dependent viscosity indicates the following function: f η (c) = A * e Bc , where f η is a function that correlates the viscosity of the protein solution with the protein concentration; c is the protein concentration; A and B are constants.
8. The method according to claim 7, wherein the concentration-dependent viscosity of the protein solution is represented by at least one of the constants A and B.
9. The method according to any one of claims 1 to 8, further comprising the step of providing the neural network to be trained, the neural network to be trained comprises: An input layer (12), the input layer having a plurality of input nodes (IN), each of the plurality of input nodes receiving an input parameter (IP); At least one hidden layer (14), in particular only one hidden layer (14), the at least one hidden layer having a plurality of hidden nodes (HN), in particular four or more hidden nodes (HN); And An output layer (16), the output layer having at least one output node (ON), the at least one output node for providing at least one output parameter (OP) indicating the concentration-dependent viscosity.
10. The method according to claim 9, wherein the hidden nodes (HN) use tan h as the activation function.
11. The method according to claim 9 or 10, wherein the at least one output parameter provided by the at least one output node (ON) indicates: The viscosity value (η cs ) that correlates the viscosity of the protein solution with the specific protein concentration of the protein solution; or At least one of the constants A and B of the following function: f η (c) = A * e Bc , where f η is a function that correlates the viscosity of the protein solution with the protein concentration; c is the protein concentration; and A and B are constants.
12. The method according to any one of claims 1 to 11, wherein the input parameter (IP) indicates protein-protein interaction.
13. The method according to any one of claims 1 to 12, wherein the neural network is an artificial neural network.
14. The method according to any one of claims 1 to 13, further comprising the step of validating the trained neural network (10) based on a plurality of validation data sets (S V ), each of the plurality of validation data sets including a plurality of input parameters (IP) and at least one associated output parameter (OP).
15. A computer-implemented method for predicting the concentration-dependent viscosity of a protein solution by using a neural network, wherein the neural network is configured to predict the concentration-dependent viscosity of the protein solution based on a plurality of input parameters associated with the protein solution, and wherein the input parameters (IP) include at least one of the following: i) Experimental data; ii) Computational data; and iii) Computer simulation data.
16. The method according to any one of claims 1 to 15, wherein the input parameter (IP) comprises at least one of the following: i) experimental data, ii) computational data, or iii) computer simulation data, wherein the experimental data i) is selected from apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), second virial coefficient (A2), and combinations thereof; the computational data ii) is selected from the isoelectric point (pI) of the protein, variable fragment (Fv)-charge (Fv-charge), and combinations thereof; and the computer simulation data iii) is selected from hydrophobic and charged patch sizes, in particular positive Fv total score, positive Fv total size, positive Fv total count, negative Fv total score, negative Fv total size, negative Fv total count, hydrophobic Fv total score, hydrophobic Fv total size, and hydrophobic Fv total count.
17. The method according to any one of claims 1 to 14 or 16, wherein the input parameter (IP) is selected from the ii) computational data and the iii) computer simulation data.
18. A method for determining the concentration-dependent viscosity of a protein solution, the method comprising: predicting the concentration-dependent viscosity of the protein solution by using a computer-implemented neural network (10); and comparing the concentration-dependent viscosity with a target viscosity.
19. The method according to claim 18, wherein the step of comparing the concentration-dependent viscosity with the target viscosity is performed to determine the suitability of the protein solution as a pharmaceutical product.
20. A method for providing a pharmaceutical product comprising a protein solution, the method comprising: predicting the concentration-dependent viscosity of the protein solution by using a computer-implemented neural network (10) (step S2); comparing the concentration-dependent viscosity with a target viscosity to determine the suitability of the protein solution as a pharmaceutical product (step S3); and optionally preparing the pharmaceutical product if the suitability of the protein solution is determined (step S5).
21. A device for predicting the concentration-dependent viscosity of an experimental protein solution, the device comprising: an experimental protein solution; a computer component comprising a neural network configured to predict the viscosity of the experimental protein solution: wherein the neural network is trained using a plurality of predetermined data sets; wherein each of the plurality of predetermined data sets comprises (i) a plurality of input parameters indicative of at least one physical property of a predefined protein solution, and (ii) at least one output parameter indicative of the concentration-dependent viscosity of the predefined protein solution; and wherein the plurality of input parameters comprises at least one of experimentally derived data, computationally derived data, and computer-simulated data.
22. A system for determining the concentration-dependent viscosity of a protein solution, the system comprising: a protein solution; and A neural network configured to determine the viscosity of the protein solution, wherein determining the viscosity of the protein solution depends on a plurality of input parameters (IP) associated with the protein solution, and further wherein the input parameters (IP) include at least one selected from the group consisting of: (i) experimental data; (ii) computational data; and (iii) computer simulation data.