A method for predicting the binding rate of a drug protein
By combining two-dimensional near-infrared spectroscopy and machine learning, the problem of insufficient accuracy and resolution in drug-protein binding rate determination methods has been solved, achieving high-precision prediction of drug-protein binding rate and meeting the needs of drug application and generic drug consistency evaluation.
Patent Information
- Application Number
- CN202511174452.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-08-21
AI Technical Summary
Existing methods for determining drug protein binding rates cannot meet the requirements for drug quality evaluation in terms of accuracy and resolution, especially for the determination of the same drug product, and model prediction methods cannot predict the binding rate of salt-type proteins.
Two-dimensional near-infrared correlation spectroscopy was used to characterize the interaction between drugs and proteins in a mixed solution. A population search algorithm was used to find the relationship between cross peaks and protein binding rate. A predictive model was established by combining machine learning methods and key influencing factors such as protein concentration, drug salt formation rate and solvent pH were introduced for correction.
It achieves high-precision prediction of drug-protein binding rate in a simulated physiological environment, can distinguish the protein structure of different drugs, solves the needs of in vitro drug evaluation and generic drug consistency evaluation, and improves the efficiency and accuracy of the assay.
Smart Images

Figure CN120741402B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of detection, and particularly relates to a drug protein binding rate prediction method. BACKGROUND
[0002] The drug protein binding rate is closely related to the clinical curative effect, is an important index for the consistency evaluation of generic drugs, and the binding degree influences the absorption, distribution, metabolism and excretion of free drugs in the body, further influences the change of corresponding kinetic parameters, and finally influences the drug efficacy. Therefore, the determination of the drug protein binding rate has important significance in drug research and development and actual clinical application.
[0003] At present, the determination methods of the drug protein binding rate mainly include in-vitro determination methods (equilibrium dialysis method, ultrafiltration method, high-performance affinity chromatography method) and model prediction methods (such as GastroPlus). In the in-vitro determination methods, the equilibrium dialysis method can maximize the reaction of the in-vivo condition, but the experimental process is too long. The ultrafiltration method has disadvantages such as Gibbs-Donna effect, non-specific dialysis equipment surface drug adsorption effect and protein leakage. The ultra-speed centrifugation method uses expensive instruments, and the sinking, reverse diffusion and other physical phenomena influence the accuracy of the determination. The common high-performance affinity chromatography in the chromatography method may change the conformation of the protein and the original binding behavior, and further interfere with the detection of the drug.
[0004] Although the above methods can realize the determination of the drug protein rate to a certain extent, in actual work, the accuracy of these methods can meet the determination of the drug protein rate of different varieties of drugs; and in the drug quality evaluation, the drug protein binding rates of different products of the same variety usually need to be determined, and the results can be used for the consistency evaluation of generic drugs, the exploration of key process control points, the individual difference research of the drug protein binding rate and the like. At this time, the in-vitro determination method cannot meet these needs in terms of the resolving power, repeatability and especially the resolving power of different products. In addition to the resolving power, the model prediction method cannot predict the protein binding rate of the salt type structure at present, and cannot completely replace the in-vitro determination method. SUMMARY
[0005] To solve the above technical problems, the application provides a drug protein binding rate prediction method.
[0006] The technical scheme adopted by the application is that a drug protein binding rate prediction method is provided, the drug and the protein are mixed, the influence of the drug protein binding on the protein structure change in the mixed solution is characterized by two-dimensional near-infrared correlation spectroscopy, a prediction model is established by using the characteristic spectrum area related to the drug protein binding, and the drug protein binding rate is predicted.
[0007] Preferably, the mixed solution is characterized by two-dimensional near-infrared correlation spectroscopy under the condition of temperature disturbance.
[0008] Preferably, the population search algorithm is used to find the relationship between the cross peaks in the synchronous spectrum and the asynchronous spectrum and the protein binding rate in the two-dimensional near-infrared spectrum.
[0009] Preferably, a prediction model based on machine learning is established by taking the characteristic signal of the obtained drug-target protein interaction as a variable and taking the theoretical drug protein binding rate value as a reference value.
[0010] Preferably, the machine learning is partial least squares regression (PLSR), support vector machine regression (SVMR), or deep neural network (DL).
[0011] Preferably, the specific steps are as follows:
[0012] Step 1: Characterize the target protein that can specifically bind to the drug to be analyzed, set temperature perturbation conditions, and detect near-infrared spectra under different temperature conditions; use two-dimensional correlation analysis to separate the structural information of the target protein, obtain the synchronous spectrum and the asynchronous spectrum;
[0013] Step 2: Mix the drug to be analyzed with the target protein, detect the near-infrared spectrum under temperature perturbation conditions, separate the structural information of the mixture by two-dimensional correlation analysis, compare the characteristic changes before and after the drug binds to the target protein, and screen the characteristics that change due to drug binding; use the population search algorithm to analyze the relationship between the cross peaks in the synchronous spectrum and the asynchronous spectrum and the protein binding rate, and obtain the characteristic signal of the drug-target protein interaction;
[0014] Step 3: Take the characteristic signal of the drug-target protein interaction as a variable, and take the theoretical drug protein binding rate value as a reference value to establish a prediction model based on machine learning.
[0015] Preferably, the drug is a class of drugs or a specific drug, and the protein is a protein that specifically binds to the drug.
[0016] Preferably, when the drug is a specific drug, constructing a high-precision prediction model further includes Step 4, which is as follows:
[0017] Step 4: On the basis of the quantitative model established in Step 3, introduce key influencing factors that affect the protein binding rate of the specific drug, modify the model, and establish a high-precision prediction model; wherein the key influencing factors are one or more of protein concentration, drug salt formation rate, and solvent system pH.
[0018] Application of the drug protein binding rate prediction method in drug evaluation.
[0019] Application of the drug protein binding rate prediction method in generic drug consistency evaluation.
[0020] The application has the advantages and positive effects that: the drug and protein binding process is characterized by two-dimensional correlation spectroscopy, two-dimensional correlation spectroscopy and machine learning are organically combined, a prediction model of drug protein binding rate in a simulated physiological environment is established from three aspects of structure characterization signal extraction, separation and quantitative analysis, further, the prediction accuracy of the model can be greatly improved by introducing key influencing factors, and an effective technical means is provided for drug protein interaction and related research;
[0021] In actual operation, the sample processing method is simple, time-consuming and low-cost, suitable for various related scientific research activities, not only can realize the protein binding rate prediction of one kind of drug (including acid / base type, salt type), but also can distinguish the protein binding rate difference of different products of the same drug, which has important significance for drug in vitro evaluation, generic drug consistency evaluation, pharmacodynamics research and the like. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 Fig. 1 is a two-dimensional correlation spectrum of human blood albumin under temperature disturbance; (a) is a synchronous spectrum of HSA solution, and (b) is an asynchronous spectrum of HSA solution;
[0023] Figure 2 Fig. 2 is a near-infrared two-dimensional correlation spectrum analysis synchronous spectrum; (a) is an HSA solution, (b) is a cefradine-HSA solution, (c) is a cefoperazone, (d) is a cefazolin sodium-HSA solution, (e) is a ceftizoxime-HSA solution, and (f) is a ceftriaxone sodium-HSA solution;
[0024] Figure 3 Fig. 3 is a two-dimensional correlation spectrum of five kinds of cephalosporin drugs mixed with HSA solution in a synchronous spectrum slice spectrum 4560-4640 cm -1 Waveband response radar chart;
[0025] Figure 4 Fig. 4 is a protein binding rate prediction model based on a slice spectrum at 4600 cm -1 in a synchronous spectrum;
[0026] Figure 5 Fig. 5 is a range and linearity investigation of a cephalosporin drug protein binding rate prediction model;
[0027] Figure 6: Ceftriaxone sodium-HSA solution near infrared spectroscopy two-dimensional correlation analysis synchronous spectrum; (a): protein binding rate 86.25%; (b): protein binding rate 87.54%; (c): protein binding rate 88.82%; (d): protein binding rate 90.11%; (e): protein binding rate 91.40%; (f): protein binding rate 92.69%; (g): protein binding rate 93.98%; (h): protein binding rate 95.26%; (i): protein binding rate 96.55%;
[0028] Figure 7 : 9 groups of ceftriaxone sodium-HAS two-dimensional correlation synchronous spectrum 4500 cm -1 Spectrum at 9 groups of ceftriaxone sodium-HAS two-dimensional correlation synchronous spectrum 4500 cm
[0029] Figure 8 : Ceftriaxone sodium salt rate-protein binding rate prediction model;
[0030] Figure 9 : Ceftriaxone sodium protein binding rate external validation results;
[0031] Figure 10 : 4 manufacturers of ceftriaxone sodium product drug protein rate distribution results;
[0032] Figure 11 : Ceftriaxone sodium generic drug population mass distribution chart. DETAILED DESCRIPTION
[0033] The embodiments of the present application will be described below with reference to the accompanying drawings.
[0034] The present application relates to a drug protein binding rate prediction method. First, the influence of drug protein binding on protein structure change is characterized by two-dimensional near-infrared correlation spectroscopy, and then a quantitative model is established to predict the drug protein binding rate by using the characteristic spectrum area related to drug protein binding.
[0035] For the mixed solution of the drug to be analyzed and the target protein, the structural changes caused by the binding of the drug and the protein are detected. By capturing the kinetic process of the mixed solution under temperature disturbance, the near-infrared spectrum is expanded from one dimension to two dimensions, the resolution and the specificity information extraction efficiency are improved, and the characterization and analysis of the protein characteristic structure are realized; the structural changes caused by the binding of the drug and the protein are used to obtain the characteristic spectrum area related to the binding rate value; further, a prediction model can be established by using the characteristic spectrum information, so as to predict the binding rate of the drug and the protein. The specific steps are as follows:
[0036] Step one: obtaining the characteristic spectrum of the target protein; characterizing the target protein capable of specific binding with the drug to be analyzed; preparing a target protein solution, setting temperature perturbation conditions, detecting near-infrared spectrum under different temperature conditions; performing spectrum quality inspection by regulating the spectrum intensity, signal-to-noise ratio, maximum and minimum noise in a certain spectrum range to ensure the accuracy and reliability of the spectrum for identification and structure characterization.
[0037] In some embodiments of the present application, the denaturation temperature of albumin is generally about 65℃, and the temperature perturbation condition can be set to a temperature range of 57-61℃ with a variation gradient of 1℃. The temperature variation process includes a warming process and a cooling process, and each temperature point is balanced for 5 minutes. The near-infrared spectrum determination conditions are as follows: wavelength range 10000-4000cm -1 , transmission sampling mode, scanning number 32 (background / sample), and resolution 4cm -1 . The obtained original spectrum is pretreated, and the original spectrum is denoised and baseline corrected by second-order 15-point SG filtering.
[0038] The target protein structure information is separated by two-dimensional correlation, and the calculation method and formula are as follows:
[0039] Analysis of spectrum change under the action of external perturbation variable , wherein t represents the external perturbation variable temperature (℃), is the wave number (cm -1 ). When t changes between and , the measured spectrum intensity y, i.e. the original spectrum set, is defined as :
[0040] (formula 1);
[0041] , wherein is the average spectrum:
[0042] (formula 2);
[0043] Two-dimensional correlation synchronous spectrum intensity is the vector product of the dynamic spectrum intensity y at different wave numbers :
[0044] (formula 3);
[0045] Two-dimensional correlation asynchronous spectrum intensity is the Hilbert-Noda matrix ( ) vector product of the dynamic spectrum intensity y at different wave numbers :
[0046] Equation 4
[0047] wherein, As shown in Equation 5:
[0048] Equation 5
[0049] The characteristics of the target protein are obtained by analyzing the dynamic process of the target protein structure changing with temperature using the synchronous spectrum analysis in two-dimensional correlation. The orthogonal cross peak of the synchronous spectrum represents the direction of the dynamic change of the albumin, and the asynchronous spectrum represents the rate of the dynamic change of the albumin with temperature.
[0050] Step 2: Characterization of the interaction between the drug and the target protein. The target protein is used as a probe, and the target protein solution simulating the physiological environment is mixed with the target drug at a clinical concentration (range) to collect near-infrared spectra with temperature as the disturbance, and the structural change information of the target drug and the target protein when they interact is characterized.
[0051] Specifically, the mixed solution of the target drug and the target protein is prepared, and the near-infrared spectrum is collected with temperature as the disturbance. The temperature disturbance condition can be selected in the range of 40-65 ℃, and the gradient can be 5 ℃. The synchronous spectrum and the asynchronous spectrum are obtained by two-dimensional correlation analysis according to the method of step 1; the characteristic signals of the target protein after binding with the drug are compared, and the characteristic signals changed due to the binding of the drug are selected to perform quantitative analysis of the drug-protein binding rate.
[0052] The two-dimensional spectral characteristics include the cross peaks in the synchronous spectrum and the asynchronous spectrum. In order to find the relationship between the cross peak and the protein binding rate, an improved population search algorithm is used, and the specific steps are as follows:
[0053] 1. Establish the mapping relationship of the two-dimensional spectral characteristic peaks and the protein molecule bond A random number sequence is generated using Tent chaotic mapping, which is easy to maintain the diversity of the population. The Tent chaotic mapping is shown in the following formula:
[0054] Equation 6
[0055] 2. Define the spectral characteristic peak matrix as a producer, the spectral peak position is updated, and the producer is explored in multiple positions. The optimal position is selected from multiple positions as the updated position of the producer. The updated position is determined as the structural information in the protein structure matrix. The update formula is:
[0056] Equation 7
[0057] wherein, the number of searching positions for the ith producer, the position of the ith producer, the fitness value of the ith producer, m is a constant for adjusting the number of searching, is the fitness value of the worst individual in the current population, the radius of the search range of the ith producer, the maximum search radius, the fitness value of the best individual in the current population, is a minimum constant to avoid the denominator being zero.
[0058] In the explosion process, for the ith producer, a total of explosion positions are generated, and the calculation of each position is shown in equation (8):
[0059] Equation 8;
[0060] wherein, represents a position generated by one explosion, indicates the value of the ith producer in the jth dimension in the tth iteration, indicates a random number between -1 and 1, which is used together with to determine the size and direction of movement.
[0061] For the explosion positions generated by the explosion operator, M positions are randomly selected and dimensions are randomly selected for mutation operation, as shown in equation (9):
[0062] Equation 9;
[0063] wherein, represents a mutation position generated after the explosion position is subjected to Gaussian mutation, , represents a random number subject to Gaussian distribution with mean and variance both being 1.
[0064] Suppose that for the ith producer, the firework optimization strategy generates a total of explosion positions and Gaussian positions, together with one position obtained from equation (8), the position with the optimal fitness value is found among the positions as the updated position of the ith producer.
[0065] After the sparrow position is updated in each iteration, a certain number of sparrow populations are randomly selected for Levy flight search, and a probability threshold is set, and the updated sparrow position is shown in equation (10):
[0066] (Formula 10);
[0067] wherein, represents the current global optimal position, s is a Levy flight step length, and r is a random number in [0, 1] that conforms to a uniform distribution.
[0068] According to the updated sparrow position, the protein structure corresponding to the characteristic peak is determined.
[0069] Step three: Establishment of a drug-protein binding rate prediction model. Using the obtained characteristic signals (spectral segments) of the interaction between the drug and the target protein, one or more segments that have a quantitative relationship with the protein binding rate are selected as variables; the drug-protein binding rate value in the drug theory or the classical pharmacokinetic parameters is used as the reference value, and a quantitative model based on machine learning is established.
[0070] wherein, the machine learning includes but is not limited to partial least squares regression (PLSR), support vector machine regression (SVMR), deep neural network (DL), etc. The model is verified by external samples, and parameters such as linearity, accuracy, and precision are evaluated: the linearity of the model is evaluated by the determination coefficient R2 of the predicted value and the theoretical value, the accuracy of the model is evaluated by the root mean square error (RMSEP), and the precision of the model is evaluated by the relative standard deviation (RSD) of repeated measurements. R2 should be greater than 0.9, and the RSD value should be within 5%. RMSEP should be within 10.
[0071] The constructed model can be used to predict the protein binding rate of each drug in the drug.
[0072] The above method combines two-dimensional correlation spectroscopy and machine learning organically, and establishes a prediction model of drug-protein binding rate in a simulated physiological environment from three aspects of structural characterization signal extraction, separation, and quantitative analysis. Among them, the target drug can be a class of drugs, and "a class of drugs" refers to a class of drugs with the same drug-protein interaction mechanism in pharmacological classification, and the same dissolution system (such as aqueous solution, tris buffer) is used in the determination. For example, cephalosporins mainly bind to human serum albumin in vivo, and different cephalosporins have different protein binding rates, so a prediction model of cephalosporin protein binding rate can be established to predict the protein binding rates of different types of cephalosporins.
[0073] In some other embodiments of the present application, the target drug can also be a certain specific drug, and accordingly, a high-precision prediction model for the specific drug can be established. For example, a prediction model for the protein binding rate of a certain cephalosporin. The construction method of the high-precision prediction model for a certain specific drug further includes the following steps on the basis of steps one, two, and three described above:
[0074] Step four: establishment of a high-precision prediction model for the protein binding rate of a specific drug. To further meet the needs of evaluating the differences in the protein binding rate of different products of a specific drug, the present application introduces key influencing factors affecting the protein binding rate of the specific drug to modify the model and establish a high-precision prediction model, thereby achieving high-precision prediction of the protein binding rate of different products of the drug.
[0075] The key influencing factors can be the concentration of target protein, drug salification rate, solvent system pH, etc.
[0076] The model is verified by external samples, and parameters such as linearity, accuracy and precision are evaluated: the linearity of the model is evaluated by the determination coefficient R2 of the predicted value and the theoretical value, the accuracy of the model is evaluated by the root mean square error (RMSEP), and the precision of the model is evaluated by the relative standard deviation (RSD) of repeated measurements. R2 should be greater than 0.9, and the RSD value should be within 5%. RMSEP should be within 10.
[0077] The two-dimensional spectral feature slice spectrum is used to establish a quantitative basic model; the key influencing factors are introduced into the basic model to improve the prediction accuracy of the model, and a high-precision drug protein binding rate model meeting the determination requirements is obtained. This method has strong structure analysis capability, can not only determine the protein binding rate of different varieties of drugs and realize the prediction of the protein binding rate of one type of drug (including acid / alkali type and salt type), but also improve the signal resolution by using two-dimensional correlation method, improve the model accuracy and the resolution ability of different products, and distinguish the differences in the protein binding rate of different products of the same drug, thereby providing an effective technical means for generic drug consistency evaluation, key process control point exploration, drug protein binding rate individual difference research, etc.
[0078] The present application will be described below in conjunction with the accompanying drawings, wherein the experimental methods not specifically described in the operation steps are performed according to the corresponding product instructions. The instruments, reagents and consumables used in the examples can be purchased from commercial companies unless otherwise specified.
[0079] Example 1: Prediction of the binding rate of cephalosporin drugs to human serum albumin
[0080] 1.1 Acquisition of characteristic spectrum of human serum albumin
[0081] The near-infrared spectrum of human serum albumin solution under temperature disturbance is obtained. The denaturation temperature of albumin is usually around 65℃, and the temperature range is set to 57~61℃ with a change gradient of 1℃. The temperature change process includes the heating process and the cooling process, and each temperature point is balanced for 5 minutes. The near-infrared spectrum determination conditions are as follows: wavelength range: 10000-4000cm -1Transmission sampling mode, 32 scans (background / sample), 4 cm resolution -1 .
[0082] Near infrared spectra were collected under temperature perturbation condition with an optical path of 1 mm. The spectral quality was checked by the spectral intensity, signal-to-noise ratio, maximum and minimum noise in a specified spectral range to ensure the accuracy and reliability of the spectra for identification and structural characterization. The original spectra were denoised by second-order 15-point SG filtering and baseline correction.
[0083] Two-dimensional correlation was used to separate the structural information of the target protein. The function of spectral change under the action of external perturbation variable was analyzed , where t represents the external perturbation variable temperature (°C), is the wave number (cm -1 ). When t changes between and , the measured spectral intensity y, i.e. the original spectrum set, is defined as :
[0084] (Formula 1);
[0085] wherein is the average spectrum:
[0086] (Formula 2);
[0087] The two-dimensional correlation synchronous spectral intensity is the vector product of the dynamic spectral intensity y at different wave numbers :
[0088] (Formula 3);
[0089] The two-dimensional correlation asynchronous spectral intensity is the Hilbert-Noda matrix (H) vector product of the dynamic spectral intensity y at different wave numbers :
[0090] (Formula 4);
[0091] wherein is shown in Formula 5:
[0092] (Formula 5);
[0093] The characteristic signal of the target protein was obtained by analyzing the characteristics of the kinetic process of the structural change of the target protein with temperature change using the synchronous spectrum in two-dimensional correlation. The characteristic spectrum of human blood albumin was obtained by two-dimensional spectrum under temperature perturbation, as shown in Figure 1 .
[0094] 1.2 Obtaining characteristic spectrum of cephalosporins after binding with human serum albumin
[0095] Add effective blood drug concentration of cephalosporins, including ceftriaxone sodium, cefathiamidine, cefazolin sodium, cephapiram or cefradine, to human serum albumin solution in simulated physiological environment, and the theoretical protein binding rate of these cephalosporins is 90-10%.
[0096] Collect near-infrared spectrum under temperature disturbance, wherein the temperature interval is selected as 40-65 ℃, and the gradient is 5 ℃. According to the aforementioned analysis steps, analyze the two-dimensional correlation spectrum under temperature disturbance to obtain the synchronous spectrum.
[0097] The results are shown in Figure 2 The autocorrelation peaks of the mixed solution of cefradine and cephapiram with low protein binding rate and human serum albumin are uniformly distributed at 4604 cm -1 The differences between the five synchronous spectra in Figure 2 (b)-(f) are mainly reflected in the distribution of the correlation peaks near 4560-4640 cm -1 After adding cephalosporins, the autocorrelation peak of human serum albumin at 4604 cm -1 is also shifted to high wavenumber, which indicates that under the same temperature disturbance, the response of the secondary structure of human serum albumin solution and the secondary structure of the cephalosporin-human serum albumin combination to temperature increase is different, and the difference increases with the increase of the drug protein binding rate. The characteristic peak that changes corresponds to the -helix information in the secondary structure.
[0098] Analyze the slice spectrum 4560-4640 cm -1 of the synchronous spectrum of the two-dimensional correlation spectrum, and the radar chart of the slice spectrum 4560-4640 cm -1 of the five cephalosporin-HSA mixed solutions is shown in Figure 3 The spectrum region near 4600 cm -1 can most effectively characterize the amount of cephalosporin-human serum albumin combination.
[0099] 1.3 Establishment of quantitative model
[0100] Characterize the structural changes of the target drug and the target protein when they interact. The two-dimensional spectral characteristics include cross peaks in synchronous spectrum and asynchronous spectrum. In order to find the relationship between cross peaks and protein binding rate, an improved population search algorithm is used, as follows:
[0101] Establish the two-dimensional spectral characteristic peaks and the protein molecular bonds The mapping relationship of the Tent chaotic mapping is used to generate random number sequence, which is easy to maintain the diversity of the population. The Tent chaotic mapping is shown in the following formula:
[0102] (Formula 6);
[0103] Define the spectral feature peak matrix For the producer, the spectral peak position is updated, the producer is explored in multiple positions, and the optimal position is selected from multiple positions as the updated position of the producer. The updated position is determined as the structure information in the protein structure matrix. The update formula is as follows:
[0104] (Formula 7);
[0105] wherein, is the number of search positions of the i th producer, is the position of the i th producer, represents the fitness value of the i th producer, and m is a constant for adjusting the search number, is the fitness value of the worst individual in the current population, represents the search radius of the i th producer, represents the maximum search radius, represents the fitness value of the best individual in the current population, is a minimum constant to avoid zero denominator.
[0106] During the explosion process, the i th producer generates explosion positions, and the calculation of each position is shown in the following formula (8):
[0107] (Formula 8);
[0108] wherein, represents the position generated by one explosion, represents the value of the i th producer in the j th dimension in the t th iteration, represents a random number between-1 and 1, and is used together with to determine the distance size and direction of movement.
[0109] M positions are randomly selected from the explosion positions generated by the explosion operator, and a dimension is randomly selected for mutation operation, as shown in the following formula (9):
[0110] (Formula 9);
[0111] wherein, represent the explosion position after Gaussian variation, and the variation position is generated, represent a random number subject to Gaussian distribution with mean and variance of 1.
[0112] Suppose that for the ith producer, the firework optimization strategy generates explosion positions and Gaussian positions together with 1 position obtained by formula (8) in total, and the position with the optimal fitness value is found in the positions as the updated position of the ith producer.
[0113] After the sparrow position is updated in each iteration, a certain number of sparrow populations are randomly selected for Levy flight search, a probability threshold is set , and the updated sparrow position is shown in formula (10):
[0114] Formula 10
[0115] wherein, represents the current global optimal position, s is the Levy flight step length, and r is a random number conforming to uniform distribution in [0, 1].
[0116] According to the updated sparrow position, the protein structure corresponding to the characteristic peak is determined.
[0117] The 4600 cm -1 of the five cephalosporin drugs was used to establish a quantitative model based on partial least squares regression (PLSR), with the pharmacokinetic parameters of the peak concentration of cephalosporin drugs as the reference value. The known binding rate of the peak concentration of cephalosporin drugs is as follows: ceftriaxone sodium protein binding rate 90%, ceftizoxime protein binding rate 85%, cefazolin sodium protein binding rate 70%, cefoperazone protein binding rate 22%, and cefradine protein binding rate 10%.
[0118] 1.4 Model verification
[0119] The above model was verified, and the correction and verification results are shown in Figure 4 The R 2 of the predicted value and the true value of the training set is 0.95, the root mean square error RMSEC of prediction is 7.4, the R 2 of the predicted value and the true value in 5-fold cross-validation is 0.92, and the root mean square error RMSECV of cross-validation is 11.8.
[0120] The above five kinds of cephalosporin injections were used as external verification samples, and the results are shown in Figure 5 As shown, the linear fitting correlation coefficient r of the external validation set predicted value and the theoretical value is 0.97, and the linearity is good. The external validation RMSEP is 8.02, and the model accuracy is good. Further, the model durability is investigated, and the results are shown in Table 1, indicating that the RSD is within 0.42%, and thus the quantitative model has good durability.
[0121] Table 1 Durability investigation of the cefalosporin protein binding rate prediction model
[0122]
[0123] The above model can be used to predict the protein binding rate of unknown cefalosporin drugs, which can provide important reference for preclinical research of new drugs, pharmacodynamics, pharmacokinetics, and population pharmacokinetics.
[0124] Example 2: Accurate prediction of albumin binding rate of ceftriaxone sodium for injection
[0125] Clinical investigations show that the difference in clinical efficacy between the original drug and the generic drug of ceftriaxone sodium for injection mainly reflects the speed of taking effect, which is mainly related to the salting rate of the product. Therefore, the salting rate is a key factor affecting the protein binding rate of ceftriaxone sodium. In this embodiment, ceftriaxone sodium solution with a salting rate of 98%-102% of the theoretical value (92.70%) of ceftriaxone sodium solution (i.e. 90.85%-94.55%) is prepared, as shown in Table 2. Nine groups of ceftriaxone sodium-HSA solutions are prepared by using the ratio of drug and protein in the simulated human physiological environment, and the actual salting rate of each group of solution is detected, and the protein binding rate of ceftriaxone sodium is corrected according to the dose-effect conversion relationship.
[0126] Table 2 Experimental design of ceftriaxone sodium-HSA solutions with different salting rates
[0127]
[0128] 2.1 Model establishment
[0129] Each of the nine groups of ceftriaxone sodium-HSA solutions is detected, and the high expression profile area of the protein binding rate of ceftriaxone is analyzed; the near-infrared spectra of each group of solutions under temperature disturbance conditions are obtained, and the disturbance conditions and analysis steps are the same as those in Example 1, and the two-dimensional correlation analysis synchronous spectrum is as shown in Figure 6 The results show that the change of the self-correlation peak in the wave band of 4460-4540 cm -1 is obvious, and the self-correlation peak in the wave band near 4500 cm -1 shifts to high wave number while the peak intensity decreases.
[0130] The 4500 cm -1 in the synchronous spectrum of the near-infrared two-dimensional correlation analysis of the nine groups of samples Figure 6The slice spectrum of the 4460-4540 cm Figure 7 band, as shown in the figure, it can be seen that as the salt formation rate decreases, the protein binding rate increases, the 4460-4540 cm -1 band gradually decreases in intensity, therefore, the synchronous spectrum 4500 cm -1 in the slice spectrum is the high expression spectrum area of the protein binding rate of ceftriaxone sodium. -1
[0131] On the basis of the quantitative model established in Example 1, the model parameters are corrected. The high expression spectrum area of the protein binding rate of ceftriaxone sodium in the aforementioned 4460-4540 cm -1 band is taken as the calculation spectrum area, and the salt formation rate is introduced as an influencing factor, and the theoretical conversion protein binding rate of ceftriaxone at different salt formation rates is used to establish a high-precision prediction model for the protein binding rate of ceftriaxone sodium for injection.
[0132] 2.2 Model verification
[0133] The model constructed is verified, and the correction and verification results of the prediction model are shown in Table 1. Figure 8 The R 2 of the predicted value and the true value of the training set is 0.99, and the root mean square error RMSEC value is 0.38. The R 2 of the predicted value and the true value in 5-fold cross-validation is 0.97, and the cross-validation root mean square RMSECV value is 0.60. The performance of the model is significantly improved compared with the model in Example 1.
[0134] The external verification results are shown in Table 2. Figure 9 The results show that the linear fitting correlation coefficient r of the predicted value and the theoretical value of the external validation set is 0.99. The root mean square error RMSEP of the external validation is 0.72. The linearity and accuracy of the model are greatly improved compared with the model in Example 1.
[0135] The model robustness verification results are shown in Table 3. The results show that the RSD is within 0.14%, and the accuracy of the quantitative model is good. The model robustness is also improved compared with the model in Example 1.
[0136] Table 3. Accuracy verification results of the ceftriaxone sodium protein binding rate prediction model
[0137]
[0138] Example 3: Difference in protein binding rate between ceftriaxone sodium original drug and generic drug
[0139] The prediction model constructed in Example 2 is used to analyze the difference in protein binding rate between ceftriaxone sodium original drug and generic drug.
[0140] The samples of representative manufacturers with high, medium and low salt formation rate distribution were selected: 3 batches of M1 (original drug), 3 batches of M2 (generic drug), 3 batches of M3 (generic drug), and 3 batches of M4 (generic drug). The spectra were collected and two-dimensional correlation analysis was performed according to the method of Example 1, and the prediction model of Example 2 was input with the high expression slice spectrum for prediction, and the prediction results are shown in Table 4.
[0141] Table 4 Information of ceftriaxone sodium generic drugs and prediction results of protein binding rate
[0142]
[0143] The protein binding rate results of generic drugs and original drugs were analyzed by box plot, as shown in Figure 10 , according to the prediction results, it can be known that the protein binding rate prediction value of ceftriaxone sodium produced by the original drug manufacturer M1 is the smallest, indicating that the protein binding rate of the original ceftriaxone sodium is lower than that of the generic ceftriaxone sodium, and the low protein binding rate can make ceftriaxone sodium widely distribute in the free state in various tissues and body fluids, thereby releasing the drug effect more quickly, and in clinical application, it can improve the efficacy at a lower serum total drug concentration level in the human body. According to the protein binding rate prediction value results of ceftriaxone sodium produced by different manufacturers, the differences between batches of manufacturers can be further characterized, and it can be seen from the figure that the batch consistency of M1 manufacturer is the best, the batch consistency of M2 manufacturer is poor, and among the generic drugs, the protein binding rate of ceftriaxone sodium produced by M3 manufacturer is the lowest, and the protein binding rate of ceftriaxone sodium produced by M2 manufacturer is higher. The results are consistent with the results of the population quality distribution of domestic generic ceftriaxone drugs Figure 11 .
[0144] The experimental results show that the original drug M1 not only has high salt formation rate, low protein binding rate, but also has very small batch difference; although the generic drug M4 representing the domestic high-quality process has obvious advantages over other distributed area generic drugs, it is also seen that it has obvious gap with the original drug in terms of salt formation rate and batch difference, which also explains the main reason for the difference in clinical efficacy between generic drugs and original drugs. As for the selected individual generic drugs, the main problem of M3 is the salt formation process and batch difference, and the main problem of M2 may also exist in the impurity control in addition to the salt formation process. It can be seen that the method can well characterize the differences between the protein binding rates of the original drug and different generic drugs, which is consistent with the results of the difference in clinical efficacy, and has important significance for the consistency evaluation of generic drugs.
[0145] The embodiments of the present application are described in detail above, but the content described is only the preferred embodiments of the present application, and cannot be considered as limiting the scope of the implementation of the present application. Any equivalent changes and improvements made within the scope of the present application shall still belong to the patent coverage of the present application.
Claims
1. A method for predicting the binding rate of a drug protein, characterized by: The drug and the protein are mixed, the drug-protein binding effect on the protein structure change in the mixed solution is characterized by two-dimensional near-infrared correlation spectroscopy, a prediction model is established by using the characteristic spectrum area related to the drug-protein binding, and the drug-protein binding rate is predicted. The specific steps are as follows: Step one: the target protein capable of specifically binding with the drug to be analyzed is characterized, a temperature disturbance condition is set, near-infrared spectra are detected under different temperature conditions, the near-infrared spectra are extended from one dimension to two dimensions by capturing the kinetic process of the mixed solution under the temperature disturbance condition, the target protein structure information is separated by two-dimensional correlation analysis, and synchronous spectrum and asynchronous spectrum are obtained; Step two: the drug to be analyzed is mixed with the target protein, near-infrared spectra are detected under the temperature disturbance condition, the mixture structure information is separated by two-dimensional correlation analysis, the relationship between the cross peaks in the synchronous spectrum and the asynchronous spectrum and the protein binding rate is analyzed by using the population search algorithm, and the characteristic signal of the interaction between the drug and the target protein is obtained; Step three: the characteristic signal of the interaction between the drug and the target protein is used as a variable, and the theoretical drug-protein binding rate value is used as a reference value, a prediction model based on machine learning is established, and the model obtained can be used for the prediction of the drug-protein binding rate of each drug in the drug.
2. The drug protein binding rate prediction method according to claim 1, characterized by: The machine learning is partial least squares regression, support vector machine regression or deep neural network.
3. The method of claim 1, wherein: The drug is a class of drugs or a specific drug, the class of drugs refers to a class of drugs with the same drug-protein interaction mechanism in pharmacological classification, and the same dissolution system is used in the determination; the protein is a protein specifically bound with the drug.
4. The method of claim 3, wherein: When the drug is a specific drug, the construction of a high-precision prediction model further includes step four, which is as follows: Step four: on the basis of the quantitative model established in step 3, the key influencing factors affecting the protein binding rate of the specific drug are introduced, the model is modified, and a high-precision prediction model is established; wherein the key influencing factors are one or more of protein concentration, drug salting rate and solvent system pH.
5. The drug-protein binding rate prediction method according to any one of claims 1-3 is applied in drug evaluation.
6. The drug-protein binding rate prediction method according to claim 4 is applied in generic drug consistency evaluation.