Drug protein binding rate prediction method

By combining two-dimensional near-infrared spectroscopy and machine learning, a high-precision drug-protein binding rate prediction model was established, which solved the problems of insufficient accuracy and repeatability of the measurement methods in the existing technology and achieved efficient prediction and differentiation of drug-protein binding rates.

CN120741402AActive Publication Date: 2025-10-03NAT INST FOR FOOD & DRUG CONTROL
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202511174452.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-10-03
Estimated Expiration
2045-08-21

AI Technical Summary

Technical Problem

Existing methods for determining drug protein binding rates have deficiencies in accuracy and repeatability, especially in drug quality evaluation, where they cannot meet the determination requirements for different products of the same variety, and model prediction methods cannot predict the protein binding rates of salt structures.

Method used

Two-dimensional near-infrared correlation spectroscopy is used to characterize the binding process of drugs and proteins. A prediction model is established by combining population search algorithm and machine learning methods. Through spectral analysis under temperature perturbation conditions, the characteristic signals of drug-protein interaction are extracted to establish a high-precision drug-protein binding rate prediction model.

Benefits of technology

It achieves high-precision prediction of drug-protein binding rate in a simulated physiological environment and can distinguish the differences in binding rates of different products. It is suitable for drug evaluation, generic drug consistency evaluation and pharmacodynamic research, with simple operation and low cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120741402A_ABST
    Figure CN120741402A_ABST
Patent Text Reader

Abstract

The invention relates to a drug protein binding rate prediction method. The method comprises the following steps: firstly, characterizing the influence of the drug protein binding effect on the protein structure change by using a two-dimensional near-infrared correlation spectroscopy, and then establishing a quantitative model by using a characteristic spectrum region related to the drug protein binding effect to predict the drug protein binding rate. A two-dimensional correlation spectrum and machine learning are organically combined, and a prediction model for simulating the drug protein binding rate in the physiological environment is established from three aspects of structural characterization signal extraction, separation and quantitative analysis; in actual operation, the technology is simple in sample treatment method, less in consumed time, lower in cost and suitable for various related scientific research activities, not only can predict the protein binding rate of one type of drugs (including acid / alkali type and salt type), but also can distinguish the difference of the protein binding rates of different products of the same drug, and has a good application prospect. The method is of great significance to in-vitro evaluation of drugs, consistency evaluation of imitated drugs, pharmacodynamic research and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of detection technology, and in particular relates to a method for predicting drug protein binding rate. Background Art

[0002] Drug protein binding is closely related to clinical efficacy and is an important indicator for generic drug consistency evaluation. The degree of drug protein binding affects the absorption, distribution, metabolism, and excretion of free drugs in the body, thereby affecting the changes in corresponding kinetic parameters and ultimately affecting the drug's efficacy. Therefore, measuring drug protein binding is of great significance in drug development and actual clinical application.

[0003] Currently, the main methods for determining drug-protein binding rates are in vitro assays (equilibrium dialysis, ultrafiltration, and high-performance affinity chromatography) and model prediction methods (such as GastroPlus). Among in vitro assays, equilibrium dialysis can best reflect in vivo conditions, but the experimental process is time-consuming. Ultrafiltration has disadvantages such as the Gibbs-Donna effect, nonspecific drug adsorption on the dialysis device surface, and protein leakage. Ultracentrifugation uses expensive instruments, and physical phenomena such as sedimentation and back diffusion affect measurement accuracy. High-performance affinity chromatography, a common chromatographic method, may alter the protein's conformation and binding behavior during detection, thereby interfering with drug detection.

[0004] While the aforementioned methods can measure drug-protein binding to a certain extent, in practice, their accuracy is limited to measuring the protein binding of different drug varieties. However, in drug quality evaluation, it is often necessary to measure the protein binding of different products within the same product. The results can be used for generic drug consistency evaluation, critical process control point investigation, and individual differences in drug-protein binding. In these cases, in vitro assays cannot meet these requirements in terms of resolution, reproducibility, and, especially, the ability to distinguish between different products. Furthermore, model prediction methods, aside from their resolution, currently cannot predict the protein binding of salt structures and cannot fully replace in vitro assays. Summary of the Invention

[0005] In order to solve the above technical problems, the present invention provides a method for predicting drug protein binding rate.

[0006] The technical solution adopted by the present invention is: a method for predicting drug-protein binding rate, which mixes drugs and proteins, characterizes the effect of drug-protein binding on protein structure changes in the mixed solution through two-dimensional near-infrared correlation spectroscopy, and establishes a prediction model using characteristic spectral regions related to drug-protein binding to predict the drug-protein binding rate.

[0007] Preferably, the mixed solution is characterized by two-dimensional near-infrared correlation spectroscopy under temperature perturbation conditions.

[0008] Preferably, a population search algorithm is used to find the relationship between the cross peaks in the synchronous spectrum and the asynchronous spectrum in the two-dimensional near-infrared spectrum and the protein binding rate.

[0009] Preferably, a prediction model based on machine learning is established using the acquired characteristic signal of the interaction between the drug and the target protein as a variable and the drug-protein binding rate value in drug theory or classical pharmacokinetic parameters as a reference value.

[0010] Preferably, the machine learning is partial least squares regression (PLSR), support vector machine regression (SVMR) or deep neural network (DL).

[0011] Preferably, the specific steps are as follows:

[0012] Step 1: Characterize the target protein that specifically binds to the drug to be analyzed, set temperature perturbation conditions, and detect near-infrared spectra under different temperature conditions; use two-dimensional correlation analysis to separate the target protein structural information and obtain synchronous and asynchronous spectra;

[0013] Step 2: The drug to be analyzed is mixed with the target protein. The near-infrared spectrum is measured under temperature perturbation conditions. Two-dimensional correlation analysis is used to separate the structural information of the mixture. The characteristic changes of the target protein before and after drug binding are compared, and the characteristics that change due to drug binding are screened. A population search algorithm is used to analyze the relationship between the cross-peaks in the synchronous and asynchronous spectra and the protein binding rate to obtain the characteristic signals of the interaction between the drug and the target protein.

[0014] Step 3: Using the characteristic signal of the interaction between the drug and the target protein as the variable and the theoretical drug-protein binding rate as the reference value, a prediction model based on machine learning is established.

[0015] Preferably, the drug is a class of drugs or a specific drug, and the protein is a protein to which the drug specifically binds.

[0016] Preferably, when the drug is a specific drug, building a high-precision prediction model further includes step 4, which is as follows:

[0017] Step 4: Based on the quantitative model established in step 3, key influencing factors that affect the protein binding rate of a specific drug are introduced, the model is modified, and a high-precision prediction model is established; among which the key influencing factors are one or more of protein concentration, drug salt formation rate, and solvent system pH.

[0018] Application of drug protein binding rate prediction methods in drug evaluation.

[0019] Application of drug protein binding rate prediction method in consistency evaluation of generic drugs.

[0020] The advantages and positive effects of the present invention include: characterizing the drug-protein binding process through two-dimensional correlation spectroscopy, organically combining two-dimensional correlation spectroscopy with machine learning, and establishing a prediction model for drug-protein binding rate in a simulated physiological environment from the perspectives of structural characterization signal extraction, separation, and quantitative analysis. Furthermore, by introducing key influencing factors, the prediction accuracy of the model can be greatly improved, providing an effective technical means for drug-protein interaction and related research.

[0021] In actual operation, this technology has a simple sample processing method, takes less time, and has low cost, making it suitable for various related scientific research activities. It can not only predict the protein binding rate of a class of drugs (including acid / base and salt types), but also distinguish the differences in drug protein binding rates of different products of the same drug. It is of great significance for in vitro drug evaluation, generic drug consistency evaluation, and pharmacodynamic research. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 : Two-dimensional correlation spectrum of human serum albumin under temperature perturbation; (a): Synchronous spectrum of HSA solution (b): Asynchronous spectrum of HSA solution;

[0023] Figure 2 : Synchronous spectrum of near-infrared two-dimensional correlation spectroscopy analysis; (a): HSA solution (b): Cefradine-HSA solution (c): Cefpiramide (d): Cefazolin sodium-HSA solution (e): Cefathiamidine-HSA solution (f) Ceftriaxone sodium-HSA solution;

[0024] Figure 3 :Slice spectrum 4560-4640cm in two-dimensional correlation spectroscopy of mixed solutions of five cephalosporins and HSA -1 Response radar chart of the band;

[0025] Figure 4 :Based on two-dimensional correlation analysis of the 4600 cm -1 Protein binding rate prediction model based on slice profile;

[0026] Figure 5 :Investigation on the range and linearity of the prediction model of protein binding rate of cephalosporins;

[0027] Figure 6: Synchronous spectrum of near-infrared spectroscopy and two-dimensional correlation analysis of ceftriaxone sodium-HSA solution; (a): protein binding rate 86.25%; (b): protein binding rate 87.54%; (c): protein binding rate 88.82%; (d): protein binding rate 90.11%; (e): protein binding rate 91.40%; (f): protein binding rate 92.69%; (g): protein binding rate 93.98%; (h): protein binding rate 95.26%; (i): protein binding rate 96.55%;

[0028] Figure 7 : 9 groups of ceftriaxone sodium-HAS two-dimensional correlation synchronization spectra 4500 cm -1 Slice spectrum;

[0029] Figure 8 ; Ceftriaxone sodium salt formation rate-protein binding rate prediction model;

[0030] Figure 9 :External validation results of ceftriaxone sodium protein binding rate;

[0031] Figure 10 : Drug protein rate distribution results of ceftriaxone sodium products from 4 manufacturers;

[0032] Figure 11 :Group quality distribution of domestically produced generic ceftriaxone. DETAILED DESCRIPTION

[0033] The embodiments of the present invention are described below with reference to the accompanying drawings.

[0034] The present invention relates to a method for predicting drug-protein binding rate. First, two-dimensional near-infrared correlation spectroscopy is used to characterize the effect of drug-protein binding on protein structural changes, and then a quantitative model is established using characteristic spectral regions related to drug-protein binding to predict the drug-protein binding rate.

[0035] For a mixed solution of the drug to be analyzed and the target protein, the structural changes caused by the binding of the drug to the protein are detected. By capturing the kinetic process of the mixed solution under temperature perturbations, the near-infrared spectrum is expanded from one dimension to two dimensions, improving the resolution and efficiency of specific information extraction, and achieving the characterization and analysis of the protein's characteristic structure. The structural changes caused by the binding of the drug to the protein are used to obtain characteristic spectral regions related to the binding rate value. Furthermore, a prediction model can be established based on the characteristic spectral information to predict the binding rate of the drug and protein. The specific steps are as follows:

[0036] Step 1: Acquire the characteristic spectral region of the target protein; characterize the target protein that can specifically bind to the drug to be analyzed; prepare the target protein solution, set the temperature perturbation conditions, and detect the near-infrared spectrum under different temperature conditions; perform spectral quality inspection by specifying the spectral intensity, signal-to-noise ratio, maximum and minimum noise within a certain spectral range to ensure the accuracy and reliability of the spectrum for identification and structural characterization.

[0037] In certain embodiments of the present invention, the denaturation temperature of albumin is generally around 65°C. The temperature perturbation condition can be set to a temperature range of 57-61°C with a gradient of 1°C. The temperature change process includes a heating process and a cooling process, wherein each temperature point is balanced for 5 minutes. The near-infrared spectroscopy measurement conditions are a wavelength range of 10,000-4,000 cm -1 , transmission sampling mode, 32 scans (background / sample), resolution 4cm -1 The original spectrum obtained by detection was preprocessed, and the original spectrum was subjected to second-order 15-point SG filtering for denoising and baseline correction.

[0038] Use two-dimensional correlation to separate the target protein structure information. The calculation method and formula are as follows:

[0039] Function for analyzing spectral changes under the influence of external disturbance variables , where t represents the external disturbance variable temperature (℃), is the wave number (cm -1 ). When t and When the measured spectral intensity y, that is, the original spectrum set, is defined as :

[0040] (Formula 1);

[0041] in, is the average spectrum:

[0042] (Formula 2);

[0043] Two-dimensional correlation synchronization spectral intensity For different wave numbers The vector product of the dynamic spectral intensity y is:

[0044] (Formula 3);

[0045] Two-dimensional correlated asynchronous spectral intensity For different wave numbers The Hilbert-Noda matrix of the dynamic spectral intensity y is ( ) Vector product:

[0046] (Formula 4);

[0047] in, As shown in Formula 5:

[0048] (Formula 5);

[0049] The synchronous spectrum in the two-dimensional correlation method is used to analyze the characteristics of the target protein's structure dynamics as it changes with temperature, and obtain the target protein's characteristic signal. The positive cross peaks of the synchronous spectrum represent the direction of the albumin dynamics change, and the asynchronous spectrum represents the rate of albumin dynamics change with temperature.

[0050] Step 2: Characterization of the interaction between the drug and the target protein. Using the target protein as a probe, a target protein solution simulating a physiological environment is mixed with the target drug at a clinical concentration (range). Near-infrared spectra are collected using temperature as a perturbation to characterize the structural changes that occur when the test drug interacts with the target protein.

[0051] Specifically, a mixed solution of the drug to be tested and the target protein is prepared, and near-infrared spectra are collected using temperature as a perturbation. The temperature perturbation condition can be selected from 40-65°C, with a gradient of 5°C. Synchronous and asynchronous spectra are obtained using two-dimensional correlation analysis according to the method in step 1. The characteristic signals of the target protein after binding with the drug are compared, and the characteristic signals that change due to drug binding are screened. From these, signals related to the magnitude are selected for quantitative analysis of the drug-protein binding rate.

[0052] The two-dimensional spectral features include cross-peaks in synchronous and asynchronous spectra. To find the relationship between cross-peaks and protein binding rate, an improved population search algorithm was used. The specific steps are as follows:

[0053] 1. Establishing two-dimensional spectral characteristic peaks Bonds to protein molecules The mapping relationship between the two is used. The Tent chaotic mapping tool is used to generate a random number sequence, which is easy to maintain the diversity of the population. The Tent chaotic mapping is shown as follows:

[0054] (Formula 6);

[0055] 2. Define the spectral characteristic peak matrix The producer is the one in which the spectral peak position is updated, and the producer is explored in multiple positions. The optimal position is selected from multiple positions as the updated position of the producer, and the updated position is determined to be the structural information in the corresponding protein structure matrix. The update formula is:

[0056] (Formula 7);

[0057] in, The number of times the position is searched for the i-th producer, is the position of the i-th producer, Represents the fitness value of the i-th producer, m is a constant used to adjust the number of searches, is the fitness value of the individual with the worst fitness function in the current population, represents the radius of the search range of the i-th producer, represents the maximum search radius, Represents the fitness value of the individual with the best fitness value in the current population, is a very small constant that prevents the denominator from being zero.

[0058] During the explosion, for the i-th producer, a total of explosion positions, and the calculation of each position is shown in formula (8):

[0059] (Formula 8);

[0060] in, Represents the location where an explosion occurred. Represents the i-th producer The value of the jth dimension in the tth iteration, Represents a random number between -1 and 1, and Together they determine the distance, size, and direction of movement.

[0061] Generated by the explosion operator Randomly select M locations from the explosion locations and randomly select dimensions to perform mutation operations on them, as shown in formula (9):

[0062] (Formula 9);

[0063] in, Represents the mutation position generated by Gaussian mutation of the explosion position. , represents a random number that follows a Gaussian distribution with a mean and variance of 1.

[0064] Assume that for the i-th producer, the fireworks optimization strategy generates explosion positions and Gaussian positions. Together with the 1 position obtained by formula (8), the position with the best fitness value is found among the positions as the updated position of the i-th producer.

[0065] After the sparrow positions are updated in each iteration, a certain number of sparrow populations are randomly selected for Lévy flight search, and the probability threshold is set. , the updated sparrow position is shown in formula (10):

[0066] (Formula 10);

[0067] in, represents the current global optimal position, s is the Levy flight step, and r is a random number in [0,1] that conforms to the uniform distribution.

[0068] According to the updated sparrow position, the protein structure corresponding to the characteristic peak is determined.

[0069] Step 3: Establish a drug-protein binding rate prediction model. Using the characteristic signals (spectral segments) of the interaction between the drug and the target protein, select one or more segments that have a quantitative relationship with the protein binding rate as variables. Using the drug-protein binding rate values ​​in drug theory or classical pharmacokinetic parameters as reference values, establish a quantitative model based on machine learning.

[0070] Machine learning-based methods include, but are not limited to, partial least squares regression (PLSR), support vector machine regression (SVMR), and deep neural networks (DL). Models should be validated using external samples and evaluated using parameters such as linearity, accuracy, and precision. Model linearity is evaluated using the coefficient of determination (R²) between the predicted and theoretical values, model accuracy is evaluated using the root mean square error (RMSEP), and model precision is evaluated using the relative standard deviation (RSD) of repeated measurements. R² should be greater than 0.9, and RSD values ​​should be within 5%. RMSEP should be within 10.

[0071] The constructed model can be used to predict the protein binding rate of each drug in this class of drugs.

[0072] This method combines two-dimensional correlation spectroscopy with machine learning to develop a predictive model for drug-protein binding in a simulated physiological environment, focusing on structural characterization signal extraction, separation, and quantitative analysis. The target drug can be a Class I drug, which, according to pharmacological classification, shares a common mechanism of action with proteins and is measured using the same solubility system (e.g., aqueous solution, tris buffer). For example, cephalosporins primarily bind to human serum albumin in vivo, and different cephalosporins have varying protein binding rates. Therefore, a predictive model for cephalosporin protein binding can be developed to predict the protein binding rates of different cephalosporin classes.

[0073] In other embodiments of the present invention, the target drug is suspected to be a specific drug. Accordingly, a high-precision prediction model for the specific drug can be established. For example, a prediction model for the protein binding rate of a cephalosporin drug can be established. The method for constructing a high-precision prediction model for a specific drug, based on steps one, two, and three above, further includes the following steps:

[0074] Step 4: Establishment of a high-precision prediction model for the protein binding rate of a specific drug. To further address the differential evaluation of the protein binding rates of different products of a specific drug, the present invention, based on the quantitative model established in Step 3, introduces key factors affecting the protein binding rate of the specific drug, modifies the model, and establishes a high-precision prediction model to achieve high-precision prediction of the protein binding rates of different drug products.

[0075] Among them, the key influencing factors may be target protein concentration, drug salt formation rate, solvent system pH, etc.

[0076] Validate the model using external samples and evaluate parameters such as linearity, accuracy, and precision. Model linearity is evaluated using the coefficient of determination (R²) between the predicted and theoretical values. Model accuracy is evaluated using the root mean square error (RMSEP). Model precision is evaluated using the relative standard deviation (RSD) of replicate measurements. R² should be greater than 0.9, and RSD values ​​should be within 5%. RMSEP should be within 10.

[0077] A quantitative basic model is established using two-dimensional spectral characteristic slice spectra; key influencing factors are introduced into the basic model to improve the model's prediction accuracy, resulting in a high-precision drug-protein binding rate model that meets the measurement requirements. This method has strong structural analysis capabilities and can not only measure the protein binding rates of different drug varieties and predict the protein binding rates of a class of drugs (including acid / base and salt types), but also utilize two-dimensional correlation methods to improve signal resolution. By introducing key influencing factors, the model's accuracy and the ability to distinguish different products can be improved, and the differences in drug-protein binding rates of different products of the same drug can be distinguished. This provides an effective technical means for the consistency evaluation of generic drugs, the exploration of critical process control points, and the study of individual differences in drug-protein binding rates.

[0078] The present invention is described below with reference to the accompanying drawings. Experimental methods without specific operating steps are carried out in accordance with the corresponding product specifications. Unless otherwise specified, the instruments, reagents, and consumables used in the examples can be purchased from commercial companies.

[0079] Example 1: Prediction of the binding rate of cephalosporins to human serum albumin

[0080] 1.1 Acquisition of the characteristic spectrum of human serum albumin

[0081] Obtain near-infrared spectra of human albumin solutions under temperature-perturbed conditions. The denaturation temperature of albumin is typically around 65°C. The temperature range was set between 57 and 61°C, with a gradient of 1°C. The temperature change process included both a ramp-up and ramp-down phase, with each temperature point equilibrated for 5 minutes. Near-infrared spectroscopy conditions: Wavelength range: 10,000-4,000 cm -1, transmission sampling mode, 32 scans (background / sample), resolution 4cm -1 .

[0082] Near-infrared spectra were collected under various temperature perturbation conditions, using a 1mm optical pathlength. Spectral quality was checked by specifying spectral intensity, signal-to-noise ratio, and maximum and minimum noise within a specified spectral range to ensure the accuracy and reliability of the spectra for identification and structural characterization. Raw spectra were subjected to second-order 15-point SG filtering for denoising and baseline correction.

[0083] Use two-dimensional correlation to separate the target protein structure information. Analyze the function of spectral changes under the influence of external perturbations , where t represents the external disturbance variable temperature (℃), is the wave number (cm -1 ). When t and When the measured spectral intensity y, that is, the original spectrum set, is defined as :

[0084] (Formula 1);

[0085] in, is the average spectrum:

[0086] (Formula 2);

[0087] Two-dimensional correlation synchronization spectral intensity For different wave numbers The vector product of the dynamic spectral intensity y is:

[0088] (Formula 3);

[0089] Two-dimensional correlated asynchronous spectral intensity For different wave numbers The Hilbert-Noda matrix of the dynamic spectral intensity y is ( ) Vector product:

[0090] (Formula 4);

[0091] in, As shown in Formula 5:

[0092] (Formula 5);

[0093] The characteristics of the target protein structure dynamics process with temperature change are analyzed by using the synchronous spectrum in the two-dimensional correlation to obtain the characteristic signal of the target protein. The characteristic spectrum of human serum albumin is obtained by using the two-dimensional spectrum of temperature perturbation, such as Figure 1 shown.

[0094] 1.2 Obtaining characteristic spectra after cephalosporins bind to human serum albumin

[0095] Cephalosporins at effective blood concentrations, including ceftriaxone sodium, cefathiamidine, cefazolin sodium, cefpiramide or cefradine, are added to a human albumin solution simulating a physiological environment. The theoretical protein binding rates of these cephalosporins are between 90% and 10%.

[0096] The near-infrared spectrum was collected with temperature as the perturbation condition, wherein the temperature range was selected as 40-65°C with a gradient of 5°C. The two-dimensional correlation spectrum under temperature perturbation was analyzed according to the above analysis steps to obtain the synchronous spectrum.

[0097] The results are as follows Figure 2 As shown in the figure, the autocorrelation peak of the mixed solution of cefradine, cefpiramide and human serum albumin with low protein binding rate is at 4604 cm -1 The distribution is consistent, for Figure 2 The differences between the five synchronous spectra in (b)-(f) are mainly manifested in the 4560-4640 cm -1 The distribution of the relevant peaks near 4604 cm -1 The autocorrelation peak at 100 nm also shifts toward the high wavenumber direction. This change indicates that under the same temperature disturbance, the response of human albumin solution and the binding of cephalosporins to the secondary structure of human albumin with increasing temperature is different, and the difference increases with the increase of drug protein binding rate. The characteristic peak that changes corresponds to the secondary structure of -Spiral information.

[0098] The slice spectrum 4560-4640cm in the two-dimensional correlation spectrum synchronization spectrum -1 The wavelength bands were analyzed, and the slice spectrum of the mixed solution of five cephalosporins and HSA was 4560-4640cm -1 The radar image of the band is as follows Figure 3 As shown, 4600 cm -1 The nearby spectral region can most effectively characterize the magnitude of the binding interaction between cephalosporins and human serum albumin.

[0099] 1.3 Establishment of quantitative model

[0100] The structural changes when the target drug interacts with the target protein are characterized. The two-dimensional spectral characteristics include cross-peaks in the synchronous spectrum and asynchronous spectrum. In order to find the relationship between the cross-peaks and the protein binding law, an improved population search algorithm is used, as follows:

[0101] Establish two-dimensional spectral characteristic peaks Bonds to protein molecules The mapping relationship between the two is used. The Tent chaotic mapping tool is used to generate a random number sequence, which is easy to maintain the diversity of the population. The Tent chaotic mapping is shown as follows:

[0102] (Formula 6);

[0103] Define the spectral characteristic peak matrix The producer is the one in which the spectral peak position is updated, and the producer is explored in multiple positions. The optimal position is selected from multiple positions as the updated position of the producer, and the updated position is determined to be the structural information in the corresponding protein structure matrix. The update formula is:

[0104] (Formula 7);

[0105] in, The number of times the position is searched for the i-th producer, is the position of the i-th producer, Represents the fitness value of the i-th producer, m is a constant used to adjust the number of searches, is the fitness value of the individual with the worst fitness function in the current population, represents the radius of the search range of the i-th producer, represents the maximum search radius, Represents the fitness value of the individual with the best fitness value in the current population, is a very small constant that prevents the denominator from being zero.

[0106] During the explosion, for the i-th producer, a total of explosion positions, and the calculation of each position is shown in formula (8):

[0107] (Formula 8);

[0108] in, Represents the location where an explosion occurred. Represents the i-th producer The value of the jth dimension in the tth iteration, Represents a random number between -1 and 1, and Together they determine the distance, size, and direction of movement.

[0109] Generated by the explosion operator Randomly select M locations from the explosion locations and randomly select dimensions to perform mutation operations on them, as shown in formula (9):

[0110] (Formula 9);

[0111] in, Represents the mutated position generated by Gaussian mutation of the explosion position. , represents a random number that follows a Gaussian distribution with a mean and variance of 1.

[0112] Assume that for the i-th producer, the fireworks optimization strategy generates explosion positions and Gaussian positions. Together with the 1 position obtained by formula (8), the position with the best fitness value is found among the positions as the updated position of the i-th producer.

[0113] After the sparrow positions are updated in each iteration, a certain number of sparrow populations are randomly selected for Lévy flight search, and the probability threshold is set. , the updated sparrow position is shown in formula (10):

[0114] (Formula 10);

[0115] in, represents the current global optimal position, s is the Levy flight step, and r is a random number in [0,1] that conforms to the uniform distribution.

[0116] According to the updated sparrow position, the protein structure corresponding to the characteristic peak is determined.

[0117] Using 4600cm of five cephalosporins -1 A quantitative model based on partial least squares regression (PLSR) was established using the slice spectra at the peak concentration of cephalosporins as reference values ​​for the pharmacokinetic parameters of cephalosporins. The known protein binding rates of the cephalosporins at the peak concentrations were: ceftriaxone sodium (90%), cefathiamidine (85%), cefazolin sodium (70%), cefpiramide (22%), and cefradine (10%).

[0118] 1.4 Model Validation

[0119] The above model was verified, calibrated and verified. Figure 4 As shown. The R of the training set prediction value and the true value 2 is 0.95, the root mean square error of prediction RMSEC value is 7.4, and the R 2 The root mean square of the interaction validation is 0.92, and the RMSECV value is 11.8.

[0120] The above five cephalosporin injections were used as external validation samples. The results were as follows: Figure 5As shown in Table 1, the linear fit correlation coefficient r between the predicted and theoretical values ​​in the external validation set was 0.97, indicating good linearity. The external validation RMSEP was 8.02, indicating good model accuracy. Furthermore, the model robustness was examined. The results, shown in Table 1, indicate that the RSD was within 0.42%, indicating good robustness of the quantitative model.

[0121] Table 1 Durability assessment of the prediction model for protein binding of cephalosporins

[0122]

[0123] The above model can be used to predict cephalosporins with unknown drug protein binding rates, and can provide important reference for preclinical research of new drugs, as well as pharmacodynamics, pharmacokinetics, and population pharmacokinetics.

[0124] Example 2: Accurate prediction of albumin binding rate of ceftriaxone sodium injection

[0125] Clinical investigations have shown that the difference in clinical efficacy between the original and generic ceftriaxone sodium injections lies primarily in the speed of onset of action, which is primarily related to the product's salt formation rate. Therefore, the salt formation rate is a key factor influencing the protein binding rate of ceftriaxone sodium. In this example, ceftriaxone sodium solutions with salt formation rates ranging from 98% to 102% (i.e., 90.85% to 94.55%) of the theoretical value (92.70%) of ceftriaxone sodium were prepared, as shown in Table 2. Nine ceftriaxone sodium-HSA solutions were prepared using drug-to-protein ratios simulated in a human physiological environment. The actual salt formation rate of each solution was measured, and the protein binding rate of ceftriaxone sodium was corrected based on the dose-effect conversion relationship.

[0126] Table 2 Experimental design of ceftriaxone sodium HSA solution with different salt formation rates

[0127]

[0128] 2.1 Model Building

[0129] 9 groups of ceftriaxone sodium-HSA solutions were tested respectively to analyze the high expression spectrum area of ​​ceftriaxone protein binding rate; the near infrared spectrum of each group of solutions under temperature perturbation conditions was obtained. The perturbation conditions and analysis steps were the same as those in Example 1. The two-dimensional correlation analysis synchronization spectrum was as follows: Figure 6 The results show that 4460-4540 cm -1 The change of the autocorrelation peak in the band is obvious, 4500 cm -1 The autocorrelation peaks in nearby bands shift to higher wavenumbers while the peak intensities decrease.

[0130] Take 9 groups of samples and analyze the synchronous spectrum at 4500 cm in the near infrared two-dimensional correlation analysis. -1 Department ( Figure 6The slice spectrum of the black dashed line in the middle is as follows: Figure 7 As shown in the figure, it can be seen that as the salt formation rate decreases, the protein binding rate increases. -1 The band correlation strength gradually decreases, so the synchronous spectrum 4500 cm -1 4460-4540 cm in the slice spectrum -1 The band is the high expression spectrum area of ​​ceftriaxone sodium protein binding rate.

[0131] Based on the quantitative model established in Example 1, the model parameters were modified. -1 The high expression spectrum area of ​​protein binding rate of ceftriaxone sodium in the band was used as the calculation spectrum area. At the same time, the influencing factor of salt formation rate was introduced. The protein binding rate of ceftriaxone with different salt formation rates was theoretically converted, as shown in Table 2. A high-precision prediction model for protein binding rate of ceftriaxone sodium for injection was established.

[0132] 2.2 Model Validation

[0133] The constructed model was verified, and the prediction model correction and verification results were as follows: Figure 8 As shown. The R of the training set prediction value and the true value 2 The R value of the predicted value and the true value in the 5-fold cross validation is 0.99, and the root mean square error RMSEC value is 0.38. 2 The performance of the model is significantly improved compared with the model in Example 1.

[0134] External verification results such as Figure 9 As shown, the results show that the linear fit correlation coefficient r between the external validation set prediction value and the theoretical value is 0.99. The external validation root mean square error RMSEP is 0.72. The model linearity and accuracy are significantly improved compared with the model in Example 1.

[0135] The results of the model robustness validation are shown in Table 3. The results indicate that the RSD is within 0.14%, indicating that the quantitative model has good accuracy. The model robustness is also improved compared to the model in Example 1.

[0136] Table 3 Results of accuracy validation of the prediction model for protein binding rate of ceftriaxone sodium

[0137]

[0138] Example 3: Difference in protein binding between the original and generic ceftriaxone sodium drugs

[0139] The prediction model constructed in Example 2 was used to analyze the difference in protein binding rate between the original drug and the generic drug for ceftriaxone sodium injection.

[0140] Samples from representative manufacturers with a high, medium, and low salt formation rate distribution were selected: three batches of M1 (original drug), three batches of M2 (generic drug), three batches of M3 (generic drug), and three batches of M4 (generic drug). Spectra were collected and two-dimensional correlation analysis was performed according to the method in Example 1. The high-expression slice spectra were input into the prediction model of Example 2 for prediction. The prediction results are shown in Table 4.

[0141] Table 4 Ceftriaxone sodium generic drug information and protein binding rate prediction results

[0142]

[0143] Use box plots to analyze the protein binding results of generic drugs and original drugs, such as Figure 10 As shown in the figure, according to the prediction results, the protein binding rate prediction value of ceftriaxone sodium produced by the original drug manufacturer M1 is the smallest, indicating that the protein binding rate of the original ceftriaxone sodium is lower than that of the generic ceftriaxone sodium. The low protein binding rate can make ceftriaxone sodium widely distributed in various tissues and body fluids in a free state, thereby releasing the drug efficacy faster. In clinical applications, it can improve the efficacy at a lower total serum drug concentration level in the human body. According to the protein binding rate prediction value results of ceftriaxone sodium produced by different manufacturers, the differences between batches of manufacturers can be further characterized. It can be seen from the figure that the consistency between batches of manufacturer M1 is the best, and the consistency between batches of manufacturer M2 is poor. Among the generic drugs, the protein binding rate of ceftriaxone sodium produced by manufacturer M3 is the lowest, and the protein binding rate of ceftriaxone sodium produced by manufacturer M2 is higher. This result is consistent with the population mass distribution of domestic generic drugs of ceftriaxone ( Figure 11 ) and the results were consistent.

[0144] Experimental results showed that the original drug M1 not only had a high salt formation rate and low protein binding rate, but also had minimal batch-to-batch variability. While the generic drug M4, representing high-quality domestic technology, had significant advantages over generic drugs from other regions, it also showed significant differences in both salt formation rate and batch-to-batch variability compared to the original drug, explaining the primary reason for the difference in clinical efficacy between generic and original drugs. For the individual generic drugs selected, the main issues with M3 lay in the salt formation process and batch-to-batch variability, while the main problem with M2, in addition to the salt formation process, may also be impurity control issues. This shows that the developed method can effectively characterize the differences in protein binding rates between the original drug and different generic drugs, which is consistent with the results of clinical efficacy differences and is of great significance for the consistency evaluation of generic drugs.

[0145] The embodiments of the present invention are described in detail above, but the contents described are only preferred embodiments of the present invention and should not be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.

Claims

1. A method for predicting drug protein binding rate, characterized by: The drug and protein are mixed, and the effect of drug-protein binding in the mixed solution on protein structure changes is characterized by two-dimensional near-infrared correlation spectroscopy. A prediction model is established using the characteristic spectral regions related to drug-protein binding to predict the drug-protein binding rate.

2. The method for predicting drug protein binding rate according to claim 1, wherein: The mixed solution was characterized by two-dimensional near-infrared correlation spectroscopy under temperature perturbation conditions.

3. The method for predicting drug protein binding rate according to claim 2, wherein: A population search algorithm was used to find the relationship between the cross-peaks in the synchronous and asynchronous spectra and the protein binding rate in two-dimensional near-infrared spectroscopy.

4. The method for predicting drug protein binding rate according to claim 3, wherein: A prediction model based on machine learning is established using the characteristic signals of the interaction between the drug and the target protein as variables and the drug-protein binding rate values ​​in drug theory or classical pharmacokinetic parameters as reference values.

5. The method for predicting drug protein binding rate according to claim 4, wherein: Machine learning is partial least squares regression, support vector machine regression, or deep neural network.

6. The method for predicting drug protein binding rate according to any one of claims 1 to 5, characterized in that: The specific steps are as follows: Step 1: Characterize the target protein that specifically binds to the drug to be analyzed, set temperature perturbation conditions, and detect near-infrared spectra under different temperature conditions; use two-dimensional correlation analysis to separate the target protein structural information and obtain synchronous and asynchronous spectra; Step 2: The drug to be analyzed is mixed with the target protein, and the near-infrared spectrum is detected under temperature perturbation conditions. The structural information of the mixture is separated using two-dimensional correlation analysis. A population search algorithm is used to analyze the relationship between the cross-peaks in the synchronous and asynchronous spectra and the protein binding rate to obtain the characteristic signal of the interaction between the drug and the target protein. Step 3: Using the characteristic signal of the interaction between the drug and the target protein as the variable and the theoretical drug-protein binding rate as the reference value, a prediction model based on machine learning is established.

7. The method for predicting drug protein binding rate according to claim 6, wherein: The drug is a class of drugs or a specific drug, and the protein is the protein to which the drug specifically binds.

8. The method for predicting drug protein binding rate according to claim 7, wherein: When the drug is a specific drug, building a high-precision prediction model also includes step four, as follows: Step 4: Based on the quantitative model established in step 3, key influencing factors that affect the protein binding rate of a specific drug are introduced, the model is modified, and a high-precision prediction model is established; among which the key influencing factors are one or more of protein concentration, drug salt formation rate, and solvent system pH.

9. Use of the drug protein binding rate prediction method according to any one of claims 1 to 7 in drug evaluation.

10. Application of the drug protein binding rate prediction method according to claim 8 in the consistency evaluation of generic drugs.

Citation Information

Patent Citations

  • Method for drug storage data processing by correlation coefficients

    CN106979934A

  • Cassia twig active ingredient potential application evaluation and production process monitoring method

    CN108535213A

  • Method for predicting ligand-protein interaction based on quantum calculation

    CN114446383A

  • Method for predicting protein content in feed based on two-dimensional correlation spectrum

    CN115541531A

  • Establishment method and application method of database for rapidly detecting proteins in tissue proteins based on near infrared

    CN117589708A