Method, device and program product for calculating modification rate of PEG single-modified recombinant protein
Through Lys-C enzyme hydrolysis and high-performance liquid phase analysis, the modification rates of each site of recombinant protein and PEG single modified recombinant protein were calculated, which solved the problem that the modification rates of different sites could not be analyzed in the prior art, and achieved the accuracy of drug quality control.
Patent Information
- Application Number
- CN202411149298.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-21
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-08-21
AI Technical Summary
The prior art is difficult to analyze the PEG modification rate of recombinant proteins at different sites, and cannot perform quality control and detection.
The recombinant protein and PEG single-modified recombinant protein were hydrolyzed by Lys-C enzyme. The modification rate of each site was calculated through high-performance liquid phase analysis, and the modification rate of PEG modification site was calculated using analytical data and relative peak area conversion.
Quantitative analysis of the modification rate of non-site single PEG modified recombinant proteins at each site was achieved, and the accuracy of drug quality control was improved.
Smart Images

Figure CN118866084B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent medicine, and more specifically, to a method, device, medium and program product for calculating the modification rate of PEG modification sites. Background Art
[0002] Recombinant proteins are an important category of biopharmaceuticals and are widely used clinically, such as recombinant human erythropoietin (EPO), etc. Recombinant proteins have many advantages when used as drugs, such as having the same composition as human components and producing effects similar to those of inherent substances in the human body. However, they also have some disadvantages, such as a short half-life and certain immunogenicity. To solve these problems, some biopharmaceuticals have been redeveloped, and polyethylene glycol (PEG)ylation is one of the main directions.
[0003] The basic structure of PEG is H-(OCH2CH2)n-OH, and this compound has characteristics such as good water solubility, non-volatility, tastelessness, no charge, and weak immunogenicity.
[0004] Activated PEG can be covalently linked to the free amino groups on the surface of recombinant protein drugs. The ligation sites of different recombinant proteins have different efficiencies, that is, the modification rates of each site are different. As a drug, it is necessary to conduct quality control and detection on the modification rates of each site of the modified protein. In the prior art, other proteases are used to hydrolyze recombinant proteins and PEG-modified recombinant proteins, and the proportion of N-terminal peptide segment deletion is analyzed, which is relatively effective for recombinant proteins with site-directed modification at the N-terminus. The disadvantage is that the modification conditions of other sites cannot be analyzed. The prior art can perform fluorescence amine color development after quantifying recombinant proteins and PEG-modified recombinant proteins, and the overall PEG modification rate can be analyzed. The disadvantage is that the modification conditions of other sites cannot be analyzed. Summary of the Invention
[0005] To solve the problems of quality control and detection of the modification rates of each site of the modified protein, the present invention provides a method for calculating the modification rate of PEG modification sites. Lys-C enzyme is used to hydrolyze recombinant proteins and PEG-monomodified recombinant proteins, and high-performance liquid analysis is performed respectively. By integrating the peak areas of each structure-confirmed peak and converting the relative peak areas, the modification rates of each site of the PEG recombinant protein are analyzed and calculated.
[0006] The present application (in the first aspect) discloses a method for calculating the modification rate of PEG modification sites, including:
[0007] S101: Obtain the analysis data and modification sites of the enzyme-digested peptides of recombinant proteins and PEG-monomodified recombinant proteins, where the analysis data is the percentage data of n peptides obtained after enzyme-digesting the peptides at n said modification sites;
[0008] S102: Calculate the modification rate of each PEG modification site based on the analysis data of the recombinant protein and the analysis data of the PEG mono-modified recombinant protein. The calculation basis is: the reduction ratio of the percentage of the same peptide segment in the analysis data of the PEG mono-modified recombinant protein compared to the analysis data of the recombinant protein is the sum of the modification rates of the modification sites at both ends of the same peptide segment.
[0009] Further, the calculation is to search for the optimal solution in the solution space of size 101 n , and the optimal solution represents the modification rate of each PEG modification site;
[0010] Optionally, the percentage of the i-th peptide segment in the analysis data of the recombinant protein is expressed as K i , and the percentage of the corresponding peptide segment in the analysis data of the PEG mono-modified recombinant protein is expressed as Kp i , where i ∈ [1, n];
[0011] Optionally, the calculation first generates a sub-solution space that meets in the solution space, and iteratively searches for the optimal solution that minimizes the error in the sub-solution space. The error is based on the above calculation basis and is expressed as the sum of the squared residuals of all peptide segments. The residual of any peptide segment is the difference between the reduction ratio of the percentage of the same peptide segment in the analysis data of the PEG mono-modified recombinant protein compared to the analysis data of the recombinant protein and the sum of the modification rates of the modification sites at both ends of the same peptide segment.
[0012] Further, the formula representing the calculation basis is:
[0013]
[0014] Among them, x i represents the modification rate of the i-th site that PEG can modify, n represents the total number of sites that PEG can modify, K i represents the percentage data of the i-th peptide segment in the analysis data of the recombinant protein; Kp i represents the percentage data of the i-th peptide segment in the analysis data of the PEG mono-modified recombinant protein;
[0015] Optionally, the error is the error of the least squares method. The error of the least squares method is based on the above calculation basis and is expressed by the formula:
[0016]
[0017] Among them, MSE represents the error of the least squares method, x i represents the modification rate of the i-th site that PEG can modify, n represents the total number of sites that PEG can modify, K iIt represents the percentage data of the i-th peptide segment in the analysis data of the recombinant protein; Kp i It represents the percentage data of the i-th peptide segment in the analysis data of the PEG mono-modified recombinant protein; x n It represents the modification rate of the last site that PEG can modify, K n It represents the percentage data of the last peptide segment in the analysis data of the recombinant protein; Kp i It represents the percentage data of the last peptide segment in the analysis data of the PEG mono-modified recombinant protein;
[0018] Optionally, the error includes the error of the least squares method and the overall error of all modification sites;
[0019] Optionally, the error is expressed as:
[0020]
[0021] Among them, Loss represents the error of iterative search, MSE represents the error of the least squares method, and λ represents the coefficient of the overall error of all modification sites; x i It represents the modification rate of the i-th site that PEG can modify, n represents the total number of sites that PEG can modify, K i It represents the percentage data of the i-th peptide segment in the analysis data of the recombinant protein; Kp i It represents the percentage data of the i-th peptide segment in the analysis data of the PEG mono-modified recombinant protein; x n It represents the modification rate of the last site that PEG can modify, K n It represents the percentage data of the last peptide segment in the analysis data of the recombinant protein; Kp i It represents the percentage data of the last peptide segment in the analysis data of the PEG mono-modified recombinant protein.
[0022] Furthermore, the specific steps of generating a sub-solution space that meets in the solution space and iteratively searching in the sub-solution space to obtain the optimal solution that minimizes the least squares error include: First, use a nested loop of forward search to generate a solution that meets , calculate the error of the generated solution, if the iterative stop condition is not reached, then determine the site that contributes the most to the error in the solution and perform priority iterative update to make the updated solution quickly approach a basically reasonable solution; perform an accurate search in the neighborhood range of the basically reasonable solution to obtain the optimal solution; a solution is expressed as (x1, x2, x3,..., x i ,..., x n ), where i ∈ [1, n], and x i represents the modification rate at the i-th site, and the forward search is that the loop of x1 is in the outermost layer, and the loops of x2,..., x nThereby search the sub-solution space;
[0023] Further, the specific steps of the forward search include: forward search the sub-solution space, calculate the current error based on the current solution, and compare it with the saved minimum error. If the current error is less than the minimum error, update the global optimal solution to the current solution; calculate the site with the largest contribution to the error in the outer modification rate of the current solution. First, update the modification rate of the site with the largest contribution to obtain the updated solution, calculate the updated error using the updated solution, and compare it with the saved minimum error. After iterative cycling, a basically reasonable solution for the forward search is obtained. The outer modification rate includes (x1, x2, …, x k , …, x half ), represents the floor function;
[0024] Optionally, the step of obtaining the optimal solution by precisely searching the neighborhood range of the basically reasonable solution includes: generating a local solution space based on the basically reasonable solution, and performing a backward search on the local solution space to obtain the optimal solution. The backward search is a loop for x n in the outermost layer, and sequentially loop inward for x n-1 , …, x1; during the iterative process, if the current error calculated from the current solution is less than the minimum error, update the local optimal solution to the current solution; sequentially loop and iterate until the backward search is completed to obtain the local optimal solution as the optimal solution. Each solution of the optimal solution represents the modification rate of each site.
[0025] Further, when updating the current solution to obtain the updated solution, determine the solution x i that contributes the most to the mean square error, update x i to obtain the updated solution x′ i , and use the updated solution x′ i to replace the solution x i in the solution to obtain the updated solution;
[0026] Optionally, the process of updating x i is expressed as:
[0027] x′ i = x i + α
[0028] where α represents the update step size;
[0029] Optionally, the enzymatic digestion is performed using Lys-C protease for hydrolysis;
[0030] Optionally, the upper limit of the modification rate of the i-th site during the forward search is
[0031]
[0032] Among them, n i represents the traversal upper limit of the i-th site, the sum of modification rates is 1, and ε represents the extended search range of the sum of modification rates;
[0033] Optionally, the upper limit of the modification rate of the i-th site during the forward search is
[0034]
[0035] Among them, n i represents the traversal upper limit of the i-th site, the sum of modification rates is 1, and ε represents the extended search range of the sum of modification rates, represents the extended expected search window, and the extended expected search window is determined based on the protein purity requirement;
[0036] Optionally, in the innermost layer of the nested loop of the forward search, it is judged whether the sum of modification rates is 1;
[0037] Optionally, the sum of modification rates being 1 can be extended to the sum of modification rates within the range of 1±ε, where ε represents the extended search range of the sum of modification rates;
[0038] Optionally, the analytical data is determined by liquid phase analysis;
[0039] Optionally, the percentage data of the peptide segments is the peak area percentage based on the results of liquid phase analysis.
[0040] Furthermore, the recombinant protein is EPO, the PEG single-modified recombinant protein is PEG-EPO, n = 9, and the 9 modification sites are the N-terminal α-amino group and the lysine side chain amino groups at the 20th, 45th, 52nd, 97th, 116th, 140th, 152nd, and 154th positions. Using mn, m20, m45, m52, m97, m116, m140, m152, m154 to represent the modification rates of the 9 modification sites, the calculation basis is expressed as:
[0041]
[0042] Among them, K i represents the percentage data of the i-th peptide segment in the analytical data of EPO; Kp i represents the percentage data of the i-th peptide segment in the analytical data of PEG-EPO;
[0043] Optionally, the error of the least squares method is expressed as:
[0044]
[0045] Among them, K i represents the percentage data of the i-th peptide segment in the analytical data of EPO, Kpi It represents the percentage data of the i-th peptide segment in the parsed data of PEG-EPO, where i ∈ [1, 9], and mn, m20, m45, m52, m97, m116, m140, m152, m154 represent the modification rates of the 9 modification sites;
[0046] Optionally, the outermost loop of the forward search is the search for the value range of mn, and the inner loops sequentially search for the value ranges of m20, m45, m52, m97, m116, m140, m152, m154;
[0047] Optionally, the error includes the least squares error and the overall error of all modification sites, expressed as:
[0048]
[0049] Among them, Loss represents the error of iterative search, MSE represents the least squares error, and λ represents the coefficient of the overall error of all modification sites; mn, m20, m45, m52, m97, m116, m140, m152, m154 represent the modification rates of the 9 modification sites;
[0050] Optionally, the value range is [0, 100], and in the innermost loop of the forward search, it is checked whether the sum of the solutions is 1 to generate a solution that meets the requirements;
[0051] Optionally, the upper limit of the value range is:
[0052] where n i represents the traversal upper limit of the i-th site, the sum of the modification rates is 1, ε represents the extended search range of the sum of the modification rates, represents the extended expected search window, and the extended expected search window is determined based on the protein purity requirement;
[0053] Optionally, the outermost loop of the reverse search is the search for the value range of m154, and the inner loops sequentially search for the value ranges of m152, m140, m116, m97, m52, m45, m20, mn.
[0054] The second aspect of the present application discloses a system for calculating the modification rate of PEG modification sites, including:
[0055] An acquisition module 201: used to acquire the parsed data and modification sites of the digested peptide segments of the recombinant protein and the PEG single-modified recombinant protein, and the parsed data is the percentage data of n peptide segments obtained after digesting the peptide segments at n said modification sites;
[0056] Calculation module 202: It is used to calculate the modification rate of each PEG modification site based on the analysis data of the recombinant protein and the analysis data of the PEG single-modified recombinant protein. The calculation basis is: the reduction ratio of the percentage of the same peptide segment in the analysis data of the PEG single-modified recombinant protein compared to the analysis data of the recombinant protein is the sum of the modification rates of the modification sites at both ends of the same peptide segment.
[0057] The third aspect of the present application discloses a computer device, which includes: a memory and a processor; the memory is used to store program instructions; the processor is used to call the program instructions, and when the program instructions are executed, it is used to execute the steps of the above method.
[0058] The fourth aspect of the present application discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it realizes the steps of the above method.
[0059] The fifth aspect of the present application discloses a computer program product, including a computer program, and when the computer program is executed by a processor, it realizes the steps of the above method.
[0060] The present application has the following beneficial effects:
[0061] 1. The present application can realize the quantitative analysis of the modification rate of each site of the non-site-specific single PEG-modified recombinant protein to achieve drug quality control;
[0062] 2. The calculation process adopts a strategy combining fast approximation and precise search for solution search. Different from searching and judging one by one in sequence, by the different contributions of the solutions of each site in the generated solutions (or solution combinations) to the error, most non-optimal solution combinations are quickly skipped to achieve a fast approach to the optimal solution combination.
[0063] 3. For the fast approximation algorithm, that is, by judging the composition of the residuals, calculate the modification rate of the site that contributes the most to the residuals among each single modification site, and then preferentially adjust the solution of this site to make the solution combination quickly approach a reasonable solution combination. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0065] Figure 1 It is a schematic flowchart of the method provided in the first aspect of the embodiment of the present invention;
[0066] Figure 2It is a schematic diagram of a program product provided in the second aspect of the embodiments of the present invention;
[0067] Figure 3 It is a schematic diagram of a computer device provided in the embodiments of the present invention;
[0068] Figure 4 It is a schematic diagram of the architecture of an exemplary computing device provided in the embodiments of the present invention;
[0069] Figure 5 It is a schematic diagram of a storage medium provided in the embodiments of the present invention;
[0070] Figure 6 It is the identification result of each peptide segment in the peptide maps of EPO(A) and PEG-EPO(B) provided in the embodiments of the present invention;
[0071] Figure 7 It is a comparison of the analysis results of PEG-EPO from different sources provided in the embodiments of the present invention;
[0072] Figure 8 It is a schematic diagram of the digested peptide segments of EPO and PEG-EPO provided in the embodiments of the present invention. Detailed implementation manners
[0073] In order to enable those skilled in the art of the present technology to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0074] In some processes described in the specification, claims and the above-mentioned accompanying drawings of the present invention, a plurality of operations that appear in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as S101, S102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.
[0075] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0076] The relevant terms of the present invention are listed as follows:
[0077] Recombinant protein: A long-chain polymer formed by the condensation of amino acids, which forms a spatial structure after curling.
[0078] Lys-C enzyme: A protein that can cleave recombinant proteins after lysine residues and cannot recognize lysine modified by PEG.
[0079] Ultra-high performance liquid chromatograph: A device that can separate the fragments of recombinant proteins after enzymatic digestion.
[0080] PDA ultraviolet detector: When used in combination with an ultra-high performance liquid chromatograph, it can quantify the fragments of recombinant proteins.
[0081] Xevo G2-S mass spectrometer: A device for detecting the mass numbers of peptide segments.
[0082] Figure 1 It is a schematic flow chart of a method for evaluating and calculating the modification rate of PEG modification sites provided by an embodiment of the present invention. Specifically, the method includes the following steps:
[0083] S101: Obtain the analytical data and modification sites of the enzymatically digested peptides of the recombinant protein and the PEG single-modified recombinant protein. The analytical data is the percentage data of n peptides obtained after digesting the peptides at n modification sites.
[0084] In some embodiments, an ultra-high performance liquid chromatograph (ACQUITY, equivalent to other devices), a PDA ultraviolet detector, a bench-top high-speed centrifuge (Eppendorf, equivalent to other devices); a metal bath (GINGKO,, equivalent to other devices), and a water purification machine (Milli-Q, equivalent to other devices) are used to measure the liquid chromatography-mass spectrometry peptide map.
[0085] In some embodiments, the reagents used for digesting the recombinant protein and the PEG-recombinant protein include: acetonitrile (Fisher Scientific, equivalent to other reagents); trifluoroacetic acid (Alfa Aesar, equivalent to other reagents); Tris, guanidine hydrochloride (Amresco, equivalent to other reagents); analytical pure hydrochloric acid and ammonium bicarbonate (Sinopharm Chemical Reagent Co., Ltd., equivalent to other reagents); dithiothreitol (Promega, equivalent to other reagents); iodoacetamide (Sigma, equivalent to other reagents); Lys-C protease (Wako, equivalent to other reagents).
[0086] In some embodiments, a multi-functional microplate reader (Molecular Devices plus5, equivalent to other devices) is used to measure the liquid chromatography-mass spectrometry peptide map.
[0087] S102: Calculate the modification rate of each PEG modification site based on the analysis data of the recombinant protein and the analysis data of the PEG mono-modified recombinant protein. The calculation basis is: the reduction ratio of the percentage of the same peptide segment in the analysis data of the PEG mono-modified recombinant protein compared to the analysis data of the recombinant protein is the sum of the modification rates of the modification sites at both ends of the same peptide segment.
[0088] The percentage of the i-th peptide segment in the analysis data of the recombinant protein is denoted as K i , and the percentage of the corresponding peptide segment in the analysis data of the PEG mono-modified recombinant protein is denoted as Kp i , where i ∈ [1, n];
[0089] A solution is denoted as (x1, x2, x3, …, x i , …, x n ), where i ∈ [1, n], and x i represents the modification rate at the i-th site. Since the value range of x i is 0 to 100, that is, each x i has 101 possible values, and there are n modification sites in total. Therefore, the solution space contains 101 n solutions.
[0090] In some embodiments, it is necessary to calculate the modification rate of each site of the mono-modified polyethylene glycolylated human erythropoietin (PEG-EPO) to achieve further quality control.
[0091] Overall scheme:
[0092] First step: Digest the human erythropoietin (EPO) and PEG-EPO samples with Lys-C protease respectively, and perform liquid chromatography-mass spectrometry (LC-MS) to measure the LC-MS peptide map under a mobile phase system containing 0.1% trifluoroacetic acid; identify each peptide segment by mass spectrometry, and integrate and calculate the peak area percentage of each peptide segment in the peptide map collected at a wavelength of 215 nm; calculate the modification rate of each site using software based on the change in the peak area percentage of the corresponding peptide segments (i.e., the peptide segments with the same retention time on the high-performance liquid chromatography map) in the EPO and PEG-EPO peptide maps.
[0093] There are 9 possible sites for PEG modification on EPO, which are the N-terminal α-amino group and the side-chain amino groups of lysine at positions 20, 45, 52, 97, 116, 140, 152, and 154.
[0094] PEG modification of the side-chain amino group of Lys will cause Lys-C protease to be unable to hydrolyze the peptide bond at this site, resulting in a decrease in the relative peak areas of the Lys enzymatic hydrolysis peptides on both sides of the PEG-modified lysine. If PEG is modified at the N-terminal α-amino group, the proportion of the first peptide at the N-terminus will decrease. For the PEGylated peptide, the retention time shifts to after all peptides, but it is also included in the area calculation. The peaks of the PEGylated peptides are basically together, forming a mixed peak cluster.
[0095] If PEG is modified at the side-chain amino group of the 20th lysine, Lys-C protease cannot hydrolyze the peptide bond at this site, which will lead to a decrease in the relative peak areas of the two-sided peptides K1 and K2, and the decreasing proportion is exactly the proportion of Kp1;
[0096] After hydrolysis of EPO by Lys-C protease, there are a total of 9 peptides (one is relatively small: the K8 peptide is relatively small and cannot be detected during the actual detection of EPO). Starting from the N-terminus, they are respectively recorded as K1, K2, K3, K4, K5, K6, K7, K8, K9, and at the same time, the percentage of each peptide is represented;
[0097] After hydrolysis of PEG-EPO by Lys-C protease, these 9 peptides can also be obtained. To distinguish from the enzymatic digestion results of EPO, these 9 peptides are respectively recorded as Kp1, Kp2, Kp3, Kp4, Kp5, Kp6, Kp7, Kp8, Kp9 starting from the N-terminus (as Figure 8 shown), and at the same time, the percentage of each peptide is represented; According to the position of the modification site, each modification site is defined, the N-terminal modification site (mn), the side-chain modification site of the 20th residue (m20), and so on, m45, m52, m97, m116, m140, m152, m154, and at the same time, the modification rate percentage of each modification site is represented.
[0098] Step 2: According to the peptide determination results after enzymatic digestion of EPO and PEG-EPO samples (as Figure 6 shown), in the peptide determination results of EPO enzymatic digestion, K1 + K2 + K3 + K4 + K5 + K6 + K7 + K8 + K9 ≈ 1;
[0099] For the single-modified PEG-EPO, the site of the modified site cannot be cleaved by Lys enzyme, so there is:
[0100] Kp1 + Kp2 + Kp3 + Kp4 + Kp5 + Kp6 + Kp7 + Kp8 + Kp9 <
[0101] K1 + K2 + K3 + K4 + K5 + K6 + K7 + K8 + K9;
[0102] The reduction relationship between PEG-EPO and the corresponding peptide segments of EPO is such that mn + m20 + m45 + m52 + m97 + m116 + m140 + m152 + m154 ≈ 1;
[0103] (K1 - Kp1) / K1 ≈ mn + m20; (1.1)
[0104] (K2 - Kp2) / K2 ≈ m20 + m45; (1.2)
[0105] (K3 - Kp3) / K3 ≈ m45 + m52; (1.3)
[0106] (K4 - Kp4) / K4 ≈ m52 + m97; (1.4)
[0107] (K5 - Kp5) / K5 ≈ m97 + m116; (1.5)
[0108] (K6 - Kp6) / K6 ≈ m116 + m140; (1.6)
[0109] (K7 - Kp7) / K7 ≈ m140 + m152; (1.7)
[0110] (K8 - Kp8) / K8 ≈ m152 + m154; (1.8)
[0111] (K9 - Kp9) / K9 ≈ m154; (1.9)
[0112] Step 3: The calculation process of the modification rate:
[0113] Since the value ranges of mn, mn20, m45, m52, m97, m116, m140, m152, and m154 are all from 0 to 100, the size of the solution space is 101 9 。
[0114] First, generate solutions that meet K1 + K2 + K3 + K4 + K5 + K6 + K7 + K8 + K9 ≈ 1, and then iteratively search for the optimal solution through the least squares method.
[0115] Among them, the least squares method obtains the optimal solution combination by making the error of the least squares method approach "0", and the error is defined as:
[0116] Loss = MSE + λ
[0117] ((K1 - Kp1) / K1 - mn - m20) 2 + ((K2 - Kp2) / K2 - m20 - m45) 2 + ((K3 - Kp3) / K3 - m45 - m52) 2 + ((K4 - Kp4) / K4 - m52 - m97)2 +((K5 - Kp5) / K5 - m97 - m116) 2 +((K6 - Kp6) / K6 - m116 - m140) 2 +((K7 - Kp7) / K7 - m140 - m152) 2 +((K8 - Kp8) / K8 - m152 - m154) 2 +((K9 - Kp9) / K9 - m152) 2
[0118] In some embodiments, λ = 1 / 3.87.
[0119] Wherein, the regularization term after MSE is the square error between the sum on the left side and the sum on the right side of equations (1.1) to (1.9).
[0120] In view of the large amount of computation and the long time to obtain the optimal solution combination, the present invention proposes an improved optimization algorithm, which adopts a strategy combining fast approach and precise search to quickly obtain the optimal solution combination, specifically including:
[0121] A solution generator with 9 - fold loop nesting generates solution combinations whose sum is 1 (or within an acceptable range of experimental error, such as 0.99 - 1.01), performs operation evaluation and temporarily stores the records.
[0122] The evaluation of the solution combination runs in the above - mentioned loop nesting. For the deepest several - fold loop nesting among the solutions generated by all solution generators, early evaluation and adjustment (+0.01) are performed, and the outer - layer nesting is used for balancing to generate a new adjusted solution. Then, the errors of the original solution (the solution of the current iteration) and the new solution (the solution of the next iteration) are calculated respectively by the least - squares method, and the solution with a smaller error in the least - squares method is retained, that is, if the error after adjustment is smaller than the original error, this adjustment is executed; if the error becomes larger, this adjustment is discarded and the next round of solution generation is entered, and the evaluation and adjustment of the solution are always in progress.
[0123] In the evaluation process of the obtained solutions, it is possible to perform early adjustment on the sites where the error is reduced after adjustment of the deep - layer sites.
[0124] After one run of the deepest - layer loop nesting is completed, intensive solution search is carried out within the range of running the deepest - layer loop nesting completely 2 more times.
[0125] Based on the obtained solutions, reverse the order of the loop nesting, appropriately expand the search range, and perform it once again to find the solution most suitable for inputting data and given conditions.
[0126] In some embodiments, the least - squares method is used for iterative search of the optimal solution, and the end - point error is controlled within 2%. If there is no solution in the control range, the solution with the smallest error is selected as the optimal solution.
[0127] In some embodiments, the number n of modification sites of the recombinant protein is obtained; the percentage Ki of each of the n peptide segments obtained by digesting the recombinant protein with Lys-C is obtained i , where i ∈ [1, n]; the percentage Kpi of each of the n combined peptide segments obtained by digesting the PEGylated drug with the same enzyme is obtained i ; where Ki i and Kpi i are determined using an ultra-high performance liquid chromatograph
[0128] In the solution space, subspace solutions are generated by forward search to obtain solutions that sum to 1; a solution is represented as (x1, x2, x3, …, xi, …, xn), where i ∈ [1, n], and xi represents the modification rate at the ith site. The forward search is such that the loop over x1 is on the outermost layer, and the loops over x2, …, xn are nested inward in sequence i , …, xn n ), where i ∈ [1, n], and xi i represents the modification rate at the ith site. The forward search is such that the loop over x1 is on the outermost layer, and the loops over x2, …, xn are nested inward in sequence n so as to search for the subspace solutions;
[0129] An error is calculated based on the reduction ratio of each peptide segment after digesting the PEGylated recombinant protein with Lys-C and the modification rate at the corresponding site. Among the combinations of the subspace solutions, an iterative calculation is performed to find the combination of solutions that minimizes the error, and the combination of solutions of the optimal solution combination represents the modification rates at each site.
[0130] In some embodiments, (1) during forward search, the upper limit of traversal of the ith site is:
[0131]
[0132] where n i represents the upper limit of traversal of the ith site, the sum of the modification rates is 1, ε represents the extended search range of the sum of the modification rates, ε = 0.01, represents the extended expected search window, and a 4% extension is selected based on the actual requirement that the protein purity is not less than 95%; the sum of the modification rates at each site is 99% - 101%.
[0133] (2) The lower limit of traversal of the ith site is:
[0134]
[0135] n i represents the upper limit of traversal of the ith site,
[0136] m i represents the lower limit of traversal of the ith site
[0137]
[0138] (3) Pseudo-code for forward search
[0139] Table 1 Pseudo-code for forward search
[0140]
[0141]
[0142] Backward search
[0143] Table 2 Pseudo-code for backward search
[0144]
[0145]
[0146] In some embodiments, to improve accuracy and reduce bias, the errors of all sites are evaluated overall, and at the same time, they need to be adjusted to a level equivalent to the contributions of each site. Therefore, the error is adjusted as follows:
[0147] Loss = MSE + [(K1 - Kp1) / K1 + (K2 - Kp2) / K2 + (K3 - Kp3) / K3 + (K4 - Kp4) / K4 + (K5 - Kp5) / K5 + (K6 - Kp6) / K6 + (K7 - Kp7) / K7 + (K8 - Kp8) / K8 + (K9 - Kp9) / K9 - mn - 2(m20 + m45 + m52 + m97 + m116 + m140 + m152 + m154)] 2
[0148] Figure 7 The modification rate results of PEG-EPO from different sources obtained by using the method of the present invention are shown as follows.
[0149] Figure 3 It is a schematic diagram of a computer device provided by an embodiment of the present invention. As Figure 3 shown, the device may include: one or more processors, and one or more memories; wherein, computer-readable code is stored in the memory, and when the computer-readable code is run by the one or more processors, the method described above can be executed.
[0150] The processor in this embodiment may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, operations and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc., and may be of the X86 architecture or the ARM architecture.
[0151] Generally speaking, the various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, a microprocessor or other computing devices. When the aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts or using some other graphical representation, it will be understood that the blocks, devices, systems, technologies or methods described herein may be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.
[0152] For example, the method or device according to the embodiments of the present disclosure may also be implemented by means of Figure 4 the architecture of the computing device 3000 shown. As Figure 4 shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for the processing and / or communication of the method provided by the present disclosure and the program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 4 the architecture shown is only exemplary, and when implementing different devices, one or more components shown in the Figure 4 computing device may be omitted according to actual needs.
[0153] The embodiments of the present invention also provide a computer-readable storage medium, such as Figure 5As shown, it is a schematic diagram of a storage medium provided by an embodiment of the present invention. Computer-readable instructions 4010 are stored on the computer storage medium 4020. When the computer-readable instructions 4010 are run by a processor, the methods according to the embodiments of the present disclosure described with reference to the above figures can be executed. The computer-readable storage medium in the embodiments of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memories of the methods described herein are intended to include, but are not limited to, these and any other suitable types of memories. It should be noted that the memories of the methods described herein are intended to include, but are not limited to, these and any other suitable types of memories.
[0154] The embodiments of the present disclosure also provide a computer program product or a computer program. When the computer program is executed by a processor, the steps of the above method are implemented, as Figure 2 shown. The computer program product or the computer program includes:
[0155] An acquisition module 201: configured to acquire the analysis data and modification sites of the digested peptides of the recombinant protein and the PEG single-modified recombinant protein, where the analysis data is the percentage data of n peptides obtained after digesting the peptides at n modification sites;
[0156] A calculation module 202: configured to calculate the modification rate of each PEG modification site based on the analysis data of the recombinant protein and the analysis data of the PEG single-modified recombinant protein. The calculation basis is: the reduction ratio of the percentage of the same peptide segment in the analysis data of the PEG single-modified recombinant protein compared to the analysis data of the recombinant protein is the sum of the modification rates of the modification sites at both ends of the same peptide segment.
[0157] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0158] In general, the various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices. When aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.
[0159] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above may refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0160] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods may be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other may be indirect couplings or communication connections through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0161] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0162] In addition, in each embodiment of the present invention, each functional unit may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0163] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art should understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.
Claims
1. A method for calculating the modification rate of a single PEG-modified recombinant protein, characterized in that: The method comprises: S101: Obtaining analytical data and modification sites of enzyme-cleaved peptides of the recombinant protein and the recombinant protein modified with PEG, wherein the analytical data is percentage data of n peptides obtained after the peptides are cleaved at n modification sites; S102: Calculate the modification rate of each PEG modification site based on the analytical data of the recombinant protein and the analytical data of the PEG-modified recombinant protein. The calculation basis is: the reduction ratio of the same peptide segment percentage of the analytical data of the PEG-modified recombinant protein compared with the analytical data of the recombinant protein is the sum of the modification rates of the modification sites at both ends of the same peptide segment. The calculation steps include: first, generate a nested loop that meets the requirements of the forward search. The error of the generated solution is calculated. If the iteration stop condition is not met, the site that contributes the most to the error in the solution is determined to be updated iteratively first, so that the updated solution quickly approaches the basically reasonable solution; a local solution space is generated based on the basically reasonable solution, and a reverse search is performed on the local solution space to obtain the optimal solution. The reverse search is based on The cycle is in the outermost layer, and then circulates inwards ; During the iteration process, if the current error calculated by the current solution is less than the minimum error, the local optimal solution is updated as the current solution; the iteration is repeated until the reverse search is completed and the local optimal solution is obtained as the optimal solution. Each solution of the optimal solution represents the modification rate of each point. A solution is expressed as ,in, , Represents the modification rate at the i-th site, and the forward search is The cycle is in the outermost layer, and then circulates inwards Thus, the sub-solution space is searched, and the percentage of the i-th peptide in the analytical data of the recombinant protein is expressed as ,conform to The solutions of constitute the sub-solution space.
2. The method for calculating the modification rate of a PEG-modified recombinant protein according to claim 1, wherein: The calculation is to search for the optimal solution in the solution space, and each value of the optimal solution represents the modification rate of each PEG modification site.
3. The method for calculating the modification rate of a PEG-modified recombinant protein according to claim 1, wherein: The percentage of the corresponding peptide fragments in the analytical data of PEG-modified recombinant protein is expressed as ,in, .
4. The method for calculating the modification rate of a PEG-modified recombinant protein according to claim 3, wherein: The error is obtained based on the calculation basis, and the error is expressed as the sum of the squares of all peptide residuals, wherein the residual of any peptide is the difference between the reduction ratio of the same peptide percentage of the analytical data of the PEG-modified recombinant protein compared to the analytical data of the recombinant protein and the sum of the modification rates of the modification sites at both ends of the same peptide.
5. The method for calculating the modification rate of a PEG-modified recombinant protein according to claim 1, wherein: The calculation basis is expressed as follows: . The left side of the formula represents the percentage reduction of the same peptide segment in the analytical data of the PEG-modified recombinant protein compared to the analytical data of the recombinant protein, and the right side of the formula represents the sum of the modification rates of the modification sites at both ends of the same peptide segment. represents the modification rate of the i-th site that can be modified by PEG, n represents the total number of sites that can be modified by PEG, Represents the percentage data of the i-th peptide segment in the analysis data of the recombinant protein; It represents the percentage data of the i-th peptide segment in the analysis data of the PEG-modified recombinant protein; wherein, one end of the n-th peptide segment is the n-th modification site, and the other end is the C-terminus of the recombinant protein.
6. The method for calculating the modification rate of a PEG-modified recombinant protein according to claim 4, characterized in that: The error is a least squares error, which is obtained based on the calculation basis and is expressed as follows: Where MSE stands for least squares error, represents the modification rate of the i-th site that can be modified by PEG, n represents the total number of sites that can be modified by PEG, Represents the percentage data of the i-th peptide segment in the analysis data of the recombinant protein; It represents the percentage data of the i-th peptide segment in the analytical data of the PEG-modified recombinant protein; represents the modification rate of the last site that can be modified by PEG, It represents the percentage data of the last peptide in the analysis data of the recombinant protein; It represents the percentage data of the last peptide segment in the analysis data of the PEG-modified recombinant protein.
7. The method for calculating the modification rate of a PEG-modified recombinant protein according to claim 4, characterized in that: The error includes the least squares error and the overall error of all modified sites.
8. The method for calculating the modification rate of a PEG-modified recombinant protein according to claim 7, characterized in that: The error is expressed as: Among them, Loss represents the error of iterative search, MSE represents the least squares error, The coefficient representing the overall error of all modification sites; represents the modification rate of the i-th site that can be modified by PEG, n represents the total number of sites that can be modified by PEG, Represents the percentage data of the i-th peptide segment in the analysis data of the recombinant protein; Represents the percentage data of the i-th peptide segment in the analysis data of the PEG-modified recombinant protein.
9. The method for calculating the modification rate of a PEG-modified recombinant protein according to claim 1, wherein: The specific steps of the forward search include: forward searching the sub-solution space, calculating the current error based on the current solution, and comparing it with the saved minimum error. If the current error is less than the minimum error, updating the global optimal solution as the current solution; calculating the site whose outer modification rate in the current solution contributes the most to the error, first updating the modification rate of the site with the largest contribution to obtain an updated solution, using the updated solution to calculate an updated error, and comparing it with the saved minimum error. After cyclic iteration, a basically reasonable solution for the forward search is obtained. The outer modification rate includes , ,floor() represents the floor rounding function.
10. The method for calculating the modification rate of a PEG-modified recombinant protein according to claim 1, characterized in that: When the updated solution is obtained based on the current solution, determine the solution that contributes the most to the mean square error ,renew Get updated solution , using the updated solution Replace the solution in the solution Get an updated solution.
11. The method for calculating the modification rate of a PEG-modified recombinant protein according to claim 10, characterized in that: renew The process is expressed as: Among them, α represents the update step size.
12. The method for calculating the modification rate of a PEG-modified recombinant protein according to claim 1, characterized in that: The enzymatic cleavage is performed by hydrolysis using Lys-C protease.
13. The method for calculating the modification rate of a PEG-modified recombinant protein according to claim 1, characterized in that: The upper limit of the modification rate of the ith site during the forward search is: in, represents the upper limit of the traversal of the i-th site, and the sum of the modification rates is 1. represents the expanded search range of the modification rate and .
14. The method for calculating the modification rate of a single PEG-modified recombinant protein according to claim 13, characterized in that: The upper limit of the modification rate of the ith site in the forward search is in, represents the upper limit of the traversal of the i-th site, and the sum of the modification rates is 1. represents the extended search range of the modification rate and, represents an expanded expected search window, which is determined based on the protein purity requirement.
15. The method for calculating the modification rate of a PEG-modified recombinant protein according to claim 1, characterized in that: In the innermost layer of the nested loop of the forward search, determine whether the sum of the modification rates is 1.
16. The method for calculating the modification rate of a single PEG-modified recombinant protein according to claim 14, characterized in that: The sum of the modification rates is 1 and can be expanded to the sum of the modification rates within 1± Within the range, represents the expanded search range of the modification rate and .
17. The method for calculating the modification rate of a single PEG-modified recombinant protein according to claim 1, characterized in that: The analytical data were determined using liquid analysis.
18. The method for calculating the modification rate of a single PEG-modified recombinant protein according to claim 1, characterized in that: The percentage data of the peptide segments are the peak area percentages based on the results of liquid phase analysis.
19. The method for calculating the modification rate of a PEG-modified recombinant protein according to claim 6, characterized in that: The recombinant protein is EPO, the PEG-modified recombinant protein is PEG-EPO, n is 9, and the 9 modification sites are the N-terminal α-amino group and the 20th, 45th, 52nd, 97th, 116th, 140th, 152nd and 154th lysine side chain amino groups, and mn, m20, m45, m52, m97, m116, m140, m152 and m154 are used to represent the modification rates of the 9 modification sites. The calculation basis is expressed as follows: in, Represents the percentage data of the i-th peptide segment in the parsed data of EPO; It represents the percentage of the same peptide in the analytical data of PEG-EPO.
20. The method for calculating the modification rate of a single PEG-modified recombinant protein according to claim 19, characterized in that: The least squares error is expressed as: in, Indicates the percentage data of the i-th peptide segment in the EPO analysis data, Indicates the percentage data of the same peptide segment in the analytical data of PEG-EPO. , mn, m20, m45, m52, m97, m116, m140, m152, and m154 represent the modification rates of the 9 modification sites.
21. The method for calculating the modification rate of a single PEG-modified recombinant protein according to claim 19, characterized in that: The outermost loop of the forward search searches the value range of mn, and the inner loops search the value ranges of m20, m45, m52, m97, m116, m140, m152, and m154 in turn.
22. The method for calculating the modification rate of a single PEG-modified recombinant protein according to claim 19, characterized in that: The error includes the least squares error and the overall error of all modified sites, which is expressed as: Among them, Loss represents the error of iterative search, MSE represents the least squares error, represents the coefficient of the overall error of all modification sites; mn, m20, m45, m52, m97, m116, m140, m152, and m154 represent the modification rates of the 9 modification sites.
23. The method for calculating the modification rate of a single PEG-modified recombinant protein according to claim 21, characterized in that: The value range is [0,100]. In the innermost loop of the forward search, it is checked whether the sum of the solutions is 1 to generate a solution that meets the requirements. The solution.
24. The method for calculating the modification rate of a single PEG-modified recombinant protein according to claim 21, characterized in that: The upper limit of the value range is: in, represents the upper limit of the traversal of the i-th site, and the sum of the modification rates is 1. represents the expanded search range of the modification rate and .
25. The method for calculating the modification rate of a single PEG-modified recombinant protein according to claim 24, characterized in that: The upper limit of the value range is: in, represents the upper limit of the traversal of the i-th site, and the sum of the modification rates is 1. represents the extended search range of the modification rate and, represents an expanded expected search window, which is determined based on the protein purity requirement.
26. The method for calculating the modification rate of a single PEG-modified recombinant protein according to claim 19, characterized in that: The outermost loop of the reverse search searches the value range of m154, and the inner loops search the value ranges of m152, m140, m116, m97, m52, m45, m20, and mn in turn.
27. A computer device, characterized in that: The device comprises: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 26.
28. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 26 are implemented.
29. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 26 are implemented.