Polymer classification device, polymer classification method, and polymer classification program
The polymer classification device addresses the limitation of one-sided charge change calculations by using random numbers to determine modification presence and calculate single-molecule charge changes, facilitating comprehensive polymer profiling and molecular variant estimation.
Patent Information
- Application Number
- PCT/JP2025/007232
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-01
- Filing Date
- 2025-02-28
- Publication Date
- 2025-09-04
AI Technical Summary
Conventional methods for calculating polymer modification rates are limited to one-sided charge changes, failing to account for molecular variants with both negative and positive charges, making it difficult to derive a comprehensive profile of charge variants.
A polymer classification device and method that includes a memory unit storing modification correspondence data and analysis data, using random numbers to determine modification presence and calculate single-molecule charge changes, enabling classification into multiple groups based on properties and three-dimensional structure data.
Enables calculation of abundance rates of all possible molecular variants, estimation of charge variant analysis results from peptide mapping, and reconstruction of molecular variants, providing a detailed profile of polymers.
Smart Images

Figure JP2025007232_04092025_PF_FP_ABST
Abstract
Description
Polymer classification device, polymer classification method, and polymer classification program
[0001] The present invention relates to a polymer classification device, a polymer classification method, and a polymer classification program.
[0002] Non-Patent Document 1 discloses a technique for calculating the profile of charge variants of an entire antibody from the results of antibody peptide mapping, in which the value obtained by adding up the modification rates for each amino acid residue of modification species that cause a negative charge change is defined as the abundance ratio of negative charge variants in one antibody heavy chain and one antibody light chain combined, and this abundance ratio is corrected based on a binomial distribution to calculate the abundance ratio of negative charge variants in the entire antibody (two antibody heavy chains and two antibody light chains).
[0003] F. Yang. , MABS, 15, 2197668 (2023)
[0004] However, in conventional inventions, the calculation of modification rates is limited to only one-sided charge changes, and therefore, although molecular variants with both negative and positive charges actually exist, there was a problem in that it was not possible to derive a profile of charge variants that took this into account.
[0005] The present invention has been made in consideration of the above-mentioned problems, and aims to provide a polymer classification device, a polymer classification method, and a polymer classification program that can classify polymers taking into account molecular variants with various properties.
[0006] In order to solve the above-mentioned problems and achieve the objectives, a polymer classification device is provided which includes a memory unit and a control unit, wherein the memory unit includes a polymer storage means which stores modification correspondence data which links the type of modification applied to a polymer and the property change value caused by the modification, and analysis data which links the type of modification of polymer fragments obtained by fragmenting the polymer and the modification rate, and the control unit includes a classification means which acquires classification data which classifies the polymer into multiple groups according to its properties based on the modification correspondence data and the analysis data.
[0007] Furthermore, in the polymer classification device according to the present invention, the polymer is a molecule that is a protein, peptide, nucleic acid, or sugar chain, an artificially modified form of such a molecule, a complex of such a molecule and / or such an artificially modified form, a biopharmaceutical, a manufacturing aid, a nucleic acid for therapeutic use, or a carrier protein.
[0008] In addition, in the polymer classification device according to the present invention, the polymer storage means further stores three-dimensional structure data that sets three-dimensional structural changes of the polymer in accordance with the modification, and the classification means acquires the classification data that classifies the polymer into a plurality of groups in accordance with the properties based on the three-dimensional structure data, the modification correspondence data, and the analysis data.
[0009] In addition, in the polymer classification device of the present invention, the classification means assigns a random number to each position where each modification occurs in one polymer molecule, and if the random number is within a range set by the modification rate of the modification corresponding to the random number based on the analysis data, performs a modification determination process for each position in one polymer molecule where each modification occurs to determine whether or not the one polymer molecule contains the modification, and performs a property determination process for a predetermined number of polymers to determine the properties resulting from the presence or absence of each modification in the one polymer molecule based on the modification correspondence data, thereby obtaining the classification data in which the predetermined number of polymers are classified into multiple groups according to the properties.
[0010] In addition, in the polymer classification device of the present invention, the classification means assigns a random number to each position where each modification occurs on one polymer molecule, and if the random number is within a range set by the modification rate of the modification corresponding to the random number based on the analysis data, performs the modification determination process for each position in one polymer molecule where each modification occurs, determining whether or not the one polymer molecule contains the modification, and performs the property determination process for the predetermined number of polymers, determining the single-molecule charge change, which is the sum of the charge change values, which are the property change values caused by each of the modifications contained in the one polymer molecule, based on the modification correspondence data, thereby obtaining the classification data in which the predetermined number of polymers are classified into multiple groups according to the single-molecule charge change.
[0011] In addition, in the polymer classification device of the present invention, the polymer is a protein, the polymer fragment is a peptide, the property change value is a charge change value, the property caused by the presence or absence of each modification in one polymer molecule is a single-molecule charge change obtained by adding up the charge change values caused by each modification contained in one protein molecule, and the classification data is a profile of charge variants.
[0012] In the polymer classification apparatus according to the present invention, the charge change per molecule is calculated using Equation 1. (s: number of molecule to be generated, i: number of position where modification occurs (i = 1, 2, 3, ..., m), j: number of modification species occurring at position i (j = 1, 2, 3, ..., l), ΔZ s : Charge change value of molecule s1, R s,i : a random number corresponding to the modification position i of the molecule s1, p i,j : modification rate of modification type j at modification position i, Δz i,j : charge change value when modification species j occurs at modification position i)
[0013] In the polymer classification device according to the present invention, the classification data includes aggregate data in which the proteins are aggregated by valence.
[0014] Furthermore, in the polymer classification device of the present invention, the classification means calculates the abundance ratio of each molecular variant that may be present in the polymer according to the combination of modifications based on the analysis data, identifies the properties of each molecular variant based on the modification correspondence data, and calculates the abundance ratio of each property based on the abundance ratio of each molecular variant and the property of each molecular variant, thereby obtaining the classification data in which the polymer is classified into multiple groups according to the properties.
[0015] Furthermore, in the polymer classification device of the present invention, the classification means calculates the abundance ratio of each molecular variant that may be present in the polymer according to the combination of the type of modification and the modification rate based on the analysis data, identifies a single-molecule charge change that is the sum of the charge change values, which are the property change values caused by each modification contained in each molecular variant, based on the modification correspondence data, and calculates the abundance ratio of each single-molecule charge change based on the abundance ratio of each molecular variant and the single-molecule charge change of each molecular variant, thereby obtaining the classification data that classifies the polymer into multiple groups according to the single-molecule charge change.
[0016] In the polymer classification device according to the present invention, the abundance ratio of each molecular variant is calculated using Equation 2 and Equation 3. (t: number of the type of molecular variant, i: number of the position where the modification occurs (i = 1, 2, 3, ..., m), j: number of the modification species occurring at position i (j = 1, 2, 3, ..., l), P t : the abundance ratio of molecular variant t, c t,i,j : 1 or 0 (if molecular variant t has modification j at position i = 1, if molecular variant t does not have modification j at position i = 0 (note that there is only one type of modification at position i of molecular variant t, so c t,i,j The sum of j is 1 or 0), p i,j : modification rate of modification type j at modification position i) (t: number of the type of molecular variant, i: number of the position where the modification occurs (i = 1, 2, 3, ..., m), j: number of the modification species occurring at position i (j = 1, 2, 3, ..., l), ΔZ t : charge change value of molecular variant t, c t,i,j : 1 or 0 (if molecular variant t has modification j at position i = 1, if molecular variant t does not have modification j at position i = 0 (note that there is only one type of modification at position i of molecular variant t, so c t,i,j The sum of j is 1 or 0), Δz i,j : the charge change value caused by modification species j at modification position i)
[0017] In addition, in the polymer classification device according to the present invention, the polymer is a protein, the polymer fragment is a peptide, the property change value is a charge change value, and the classification data is a charge profile.
[0018] In addition, in the polymer classification device according to the present invention, the analytical data is characterized in that it is an analytical result obtained by a method of fragmenting the protein into peptides and obtaining the type of modification and the modification rate on the protein.
[0019] In the polymer sorting apparatus according to the present invention, the protein is an antibody, an antibody-drug conjugate, or an antibody-like molecule.
[0020] In addition, in the polymer classification device according to the present invention, the modification is one or a combination selected from the group consisting of deamidation of asparagine, cyclization of N-terminal glutamine, cyclization of N-terminal glutamic acid, deletion of C-terminal lysine, remaining of a sequence derived from an additional sequence at the N-terminus, acetylation of an amino acid at the N-terminus, and amidation of an amino acid at the C-terminus.
[0021] In addition, in the polymer classification device according to the present invention, the modification is one or a combination selected from the group consisting of glycation of lysine, glycosylation of asparagine, succinimidation of asparagine, oxidation of asparagine to isoaspartate, succinimidation of aspartic acid, isomerization of aspartic acid, glycosylation of threonine, glycosylation of serine, and deamidation of glutamine.
[0022] In addition, in the polymer classification device according to the present invention, the modification is one or a combination selected from the group consisting of advanced glycation end products of lysine, succinylation of lysine, remaining sequences derived from additional sequences at the C-terminus, oxidation of cysteine, phosphorylation of serine, phosphorylation of threonine, and phosphorylation of tyrosine.
[0023] In addition, in the polymer classification device of the present invention, the modification is characterized by being one or a combination selected from the group consisting of mutation of a monomer that is a constituent unit of the polymer, deletion of a monomer, introduction of a non-natural monomer, and introduction of an artificial modification.
[0024] Furthermore, the polymer classification method of the present invention is a polymer classification method to be executed by a polymer classification device equipped with a memory unit and a control unit, wherein the memory unit comprises a polymer storage means for storing modification correspondence data that links the type of modification on a polymer to the property change value resulting from the modification, and analytical data that links the type of modification of polymer fragments obtained by fragmenting the polymer to the modification rate, and is characterized by including a classification step that is executed by the control unit to obtain classification data that classifies the polymer into multiple groups according to its properties based on the modification correspondence data and the analytical data.
[0025] Furthermore, the polymer classification program of the present invention is a polymer classification program to be executed by a polymer classification device equipped with a memory unit and a control unit, wherein the memory unit comprises a polymer memory means for storing modification correspondence data set by linking the type of modification on a polymer and the property change value caused by the modification, and analysis data set by linking the type of modification of polymer fragments obtained by fragmenting the polymer and the modification rate, and the control unit is characterized by executing a classification step in which, based on the modification correspondence data and the analysis data, the program acquires classification data that classifies the polymer into multiple groups according to its properties.
[0026] The present invention has the advantage of being able to calculate the abundance rates of all possible molecular variants of a polymer. Furthermore, the present invention has the advantage of being able to estimate the results of conventional charge variant analysis methods from the results of peptide mapping. Furthermore, the present invention has the advantage of being able to reconstruct molecular variants (e.g., charge mutants) from the results of peptide mapping. Furthermore, the present invention has the advantage of being able to calculate the profile of the entire polymer from the modification types and modification rates of the polymer obtained using mass spectrometry or the like.
[0027] FIG. 1 is a block diagram showing an example of the configuration of a polymer classification device according to this embodiment. FIG. 2 is a flowchart showing an example of polymer classification processing according to this embodiment. FIG. 3 is a diagram showing an example of polymer classification processing according to this embodiment. FIG. 4 is a diagram showing an example of polymer classification processing for a virtual protein according to this embodiment. FIG. 5 is a diagram showing an example of an input list according to this embodiment. FIG. 6 is a diagram showing an example of analysis results according to this embodiment. FIG. 7 is a diagram showing an example of a modified species according to this embodiment. FIG. 8 is a diagram showing an example of an output according to this embodiment. FIG. 9 is a diagram showing an example of an antibody according to this embodiment. FIG. 10 is a diagram showing an example of LC / MS / MS analysis conditions according to this embodiment. FIG. 11 is a diagram showing an example of LC analysis conditions according to this embodiment. FIG. 12 is a diagram showing an example of capillary isoelectric focusing analysis conditions according to this embodiment. FIG. 13 is a diagram showing an example of capillary zone electrophoresis conditions according to this embodiment. FIG. 14 is a diagram showing an example of analysis results according to this embodiment. FIG. 15 is a diagram showing an example of analysis results of a protein sample having a molecular weight of approximately 48,000 according to this embodiment. FIG. 16 is a diagram showing an example of analysis results of a peptide sample according to this embodiment.
[0028] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS An embodiment of the present invention will be described in detail with reference to the accompanying drawings. However, the present invention is not limited to this embodiment.
[0029] [1. Overview] First, an overview of the present invention will be described.
[0030] Traditionally, proteins such as biopharmaceuticals are highly heterogeneous, and various molecular variants exist even in the same protein solution. This can affect efficacy and safety, making this an important control point in quality assessment. Therefore, biopharmaceuticals have traditionally been managed using analytical methods tailored to each molecular variant. For example, to evaluate the content of variants with different charges (i.e., charge variants, charge mutants, charge variants, charge isomers, or charge variants), "charge variant analysis," which separates and analyzes proteins based on their charge state, has been used. Meanwhile, in recent years, a more efficient method known as "peptide mapping," which fragments and analyzes proteins, has attracted attention as a method known as the Multi-Attribute Method (MAM), which simultaneously analyzes multiple quality attributes. However, peptide mapping involves analyzing the original protein by digesting it into peptides using a digestive enzyme, which results in the loss of information about the entire original protein profile (e.g., the number of charge variants present). In other words, in conventional peptide mapping, proteins are digested and fragmented into peptides for analysis, the original proteins are heterogeneous, and it is not known which original mutant protein the analyzed peptides are derived from. This makes it difficult to reconstruct information about the original protein from the results of peptide mapping, and makes it difficult to bridge the results of peptide mapping with information about known proteins.
[0031] Therefore, in this embodiment, a mechanism is provided for estimating the profile of charge variants obtained by charge variant analysis from the results of peptide mapping, for example.
[0032] [2. Configuration of the polymer classification apparatus 100] The polymer classification apparatus 100 according to this embodiment can be configured by functionally or physically distributing or integrating any unit (either as a stand-alone type or a system type). In this embodiment, an example of the configuration of the polymer classification apparatus 100 will be described with reference to Fig. 1. Fig. 1 is a block diagram showing an example of the configuration of the polymer classification apparatus 100 according to this embodiment.
[0033] 1 , the polymer classification apparatus 100 may be an information processing device such as a personal computer or a workstation. The polymer classification apparatus 100 includes a control unit 102, a storage unit 106, and an input / output unit 112, and the various units included in the polymer classification apparatus 100 are communicatively connected via any communication path. For example, the polymer classification apparatus 100 may include the control unit 102 and / or the storage unit 106 as an analysis environment on the cloud. The polymer classification apparatus 100 is communicatively connected to other devices via a network 300.
[0034] The input / output unit 112 may have a function for inputting and outputting data (I / O). Here, the input / output unit 112 may be, for example, a key input unit, a touch panel, a control pad (e.g., a touch pad, a game pad, etc.), a mouse, a keyboard, a microphone, etc. The input / output unit 112 may also be a display unit (e.g., a display, monitor, and touch panel made of liquid crystal or organic electroluminescence, etc.) that displays (input / output) information of application software, etc. The input / output unit 112 may also be an audio output unit (e.g., a speaker, etc.) that outputs audio information as audio. The input / output unit 112 may also be an image input unit (e.g., a camera, etc.) that records images (still images and videos) captured by an imaging element such as a CCD image sensor or a CMOS image sensor as digital data. The input / output unit 112 may also be a fingerprint sensor, a camera (e.g., an infrared camera, etc.) that can be used for iris authentication or face authentication, etc., and / or a biometric sensor such as a vein sensor.
[0035] The storage unit 106 stores various databases, tables, and / or files. The storage unit 106 stores computer programs that cooperate with an operating system (OS) to issue commands to a central processing unit (CPU) to perform various processes. The storage unit 106 may be, for example, a random access memory (RAM), a read-only memory (ROM), a hard disk drive (HDD), and / or a solid state drive (SSD). The storage unit 106 may store image data recorded by the input / output unit 112, data received via the network 300, and / or input data input via the input / output unit 112. The storage unit 106 conceptually includes a polymer database 106a.
[0036] The polymer database 106a stores polymer data. The polymer database 106a may store modification correspondence data linking the type of modification of a polymer with the property change value resulting from the modification, as well as analytical data linking the type of modification of a polymer fragment obtained by fragmenting the polymer with the modification rate. The polymer may be a protein, peptide, nucleic acid, or sugar chain molecule, an artificially modified version of the molecule, a complex of the molecule and / or the artificially modified version, a biopharmaceutical, a manufacturing aid (e.g., growth factors in regenerative medicine products or cell therapy products), a nucleic acid for treatment (e.g., gene therapy products), or a carrier protein (e.g., capsid). The polymer database 106a may also store three-dimensional structure data specifying a three-dimensional conformational change of the polymer depending on the modification. The polymer may be a protein, the polymer fragment may be a peptide, and the property change value may be a charge change value. The analytical data may also be the results of an analysis performed by a method of fragmenting a protein into peptides and acquiring the type of modification and the modification rate of the protein. The protein may be an antibody, an antibody-drug conjugate, or an antibody-like molecule. The modification may be one or a combination selected from the group consisting of asparagine deamidation, N-terminal glutamine cyclization, N-terminal glutamic acid cyclization, C-terminal lysine deletion, remaining sequence derived from the N-terminal additional sequence, N-terminal amino acid acetylation, and C-terminal amino acid amidation. The modification may be one or a combination selected from the group consisting of lysine glycation, asparagine glycosylation, asparagine succinimidation, asparagine isoaspartic acid succinimidation, aspartic acid isomerization, threonine glycosylation, serine glycosylation, and glutamine deamidation. The modification may be one or a combination selected from the group consisting of lysine advanced glycation end products, lysine succinylation, remaining sequence derived from the C-terminal additional sequence, cysteine oxidation, serine phosphorylation, threonine phosphorylation, and tyrosine phosphorylation.The modification may be one or a combination selected from the group consisting of a mutation of a monomer that is a structural unit of a polymer, a deletion of a monomer, the introduction of a non-natural monomer, and the introduction of an artificial modification. The polymer database 106a may also store classification data that classifies polymers into a plurality of groups.
[0037] The control unit 102 is a CPU or the like that performs overall control of the polymer classification apparatus 100. The control unit 102 has an internal memory for storing control programs such as an OS, programs that define various processing procedures, required data, etc., and executes various information processing operations based on these stored programs. Functionally, the control unit 102 is equipped with an analysis and acquisition unit 102a and a classification unit 102b.
[0038] The analysis acquisition unit 102a acquires analytical data, and may register the analytical data of the polymer and / or polymer fragment in the polymer database 106a.
[0039] The classification unit 102b acquires classification data in which polymers are classified into a plurality of groups. Here, the classification unit 102b may acquire classification data in which polymers are classified into a plurality of groups according to their properties. Furthermore, the classification unit 102b may acquire classification data in which polymers are classified into a plurality of groups according to their properties based on modification correspondence data and analytical data. Furthermore, the classification unit 102b may acquire classification data in which polymers are classified into a plurality of groups according to their properties based on three-dimensional structure data, modification correspondence data, and analytical data. Furthermore, the classification unit 102b may acquire classification data in which polymers are classified into a plurality of groups according to their properties by assigning a random number to each position at which each modification occurs in a single polymer molecule, and, if the random number falls within a range set by the modification rate of the modification corresponding to the random number based on the analytical data, performing a modification determination process for each position at which each modification occurs in the single polymer molecule to determine whether or not the single polymer molecule contains the modification, and performing a property determination process for a predetermined number of polymers to determine the properties resulting from the presence or absence of each modification in the single polymer molecule based on the modification correspondence data, thereby acquiring classification data in which a predetermined number of polymers are classified into a plurality of groups according to their properties. The classification unit 102b may also assign a random number to each position where each modification occurs in a single polymer molecule, and if the random number is within a range set by the modification rate of the modification corresponding to the random number based on the analysis data, perform a modification determination process for each position where each modification occurs in the single polymer molecule to determine whether or not the single polymer molecule contains the modification. Based on the modification correspondence data, the classification unit 102b may also perform a property determination process for a predetermined number of polymers to determine a single-molecule charge change, which is a property change value resulting from each modification contained in the single polymer molecule, by adding up charge change values. In this case, the property resulting from the presence or absence of each modification in a single polymer molecule is a single-molecule charge change, which is a sum of charge change values resulting from each modification contained in a single protein molecule, and the classification data may be a profile of charge variants. Furthermore, the single-molecule charge change may be calculated using Equation 1. (s: number of molecule to be generated, i: number of position where modification occurs (i = 1, 2, 3, ..., m), j: number of modification species occurring at position i (j = 1, 2, 3, ..., l), ΔZ s : Charge change value of molecule s1, R s,i : a random number corresponding to the modification position i of the molecule s1, p i,j : modification rate of modification type j at modification position i, Δz i,j : charge change value when modification species j occurs at modification position i)
[0040] The classification data may also include aggregate data that aggregates proteins by valency. The classification unit 102b may calculate, based on the analytical data, the abundance ratio of each molecular variant corresponding to a combination of modifications that may be present in the polymer, identify the properties of each molecular variant based on the modification correspondence data, and calculate the abundance ratio of each molecular variant and the abundance ratio of each property based on the properties of each molecular variant, thereby obtaining classification data that classifies the polymer into multiple groups according to its properties. The classification unit 102b may also calculate, based on the analytical data, the abundance ratio of each molecular variant corresponding to a combination of the type of modification and the modification rate that may be present in the polymer, identify a single-molecule charge change that is the sum of the charge change values, which are property change values caused by each modification contained in each molecular variant, based on the modification correspondence data, and calculate the abundance ratio of each molecular variant and the abundance ratio of each single-molecule charge change based on the single-molecule charge change of each molecular variant, thereby obtaining classification data that classifies the polymer into multiple groups according to the single-molecule charge change. The classification data may also be a charge variant profile. The abundance ratio of each molecular variant may also be calculated using Equation 2 and Equation 3. (t: number of the type of molecular variant, i: number of the position where the modification occurs (i = 1, 2, 3, ..., m), j: number of the modification species occurring at position i (j = 1, 2, 3, ..., l), P t : the abundance ratio of molecular variant t, c t,i,j : 1 or 0 (if molecular variant t has modification j at position i = 1, if molecular variant t does not have modification j at position i = 0 (note that there is only one type of modification at position i of molecular variant t, so c t,i,jThe sum of j is 1 or 0), p i,j : modification rate of modification type j at modification position i) (t: number of the type of molecular variant, i: number of the position where the modification occurs (i = 1, 2, 3, ..., m), j: number of the modification species occurring at position i (j = 1, 2, 3, ..., l), ΔZ t : charge change value of molecular variant t, c t,i,j : 1 or 0 (if molecular variant t has modification j at position i = 1, if molecular variant t does not have modification j at position i = 0 (note that there is only one type of modification at position i of molecular variant t, so c t,i,j The sum of j is 1 or 0), Δz i,j : the charge change value caused by modification species j at modification position i)
[0041] 3. Polymer Classification Processing An example of the polymer classification processing according to this embodiment will be described with reference to Fig. 2 to Fig. 14. Fig. 2 is a flowchart showing an example of the polymer classification processing according to this embodiment.
[0042] As shown in FIG. 2, when a user inputs analysis results obtained by a method of fragmenting a polymer and obtaining the type and modification rate of the polymer via the input / output unit 112, the analysis acquisition unit 102a registers the analysis results as analysis data in the polymer database 106a (step SA-1).
[0043] Then, the classification unit 102b obtains classification data in which the polymers are classified into a plurality of groups according to their properties based on the three-dimensional structure data, modification correspondence data, and analysis data (step SA-2).
[0044] Then, the classification unit 102b displays the classification data on the input / output unit 112 (step SA-3), and ends the process.
[0045] Here, a specific example of the polymer classification process in this embodiment will be described with reference to Fig. 3 to Fig. 6. Fig. 3 is a diagram showing an example of the polymer classification process in this embodiment. Fig. 4 is a diagram showing an example of the polymer classification process for a virtual protein in this embodiment. Fig. 5 is a diagram showing an example of an input list in this embodiment. Fig. 6 is a diagram showing an example of an analysis result in this embodiment.
[0046] As shown in Figure 3, in the past, proteins were analyzed as they were and charge variant profiles were calculated. However, in this embodiment, the core of the method is a calculation method in which a charge variant profile (main peak, acidic peak, basic peak) of the entire protein is reconstructed from the abundance ratio of each modification based on a list of the types and modification rates of modifications for each amino acid residue or peptide.
[0047] 4, this embodiment includes Method 1, in which random numbers are used to generate a large number of protein molecules having various modifications based on the modification rates obtained by peptide mapping, and the number of protein molecules is counted for each charge change value, and Method 2, in which the abundance ratios of all molecular variants are theoretically calculated and the ratios for each charge change value are calculated. Here, in Method 1, for example, the charge change value of one molecule of a specific protein may be calculated, and this may be repeated 10,000 times to generate 10,000 pseudo-protein molecules.
[0048] 4, in this embodiment, the rate of charge change for the entire protein is calculated from a table of modifications (charge changes) and modification rates for each amino acid or each peptide by Method 1 or Method 2. In Method 1, unchanged molecules and molecularly modified molecules are artificially generated one by one according to Equation 1, and a random number R s,i (0.00 to 100.00%) and s,i is the corresponding modification rate p i,j If the molecule s falls within the interval set based on i,j This random numbering and modification determination is performed for all modification positions i of interest, and the charge change value Δzi,j By summing up the charge change value ΔZ of the entire molecule s s After performing this process multiple times (n times), the number of molecules generated for each charge change value ΔZ is counted and divided by n to calculate the profile of charge variants of the entire protein. Note that performing the process n times is equivalent to generating n molecular variants artificially. In Method 2, the abundance ratio P t The charge change value ΔZ of each molecular variant is calculated theoretically according to Equation 3. t is calculated, and the abundance ratios are summed for each charge change value ΔZ to calculate the profile of charge variants of the entire protein.
[0049] In this embodiment, as shown in Figure 5, in the antibody pharmaceutical trastuzumab, which is a complex consisting of two light chains (214 amino acids) and two heavy chains (449 amino acids), 33 modifications were detected in one light chain and one heavy chain by peptide mapping, of which 15 modifications involved in charge change were detected. Note that, as shown in Figure 5, the table of modifications and modification rates for each amino acid or peptide may include data obtained by peptide mapping (type of modification, abundance ratio of each modification) and theoretically or empirically obtained data (charge change value when the modification occurs).
[0050] As shown in FIG. 6, in this embodiment, the simulation results and the theoretical calculation results were almost in agreement, and a high degree of agreement was also obtained between the simulation results and the theoretical calculation results and the conventional method.
[0051] In this embodiment, a classification of modifications (e.g., deamidation, etc.) is defined as a "modification type," and a modification that specifically occurs in a particular amino acid residue or a particular peptide fragment (e.g., deamidation detected in a peptide containing asparagine at position 55 or aspartic acid at position 55) is defined as a "modification."
[0052] Furthermore, in this embodiment, peptide mapping involves subjecting a fragmented protein sample to liquid chromatography / tandem mass spectrometry (LC / MS / MS), and using dedicated analytical software to create a database of precursor ions and fragment ions from the known amino acid sequence and predicted modification species. This database is then compared with actual measurement data and subjected to statistical processing, thereby enabling confirmation of the amino acid sequence of the protein and estimation and quantification of modifications.
[0053] Furthermore, in this embodiment, MAM (Multi-Attribute Method) is a quantitative analysis method based on peptide mapping technology and developed with an aim to control the quality of proteins. This is a series of analysis methods in which, based on the results of peptide mapping using a high-performance mass spectrometer, unmodified and modified peptides to be monitored as important quality characteristics, etc. are selected, and quality control is performed by monitoring them using a mass spectrometer with relatively lower performance.
[0054] Furthermore, as shown in FIG. 7, in this embodiment, there are positions (amino acid residues) in the protein where each modification occurs, and the modification causes a change in charge.
[0055] An example of protein classification processing in this embodiment will be described with reference to Figures 8 and 9. Figure 8 is a diagram showing an example of output in this embodiment. Figure 9 is a diagram showing an example of an antibody in this embodiment.
[0056] In this embodiment, Step 1: peptide fragmentation by enzymatic digestion of a protein involves digesting a hypothetical 20-residue protein "ETYICDKMHLPSNTKDILVE" with the digestive enzyme trypsin, which selectively cleaves at the C-terminal side of K or R. When digestion is complete, three fragments, "ETYICDK," "MHLPSNTK," and "DILVE," are generated. Here, in this embodiment, it is assumed that cyclization of N-terminal glutamic acid and deamidation of asparagine occur at a certain rate in this sequence. In reality, specific amino acid residues can undergo multiple types of modification. As an example, asparagine in the consensus sequence undergoes multiple types of glycosylation (N-linked type), and the charge change value, etc., varies depending on the type of glycosylation.
[0057] In this embodiment, Step 2 is performed to analyze the digest (LC / MS, LC / MS / MS). This analysis may be performed using LC, LC / MS, CE, CE / MS, or other techniques such as MALDIMS.
[0058] In this embodiment, Step 3: Identification of unmodified and modified peptides from data such as mass spectra and acquisition of quantitative values involves using parameters such as peak area, peak intensity, and number of spectra in quantification to calculate quantitative values (abundance ratio of each modification).
[0059] In this embodiment, in Step 4: Inputlist creation processing, a "modification list for each amino acid residue" is created. Here, the Inputlist may be a "modification list for each peptide." Here, in this embodiment, if data on "incompletely digested peptides" or "non-specifically cleaved peptides" in Step 1 exists, these data can also be added together and included in the list.
[0060] In this embodiment, in Step 5-1: "Theoretical Calculation (Method 2)," the charge distribution of the original protein is calculated by directly using the modification ratios obtained in the input list for each amino acid or peptide. In this embodiment, when there are few amino acid residues or modifications, calculations can be performed using simple formulas. However, when there are many amino acid residues and complex modifications, the number of combinations becomes enormous, and the formula becomes complicated. Here, in this embodiment, the "calculation based on the modification list for each peptide" calculates the abundance ratios and charge changes of all possible protein molecular variants based on the input list (for example, if peptide 1 can take two states, peptide 2 can take two states, and peptide 3 can take one state, the number of possible protein molecular variants is 2 × 2 × 1 = 4, and the abundance ratios of four types of molecular variants are calculated). In this embodiment, the charge distribution of the original protein is obtained by summing the abundance ratios of molecular variants for each charge change value from the charge changes and abundance ratios of all the obtained molecular variants.
[0061] Furthermore, in this embodiment, in Step 5-2: the process of calculating the charge distribution of the original protein by "Simulation (Method 1)," the proportion of modifications in the Inputlist for each amino acid or peptide is regarded as the probability of occurrence of that modification, one protein molecule is artificially generated based on this probability, and the charge change value of the entire protein molecule is calculated. This process is repeated n times to generate n protein molecules, and the charge change value of each is calculated. The number of protein molecules for each charge change value is then added up to obtain the charge distribution (based on the number of molecules) of the original protein. For example, in this embodiment, in the "calculation based on modifications for each peptide," when the occurrence probabilities of a peptide and its modification based on the Inputlist are "ETYICDK (0.2)," "MHLPSNTK (0.3)," and "DILVE (0)," a numerical value between 0 and 1 (random number) is randomly generated for each peptide, and if the value is equal to or less than the respective occurrence probabilities, it is determined that "one molecule of that protein has that modification" (for example, if the random numbers obtained in the first trial (= generation of the first molecule) are "0.03, 0.82, 0.11" (0.03≦0.2, 0.82>0.3, 0.11>0), the protein molecule has the modification states of "ETYICDK: modified," "MHLPSNTK: unmodified," and "DILVE: unmodified," and the charge change value of the entire molecule is −1).
[0062] 8, in this embodiment, in Step 6: Output acquisition processing, the ratio of the main peak (no charge mutation (0)) to the peaks of charge variants (components with a value of -1 or +1) is obtained and displayed in a table, list, or graph format. The charge change value of the obtained charge variants falls within the range of integer values expected from the modification detected in the target protein.
[0063] In the "Simulation (Method 1)" of this embodiment, a list including (1) modification identifiers (e.g., combinations of modification types and modification positions (positions of peptide sequences or amino acids)), (2) the proportions of each modification obtained by measurement, and (3) the charge change values that occur with each modification is input as [Input]. Note that in this embodiment, the minimum list may be, for example, an m-row x 2-column matrix in which (1) identifiers (m items) are represented by rows and (2) the proportions of each modification and (3) the charge change values are represented by columns.
[0064] In the "Simulation (Method 1)" of this embodiment, in the [Process], for [Input], (1) the "proportion of each modification obtained by measurement" is regarded as the "occurrence rate of each modification," and under the assumption that a given protein molecule has each modification probabilistically based on the independent "occurrence rate of each modification," a large number (n) of protein variant molecules (including those without modifications) are generated using random numbers, (2) the individual charge changes of the n protein variant molecules are calculated, and (3) the number of n protein variant molecules for each charge change value is calculated, and then the proportions are calculated. Specifically, in this embodiment, the generation of a given protein variant molecule and the calculation of its charge change value are performed using Equation 1 as the above (1) and (2). (s: number of molecule to be generated, i: number of position where modification occurs (i = 1, 2, 3, ..., m), j: number of modification species occurring at position i (j = 1, 2, 3, ..., l), ΔZ s : Charge change value of molecule s1, R s,i : a random number corresponding to the modification position i of the molecule s1, p i,j : modification rate of modification type j at modification position i, Δz i,j : charge change value when modification species j occurs at modification position i)
[0065] In the "Simulation (Method 1)" of this embodiment, performing the calculation of Equation 1 n times for s = 1, 2, ... is equivalent to generating (simulating) n protein variant molecules and calculating the charge change values for each. By making n sufficiently large, it is possible to obtain results that are almost identical to those in the "Theoretical Calculation (Method 2)." When using spreadsheet software, a function that generates random numbers can be created to generate a single protein variant molecule and calculate its charge change value.
[0066] Furthermore, in this embodiment, when the target protein is a complex consisting of multiple subunits, calculations are performed in either the manner of (A) or (B) below. That is, in this embodiment, (A) the "generation of protein variant molecules" is incorporated into the calculation, and in the case of a homocomplex of h subunits, the detected modifications X1, X2, X3, ... are considered to exist in h numbers each, and all are calculated as stochastically independent. In the case of a heterocomplex, the modifications of all subunits are calculated as stochastically independent sources; or (B) after calculating (1) and (2) above for each subunit, the charge change is calculated by adding up the "charge change values" between the subunits, and the corresponding ratio is calculated by multiplying the "percentage per charge change value" in (3) above between the subunits. Here, as shown in FIG. 9, in peptide mapping in this embodiment, data on modification and modification rate are obtained for each primary sequence of the subunit (light and heavy chains of an antibody), and the modification rate in the original subunit configuration is calculated to reconstruct the modification state of the entire protein (antibody).
[0067] In the "Simulation (Method 1)" of this embodiment, the charge distribution of the entire protein molecule is output as [Output]. For example, for convenience, when the charge variant with the most frequent number of protein variant molecules in [Process] (3) above is defined as the "main peak," and the peaks on the negative side of this peak are defined as acidic peaks, and the peaks on the positive side of this peak are defined as basic peaks, the output is [the proportion of all or any of the "main peak," "acidic peak," and "basic peak"], [the proportion of all or any of the "main peak," "acidic peak" for each charge change value, and "basic peak" for each charge change value].
[0068] An example of an experimental protocol in this embodiment will be described with reference to Fig. 10 to Fig. 13. Fig. 10 is a diagram showing an example of LC / MS / MS analysis conditions in this embodiment. Fig. 11 is a diagram showing an example of LC analysis conditions in this embodiment. Fig. 12 is a diagram showing an example of capillary isoelectric focusing analysis conditions in this embodiment. Fig. 13 is a diagram showing an example of capillary zone electrophoresis conditions in this embodiment.
[0069] In this embodiment, sample pretreatment, LC / MS / MS measurement, and data analysis for peptide mapping were performed using the method described by Tajiri-Tsukada, Bioengineered, 11, 984-1000 (2020), with some modifications. For sample pretreatment, 5 μL of trastuzumab solution (4 μg / μL) was mixed with 150 μL of Tris-hydrochloride buffer (250 mM, pH 7.5) containing guanidine hydrochloride (7.5 M) and 3 μL of dithiothreitol aqueous solution (500 mM). The mixture was then incubated at room temperature for 30 minutes. For sample pretreatment, 7 μL of iodoacetamide aqueous solution (500 mM) was added and incubated for 15 minutes in the dark at room temperature, followed by the addition of 4 μL of dithiothreitol aqueous solution (500 mM). For sample pretreatment, the specimen was buffer-exchanged into Tris-HCl buffer (100 mM, pH 7.5) using a gel filtration column (NAP-5 columns, Cytiva). The entire specimen was mixed with a solution containing 6 μg of trypsin and 2 μg of lysyl endopeptidase and incubated at 37°C for 30 minutes. After this, 5 μL of formic acid and 12 μL of acetonitrile were added. For sample pretreatment, the entire reaction solution was subjected to solid-phase extraction (Oasis PRiME HLB, Waters). The eluate was evaporated to dryness and then redissolved in 40 μL of 2% acetonitrile containing 0.1% formic acid. 2 μL of the redissolved specimen was subjected to LC / MS / MS analysis.
[0070] 10 , the LC / MS / MS analysis in this embodiment was performed using an instrument configuration and analytical conditions typical for peptide mapping. The resulting raw measurement data was analyzed using analysis software provided by the mass spectrometer manufacturer (BioPharma Finder ver. 3.2 and ver. 5.0, Thermo Corp.) to identify peptides and modified peptides and calculate the modification ratio for each peptide. The list of modification ratios obtained here was used for analysis. In the LC / MS / MS analysis in this embodiment, the amino acid sequences to be analyzed were the heavy chain (sequence excluding the C-terminal lysine residue) and light chain of trastuzumab. The modification species to be analyzed were pyroglutamate oxidation of the N-terminal glutamic acid residue of the heavy chain, residual lysine residue at the C-terminus of the heavy chain, deamidation (asparagine residue, glutamine residue), glycation (lysine residue), oxidation (methionine residue, tryptophan residue), and N-glycosylation (asparagine residue in the heavy chain consensus sequence). In the LC / MS / MS analysis of this embodiment, all cysteine residues in the sequence were analyzed as being carbamidomethylated.
[0071] Furthermore, as shown in FIG. 11 , in the ion exchange chromatography of this embodiment, as sample pretreatment, a 1 μg / μL trastuzumab solution was prepared by dilution with mobile phase A, and 1 μL of the solution was subjected to LC analysis using a typical apparatus configuration and analytical conditions. The obtained raw measurement data was analyzed using analysis software (Chromeleone, Thermo) provided by the LC apparatus manufacturer, and the abundance ratios of the main peak, acidic peak, and basic peak were calculated.
[0072] As shown in FIG. 12, in the capillary isoelectric focusing of this embodiment, as a sample pretreatment, 3 μL of trastuzumab solution (21 μg / μL) was mixed with 200 μL of cIEFgel (Sciex) containing 3.75 M urea, 12 μL of ampholyte (Pharmalyte) of pH 3-10, and 3-10, Cytiva), 20 μL of arginine aqueous solution (500 mM), 2 μL of iminodiacetic acid aqueous solution (200 mM), 2 μL of pImarker4.1 (Sciex), 2 μL of pImarker7.0 (Sciex), and 2 μL of pImarker10.0 (Sciex) were mixed to prepare an analytical sample. Analysis was carried out according to the standard protocol using kit reagents from the instrument manufacturer (chemical mobilization isoelectric focusing kit, Sciex). The obtained raw measurement data was analyzed using analysis software (32Krat, Sciex) provided by the capillary electrophoresis instrument manufacturer, and the abundance ratios of the main peak, acidic peak, and basic peak were calculated.
[0073] As shown in FIG. 13 , in capillary zone electrophoresis in this embodiment, sample pretreatment was performed using a trastuzumab solution (21 μg / μL) as the analysis sample and a kit reagent (CZE Rapid Charge Variant Analysis Kit, Sciex) provided by the device manufacturer in accordance with a standard protocol. The obtained raw measurement data was analyzed using analysis software (32Krat, Sciex) provided by the capillary electrophoresis device manufacturer, and the abundance ratios of the main peak, acidic peak, and basic peak were calculated.
[0074] [Examples] Examples 1-4 will be described with reference to Fig. 14. Fig. 14 is a diagram showing an example of an analysis result in this embodiment.
[0075] As shown in Figure 14, this example shows two examples in which the modifications to be analyzed in trastuzumab were reduced, and also shows the results of Example 1: [Method 1] and Example 2: [Method 2], which targeted all modifications, as well as the results of the conventional method. In this Example, in the "Example covering all modifications," calculations were performed assuming that four modification types (deamidation, glycation, non-processing of C-terminal lysine, and cyclization of N-terminal glutamic acid) were present at a total of 30 sites on the two light chains and two heavy chains. In Example 3: [Method 1], in which fewer modifications were used (only those with a modification rate of 1% or more were included in the calculations), 10 modifications across four modification types (deamidation, glycation, non-processing of C-terminal lysine, and cyclization of N-terminal glutamic acid) were used in the calculations. In Example 4: [Method 2], in which fewer modifications were used (only those with a modification rate of 2% or more were included in the calculations), 6 modifications across three modification types (deamidation, glycation, and cyclization of N-terminal glutamic acid) were used in the calculations. Note that in this Example, if the number of modifications analyzed were further reduced based on the modification rate, there would be no modifications that would result in a positive Δzi, and therefore Example 2 is the minimum Example.
[0076] Example 5 will be described with reference to Fig. 15. Fig. 15 is a diagram showing an example of the analysis results of a protein sample with a molecular weight of approximately 48,000 in this embodiment.
[0077] First, in Example 5, the differences from the analyses shown in FIGS. 10 and 11 are as follows: an aqueous iodoacetic acid solution was used instead of an aqueous iodoacetamide solution for sample pretreatment in peptide mapping; Zeba Spin Desalting Columns (7K MWCO) (Thermo) were used as the gel filtration column; and solid-phase extraction of the sample after digestion was not performed.
[0078] In Example 5, the ion exchange chromatography was performed using an Accura BioPro IEX SF column (particle size: 3 μm, inner diameter: 2.1 mm, length: 100 mm) (YMC Corporation) at a mobile phase flow rate of 0.1 mL / min, with gradient elution (linear change in composition from 0% B to 75% B over 30 minutes).
[0079] Here, as shown in FIG. 15 , in this Example 5, when ranibizumab with a molecular weight of approximately 48,000 was analyzed, the acidic peak was calculated to be 0.01% (expressed as 0.0%) by [theoretical calculation (method 2)], and the acidic peak component was detected.
[0080] Furthermore, as shown in Figure 15, Example 5 demonstrated that even trace amounts of molecular variants detected by ion exchange chromatography can be evaluated by [theoretical calculation (Method 2)], and that this method is applicable not only to antibody proteins with a large molecular weight of approximately 148,000, but also to proteins with a molecular weight of approximately 48,000.
[0081] Example 6 will now be described with reference to Figure 16. Figure 16 is a diagram showing an example of the analysis results of a peptide sample in this embodiment.
[0082] First, in Example 6, the differences from the analyses shown in Figures 10 and 11 are as follows: as sample pretreatment for peptide mapping, guanidine hydrochloride and aqueous iodoacetamide were not added to the specimen, buffer exchange using a gel filtration column was not performed, only lysilyl endopeptidase was used as the digestive enzyme, and solid-phase extraction of the specimen after digestion was not performed.
[0083] In Example 6, the ion exchange chromatography was performed using an Accura BioPro IEX SF column (particle size: 3 μm, inner diameter: 2.1 mm, length: 100 mm) (YMC Corporation), with the concentrations of mobile phase A and mobile phase B increased by 5 times, at a mobile phase flow rate of 0.1 mL / min, and gradient elution (0% B for 2 minutes after sample injection, followed by a linear change in composition to 100% B over 15 minutes).
[0084] As shown in Figure 16, in Example 6, when sermorelin, a peptide consisting of 29 amino acid residues with a molecular weight of approximately 3,300, and its degraded products were analyzed, the results of ion exchange chromatography and the results of the theoretical calculation (Method 2) were in good agreement, demonstrating that the method can be applied to the evaluation of peptides.
[0085] Therefore, Examples 1 to 6 demonstrated that the method can be applied to the evaluation of charge profiles of peptides to large proteins, and also to the evaluation of changes in charge profiles due to their degradation.
[0086] [4. Other Embodiments] In addition to the above-described embodiment, the present invention may be implemented in various different embodiments within the scope of the technical concept described in the claims.
[0087] For example, among the processes described in the embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using known methods.
[0088] Furthermore, the processing procedures, control procedures, specific names, information including parameters such as registered data and search conditions for each process, screen examples, and database configurations shown in this specification and drawings can be changed as desired unless otherwise specified.
[0089] Furthermore, with regard to the polymer sorting device 100 and the like, the components shown in the drawings are functional concepts, and do not necessarily have to be physically configured as shown in the drawings.
[0090] For example, all or any part of the processing functions of the polymer classification apparatus 100, particularly the processing functions performed by the control unit, may be implemented by a CPU and a program interpreted and executed by the CPU, or may be implemented as hardware using wired logic. The program is recorded on a non-transitory computer-readable recording medium containing programmed instructions for causing the information processing device to execute the processes described in this embodiment, and is mechanically read as needed. That is, a computer program for providing instructions to the CPU in cooperation with the OS and performing various processes is recorded in a storage unit such as a ROM or HDD (Hard Disk Drive). This computer program is executed by being loaded into RAM, and cooperates with the CPU to form the control unit.
[0091] In addition, this computer program may be stored in an application program server connected to the polymer classification device 100 or the like via any network 300, and it is also possible to download all or part of it as needed.
[0092] Furthermore, the program for executing the processes described in this embodiment may be stored in a non-transitory computer-readable recording medium, or may be configured as a program product. Here, this "recording medium" includes memory cards, USB (Universal Serial Bus) memories, SD (Secure Digital) cards, flexible disks, magneto-optical disks, ROMs, EPROMs (Erasable Programmable Read Only Memory), EEPROMs (registered trademark) (Electrically Erasable and Programmable Read Only Memory), CD-ROMs (Compact Disk Read Only Memory), MOs (Magneto-Optical disks), DVDs (Digital Versatile Disks), and more. This includes any "portable physical medium" such as a Blu-ray Disc, a DVD player, a DVD player, a Blu-ray Disc, etc.
[0093] Furthermore, a "program" is a data processing method written in any language or description method, and does not matter whether it is in the form of source code or binary code. Note that a "program" is not necessarily limited to a single structure, but also includes a structure that is distributed as multiple modules or libraries, or a structure that achieves its function by cooperating with a separate program, such as an OS. Note that the specific configuration and reading procedure for reading a recording medium in each device shown in this embodiment, as well as the installation procedure after reading, can use well-known configurations and procedures.
[0094] The various databases stored in the memory unit are storage means such as memory devices such as RAM and ROM, fixed disk devices such as hard disks, flexible disks, and optical disks, and store various programs, tables, databases, and web page files used for various processes and providing websites.
[0095] The polymer classification apparatus 100 may be configured as an information processing device such as a known personal computer or workstation, or may be configured as an information processing device connected to any peripheral device. The polymer classification apparatus 100 may also be implemented by installing software (including programs, data, etc.) that causes the device to perform the processes described in this embodiment.
[0096] Furthermore, the specific form of distribution and integration of the devices is not limited to that shown in the drawings, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various additions or functional loads. In other words, the above-mentioned embodiments can be implemented in any combination, or embodiments can be implemented selectively.
[0097] 100 Polymer classification device 102 Control unit 102a Analysis and acquisition unit 102b Classification unit 106 Storage unit 106a Polymer database 112 Input / output unit 300 Network
Claims
1. A polymer classification device comprising a memory unit and a control unit, wherein the memory unit comprises: polymer storage means for storing modification correspondence data that links the type of modification applied to a polymer with the property change value resulting from the modification, and analytical data that links the type of modification of polymer fragments obtained by fragmenting the polymer with the modification rate, and the control unit comprises: classification means for acquiring classification data that classifies the polymers into multiple groups according to their properties based on the modification correspondence data and the analytical data.
2. The polymer classification device according to claim 1, characterized in that the polymer is a molecule that is a protein, peptide, nucleic acid or sugar chain, an artificially modified form of such a molecule, a complex of such a molecule and / or such an artificially modified form, a biopharmaceutical, a manufacturing aid, a therapeutic nucleic acid, or a carrier protein.
3. The polymer classification device described in claim 1, characterized in that the polymer storage means further stores three-dimensional structure data that sets three-dimensional structural changes of the polymer according to the modification, and the classification means acquires the classification data that classifies the polymer into a plurality of groups according to the properties based on the three-dimensional structure data, the modification correspondence data, and the analysis data.
4. The polymer classification device described in claim 1, characterized in that the classification means assigns a random number to each position where each modification occurs in the polymer molecule, and if the random number is within a range set by the modification rate of the modification corresponding to the random number based on the analysis data, performs a modification determination process for each position in the polymer molecule where each modification occurs to determine whether or not the polymer molecule contains the modification, and performs a property determination process for a predetermined number of polymers to determine the property resulting from the presence or absence of each modification in the polymer molecule based on the modification correspondence data, thereby obtaining the classification data in which the predetermined number of polymers are classified into multiple groups according to the property.
5. The polymer classification device described in claim 4, characterized in that the classification means assigns a random number to each position where each modification occurs in the polymer molecule, and if the random number is within a range set by the modification rate of the modification corresponding to the random number based on the analysis data, performs the modification determination process for each position where each modification occurs in the polymer molecule to determine whether or not the polymer molecule contains the modification, and performs the property determination process for the specified number of polymers to determine the single-molecule charge change, which is the sum of the charge change values, which are the property change values caused by the modifications contained in the polymer molecule, based on the modification correspondence data, thereby obtaining the classification data in which the specified number of polymers are classified into multiple groups according to the single-molecule charge change.
6. The polymer classification device described in claim 4, characterized in that the polymer is a protein, the polymer fragment is a peptide, the property change value is a charge change value, the property caused by the presence or absence of each modification in one polymer molecule is a single molecule charge change obtained by adding up the charge change values caused by each modification contained in one protein molecule, and the classification data is a profile of charge variants.
7. The polymer classification device according to claim 5 or 6, wherein the single molecule charge change is calculated using Equation 1. (s: number of molecule to be generated, i: number of position where modification occurs (i = 1, 2, 3, ..., m), j: number of modification species occurring at position i (j = 1, 2, 3, ..., l), ΔZ s : Charge change value of molecule s1, R s,i : a random number corresponding to the modification position i of the molecule s1, p i,j : modification rate of modification type j at modification position i, Δz i,j : charge change value when modification species j occurs at modification position i) 8. The polymer classification device according to claim 6, wherein the classification data includes aggregate data that aggregates the proteins by valence.
9. The polymer classification device described in claim 1, characterized in that the classification means calculates the abundance ratio of each molecular variant that may be present in the polymer according to the combination of modifications based on the analysis data, identifies the properties of each molecular variant based on the modification correspondence data, and calculates the abundance ratio of each property based on the abundance ratio of each molecular variant and the property of each molecular variant, thereby obtaining the classification data in which the polymer is classified into multiple groups according to the properties.
10. The polymer classification device described in claim 9, characterized in that the classification means calculates the abundance ratio of each molecular variant that may be present in the polymer according to the combination of the type of modification and the modification rate based on the analysis data, identifies a single-molecule charge change that is the sum of the charge change values, which are the property change values caused by each modification contained in each molecular variant, based on the modification correspondence data, and calculates the abundance ratio of each single-molecule charge change based on the abundance ratio of each molecular variant and the single-molecule charge change of each molecular variant, thereby obtaining the classification data that classifies the polymer into multiple groups according to the single-molecule charge change.
11. The polymer classification device according to claim 9 or 10, characterized in that the abundance ratio of each molecular variant is calculated using Equation 2 and Equation 3. (t: number of the type of molecular variant, i: number of the position where the modification occurs (i = 1, 2, 3, ..., m), j: number of the modification species occurring at position i (j = 1, 2, 3, ..., l), P t : the abundance ratio of molecular variant t, c t,i,j : 1 or 0 (if molecular variant t has modification j at position i = 1, if molecular variant t does not have modification j at position i = 0 (note that there is only one type of modification at position i of molecular variant t, so c t,i,j The sum of j is 1 or 0), p i,j : modification rate of modification type j at modification position i) (t: number of the type of molecular variant, i: number of the position where the modification occurs (i = 1, 2, 3, ..., m), j: number of the modification species occurring at position i (j = 1, 2, 3, ..., l), ΔZ t : charge change value of molecular variant t, c t,i,j : 1 or 0 (if molecular variant t has modification j at position i = 1, if molecular variant t does not have modification j at position i = 0 (note that there is only one type of modification at position i of molecular variant t, so c t,i,j The sum of j is 1 or 0), Δz i,j : the charge change value caused by modification species j at modification position i) 12. The polymer classification device of claim 9, wherein the polymer is a protein, the polymer fragment is a peptide, the property change value is a charge change value, and the classification data is a charge profile.
13. The polymer classification device described in claim 6 or 12, characterized in that the analytical data is the analytical result obtained by a method of fragmenting the protein into peptides and obtaining the type of modification and the modification rate for the protein.
14. The polymer sorting device according to claim 6 or 12, wherein the protein is an antibody, an antibody-drug conjugate, or an antibody-like molecule.
15. The polymer classification device described in claim 14, characterized in that the modification is one or a combination selected from the group consisting of deamidation of asparagine, cyclization of N-terminal glutamine, cyclization of N-terminal glutamic acid, deletion of C-terminal lysine, remaining of a sequence derived from an additional sequence at the N-terminus, acetylation of an amino acid at the N-terminus, and amidation of an amino acid at the C-terminus.
16. The polymer classification device of claim 14, wherein the modification is one or a combination selected from the group consisting of glycation of lysine, glycosylation of asparagine, succinimidation of asparagine, oxidation of asparagine to isoaspartic acid, succinimidation of aspartic acid, isomerization of aspartic acid, glycosylation of threonine, glycosylation of serine, and deamidation of glutamine.
17. The polymer classification device described in claim 14, characterized in that the modification is one or a combination selected from the group consisting of advanced glycation end products of lysine, succinylation of lysine, remaining sequences derived from additional sequences at the C-terminus, oxidation of cysteine, phosphorylation of serine, phosphorylation of threonine, and phosphorylation of tyrosine.
18. The polymer classification device described in claim 1, characterized in that the modification is one or a combination selected from the group consisting of mutation of a monomer that is a building block of the polymer, deletion of a monomer, introduction of a non-natural monomer, and introduction of an artificial modification.
19. A polymer classification method to be executed by a polymer classification device having a memory unit and a control unit, wherein the memory unit comprises: polymer storage means for storing modification correspondence data that links the types of modifications made to polymers with the property change values caused by the modifications, and analysis data that links the types of modifications made to polymer fragments obtained by fragmenting the polymer with the modification rates; and the polymer classification method characterized by including: a classification step, executed by the control unit, of acquiring classification data that classifies the polymers into multiple groups according to their properties based on the modification correspondence data and the analysis data.
20. A polymer classification program to be executed by a polymer classification device having a memory unit and a control unit, wherein the memory unit comprises: a polymer storage means for storing modification correspondence data that links the type of modification on a polymer with the property change value caused by the modification, and analysis data that links the type of modification of polymer fragments obtained by fragmenting the polymer with the modification rate; and the control unit executes a classification step that acquires classification data that classifies the polymer into multiple groups according to its properties based on the modification correspondence data and the analysis data.
Citation Information
Patent Citations
Measuring mass spectrum method of
JP2005099021A
Intact mass reconstruction from peptide level data and facilitated comparison with experimental intact observation
US20210048440A1