A method, apparatus, equipment and medium for identifying sugar types
By using glycoform identification methods and database matching technology to automate the processing of mass spectrometry data, the problems of low efficiency and high cost in existing glycoform identification technologies have been solved, achieving efficient and accurate glycoform identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-14
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies suffer from low efficiency and high labor costs in identifying glycoforms, especially in the process of mass spectrometry data interpretation and glycoform pairing, where calculation errors are prone to occur, and their versatility is low.
A method for identifying glycoforms is provided. By acquiring user-input configuration parameters and combining them with a pre-established glycoform database, the method extracts the mass-to-charge ratio array of single ions of free polysaccharides and matches them in the database, thus automating the glycoform identification process. This method includes the application of liquid chromatography and mass spectrometry analysis techniques.
It improves the efficiency and accuracy of glycoform identification, significantly reduces labor costs, enables the application of automated analysis tools, and reduces the experience threshold.
Smart Images

Figure CN116242903B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of bioinformatics technology, and in particular to a method, apparatus, equipment and medium for glycoform identification. Background Technology
[0002] Glycosylation is the process by which sugars are transferred to proteins by glycosyltransferases, forming glycosidic bonds with amino acid residues on the protein. Essentially, it's the process of attaching sugars to proteins or lipids. Proteins undergoing glycosylation become glycoproteins. Glycosylation is an important post-translational modification of proteins, playing a functional role in regulating protein expression and aiding in protein folding. In the pharmaceutical field, glycosylation significantly impacts the efficacy, stability, and immunogenicity of protein drugs, serving as a critical quality attribute throughout all stages of drug development; therefore, meticulous characterization is essential during the development process.
[0003] In related technologies, there is a need for the identification and recognition of glycoforms. Currently, the mainstream strategy is mass spectrometry to determine the glycoforms present in a sample. However, in practical applications, it has been found that when analyzing mass spectrometry data to identify glycoforms, the K+ introduced during the labeling process of free polysaccharides presents challenges. + NH4 + and Na + Interference from single or multiple adducts means that even after deconvolution, the molecular weight of adduct-free molecules still needs to be manually calculated, requiring extensive manual calculations to identify and eliminate interfering peaks. With the development of bioinformatics, recombinant proteins, fusion proteins, tandem scFv, or IgM antibodies often exhibit multi-site, multi-type, multi-antenna, and structurally complex glycoforms. Traditional methods of manually interpreting mass spectrometry data and then searching online libraries are inefficient and prone to computational errors. Furthermore, they require extensive experience for counterfeit detection, resulting in high labor costs and low identification efficiency and accuracy. Summary of the Invention
[0004] The purpose of this application is to at least partially solve one of the technical problems existing in the related art.
[0005] Therefore, one object of the embodiments of this application is to provide a method, apparatus, device and medium for identifying sugar form.
[0006] To achieve the above-mentioned technical objectives, the technical solutions adopted in the embodiments of this application include:
[0007] On the one hand, embodiments of this application provide a method for identifying glycoforms, the method comprising:
[0008] The configuration parameters input by the user are obtained and associated with a pre-established glycoform database. The configuration parameters include first-type information of the labeling reagent, mass deviation data, and second-type information of the adduct. The glycoform database includes molecular weight information and glycoform information of various sugar residues.
[0009] Based on the mass spectrometry data corresponding to the free polysaccharide, extract the single ion mass-to-charge ratio array corresponding to the free polysaccharide; the single ion mass-to-charge ratio array includes several single ion mass-to-charge ratio values; the free polysaccharide refers to the free polysaccharide obtained by unlinking the polysaccharide from the protein in the target sample; the mass spectrometry data refers to the mass spectrometry data obtained by labeling and separating the free polysaccharide with labeling reagents and then analyzing it using mass spectrometry technology.
[0010] In the glycoform database, the glycoforms corresponding to the free polysaccharides are searched and matched according to the single ion mass-to-charge ratio array to determine the identification results of the glycoforms contained in the target sample.
[0011] In addition, the sugar form identification method according to the above embodiments of this application may also have the following additional technical features:
[0012] Furthermore, in one embodiment of this application, free polysaccharides are obtained by enzymatic digestion or chemical methods to decouple the polysaccharides from the proteins.
[0013] Furthermore, in one embodiment of this application, the free polysaccharides in the target sample are separated by liquid chromatography (LC) and capillary electrophoresis (CE).
[0014] Furthermore, in one embodiment of this application, the labeling reagent is any one of 2-aminobenzoamide (2-AB), 2-aminobenzoic acid (2-AA), RapiFluor-MS, InstantPC, and procainamide.
[0015] Furthermore, in one embodiment of this application, the type of the adduct is K. + NH4 + Na + Any one of them.
[0016] Furthermore, in one embodiment of this application, the extraction of the single ion mass-to-charge ratio array corresponding to the free polysaccharide includes:
[0017] Obtain the mass spectrum corresponding to the free polysaccharide based on the mass spectrometry data;
[0018] From the mass spectrum, several mass-to-charge ratio values of individual ions are selected according to their abundance to obtain the single-ion mass-to-charge ratio array; the single-ion mass-to-charge ratio array includes at least the mass-to-charge ratio of the single ion with the highest abundance and the mass-to-charge ratio values of the isotopes adjacent to the single ion with the highest abundance.
[0019] Furthermore, in one embodiment of this application, the step of extracting the single ion mass-to-charge ratio array corresponding to the free polysaccharide further includes: magnifying the contour of the mass spectrum.
[0020] Furthermore, in one embodiment of this application, the method further includes the following steps:
[0021] According to preset dimensions, the various sugar types in the sugar type database are divided to obtain multiple sugar type data tables.
[0022] Furthermore, in one embodiment of this application, the step of searching and matching the glycoform corresponding to the free polysaccharide in the glycoform database according to the single ion mass-to-charge ratio array includes:
[0023] Based on the target sample type and the cell type expressing the target sample, a target glycoform data table is determined from the plurality of glycoform data tables;
[0024] In the target glycoform data table, the glycoform corresponding to the free polysaccharide is searched and matched according to the single ion mass-to-charge ratio array.
[0025] Furthermore, in one embodiment of this application, the method further includes the following steps:
[0026] If no matching identification result is found in the target glycoform data table, a prompt message will be output;
[0027] The prompt message is used to remind the user to reselect the target sugar type data table or re-enter the configuration parameters.
[0028] On the other hand, embodiments of this application provide a glycoform identification device, the device comprising:
[0029] A configuration module is used to obtain configuration parameters input by the user and associate the configuration parameters with a pre-established glycoform database. The configuration parameters include first-type information of the labeling reagent, mass deviation data, and second-type information of the adduct. The glycoform database includes molecular weight information and glycoform information of various sugar residues.
[0030] The extraction module is used to extract the single ion mass-to-charge ratio array corresponding to the free polysaccharide based on the mass spectrometry data corresponding to the free polysaccharide; the single ion mass-to-charge ratio array includes several single ion mass-to-charge ratio values; the free polysaccharide refers to the free polysaccharide obtained by unlinking the polysaccharide from the protein in the target sample; the mass spectrometry data refers to the mass spectrometry data obtained by labeling and separating the free polysaccharide with labeling reagents and then analyzing it using mass spectrometry technology.
[0031] The matching module is used to search and match the glycoforms corresponding to the free polysaccharides in the glycoform database according to the single ion mass-to-charge ratio array, and determine the identification result of the glycoforms contained in the target sample.
[0032] Furthermore, in one embodiment of this application, the apparatus further includes:
[0033] The analysis module is used to acquire mass spectrometry data of free polysaccharides in the target sample through mass spectrometry analysis technology.
[0034] On the other hand, embodiments of this application provide a computer device, including:
[0035] At least one processor;
[0036] At least one memory for storing at least one program;
[0037] When the at least one program is executed by the at least one processor, the at least one processor implements the above-described method for identifying glycoforms.
[0038] On the other hand, embodiments of this application also provide a computer-readable storage medium storing a processor-executable program, which, when executed by a processor, is used to implement the aforementioned method for identifying glycoforms.
[0039] The advantages and beneficial effects of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application:
[0040] This application discloses a method for glycoform identification, comprising: acquiring configuration parameters input by a user and associating the configuration parameters with a pre-established glycoform database; the configuration parameters include first-type information of the labeling reagent, mass deviation data, and second-type information of the adduct; the glycoform database includes molecular weight information and glycoform information of various sugar residues; extracting a single-ion mass-to-charge ratio array corresponding to the free polysaccharide based on the mass spectrometry data corresponding to the free polysaccharide; the single-ion mass-to-charge ratio array includes several single-ion mass-to-charge ratio values; free polysaccharide refers to the free polysaccharide obtained by unlinking the polysaccharide from the protein in the target sample; the mass spectrometry data refers to the mass spectrometry data obtained by mass spectrometry analysis after labeling and separating the free polysaccharide with the labeling reagent; and searching and matching the glycoforms corresponding to the free polysaccharide in the glycoform database according to the single-ion mass-to-charge ratio array to determine the identification result of the glycoform contained in the target sample. This method can solve the work of interpreting the mass spectrometry data of free polysaccharide, matching glycoforms, and translating sugar names and structures in one stop based on automated analysis tools. This method can automate the entire glycoform identification process, improve identification efficiency and accuracy, and significantly reduce the experience threshold for free polysaccharide analysis, thereby helping to reduce labor costs. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following description is provided with accompanying drawings of the relevant technical solutions in the embodiments of this application or the prior art. It should be understood that the accompanying drawings described below are only for the purpose of clearly illustrating some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0042] Figure 1 This is a flowchart illustrating a method for identifying glycoforms provided in an embodiment of this application;
[0043] Figure 2 This is a data illustration of a glycoform database provided in an embodiment of this application;
[0044] Figure 3 This is a chromatogram for LC-MS glycoform identification provided in an embodiment of this application;
[0045] Figure 4 This is one of the embodiments provided in this application. Figure 3 The mass spectrum corresponding to the response peak with a retention time of 5.27 min;
[0046] Figure 5 This application provides an embodiment of a method for... Figure 4 The mass spectrum after magnification;
[0047] Figure 6This is a schematic diagram of the matching results of a glycoform database provided in an embodiment of this application;
[0048] Figure 7 This is a chromatogram of a completed glycoform labeling provided in an embodiment of this application;
[0049] Figure 8 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0050] The present application will be further described below with reference to the accompanying drawings and specific embodiments. The described embodiments should not be considered as limitations on the present application, and all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of the present application.
[0051] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0053] Glycosylation is the process by which sugars are transferred to proteins by glycosyltransferases, forming glycosidic bonds with amino acid residues on the protein. Essentially, it's the process of attaching carbohydrates to proteins or lipids. Proteins undergoing glycosylation become glycoproteins. Glycosylation is an important post-translational modification of proteins, playing a functional role in regulating protein expression and aiding in protein folding. In the pharmaceutical field, glycosylation significantly impacts the efficacy, stability, and immunogenicity of protein drugs, serving as a critical quality attribute throughout all stages of drug development; therefore, meticulous characterization is essential for its redevelopment.
[0054] In related technologies, there is a need for the identification and recognition of glycoforms. Currently, the mainstream strategy is to use mass spectrometry (MS) to analyze the MS data using mass spectrometry software to determine the glycoforms present in the sample. For example, commonly used MS spectrometry software such as MaxEnt3 or Bayes can be used to deconvolve free glycan peaks. However, because free glycans may introduce K+ during the labeling process... + NH4 + and Na +Single or multiple adducts mean that even after deconvolution, the molecular weight of the adduct-free component still needs to be manually calculated, requiring a large amount of manual calculation to identify and eliminate interfering peaks. Mass spectrometry software is often exclusive, with its basic package typically containing only glycosylation solutions for its own labeled reagents, making it incompatible with analytical applications using other labeled reagents, resulting in low versatility. Commonly used glycoform search software such as Glycoworkbench, Glycomod, and Glycam can match glycoform names and structures based on the molecular weight of the labeled free glycan, but due to the use of complex and obscure nomenclature and structural representations, manual translation and redrawing of glycoform structures are often required, making database searching and interpretation difficult.
[0055] In summary, the analysis shows that the traditional method of manually interpreting mass spectrometry data and then searching the database online for sugar type identification is inefficient, prone to calculation errors, and requires extensive experience in counterfeiting detection, resulting in high labor costs and low identification efficiency and accuracy.
[0056] In view of this, this application provides a method for glycoform identification. This method, based on automated analysis tools, can solve the work of interpreting free polysaccharide mass spectrometry data, glycoform pairing, sugar name and structure translation in one stop. This method can automate the entire glycoform identification process, improve identification efficiency and accuracy, and significantly reduce the experience threshold for free polysaccharide analysis, thereby helping to reduce labor costs.
[0057] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a method for identifying glycoforms provided in an embodiment of this application. (Refer to...) Figure 1 This method for identifying sugar types includes, but is not limited to:
[0058] Step 110: Obtain the configuration parameters input by the user and associate the configuration parameters with a pre-established glycoform database; the configuration parameters include the first type information of the labeling reagent, the mass deviation data, and the second type information of the adduct, and the glycoform database includes the molecular weight information and glycoform information of various sugar residues;
[0059] Step 120: Based on the mass spectrometry data corresponding to the free polysaccharide, extract the single ion mass-to-charge ratio array corresponding to the free polysaccharide; the single ion mass-to-charge ratio array includes several single ion mass-to-charge ratio values; the free polysaccharide refers to the free polysaccharide obtained by unlinking the polysaccharide from the protein in the target sample; the mass spectrometry data refers to the mass spectrometry data obtained by labeling and separating the free polysaccharide with labeling reagents and then analyzing it using mass spectrometry.
[0060] Step 130: In the sugar type database, search and match the sugar type corresponding to the free polysaccharide according to the single ion mass-to-charge ratio array to determine the identification result of the sugar type contained in the target sample.
[0061] In this application embodiment, a method for identifying sugar form is provided. This method can design and develop a corresponding sugar form database and related matching algorithm based on the Excel VBA platform, and determine the sugar form of free polysaccharides through mass spectrometry analysis technology, thereby realizing the identification of sugar form of sugar contained in the sample.
[0062] Specifically, in this embodiment, the sample to be identified can be referred to as the target sample. Mass spectrometry is an analytical method for measuring the mass-to-charge ratio (m / z) of ions. Its basic principle is to ionize the free glycans in the sample in an ion source, generating positively charged ions with different mass-to-charge ratios. These ions are then accelerated by an electric field to form an ion beam, which enters the mass analyzer to determine its molecular weight. It is understood that since sugars themselves lack chromophores, they are often labeled or derivatized to more effectively detect sugar chains during separation, purification, and structural identification. Therefore, in this embodiment, labeling reagents can be used to label the free glycans in the target sample during mass spectrometry analysis. In some embodiments, the types of labeling reagents may include, but are not limited to, 2-AB, RapiFluor-MS, InstantPC, etc. In actual use, the glycans are first decoupled from the protein and enriched, and then any labeling reagent is selected to label the free glycans. In some embodiments, enzymatic digestion is used to detach the glycan from the protein. The enzymes used for digesting the target sample include any one of N-glycoprotein deglycosylation enzymes (such as PNGase A and PNGase F) and endoglucosidases (such as Endo S, Endo H, and Endo F). Before mass spectrometry analysis, the free glycan needs to be separated. Methods for separating free glycan include, but are not limited to, liquid chromatography (LC) and capillary electrophoresis (CE).
[0063] In this embodiment, tandem mass spectrometry (LC-MS) is used to identify the glycoform of the target sample. First, the free polysaccharides in the target sample are separated by liquid chromatography, and then the target sample is sent to a mass spectrometer for mass spectrometric identification. A mass spectrometer generally includes an ion source and a mass analyzer. In this application, the specific ionization method and the method for measuring the m / z ratio of ions used in the mass spectrometry analysis are not limited; they can be implemented with reference to existing technologies. Mass spectrometry data corresponding to the target sample can be obtained through mass spectrometry analysis. This mass spectrometry data is generally in the form of a chromatogram, and a typical mass spectrometry... Figure 1The horizontal axis typically represents the m / z ratio, and the vertical axis represents the relative intensity of the peaks (expressed as a percentage). The strongest peak will have a relative intensity of 100. Based on mass spectrometry data, the ion mass-to-charge ratio of various saccharides contained in the target sample can be determined. Based on the ion mass-to-charge ratio information, the corresponding glycoform information can be further determined.
[0064] It should be noted that all elements in nature contain isotopes. When identifying glycoforms, the molecular weight of the glycoform can be calculated by deconvolving the mass-charge ratio array of isotopes of the same ion, thereby matching the glycan type. A single isotope represents the composition of the analyte. 12 C, 1 H, 16 O, 14 The precise molecular weights of these most abundant nuclides, such as N, are determined by mass spectrometry, as free polysaccharides generally have a molecular weight less than 5000 Da, and the resolution is sufficient to distinguish isotopic peaks. 12 C and 13 The natural abundance ratio of carbon (C) is approximately 99:1. However, if the molecular weight of the analyte is sufficiently large and the amount of carbon (C) is sufficiently high, the probability of obtaining a certain number of carbon atoms will vary. 13 C forms a series of isotopes with uniform spacing of 1 Da. During the ionization process in the mass spectrum, after the isotopes are tagged, a series of continuous m / z peaks are formed. Based on the spacing pattern of the m / z peaks of the isotopes, the valence state of the ion can be deduced. Deconvolution calculation can obtain the molecular weight m of the ion, thus allowing for a relatively accurate determination of the glycoform.
[0065] Of course, it should also be added that when labeling the target sample to be identified, the labeling reagent itself will also cause changes in the molecular weight of the ions corresponding to the free glycosaminoglycans. Moreover, in some cases, partial adducts may be formed, for example, the type of adduct could be K. + NH4 + Na +These factors can all lead to a difference between the molecular weight of the ions corresponding to the free polysaccharide and the ions themselves. Therefore, in this embodiment of the application, to address the complex and cumbersome matching situation, a glycoform database can be pre-established. This database can include molecular weight information of sugar residues of various known sugars, as well as the glycoform information corresponding to those sugars. Here, the glycoform information may include, but is not limited to, its name and specific structural information. Then, for labeling reagents and adducts that may vary depending on the actual situation, the user can select them during each identification and input this information into the identification system or software. That is, during each identification, the configuration parameters input by the user can be obtained. These configuration parameters may include the type information of the labeling reagent, mass deviation data, and the type information of the adduct. The type information of the labeling reagent is recorded as the first type information, and the type information of the adduct is recorded as the second type information. After obtaining the configuration parameters input by the user, these configuration parameters can be associated with a pre-established glycoform database. Then, based on the configuration parameters input by the user, the precise molecular weight of each glycoform can be obtained during mass spectrometry analysis. Specifically, this value can be obtained by summing the molecular weights of the free polysaccharide residues, H2O, labeling reagents, and adduct ions (if any) corresponding to each glycoform.
[0066] In this embodiment, tandem mass spectrometry (LC-MS) is used to obtain the chromatogram and mass spectrometry data of the target sample. Each response peak in the chromatogram may represent a free polysaccharide contained in the target sample. Therefore, in this embodiment, the mass spectrometry data of each response peak is the mass spectrometry data of each free polysaccharide in the target sample. The mass spectrometry data of each response peak can be analyzed to determine the sugar type information corresponding to each response peak.
[0067] Specifically, in this embodiment, for each response peak, a single ion mass-to-charge ratio array corresponding to that response peak can be extracted. Here, the single ion mass-to-charge ratio array includes several adjacent mass spectrometry data peaks within a specified range. It should be noted that, generally, for a response peak at a certain location, the peak data with the highest abundance in its corresponding mass spectrum is the only and most reliable glycoform matched. However, for glycoforms with larger molecular weights, due to the increased complexity of the mass spectrometry data, it may be impossible to find the strongest peak (many valence peaks on the mass spectrum have similar intensities). If a single peak data is used for matching, it may lead to matching errors. Furthermore, due to the influence of adduct ions and other factors, the actual molecular weights of different molecules of the same glycoform may have some interference. For example, if a chromatographic peak contains a 900 Da component, due to the influence of adduct ions in mass spectrometry detection, taking the 900 Da component as an example, in the mass spectrum, this glycoform component may exist not only in the standard 900 Da form but also in the form of 900+39(+K). + ), 900+68(+2K) + ), 900+17(+NH4) + The presence of multiple detection peaks in similar regions on the spectrum due to the presence of ions (900 Da + 39 (+K)) is understandable. It's understood that a standard 900 Da component represents a single ion, and similarly, 900 + 39 (+K) + It is also a single ion. In addition, carbohydrate components are often detected in the form of non-reducing end-source internal cleavage or sugar unit shedding, such as 900-203 (loss of N-acetylglucosamine) and 900-291 (loss of sialic acid), and may also show monovalent and divalent peaks at the same time. Therefore, the actual 900 Da sugar form may exist in many related forms.
[0068] Therefore, in this embodiment, for each response peak, the corresponding mass spectrometry data can be analyzed to obtain a single ion mass-to-charge ratio (MMR) array. Calculation and matching analysis using the single ion MMR array can significantly reduce the possibility of misjudgment and improve the accuracy of identification. Here, each single ion MMR array can include at least two adjacent single ion isotope MMR values, namely the MMR of the single ion with the highest abundance and the MMR of its adjacent isotope. This application does not limit the specific number of these values.
[0069] After obtaining the mass-to-charge ratio (MMR) array of a single ion, a search and matching process can be performed in the aforementioned glycoform database based on this array. It is understood that, after being refined through configuration parameters, the aforementioned glycoform database can be adjusted to obtain the molecular weights of various glycoforms and their variants that may arise under the conditions of this identification. By inputting the names and variants of each glycoform calculated from the single ion MMR array and comparing them comprehensively, the glycoform information corresponding to each response peak can be determined. By searching and matching all the response peaks, the identification results of all glycoforms contained in the target sample can be determined.
[0070] The implementation process of the sugar type identification method provided in this application will be described below with reference to a specific application embodiment.
[0071] In this embodiment, in practical application, the relevant software and algorithm tools can be developed and implemented based on the Excel platform. The glycoform database is pre-loaded with molecular weight information and glycoform information of various known glycoforms corresponding to sugar residues. This data may include, but is not limited to, a glycoform library consisting of 171 oligosaccharides, a glycoform library of 200 Oxford nomenclature sugars, a glycoform library of 39 common N-glycans in human cells, a glycoform library of 182 common N-glycans in CHO cells, and a glycoform library of 16 common O-glycans. Furthermore, the system can include molecular weight information of various labeling reagents such as 2-AB, RapiFluor-MS, and InstantPC, as well as K... + NH4 + and Na + Molecular weight information of adducts.
[0072] Specifically, to facilitate subsequent matching and ensure more scientific and reasonable identification results, in some embodiments, the glycoform database can be divided into multiple glycoform data tables according to preset dimensions. For example, in this embodiment, six glycoform data tables can be pre-edited: sheet1 (Short name used IgG glycans), sheet2 (Oligosaccharide composition), sheet3 (N-Glycan List), sheet4 (Human host N-Glycan), sheet5 (CHO host N-Glycan), and sheet6 (O-Glycan). In some cases, if it is known in advance which glycoform data table the sugars in the target sample belong to, that glycoform data table can be used directly for matching. This improves matching efficiency and reduces data processing volume; it also reduces the possibility of misjudgments, resulting in more scientific and reasonable identification results. This completes the preliminary data preparation work.
[0073] When performing glycoform identification, users can configure parameters in the system software: select appropriate labeling reagents (e.g., 2-AB, RapiFluor-MS, and InstantPC are three mainstream reagents), with InstantPC as the default; input mass tolerance data (determined based on the accuracy of the mass spectrometer and the instrument's condition, with ±0.1 Da as the default); and input adduct information (K...). + NH4 + Na + `and null` represent considering the glycoform + adduct form during glycoform pairing. Based on user actions, configuration parameters can be obtained and associated with the aforementioned glycoform database. During subsequent matching, the glycoform database can access this configuration parameter information. In this embodiment, taking "Short name used IgG glycans" as an example, after associating the configuration parameters, refer to... Figure 2 This table can record the name of the sugar type, the corresponding sugar structure diagram, and the molecular weight after the structure is labeled.
[0074] It should be noted that in some embodiments, the user can also actively select and configure a suitable glycoform database, i.e., determine the target glycoform database from multiple glycoform databases. Specifically, this can be based on the target sample type and the cell type expressing the target sample. For example, if the target sample type is an immunoglobulin and it is expressed through CHO cells, then the "Short name used IgG glycans" target glycoform database is recommended. These can be selected based on experience accumulating common glycoforms for each cell type. Furthermore, in some embodiments, the above process can be automatically implemented by the system software without user intervention, thereby improving automation and identification efficiency.
[0075] Next, the target sample can be treated with labeling reagents, and the free polysaccharides in the target sample can be separated by liquid chromatography. Mass spectrometry data of each free polysaccharide can then be obtained using mass spectrometry. (Refer to...) Figure 3 , Figure 3 A chromatogram for glycoform identification by LC-MS is shown. The horizontal axis in this chromatogram represents retention time. Generally, free polysaccharides with the same retention time belong to the same glycoform. Taking the response peak corresponding to 5.27 min as an example, refer to... Figure 4 , Figure 4 It shows Figure 3The mass spectrum corresponding to the response peak with a retention time of 5.27 min shows that the response peak includes multiple mass spectrometry data peaks with m / z ratios between 862 and 864. In this embodiment, the mass-to-charge ratio array of a single ion corresponding to the response peak can be extracted.
[0076] Specifically, in some embodiments, extracting the single ion mass-to-charge ratio array corresponding to the response peak includes:
[0077] Obtain the mass spectrum corresponding to the response peak based on the mass spectrometry data;
[0078] From the mass spectrum, several mass-to-charge ratio values of individual ions are selected according to their abundance to obtain the single-ion mass-to-charge ratio array; the single-ion mass-to-charge ratio array includes at least the mass-to-charge ratio of the single ion with the highest abundance and the mass-to-charge ratio values of the isotopes adjacent to the single ion with the highest abundance.
[0079] In this embodiment of the application, when extracting the mass-to-charge ratio array of a single ion corresponding to the response peak, the outline of the mass spectrum can be magnified first to facilitate the identification of each existing mass spectrometry data peak. For example, see... Figure 5 ,right Figure 4 By magnifying the mass spectrum, the precise mass-to-charge ratio of the isotopes of individual ions can be determined. Of course, some low peaks with extremely low relative abundance can be ignored here. Then, the mass-to-charge ratio array for individual ions is determined. Specifically, in some embodiments, the number of data points in the mass-to-charge ratio array for individual ions can be preset, for example, set to 2. Figure 5 For example, the values 862.84309 and 863.34801 can be selected as the single ion mass-to-charge ratio array. Of course, the number of data points can be flexibly set as needed, and this application does not impose any restrictions on this.
[0080] Of course, in practical applications, for each response peak, the user can zoom in on the mass spectrum outline and input the mass-to-charge ratio of the single ion isotope with the highest abundance displayed in the mass spectrum and its immediately adjacent mass-to-charge ratio. The glycoform will be automatically matched, and the matching result can be checked. Then, the second highest abundance can be found and checked, the third highest abundance can be checked, and so on, and so on. It is understandable that the more peaks of the input mass spectrometry data, the higher the accuracy of the matching result. (Refer to...) Figure 6 Taking glycoform matching as an example, where two adjacent single ion mass-to-charge ratio data (input m / z1 and input m / z2) are input, Figure 6 A schematic diagram of the matching results is shown. Based on two sets of isotope data, input m / z1 and input m / z2, the identification results can be automatically generated with high accuracy.
[0081] Here, as explained above, when matching the glycoforms corresponding to the strong peaks, a defined target glycoform data table can be used. In some embodiments, if no matching identification result is found in the target glycoform data table, "Not found" can be displayed, and a corresponding prompt message can be output. This prompt message can be used to remind the user to consider whether the target glycoform data table is not included and needs to be changed; or whether the configuration parameters are set incorrectly and need to be re-entered before matching.
[0082] In the embodiments of this application, reference is made to Figure 7 After identifying each response peak, the structural diagrams of each glycoform obtained by copying and pasting can be marked on the peaks of the chromatogram, thus completing the identification of the glycoform.
[0083] This application embodiment also provides a glycoform identification device, the device comprising:
[0084] The analysis module is used to acquire mass spectrometry data of free polysaccharides in the target sample through mass spectrometry analysis technology;
[0085] A configuration module is used to obtain configuration parameters input by the user and associate the configuration parameters with a pre-established glycoform database. The configuration parameters include first-type information of the labeling reagent, mass deviation data, and second-type information of the adduct. The glycoform database includes molecular weight information and glycoform information of various sugar residues.
[0086] The extraction module is used to extract the single ion mass-to-charge ratio array corresponding to the free polysaccharide based on the mass spectrometry data corresponding to the free polysaccharide; the single ion mass-to-charge ratio array includes several single ion mass-to-charge ratio values; the free polysaccharide refers to the free polysaccharide obtained by unlinking the polysaccharide from the protein in the target sample; the mass spectrometry data refers to the mass spectrometry data obtained by labeling and separating the free polysaccharide with labeling reagents and then analyzing it using mass spectrometry technology.
[0087] The matching module is used to search and match the glycoforms corresponding to the free polysaccharides in the glycoform database according to the single ion mass-to-charge ratio array, and determine the identification result of the glycoforms contained in the target sample.
[0088] Understandable Figure 1 The content of the saccharide form identification method embodiment shown is applicable to the saccharide form identification device embodiment. The specific functions implemented by the saccharide form identification device embodiment are the same as those shown in the embodiment. Figure 1 The illustrated method for sugar type identification is the same as the one shown in the embodiment, and the beneficial effects achieved are the same. Figure 1 The beneficial effects achieved by the illustrated method for sugar type identification are also the same.
[0089] Reference Figure 8 This application also discloses a computer device, including:
[0090] At least one processor 201;
[0091] At least one memory 202 is used to store at least one program;
[0092] When at least one program is executed by at least one processor 201, such that at least one processor 201 performs as follows: Figure 1 An example of a sugar type identification method is shown.
[0093] It is understandable that, such as Figure 1 The content of the sugar type identification method embodiment shown is applicable to the computer device embodiment, and the specific functions implemented by the computer device embodiment are the same as those shown. Figure 1 The method for identifying sugar types shown is the same as the embodiment described above, and the beneficial effects achieved are the same as those described above. Figure 1 The beneficial effects achieved by the illustrated method for sugar type identification are also the same.
[0094] This application also discloses a computer-readable storage medium storing a processor-executable program, which, when executed by a processor, is used to implement, for example... Figure 1 An example of a sugar type identification method is shown.
[0095] It is understandable that, such as Figure 1 The content of the glycoform identification method embodiment shown is applicable to the embodiment of this computer-readable storage medium. The specific functions implemented by the embodiment of this computer-readable storage medium are the same as those shown below. Figure 1 The method for identifying sugar types shown is the same as the embodiment described above, and the beneficial effects achieved are the same as those described above. Figure 1 The beneficial effects achieved by the illustrated method for sugar type identification are also the same.
[0096] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order shown in the operation diagrams. For example, depending on the functions / operations involved, two consecutively shown blocks may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order. Furthermore, the embodiments presented and described in the flowcharts of this application are provided by way of example to provide a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are contemplated in which the order of various operations is changed and sub-operations described as part of a larger operation are executed independently.
[0097] Furthermore, although this application is described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated into a single physical system and / or software module, or one or more functions and / or features may be implemented in a separate physical system or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding this application. Rather, given the properties, functions, and internal relationships of the various functional modules in the system disclosed herein, the actual implementation of the module will be understood within the scope of ordinary skill of an engineer. Therefore, those skilled in the art can implement the application set forth in the claims using ordinary techniques without excessive experimentation. It is also understood that the specific concepts disclosed are merely illustrative and not intended to limit the scope of this application, which is determined by the full scope of the appended claims and their equivalents.
[0098] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0099] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, system, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, system, or device). For the purposes of this specification, "computer-readable medium" can mean any system that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, system, or device.
[0100] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections with one or more wires (electronic systems), portable computer disk drives (magnetic systems), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic systems, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0101] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0102] In the foregoing description of this specification, the references to terms such as "one embodiment," "another embodiment," or "some embodiments," etc., indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0103] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.
[0104] The foregoing has provided a detailed description of the preferred embodiments of this application. However, this application is not limited to these embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.
[0105] In the description of this specification, the references to terms such as "one embodiment," "another embodiment," or "some embodiments," etc., indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0106] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.
Claims
1. A method for identifying sugar form, characterized in that, The method includes: The configuration parameters input by the user are obtained and associated with a pre-established glycoform database. The configuration parameters include first-type information of the labeling reagent, mass deviation data, and second-type information of the adduct. The glycoform database includes molecular weight information and glycoform information of various sugar residues. Based on the mass spectrometry data corresponding to the free polysaccharide, extract the single ion mass-to-charge ratio array corresponding to the free polysaccharide; the single ion mass-to-charge ratio array includes several single ion mass-to-charge ratio values; the free polysaccharide refers to the free polysaccharide obtained by unlinking the polysaccharide from the protein in the target sample; the mass spectrometry data refers to the mass spectrometry data obtained by labeling and separating the free polysaccharide with labeling reagents and then analyzing it using mass spectrometry technology. In the glycoform database, the glycoforms corresponding to the free polysaccharide are searched and matched according to the single ion mass-to-charge ratio array to determine the identification result of the glycoform contained in the target sample; The extraction of the single ion mass-to-charge ratio array corresponding to the free polysaccharide includes: Obtain the mass spectrum corresponding to the free polysaccharide based on the mass spectrometry data; From the mass spectrum, several mass-to-charge ratio values of individual ions are selected according to their abundance to obtain the mass-to-charge ratio array of individual ions; the mass-to-charge ratio array of individual ions includes at least the mass-to-charge ratio of the most abundant individual ion and the mass-to-charge ratio values of the isotopes adjacent to the most abundant individual ion mass-to-charge ratio.
2. The method for identifying glycoform according to claim 1, characterized in that, The labeling reagent is any one of 2-AB, 2-AA, RapiFluor-MS, InstantPC, and procainamide.
3. The method for identifying glycoform according to claim 1, characterized in that, The type of adduct is K. + NH4 + Na + Any one of them.
4. A method for identifying glycoforms according to any one of claims 1-3, characterized in that, The method further includes the following steps: According to preset dimensions, the various sugar types in the sugar type database are divided to obtain multiple sugar type data tables.
5. The method for identifying glycoform according to claim 4, characterized in that, The process of searching and matching the glycoform corresponding to the free polysaccharide in the glycoform database based on the single ion mass-to-charge ratio array includes: Based on the target sample type and the cell type expressing the target sample, a target glycoform data table is determined from the plurality of glycoform data tables; In the target glycoform data table, the glycoform corresponding to the free polysaccharide is searched and matched according to the single ion mass-to-charge ratio array.
6. The method for identifying glycoform according to claim 5, characterized in that, The method further includes the following steps: If no matching identification result is found in the target glycoform data table, a prompt message will be output; The prompt message is used to remind the user to reselect the target sugar type data table or re-enter the configuration parameters.
7. A sugar form identification device, characterized in that, The device includes: A configuration module is used to obtain configuration parameters input by the user and associate the configuration parameters with a pre-established glycoform database. The configuration parameters include first-type information of the labeling reagent, mass deviation data, and second-type information of the adduct. The glycoform database includes molecular weight information and glycoform information of various sugar residues. The extraction module is used to extract the single ion mass-to-charge ratio array corresponding to the free polysaccharide based on the mass spectrometry data corresponding to the free polysaccharide; the single ion mass-to-charge ratio array includes several single ion mass-to-charge ratio values; the free polysaccharide refers to the free polysaccharide obtained by unlinking the polysaccharide from the protein in the target sample; the mass spectrometry data refers to the mass spectrometry data obtained by labeling and separating the free polysaccharide with labeling reagents and then analyzing it using mass spectrometry technology. The matching module is used to search and match the glycoforms corresponding to the free polysaccharide in the glycoform database according to the single ion mass-to-charge ratio array, and determine the identification result of the glycoform contained in the target sample; The extraction of the single ion mass-to-charge ratio array corresponding to the free polysaccharide includes: Obtain the mass spectrum corresponding to the free polysaccharide based on the mass spectrometry data; From the mass spectrum, several mass-to-charge ratio values of individual ions are selected according to their abundance to obtain the mass-to-charge ratio array of individual ions; the mass-to-charge ratio array of individual ions includes at least the mass-to-charge ratio of the most abundant individual ion and the mass-to-charge ratio values of the isotopes adjacent to the most abundant individual ion mass-to-charge ratio.
8. A computer device, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements a glycoform identification method as described in any one of claims 1-6.
9. A computer-readable storage medium storing a processor-executable program, characterized in that: The processor-executable program, when executed by the processor, is used to implement a glycoform identification method as described in any one of claims 1-6.
Citation Information
Patent Citations
Mass spectrometry data processing device
CN109661574A
Sugar chain structure analysis device and sugar chain structure analysis program
CN112805560A