A data processing method for chromatography mass spectrometry analysis

By automatically processing chromatographic and mass spectrometry data, the problem of overlapping peaks and mass spectrometry signal processing in complex samples is solved, and the accurate analysis of compound composition is achieved, which improves analysis efficiency and accuracy.

CN114487245BActive Publication Date: 2025-05-02SUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210008618.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-06
Publication Date
2025-05-02
Estimated Expiration
2042-01-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively separate and analyze overlapping chromatographic peaks present in complex samples, and to process a large number of overlapping mass spectrometry signals, resulting in difficulty in qualitative and quantitative analysis.

Method used

Through a data processing method for chromatographic mass spectrometry analysis, chromatographic and mass spectrometry data are automatically processed, including overlapping peak fitting of chromatographic peaks and attribution correction of mass spectrometry signals, to accurately judge the composition ratio and structural composition of compounds in complex samples.

Benefits of technology

Accurate composition analysis of multiple compounds in complex samples is achieved, reducing the work burden of analysts and improving the accuracy and efficiency of the analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114487245B_ABST
    Figure CN114487245B_ABST
Patent Text Reader

Abstract

The present invention discloses a data processing method for chromatography-mass spectrometry analysis, including overlapping peak fitting of a chromatogram and attribution correction of mass spectrometry data. The present invention has the advantage that the composition ratio of multiple compounds in a complex sample and the structural composition of each component can be accurately determined by automatically processing chromatography-mass spectrometry analysis data, so that analysts are no longer required to perform the troublesome operation of confirming or comparing database retrieval results themselves, and the burden on analysts engaged in identification work can be greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing of compounds, and in particular to a data processing method for chromatography-mass spectrometry analysis, which processes data collected by a chromatography-mass spectrometry analysis device such as a liquid chromatography-mass spectrometry analysis device (LC-MS) composed of a liquid chromatograph and a mass spectrometry analysis device, a gas chromatography-mass spectrometry analysis device (GC-MS) composed of a gas chromatograph and a mass spectrometry analysis device, or a capillary electrophoresis-mass spectrometry analysis device (CE-MS) composed of a capillary electrophoresis instrument and a mass spectrometry analysis device, so as to identify or infer the structure of compounds contained in a sample. Background Art

[0002] As one of the means of separation and quantitative analysis of complex components, chromatography technology has the advantages of fast analysis speed, high separation efficiency, and small sample consumption. When using chromatography technology to separate complex components, under ideal experimental conditions, each single peak corresponds to one component. However, in fact, when two or more components have similar structures or properties, overlapping peaks are easily formed. In accurate quantitative research, traditional methods such as tangent method and vertical cutting method are used. Although they are fast, they have low accuracy, which brings difficulties to subsequent qualitative and quantitative analysis. Therefore, how to effectively separate overlapping chromatographic peaks is one of the important issues that need to be solved urgently.

[0003] The chemometric methods that have been continuously improved in the past few decades can accurately analyze overlapping chromatographic peaks. For example, the longitudinal iteration method starts fitting from the front and back edges far away from the overlapping area to correct the overlapping area of ​​the chromatographic peaks, but this method requires the overlapping peaks to have valley points; the algebraic peak separation methods mainly include the spectrum peak fitting algorithm based on Gaussian function, the spectrum peak fitting algorithm based on the least squares method, and the wavelet transform algorithm. These methods can obtain better calculation results and peak separation effects. However, in the peak separation process, they require certain parameter estimation and optimization, model selection and other steps, which are computationally intensive and time-consuming, and are not suitable for real-time online processing. Among the peak separation methods based on pattern recognition, the immune algorithm has a better separation effect, but it is only suitable for the separation of overlapping peaks of known components, and its application is limited.

[0004] Mass spectrometry is an efficient and sensitive technology that can provide structural information of compounds. However, for complex compounds, such as sugars, due to the existence of microscopic heterogeneity, a wide variety of compounds, difficult separation, multiple and overlapping mass spectrometry signals, and multiple charges, the efficient and accurate processing of a large amount of mass spectrometry data will form a bottleneck for high-throughput omics research. On the one hand, manual annotation of these analytical data is time-consuming and inefficient. The bigger problem is the lack of accuracy and standards, and peaks with low abundance and large data errors are easily missed. On the one hand, a series of software and methods for assisting the analysis of sugar structure information have emerged, which are fast and easy to operate. The establishment of a database has greatly reduced the difficulty brought by tedious data processing. The data obtained through database search and analysis are comprehensive, accurate and not missed. They will be graded according to the series of isotope peaks, which more intuitively reflects the credibility of peak attribution. These software and methods provide favorable support for the discovery of new sugar molecules, high-throughput and efficient research, and annotation and identification of sugar mass spectrometry data.

[0005] However, due to the complexity of oligosaccharide structure, existing research work still has limitations in the understanding of the formation mechanism of oligosaccharide mass spectrometry. Therefore, the accuracy of theoretical mass spectrometry prediction is not high, which affects the accuracy of the analysis results, and there are a large number of false positive results in the obtained data. One type of false positive data is due to serious overlap of mass spectrometry signals, incorrect charge recognition, and thus deconvolution errors, which ultimately lead to incorrect assigned composition. The molecular weight of this type of false positive data does not match the retention time. One type of false positive data is due to the fact that sulfated oligosaccharides easily lose sulfonic acid groups in the ion source. Therefore, for oligosaccharides with a low degree of sulfate, it is difficult to distinguish them from the fragment ion peaks generated by the loss of sulfonic acid groups in highly sulfated oligosaccharides when assigning them by mass spectrometry, and there is no way to determine the sugar chain composition during analysis.

[0006] In summary, it is very difficult to conduct qualitative and quantitative analysis of a large number of overlapping peaks in the chromatogram of a complex sample system. It is a great burden for analysts to manually analyze, confirm, judge, and identify compounds based on the results of mass spectrometry database retrieval. When analyzing pharmaceuticals or illegal drugs, especially in the process of analyzing a large number of complex compound systems with the same basic structural skeleton and slight differences in substituents, effective, accurate, and automated LC-MS analysis, retrieval, and analysis methods and tools are particularly important. Summary of the invention

[0007] The purpose of the present invention is to solve the above-mentioned technical problems and to provide a data processing method for chromatography-mass spectrometry analysis, which can accurately determine the composition ratio of multiple compounds in complex samples and the structural composition of each component by automatically processing the chromatography-mass spectrometry analysis data. Therefore, it is no longer necessary for analysts to perform the troublesome operation of confirming or comparing the database retrieval results themselves, and the burden on analysts engaged in identification work can be greatly reduced.

[0008] The technical solution of the present invention is: a data processing method for chromatographic mass spectrometry analysis, which determines the composition of multiple compounds in a complex sample by analyzing and processing chromatographic and mass spectrometric data, including overlapping peak fitting of the chromatogram, and the specific steps are as follows:

[0009] Step 1: Determine the chromatographic peak to be fitted;

[0010] Step 2: Separate the baseline-separable compounds using the same chromatographic conditions to obtain standard chromatographic peaks, and obtain the standard chromatographic peak shape parameters through the left and right standard deviations of each standard chromatographic peak;

[0011] Step 3: Perform the first fitting on the chromatographic peak to be fitted according to the standard chromatographic peak shape parameters;

[0012] Step 4: Iterate and fit the chromatographic peak of a single region, and then iterate and fit the chromatographic peak of the complete region;

[0013] Step 5: Repeat the above iterative and fitting process of the complete chromatographic peak, and superimpose the fitted peaks of the complete chromatographic peaks, and compare the fitting degree R between the superimposed peaks and the original chromatographic peak data. 2 Calculate, when the fit R 2 When the maximum is reached, stop the iterative calculation;

[0014] Step 6: The single fitted peak shape curve calculated by the last iteration is used for retention time confirmation, integration and peak shape analysis in the region to complete the separation of individual peaks from the multiple peak chromatogram.

[0015] As a preferred technical solution, in step 4, the chromatographic peaks of a single region are iterated and fitted, and the specific method is as follows:

[0016] Step 41: For n chromatographic peaks in a certain region, subtract the chromatographic peaks in other regions and the P1…P in the region from the original chromatographic peak data. n-1 After the superposition peaks, adjust P n The vertex position of and fit;

[0017] Step 42: Subtract the chromatographic peaks in other regions and P1…P in this region from the original chromatographic peak data. n-2 , P n After the superposition peaks, adjust P n-1 The vertex position of the chromatographic peaks is adjusted and fitted until the positions of the n chromatographic peaks in the region are adjusted;

[0018] Step 43: Superimpose the fitted peaks of n chromatographic peaks, compare the superimposed peaks with the original chromatographic peak data, adjust and fit the peak shape parameters according to the comparison difference, and when the local fitting degree R 2When the maximum is reached, the iteration stops.

[0019] As a preferred technical solution, in step 4, the chromatographic peaks in the complete region are iterated and fitted, and the specific method is as follows:

[0020] From left to right or from right to left, iterative fitting is performed from one region to the next region until all regional chromatographic peaks have completed iteration, that is, a complete regional chromatographic peak iteration and fitting is completed.

[0021] As a preferred technical solution, it also includes performing mass spectrometry attribution correction according to the compound retention time range, and the specific method is as follows:

[0022] Step 1: Establish an accurate molecular weight database of all possible compounds in a complex system;

[0023] Step 2: Deconvolute the mass-to-charge ratio in the mass spectrometry data to obtain the corresponding accurate molecular weight; match it with the theoretical molecular weight of the corresponding structural feature in the established database. If the deviation between the actual molecular weight and the theoretical molecular weight is less than 20ppm, the first assignment is completed;

[0024] Step 3: Perform a second assignment based on the retention time distribution range of each group of compounds in the chromatographic fitting results. If the structural characteristics of the compound are consistent with the retention time, the assignment is confirmed; if the structural characteristics of the compound are inconsistent with the retention time, the assignment is wrong, and the wrongly assigned mass spectrometric signal is re-deconvoluted according to different charge numbers. The structural characteristics of the obtained accurate molecular weight are matched with the corresponding retention time in the chromatographic fitting until all possible matches are completed. All assignments are confirmed and the mass spectrometric signals that cannot be assigned, i.e., false positive signals, are eliminated.

[0025] As a preferred technical solution, it also includes attribution correction of the real / dropped sulfate ester group, and the specific method is as follows:

[0026] The sulfate group true / dropped attribution correction was performed based on the retention time difference between the undersulfated compounds formed by the sulfate group dropping in the mass spectrometer ion source and the true undersulfated compounds and the oversulfated compounds.

[0027] As a preferred technical solution, the chromatographic analysis data is a chromatogram obtained by separation and analysis of complex compounds using chromatography, electrophoresis or other separation techniques, or a chromatogram obtained by electrophoresis conversion.

[0028] As a preferred technical solution, the mass spectrometry analysis data is the mass-to-charge ratio, kurtosis, intensity, isotope signal, total ion current (TIC) obtained when chromatography, capillary electrophoresis or other separation techniques are combined with mass spectrometry, or the compound composition data obtained by searching the above data in a database.

[0029] As a preferred technical solution, in step 1, the chromatographic peak to be fitted is determined by the first-order derivative and the second-order derivative of the chromatographic peak, and the chromatographic peak includes a normal peak, a shoulder peak and a hidden peak.

[0030] The advantages of the present invention are:

[0031] 1. The data processing method for chromatography-mass spectrometry analysis of the present invention accurately determines the composition ratio of multiple compounds in a complex sample and the structural composition of each component by automatically processing the chromatography-mass spectrometry analysis data. Therefore, it is no longer necessary for analysts to perform the troublesome operation of confirming or comparing the database search results themselves, and the burden of analysts engaged in identification operations can be greatly reduced.

[0032] 2. The present invention can perform relative quantitative analysis and mass spectrometry qualitative analysis based on the chromatogram fitting peaks, and can confirm the true compound composition in each test sample, thereby further exploring the mechanism, structure-activity relationship, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. The accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0034] The present invention will be further described below in conjunction with the accompanying drawings and embodiments:

[0035] Figure 1 A flow chart of a method for processing data for chromatography-mass spectrometry analysis of the present invention;

[0036] Figure 2 It is a characteristic schematic diagram of three types of fitting peaks of the present invention;

[0037] Figure 3 A schematic diagram of the retention time range attribution correction of the compounds of the present invention;

[0038] Figure 4 This is a schematic diagram of the correction of the true / falling attribution of sulfate ester groups in the present invention;

[0039] Figure 5 It is a schematic diagram of an application of the present invention in Example 1 for chromatogram overlapping peak fitting and compound retention time range attribution correction when the compound type is unknown and the standard peak shape is unknown;

[0040] Figure 6 It is a schematic diagram of an application of the present invention in Example 2 for overlapping peak fitting of a chromatogram with clear compound types and standard peak shapes;

[0041] Figure 7It is a schematic diagram of an application of compound retention time range attribution correction and sulfate group true / dropped attribution correction in Example 3 of the present invention. DETAILED DESCRIPTION

[0042] The above scheme is further described below in conjunction with specific examples. It should be understood that these examples are used to illustrate the present invention and are not limited to the scope of the present invention. The implementation conditions adopted in the examples can be further adjusted according to the conditions of the specific manufacturer, and the unspecified implementation conditions are usually the conditions in conventional experiments.

[0043] Example 1

[0044] Experimental purpose: To analyze the oligosaccharide sequence of enoxaparin sodium by high-resolution liquid chromatography-mass spectrometry, so as to evaluate the consistency of enoxaparin sodium samples from different animal sources.

[0045] Experimental method: Two molecular sieve columns with different pore sizes were connected in series for separation, and high-resolution mass spectrometry was used to analyze enoxaparin sodium oligosaccharides. Since enoxaparin sodium oligosaccharides have double bonds at the non-reducing end and have characteristic absorption at 232nm, the UV spectrum at 232nm was analyzed in this experiment. The database of enoxaparin sodium was established by GlycReSoft, and mass spectrometry data was obtained by database retrieval. There were a large number of false positive results in the retrieval results.

[0046] Data processing method (refer to Figure 1 ):

[0047] 1. Chromatogram overlapping peak fitting:

[0048] 1) Determine the chromatographic peaks to be fitted: Determine the chromatographic peaks to be fitted by the first-order derivative and second-order derivative of the chromatographic peaks. Determine the chromatographic peaks to be fitted based on the characteristics of the three types of peaks: normal peaks, shoulder peaks, and hidden peaks. Figure 2 As shown, the normal peak has an obvious peak apex, the first-order derivative of the peak apex is 0, and the first-order derivative on the left side of the peak apex is greater than 0, and on the right side is less than 0; for the shoulder peak, it will cause the first-order derivative to change and an extreme point to appear; for the hidden peak, it is equivalent to forming a shoulder peak on the first-order derivative, and the position of the hidden peak needs to be determined based on the second-order derivative.

[0049] 2) Obtain standard peak shape parameters: Use the same chromatographic conditions to separate baseline-separable compounds to obtain standard chromatographic peak 1 (narrow peak) and standard chromatographic peak 2 (broad peak). By formula 1: The standard chromatographic peaks were fitted to obtain the left and right standard deviations (peak shape parameters) of each standard chromatographic peak. For example, the left and right standard deviations of standard chromatographic peak 1 are σ a1 and σ a2 , the left and right standard deviations of standard chromatographic peak 2 are σ b1 and σb2 , the chromatographic peak is composed of several equally spaced points, the formula Where x, y are the time and response corresponding to each point, h is the peak height of the chromatographic peak, t is the retention time of the chromatographic peak, and σ is the standard deviation of the chromatographic peak (generally half of the peak width at 0.607 times the peak height).

[0050] 3) First fitting: According to the left and right standard deviations of the standard chromatographic peak 2, the first fitting is performed on the chromatographic peak to be fitted by substituting it into the above formula 1.

[0051] 4) Fitting peak iteration:

[0052] a) Single region iteration and fitting: among the n chromatographic peaks in a certain region, the original chromatographic peak data is used to subtract the chromatographic peaks in other regions and the region P1…P n-1 After the superposition peaks, adjust P n The vertex position is fitted; the original chromatographic peak data is used to subtract the chromatographic peaks in other regions and the region P1…P n-2 , P n After the superposition peaks, adjust P n-1 The vertex position of the n chromatographic peaks is adjusted and fitted... and so on. The positions of the n chromatographic peaks in the region are adjusted. Then the fitted peaks of the n chromatographic peaks are superimposed and the superimposed peaks are compared with the original chromatographic peak data (local fitting degree R 2 ), adjust the peak shape parameters and fit according to the contrast gap. When the local fitting degree R 2 When the maximum is reached, the iteration stops.

[0053] b) Iteration and fitting of the complete chromatographic peak: Iteration and fitting are started from the rightmost region (ie, dp2), and the next region is iterated from right to left until all regions are iterated. This is one iteration and fitting of the complete chromatographic peak.

[0054] c) Repeated iteration and fitting: Repeat the above complete chromatographic peak iteration and fitting process, superimpose the fitting peaks, and compare the superimposed peaks with the original chromatographic peak data to determine the fitting degree R 2 Calculate, when the fit R 2 When the maximum is reached, the iterative calculation stops.

[0055] 5) Peak shape quantification: The single fitted peak shape curve calculated in the last iteration is used for quantification and peak shape analysis, thereby completing the process of stripping individual peaks from the multi-peak chromatogram.

[0056] 2. Mass spectrometry data attribution correction (refer to Figure 3 As shown, the retention time range of the compound is corrected):

[0057] According to the fitting results, the retention time distribution range of each compound group was obtained. Next, the structural characteristics and retention time distribution range of the compounds were used to verify the matching oligosaccharide composition in the search results. Thus, the data were divided into two categories, one for correct attribution and the other for incorrect attribution. The data were re-deconvoluted, that is, the structural characteristics of the correct attribution were determined according to the retention time of the composition, and then re-deconvoluted according to the mass-to-charge ratio to obtain the molecular weight corresponding to different charges (1-10 charges), and then matched with the theoretical molecular weight of the corresponding structural characteristics in the established enoxaparin sodium database. If the deviation between the calculated molecular weight and the theoretical molecular weight is less than 20ppm, the attribution is corrected.

[0058] Consistency evaluation (refer to Figure 5 As shown in Figure 2), after the enoxaparin sodium from different animal sources was processed by the above data processing method, the final oligosaccharide composition of each test sample was analyzed to evaluate the consistency of the enoxaparin sodium samples. It can be seen from the PCA analysis diagram that after removing the false positive oligosaccharide data, the enoxaparin sodium from porcine intestine can be successfully distinguished from the enoxaparin sodium samples from other animal sources.

[0059] Example 2

[0060] Purpose of the experiment: There may be other impurity polysaccharides in heparin drugs. High performance liquid chromatography was used to separate different glycosaminoglycans and measure the content of other impurity polysaccharides in the production process of heparin drugs.

[0061] Experimental method: Anion exchange chromatography column was used for separation, and mixed standards of different concentrations (2 in the figure is heparin, which is the main component) were tested. The peak areas of each were obtained by fitting, and the linear relationship of different glycosaminoglycans was obtained.

[0062] Data processing method (refer to Figure 1 and Figure 6 ):

[0063] 1. Determine the chromatographic peaks that need to be fitted.

[0064] 2. Obtain standard peak shape parameters: Test each compound in the overlapping peaks separately, obtain the chromatogram of each compound, fit the standard chromatographic peaks by the formula, and obtain the left and right standard deviations of each standard chromatographic peak.

[0065] 3. First fitting: Substitute the peak shape parameters of each standard chromatographic peak into Formula 1, and fit the chromatographic peaks that need to be fitted accordingly.

[0066] 4. Fitting peak iteration: subtract the superimposed peak shape of individual peaks 2, 3, and 4 obtained by the first fitting from the peak shape of the multiple peaks, and the peak shape of individual peak 1 is obtained by the first iteration; subtract the superimposed peak shape of individual peaks 1, 3, and 4 obtained by the first fitting from the peak shape of the multiple peaks, and the peak shape of individual peak 2 is obtained by the first iteration; ...;

[0067] 5. Repeat iteration and fitting: Repeat the iteration and fitting process of 3) and 4) above, superimpose each fitting peak, and perform R 2 Calculate, when R 2 When the maximum value is reached, the iteration calculation stops.

[0068] 6. Peak shape quantification: The single fitted peak shape curve calculated from the last iteration was used for quantification. The mixed standard at 7 concentrations (0.5, 2, 4, 10, 15, 20, 50 mg / mL) was tested, and the peak areas at different concentrations of each compound were fitted to obtain the linear relationship of each compound.

[0069] 7. Quantification of actual samples: The actual samples were tested and fitted to obtain the corresponding peak areas. Substituting into the linear relationship, the concentration of compound 1 was 6.89 mg / mL, the concentration of compound 2 was 44.39 mg / mL, the concentration of compound 3 was 7.95 mg / mL, and the concentration of compound 4 was 7.43 mg / mL. Therefore, the contents of each component were 10.3%, 66.6%, 11.9% and 11.1%, respectively.

[0070] Example 3

[0071] Experimental purpose: To analyze the oligosaccharide sequences of carrageenan acid hydrolysis products at different times by high-resolution liquid chromatography-mass spectrometry (HPLC-MS / MS) and to explore the acid hydrolysis rules of carrageenan.

[0072] Experimental method: Separation was performed using two molecular sieve columns with different pore sizes in series, and oligosaccharide analysis was performed on the acid hydrolysis products of carrageenan at different times by high-resolution mass spectrometry. Since carrageenan has no characteristic ultraviolet absorption, ultraviolet chromatogram analysis was not performed in this experiment. GlycReSoft was used to establish a database of carrageenan oligosaccharides, and mass spectrometry data was obtained by database retrieval. There were also a large number of false positive results in the retrieval results.

[0073] Data processing method:

[0074] 1. Mass spectrometry data attribution correction (refer to Figure 3As shown, the attribution correction for the compound retention time range): the mass spectrum was manually analyzed to determine the time distribution range of each compound group. Next, the attribution of the matched oligosaccharide composition was verified using the relationship between the structural characteristics of the compound and the retention time. The data was thus divided into two categories, one with correct attribution and the other with incorrect attribution. The data was re-deconvoluted, that is, the structural characteristics of the correct attribution were determined according to the retention time of the composition, and then re-deconvoluted according to the mass-to-charge ratio to obtain the molecular weight corresponding to different charges, and then matched with the theoretical molecular weight of the corresponding structural characteristics in the established carrageenan database. If the deviation between the calculated molecular weight and the theoretical molecular weight is less than 20ppm, the attribution is corrected.

[0075] 2. Sulfate ester real / fallback correction (refer to Figure 4 ): First, according to the characteristics of the separation method (size exclusion chromatography, the peak with the largest molecular weight emerges first), within the range of a single structural feature, the fragment ion peak generated by the loss of the sulfonic acid group of the highly sulfated oligosaccharide has the same molecular weight as the true low-sulfated oligosaccharide, but the retention time is different; the retention time of the low-sulfated oligosaccharide formed by the drop of the sulfate ester group is close to that of the highly sulfated oligosaccharide (<0.2min), while the retention time of the true low-sulfated oligosaccharide and the highly sulfated oligosaccharide is far apart. Moreover, within the range of each structural feature, the composition of oligosaccharides with the same glycosyl composition but different degrees of sulfation is linearly related.

[0076] Acid hydrolysis law exploration, reference Figure 7 , where the dotted lines in the EIC diagram represent even-numbered oligosaccharides and the solid lines represent odd-numbered oligosaccharides:

[0077] Under the same termination pH condition (pH=1), carrageenan was acid-hydrolyzed for different times (1h, 3h, 6h, 12h), and it was found that the acid-hydrolyzed products at different acid-hydrolyzed times were mainly even-numbered sugars. The degree of polymerization of the acid-hydrolyzed products after acid-hydrolyzed for 1h was dp2-dp38, and they were mainly composed of oligosaccharides with the highest degree of sulfate; the acid-hydrolyzed products after acid-hydrolyzed for 3h, 6h and 12h all contained multiple real low-sulfated oligosaccharides.

[0078] Under the same acid hydrolysis time (3h) and different termination pH conditions (pH=1, 7, 9 and 12), it can be found that the acid hydrolysis products at pH=1 and 7 are mainly even-numbered sugars; the acid hydrolysis products at pH=9 contain both even-numbered sugars and odd-numbered sugars; the acid hydrolysis products at pH=12 are mainly odd-numbered sugars.

[0079] Therefore, it can be found that under the same pH termination conditions, as the acid hydrolysis time increases, the polymerization degree of carrageenan oligosaccharides gradually decreases; and under the same acid hydrolysis time, as the acidic conditions transition to alkaline conditions, the content of carrageenan odd-numbered oligosaccharides gradually increases.

[0080] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical concept disclosed by the present invention shall still be covered by the claims of the present invention.

Claims

1. A data processing method for chromatography-mass spectrometry analysis, which determines the composition of multiple compounds in a complex sample by analyzing and processing chromatography data and mass spectrometry data, characterized in that: Including overlapping peak fitting of the chromatogram, the specific steps are as follows: Step 1: Determine the chromatographic peak to be fitted; Step 2: Separate the baseline-separable compounds using the same chromatographic conditions to obtain standard chromatographic peaks, using formula 1: The standard chromatographic peaks were fitted, and the left and right standard deviations of each standard chromatographic peak were obtained, thereby obtaining the peak shape parameters of the standard chromatographic peaks, where x and y are the time and response corresponding to each point, h is the peak height of the standard chromatographic peak, t is the retention time of the standard chromatographic peak, and σ is the standard deviation of the standard chromatographic peak; Step 3: According to the peak shape parameters of the standard chromatographic peak, the above formula 1 is used to perform the first fitting on the chromatographic peak to be fitted; Step 4: Iterate and fit the chromatographic peak of a single region, and then iterate and fit the chromatographic peak of the complete region; Step 5: Repeat the above iteration and fitting process of the complete region chromatographic peak, and superimpose the fitting peaks of the complete chromatographic peaks, and calculate the fitting degree R2 of the superimposed peaks and the original chromatographic peak data. When the fitting degree R2 reaches the maximum, stop the iterative calculation; Step 6: The single fitted peak shape curve calculated from the last iteration is used for retention time confirmation, integration, and peak shape analysis to complete the separation of individual peaks from the multiple peak chromatogram.

2. The method for processing data for chromatography mass spectrometry analysis according to claim 1, characterized in that: In step 4, the chromatographic peaks of a single region are iterated and fitted. The specific method is as follows: Step 41: For n chromatographic peaks in a single region, after subtracting the chromatographic peaks in other regions and the superimposed peaks of P1…Pn-1 in the single region from the original chromatographic peak data, the vertex position of Pn is adjusted and fitted; Step 42: after subtracting the chromatographic peaks in other regions and the superimposed peaks of P1 ... Pn-2 and Pn in the single region from the original chromatographic peak data, the vertex position of Pn-1 is adjusted and fitted until the positions of the n chromatographic peaks in the single region are adjusted; Step 43: Superimpose the fitted peaks of n chromatographic peaks, compare the superimposed peaks with the original chromatographic peak data, adjust and fit the peak shape parameters according to the comparison difference, and stop the iteration when the local fitting degree R2 of the single area reaches the maximum.

3. The method for processing data for chromatography mass spectrometry analysis according to claim 1, characterized in that: In step 4, the chromatographic peaks in the complete region are iterated and fitted. The specific method is as follows: From left to right or from right to left, iterative fitting is performed from one region to the next region until all regional chromatographic peaks have completed iteration, that is, a complete regional chromatographic peak iteration and fitting is completed.

4. The method for processing data for chromatography-mass spectrometry analysis according to claim 1, characterized in that: It also includes mass spectrometry attribution correction based on the compound retention time range, and the specific method is as follows: Step 1: Establish an accurate molecular weight database of all possible compounds in a complex system; Step 2: Deconvolute the mass-to-charge ratio in the mass spectrometry data to obtain the corresponding accurate molecular weight; match it with the theoretical molecular weight of the corresponding structural feature in the established database. If the deviation between the actual molecular weight and the theoretical molecular weight is less than 20ppm, the first assignment is completed; Step 3: Perform a second assignment based on the retention time distribution range of each group of compounds in the chromatographic fitting results. If the structural characteristics of the compound are consistent with the retention time, the assignment is confirmed; if the structural characteristics of the compound are inconsistent with the retention time, the assignment is wrong, and the wrongly assigned mass spectrometric signal is re-deconvoluted according to different charge numbers. The structural characteristics of the obtained accurate molecular weight are matched with the corresponding retention time in the chromatographic fitting until all possible matches are completed. All assignments are confirmed and the mass spectrometric signals that cannot be assigned, i.e., false positive signals, are eliminated.

5. The method for processing data for chromatography-mass spectrometry analysis according to claim 1, characterized in that: It also includes attribution correction for sulfate ester true / drop, the specific method is as follows: The sulfate group true / dropped attribution correction was performed based on the retention time difference between the undersulfated compounds formed by the sulfate group dropping in the mass spectrometer ion source and the true undersulfated compounds and the oversulfated compounds.

6. The method for processing data for chromatography mass spectrometry analysis according to claim 1, characterized in that: The chromatographic data is a chromatogram obtained by separation and analysis of complex compounds using chromatography, electrophoresis or other separation techniques, or a chromatogram obtained by electrophoresis conversion.

7. The method for processing data for chromatography mass spectrometry analysis according to claim 1, characterized in that: The mass spectrometry data refers to the mass-to-charge ratio, kurtosis, intensity, isotope signal, total ion current obtained when chromatography, capillary electrophoresis or other separation techniques are combined with mass spectrometry, or the compound composition data obtained through database retrieval.

8. The method for processing data for chromatography mass spectrometry analysis according to claim 1, characterized in that: In step 1, the chromatographic peak to be fitted is determined by the first-order derivative and the second-order derivative of the chromatographic peak, and the chromatographic peak includes a normal peak, a shoulder peak and a hidden peak.

Citation Information

Patent Citations

  • Precisive measurement for parameter of chromatography spike and area of overlapped peak

    CN1712955A

  • Detecting peaks in two-dimensional signals

    US20100283785A1